RNA compositions for delivery of insulinotropic agent

By delivering incretin mimics via polynucleotide precursors, the problem of limited supply of incretin agents is solved, achieving broader accessibility and fewer side effects, making it suitable for the treatment of obesity and related diseases.

CN121844046APending Publication Date: 2026-04-10C·米库尔卡
View PDF 69 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
C·米库尔卡
Filing Date
2024-09-11
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The limited production and supply of existing incretin agents means that patients with obesity and related diseases such as type 2 diabetes and non-alcoholic fatty liver disease cannot receive effective treatment. Furthermore, existing peptide-based products are costly and have significant gastrointestinal side effects, leading to a high rate of treatment discontinuation.

Method used

By using polynucleotide precursors to deliver incretin mimics, including GLP1 and GIP receptor agonists, pharmacokinetic characteristics are optimized, injection volume and side effects are reduced, and accessibility and therapeutic efficacy are improved by employing half-life-extending components such as albumin-binding domains and VHH antibodies.

Benefits of technology

It provides broader treatment accessibility, reduces treatment discontinuation rates, improves pharmacokinetics, achieves longer duration of action and lower dosing frequency, and is suitable for the treatment of obesity and its sequelae such as type 2 diabetes, cardiovascular disease, kidney disease, NASH and NAFLD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844046A_ABST
    Figure CN121844046A_ABST
Patent Text Reader

Abstract

The present disclosure provides compositions (e.g., pharmaceutical compositions) and related techniques (e.g., components thereof and / or methods related thereto) for delivering an insulinotropic agent. The present disclosure provides, inter alia, methods of treatment of various diseases using polyribonucleotides encoding an insulinotropic agent.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 662,890, filed June 21, 2024, and PCT Application No. PCT / IB2023 / 059007, filed September 11, 2023, the entire contents of which are incorporated herein by reference. Background Technology

[0002] Obesity is the most prevalent chronic disease worldwide, currently affecting approximately 650 million adults. Obesity is considered a starting point and key contributing factor to prediabetes, type 2 diabetes (T2D, and its complications), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), cardiovascular disease and kidney disease, as well as premature death. It is estimated that by 2030, obesity (BMI > 30 kg / m²) will become a major health concern. 2 The number of people with this condition will exceed one billion, of whom approximately 10% will suffer from severe type III obesity (BMI > 40 kg / m²). 2 Half of all obese people live in just nine countries: the United States, China, India, Brazil, Mexico, Russia, Egypt, Germany, and Turkey. Furthermore, childhood obesity is rising dramatically worldwide. Type 2 diabetes mellitus (T2D), NAFLD, NASH, cardiovascular disease, and kidney disease are also prevalent and unrelated to obesity. There is a need to develop further therapies for the treatment and / or prevention of obesity and other related diseases. Summary of the Invention

[0003] Treatment of obesity and type 2 diabetes mellitus (T2D) with incretins and incretin mimics, particularly GLP1 and GIP receptor agonists or combinations thereof, has shown significant benefits to individuals suffering from these diseases. However, the availability of this new class of active agents is limited due to production and cost constraints, leaving millions without access to the necessary treatment. This disclosure recognizes these limitations and provides a novel therapeutic approach to eliminate the need for the self-production of incretins by delivering incretin mimics using polynucleotide precursors of the incretins. In particular, this disclosure provides the polynucleotide precursors of the incretins as molecular entities, their production, formulation, and administration for the treatment of obesity and its sequelae, including T2D, early type 1 diabetes mellitus (T1D), cardiovascular disease, kidney disease, NASH, and NAFLD. This disclosure also provides methods for using these agents to treat diseases, including obesity-independent T2D, early type 1 diabetes mellitus (T1D), cardiovascular disease, kidney disease, NASH, and NAFLD. This disclosure further provides methods for using these agents to treat sequelae of NASH, including liver fibrosis and cirrhosis.

[0004] The present disclosure also recognizes that such a therapeutic approach, i.e., delivery of an incretin polynucleotide precursor, presents additional benefits over current therapies, including but not limited to, making the product more widely accessible to the obese population who cannot access the existing products due to limited supply, high price, lack of health insurance; requiring lower injection volume of the formulation compared to the commercially available peptide-based products; lower patient discontinuation rates due to factors such as gastrointestinal side effects; and improved properties such as optimized pharmacokinetic profile.

[0005] More widely accessible to the obese population who cannot access the current products due to limited supply, high price, lack of health insurance; lower injection volume of the formulation compared to the commercially available peptide-based products; lower frequency of patient discontinuation due to factors such as gastrointestinal side effects; and improved properties such as improved pharmacokinetic profile. In some embodiments, the improved pharmacokinetic profile has the advantage of reducing the frequency of dosing with a therapeutic agent that has a longer duration of action.

[0006] In one aspect, the present disclosure provides a composition comprising a polynucleotide encoding an incretin agent. In some embodiments, the incretin agent is a GLP1 receptor agonist. In some embodiments, the incretin agent is a GIP receptor agonist. In some embodiments, the incretin agent is a GLP1 / GIP dual receptor agonist. In some embodiments, the incretin agent is a GLP1 / GCG dual receptor agonist. In some embodiments, the incretin agent is a GLP1 / GIP / GCG triple receptor agonist.

[0007] In some embodiments, the incretin agent comprises an incretin peptide having an amino acid sequence as set forth in any one of SEQ ID NOs: 5-7, 63-64, 69-70, and 74-75. In some embodiments, the incretin agent comprises an incretin peptide having an amino acid sequence as set forth in any one of SEQ ID NOs: 8-9, 62, and 72. In some embodiments, the incretin agent comprises an incretin peptide having an amino acid sequence as set forth in SEQ ID NO: 11. In some embodiments, the incretin agent comprises an incretin peptide having an amino acid sequence as set forth in any one of SEQ ID NOs: 12-14. In some embodiments, the incretin agent comprises an incretin peptide having an amino acid sequence as set forth in SEQ ID NO: 15.

[0008] In some embodiments, the incretin peptide is optionally fused to a signal peptide, optionally through the N-terminus of the incretin peptide, through a linker peptide. In some embodiments, the signal peptide has an amino acid sequence as set forth in any one of SEQ ID NOs: 16-39 and 65-67. In some embodiments, the signal peptide has an amino acid sequence as set forth in any one of SEQ ID NOs: 16-21 and 65-67. In some embodiments, the signal peptide has an amino acid sequence as set forth in SEQ ID NO: 17. In some embodiments, the signal peptide has an amino acid sequence as set forth in SEQ ID NO: 65. In some embodiments, the signal peptide has an amino acid sequence as set forth in SEQ ID NO: 66. In some embodiments, the incretin agent comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 41-45, 52-61, and 108-152.

[0009] In some embodiments, the incretin agent comprises an incretin peptide optionally fused to one or more additional incretin peptides, optionally through one or more linker peptides. In some embodiments, the one or more linker peptides comprise an amino acid sequence as set forth in any one of SEQ ID NOs: 1-5, 68, or 156. In some embodiments, the incretin agent comprises an incretin peptide fused to two or more incretin peptides. In some embodiments, the incretin agent comprises at least one GLP1 receptor agonist and at least one GIP receptor agonist. In some embodiments, the incretin agent comprises at least two GLP1 receptor agonists. In some embodiments, the incretin agent comprises at least two GIP receptor agonists.

[0010] In some embodiments, the incretin agent comprises one or more furin cleavage sites. In some embodiments, the one or more furin cleavage sites are located between adjacent incretin peptides. In some embodiments, the one or more furin cleavage sites comprise an amino acid sequence as set forth in SEQ ID NO: 153. In some embodiments, the incretin agent comprises one or more units each comprising, from N-terminus to C-terminus: GLP1 receptor agonist-linker peptide-furin cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 76, 77, 78, 79, 80, 81), two units (e.g., SEQ ID NO: 82); or four units (e.g., SEQ ID NO: 83). In some embodiments, the incretin agent comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 76-83, 94-97, 102-107.

[0011] In some embodiments, the incretin agent comprises a half-life extension moiety. In some embodiments, the half-life extension moiety comprises albumin (e.g., human serum albumin). In some embodiments, the human serum albumin comprises an amino acid sequence that is at least 90%, 95%, or 99% identical to SEQ ID NO: 159. In some embodiments, the human serum albumin comprises the amino acid sequence set forth in SEQ ID NO: 159. In some embodiments, the incretin agent comprises albumin (e.g., human serum albumin) fused to one or more units each comprising, from N-terminus to C-terminus: (i) a GLP1 receptor agonist-linker peptide (e.g., SEQ ID NO: 98); (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 100); (iii) a GLP1 receptor agonist-linker peptide- furin cleavage site (e.g., SEQ ID NO: 102); or (iv) a GLP1 receptor agonist-linker peptide- furin cleavage site-GIP receptor agonist, e.g., wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 104), two units (e.g., SEQ ID NO: 106), or four units (e.g., SEQ ID NO: 107). In some embodiments, the incretin agent comprises an amino acid sequence set forth in any of SEQ ID NOs: 98, 100, 102, 104, 106, 107, or any combination thereof.

[0012] In some embodiments, the half-life extension moiety comprises an albumin binding domain (ABD). In some embodiments, the ABD is derived from protein G of Streptococcus strain GI48 and / or protein PAB of Finegoldia magna, such as ABD035 and SA21. In some embodiments, the half-life extension moiety comprises an ABD that binds to domain II of human serum albumin and does not overlap or interfere with binding to the FcRn binding site on albumin. In some embodiments, the half-life extension moiety comprises ABD Con. In some embodiments, the half-life extension moiety comprises an albumin binding domain (ABD) derived from the bacterial protein Sso7d from the hyperthermophilic archaeon Sulfolobus solfataricus, such as M11.12 and M18.2.5. In some embodiments, the half-life extension moiety comprises a DARPin that binds albumin.

[0013] In some embodiments, the ABD comprises an immunoglobulin domain or fragment thereof that binds albumin. In some embodiments, the ABD comprises a fully human domain antibody (dAb) that binds albumin, such as an AlbudAb. In some embodiments, the ABD comprises a Fab that binds albumin, such as a dsFv CA645. In some embodiments, the ABD comprises a heavy chain-only (VHH) antibody that binds albumin, such as a nanobody. In some embodiments, the VHH antibody comprises a VHH domain having complementarity determining region (CDR) sequences HCDR1, HCDR2, and / or HCDR3 as set forth in SEQ ID NO: 191 (GFTLDYYA), SEQ ID NO: 192 (IASSGGST), and / or SEQ ID NO: 193 (AAAVLECRTVVRGYDY), respectively. In some embodiments, the VHH antibody comprises an amino acid sequence that is at least 90%, 95%, or 99% identical to SEQ ID NO: 154. In some embodiments, the VHH antibody comprises an amino acid sequence as set forth in SEQ ID NO: 154. In some embodiments, the incretin agent comprises a VHH antibody that binds albumin fused to a unit comprising, from N-terminus to C-terminus: (i) a GLP1-linker peptide (e.g., SEQ ID NO: 99); (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 101); or (iii) a GLP1 receptor agonist-linker peptide- furin-GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 103 or 105). In some embodiments, the incretin agent comprises an amino acid sequence as set forth in any one of SEQ ID NO: 99, 101, 103, 105.

[0014] In some embodiments, the half-life extension moiety does not comprise an Fc domain, such as from a human IgG, optionally from a human IgG1, IgG2, IgG3, or IgG4. In some embodiments, the half-life extension moiety comprises an Fc domain, such as from a human IgG, optionally from a human IgG1, IgG2, IgG3, or IgG4. In some embodiments, the human IgG is a human IgG4. In some embodiments, the incretin agent comprises an IgG4 Fc domain fused to a unit comprising, from N-terminus to C-terminus: (i) a GLP1 receptor agonist-linker peptide (e.g., SEQ ID NO: 10, 89, 90, 91); (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 92, 93); or (iii) a GLP1 receptor agonist-linker peptide- furin-GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 94, 95, 96, 97). In some embodiments, the IgG4 Fc domain comprises an amino acid sequence that is at least 90%, 95%, or 99% identical to SEQ ID NO: 155. In some embodiments, the IgG4 Fc domain comprises an amino acid sequence as set forth in SEQ ID NO: 155. In some embodiments, the incretin agent comprises an amino acid sequence as set forth in any one of SEQ ID NO: 10 and 89-97.

[0015] In some embodiments, the Fc domain comprises one or more mutations in one or both Fc constant domains that increase the half-life of the incretin agent and / or induce dimerization. In some embodiments, the one or more mutations comprise one or more mutations in the CH3 domain. In some embodiments, the one or more mutations that induce dimerization comprise: (i) Y349C, T366S, L368A, and / or Y407V (according to EU numbering); or (ii) S354C and / or T366W (according to EU numbering). In some embodiments, the one or more mutations comprise Y349C, T366S, L368A, and Y407V (“FcKIH-b”, according to EU numbering); or S354C and T366W (“FcKIH-a”, according to EU numbering). In some embodiments, the incretin agent comprises a first polypeptide chain and a second polypeptide chain, wherein the first polypeptide chain comprises an incretin peptide fused to a first Fc domain, wherein the first Fc domain comprises the mutations Y349C, T366S, L368A, and Y407V (“FcKIH-b”, according to EU numbering), and wherein the second polypeptide chain comprises an incretin peptide fused to a second Fc domain, wherein the second Fc domain comprises the mutations S354C and T366W (“FcKIH-a”, according to EU numbering).

[0016] In some embodiments, the one or more mutations that increase the half-life of the incretin agent comprise M428L and N434S (“LS”, according to EU numbering). In some embodiments, the incretin agent comprises an Fc domain with Fc KIH-a mutations on a first polypeptide chain and an Fc domain with Fc KIH-b mutations on a second polypeptide chain, wherein the Fc domain on each polypeptide chain is independently fused to one or more units comprising, from N-terminus to C-terminus: (i) a GLP1 receptor agonist- linker peptide (e.g., SEQ ID NO: 84, 85, 86, 87); or (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 88). In some embodiments, the incretin agent comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 84-88.

[0017] In some embodiments, the Fc domain comprises one or more mutations that ablate the effector activity (e.g., binding to Fcy receptor or Clq) of the Fc domain. In some embodiments, the one or more mutations that ablate the effector activity (e.g., binding to Fcy receptor or Clq) of the Fc domain comprise the following mutations: L234S, L235T, and G236R (“STR”, according to EU numbering). In some embodiments, the one or more mutations that ablate the effector activity (e.g., binding to Fcy receptor or Clq) of the Fc domain comprise the following mutations: L234A and L235A (“LALA”, according to EU numbering). In some embodiments, the one or more mutations that ablate the effector activity (e.g., binding to Fcy receptor or Clq) of the Fc domain comprise the following mutations: L234A / L235A / P329G (“LALAPG”, according to EU numbering).

[0018] In some embodiments, the half-life extension moiety comprises an albumin-binding VNAR. In some embodiments, the half-life extension moiety comprises an XTEN sequence.

[0019] In some embodiments, the polyribonucleotide has a ribonucleic acid sequence that is at least 90% identical to any one of SEQ ID NOs: 177-185 and 224-256. In some embodiments, the polyribonucleotide has a ribonucleic acid sequence as set forth in any one of SEQ ID NOs: 177-185 and 224-256.

[0020] In some embodiments, the polyribonucleotide comprises at least one non-coding sequence component that enhances RNA stability and / or translation efficiency. In some embodiments, the at least one non-coding sequence component comprises a 5' cap structure, a 5' UTR, a 3' UTR, and / or a polyA tail.

[0021] In some embodiments, the polynucleotide comprises, in the 5' to 3' direction: a. a 5' UTR; b. a signal peptide coding sequence; c. an enteroincretin peptide coding sequence; d. a 3' UTR; and e. a polyA tail.

[0022] In some embodiments, the polynucleotide comprises, in the 5' to 3' direction: (1) a. a 5' UTR; b. a signal peptide coding sequence; c. an enteroincretin peptide coding sequence; d. a linker peptide coding sequence; e. a half-life extension moiety coding sequence; f. a 3' UTR; and g. a polyA tail; or (2) a. a 5' UTR; b. a signal peptide coding sequence; c. a half-life extension moiety coding sequence; d. a linker peptide coding sequence; e. an enteroincretin peptide coding sequence; f. a 3' UTR; and g. a polyA tail.

[0023] In some embodiments, the enteroincretin peptide is encoded by a coding sequence that is codon-optimized and / or has an increased G / C content compared to a wild-type coding sequence, wherein the codon-optimization and / or the increase in G / C content does not change the sequence of the encoded amino acid sequence.

[0024] In some embodiments, the polynucleotide comprises at least one modified ribonucleotide. In some embodiments, the polynucleotide comprises a modified nucleoside in place of a uridine. In some embodiments, the polynucleotide comprises a modified nucleoside in place of each uridine. In some embodiments, the modified nucleoside is selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U). In some embodiments, the modified nucleoside is N1-methyl-pseudouridine (m1ψ).

[0025] In some embodiments, the polynucleotide comprises a 5' cap structure. In some embodiments, the polynucleotide comprises a 5' UTR. In some embodiments, the polynucleotide comprises a 3' UTR. In some embodiments, the polynucleotide comprises a polyA tail. In some embodiments, the polyA tail comprises at least 100 nucleotides. In some embodiments, the polynucleotide is an mRNA.

[0026] In some embodiments, the polynucleotide is formulated as a liquid, as a solid, or a combination thereof. In some embodiments, the polynucleotide is formulated for injection. In some embodiments, the polynucleotide is formulated for intraperitoneal or intravenous administration.

[0027] In some embodiments, the polyribonucleotide is formulated as or is to be formulated as a lipid particle. In some embodiments, the polyribonucleotide is formulated as or is to be formulated as a lipid nanoparticle. In some embodiments, the polyribonucleotide is encapsulated within a lipid nanoparticle. In some embodiments, the lipid nanoparticle is a pancreas- and / or intestine-targeting lipid nanoparticle. In some embodiments, the lipid nanoparticle is a cationic lipid nanoparticle.

[0028] In some embodiments, the lipids forming the lipid nanoparticle comprise a. a polymer conjugated lipid; b. a cationic lipid; and c. a neutral lipid. In some embodiments, the polymer conjugated lipid is a PEG conjugated lipid. In some embodiments, the cationic lipid is an ionizable lipidoid material.

[0029] In some embodiments, the cationic lipid has one of the following structures: X-1 X-2 X-3 X-4.

[0030] In some embodiments, the neutral lipid comprises a helper lipid such as 1,2- distearoyl-sn-glycero-3-phosphocholine (DSPC) and / or cholesterol.

[0031] In some embodiments, the cationic lipid is selected from cationic lipids X-2, X-3, or X-4, and the neutral lipid comprises a helper lipid (such as DOTAP, DOPE, or PS) and cholesterol.

[0032] In some embodiments, the polymer conjugated lipid is C14-PEG2000.

[0033] In some embodiments, the lipid nanoparticle comprises: i) about 30 mol% to about 50 mol% of a cationic lipid; ii) about 1 mol% to 5 mol% of a PEG conjugated lipid; iii) about 30 mol% to about 50 mol% of a helper lipid; and iv) about 20 mol% to about 40 mol% of cholesterol.

[0034] In some embodiments, the lipid nanoparticle comprises about 35 mol% of a cationic lipid; about 40 mol% of a helper lipid, about 22.5 mol% of cholesterol, and about 2.5 mol% of a PEG conjugated lipid.

[0035] In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipids X-2, X-3, or X-4, about 40 mol% of DOTAP, DOPE, or PS, about 22.5 mol% of cholesterol, and about 2.5 mol% of C14-PEG2000.

[0036] In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of PS; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of PS; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of PS; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

[0037] In some embodiments, the lipid nanoparticles are formulated for intraperitoneal (ip) delivery. In some embodiments, the lipid nanoparticles have an average size of about 50-150 nm. In some embodiments, the composition comprises one or more pharmaceutically acceptable carriers, diluents, and / or excipients. In some embodiments, the composition further comprises a cryoprotectant. In some embodiments, the cryoprotectant is sucrose. In some embodiments, the composition comprises an aqueous buffer solution. In some embodiments, the aqueous buffer solution comprises sodium ions.

[0038] In another aspect, this disclosure provides a method for treating a disease state in an individual in need, comprising administering to the individual a therapeutically effective amount of a composition comprising one or more polynucleotides as described herein.

[0039] In some embodiments, the method further comprises administering one or more DPP-4 inhibitors. In some embodiments, one or more DPP-4 inhibitors and the composition are administered simultaneously. In some embodiments, one or more DPP-4 inhibitors and the composition are administered sequentially. In some embodiments, one or more DPP-4 inhibitors are administered before the composition. In some embodiments, one or more DPP-4 inhibitors are administered after the composition. In some embodiments, one or more DPP-4 inhibitors comprise sitagliptin, vildagliptin, saxagliptin, linagliptin, gemigliptin, anagliptin, teneligliptin, alogliptin, trelagliptin, omarigliptin, evogliptin, gosogliptin, dutogliptin, neogliptin, retagliptin, denagliptin, cofrogliptin, fotagliptin, prusogliptin, berberine, or any combination thereof. In some embodiments, one or more DPP-4 inhibitors are administered orally.

[0040] In some embodiments, the disease state is obesity or an obesity-related condition. In some embodiments, an obesity-related condition is prediabetes, type 2 diabetes (T2D), early type 1 diabetes (T1D), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), cardiovascular (CV) disease, kidney disease, or an increased risk of premature death. In some embodiments, cardiovascular (CV) disease includes major cardiovascular events (MACE), including CV death, non-fatal myocardial infarction, non-fatal stroke, and / or heart failure with preserved ejection fraction (HFpEF).

[0041] In some embodiments, the method improves the individual's weight management. In some embodiments, the method reduces the individual's weight gain or induces weight loss. In some embodiments, the disease state is diabetes. In some embodiments, the method improves the individual's glycemic control. In some embodiments, the method lowers the individual's HbA1c. In some embodiments, the diabetes is prediabetes, type 2 diabetes (T2D), or early type 1 diabetes (T1D).

[0042] In some embodiments, the disease state is cardiovascular (CV) disease. In some embodiments, cardiovascular disease includes major cardiovascular events (MACE), including CV death, nonfatal myocardial infarction, nonfatal stroke, and / or heart failure with preserved ejection fraction (HfpEF). In some embodiments, the method improves the individual's blood pressure and / or blood lipids.

[0043] In some embodiments, the disease state is kidney disease. In some embodiments, the disease state is non-alcoholic fatty liver disease (NAFLD). In some embodiments, the disease state is non-alcoholic steatohepatitis (NASH) and optionally its sequelae, liver fibrosis and cirrhosis.

[0044] In some embodiments, administering the composition to an individual comprises administering one or more doses of the composition to the individual. In some embodiments, one or more doses of the composition are administered to the individual daily, every other day, or weekly. In some embodiments, one or more doses of the composition are administered to the individual at a frequency of less than once per week. In some embodiments, one or more doses of the composition are administered to the individual every 2, 3, or 4 weeks. In some embodiments, the composition is administered by injection. In some embodiments, the composition is administered subcutaneously, intravenously, intramuscularly, or intraperitoneally. In some embodiments, the composition is administered intraperitoneally. In some embodiments, the composition is administered noninvasively (e.g., orally or nasally). In some embodiments, administration of the composition results in the expression of an incretin in the individual. In some embodiments, the composition is administered in a volume of less than 0.5 mL.

[0045] In another aspect, this disclosure provides the use of any composition comprising one or more polynucleotides described herein for treating a disease state in an individual in need.

[0046] In another aspect, this disclosure provides a method for producing an incretin, comprising administering to cells a composition comprising the polynucleotides described herein, such that the cells express and secrete the incretin.

[0047] In another aspect, this disclosure provides an incretin agent comprising an incretin peptide fused to a signal peptide. In some embodiments, the incretin peptide is fused to the signal peptide via the N-terminus of the incretin peptide, optionally via a linker peptide. In some embodiments, the signal peptide has an amino acid sequence as shown in any of SEQ ID NO: 16-39 and 65-67. In some embodiments, the signal peptide has an amino acid sequence as shown in any of SEQ ID NO: 16-21 and 65-67. In some embodiments, the signal peptide has an amino acid sequence as shown in SEQ ID NO: 17. In some embodiments, the signal peptide has an amino acid sequence as shown in SEQ ID NO: 65. In some embodiments, the signal peptide has an amino acid sequence as shown in SEQ ID NO: 66. In some embodiments, the incretin agent comprises an incretin peptide fused to a signal peptide, comprising an amino acid sequence as shown in any of SEQ ID NO: 41-45, 52-61, and 108-152.

[0048] In some embodiments, the incretin agent comprises an incretin peptide optionally fused to one or more additional incretin peptides via one or more linker peptides. In some embodiments, the one or more linker peptides comprise an amino acid sequence as shown in any of SEQ ID NO: 1-5, 68, or 156. In some embodiments, the incretin agent comprises an incretin peptide fused to two or more incretin peptides.

[0049] In some embodiments, the incretin agent comprises at least one GLP1 receptor agonist and at least one GIP receptor agonist. In some embodiments, the incretin agent comprises at least two GLP1 receptor agonists. In some embodiments, the incretin agent comprises at least two GIP receptor agonists. In some embodiments, the incretin agent comprises one or more furin cleavage sites. In some embodiments, the one or more furin cleavage sites are located between adjacent incretin peptides. In some embodiments, the one or more furin cleavage sites comprise an amino acid sequence as shown in SEQ ID NO: 153. In some embodiments, the incretin agent comprises one or more units, each of which comprises, from the N-terminus to the C-terminus: GLP1 receptor agonist-linking peptide-furin cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 76, 77, 78, 79, 80, 81), two units (e.g., SEQ ID NO: 82), or four units (e.g., SEQ ID NO: 83). In some embodiments, the incretin agent comprises an amino acid sequence as shown in any of SEQ ID NO: 76-83, 94-97, 102-107.

[0050] In some embodiments, the incretin agent comprises a half-life extension portion. In some embodiments, the half-life extension portion comprises albumin (e.g., human serum albumin). In some embodiments, human serum albumin comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO: 159. In some embodiments, human serum albumin comprises an amino acid sequence as shown in SEQ ID NO: 159.

[0051] In some embodiments, the incretin agent comprises albumin (e.g., human serum albumin) fused to one or more units, each of which comprises, from the N-terminus to the C-terminus: (i) a GLP1 receptor agonist-linker peptide (e.g., SEQ ID NO: 98); (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 100); (iii) a GLP1 receptor agonist-linker peptide-furin protease cleavage site (e.g., SEQ ID NO: 102); or (iv) a GLP1 receptor agonist-linker peptide-furin protease cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 104), two units (e.g., SEQ ID NO: 106), or four units (e.g., SEQ ID NO: 107). In some embodiments, the incretin agent comprises an amino acid sequence as shown in any of SEQ ID NO: 98, 100, 102, 104, 106, 107, or any combination thereof.

[0052] In some embodiments, the half-life extension portion comprises an albumin-binding domain (ABD). In some embodiments, the ABD is derived from protein G of Streptococcus strain GI48 and / or protein PAB of *Gastrodinium coli*, such as ABD035 and SA21. In some embodiments, the half-life extension portion comprises an ABD that binds to domain II of human serum albumin without overlapping or interfering with binding to the FcRn binding site on albumin. In some embodiments, the half-life extension portion comprises ABDCon. In some embodiments, the half-life extension portion comprises an ABD derived from bacterial protein Sso7d from the hyperthermophilic archaea *Sulphozoa sulfolithoides*, such as M11.12 and M18.2.5. In some embodiments, the half-life extension portion comprises albumin-binding DARPin. In some embodiments, the ABD comprises an albumin-binding immunoglobulin domain or a fragment thereof. In some embodiments, the ABD comprises an albumin-binding fully human domain antibody (dAb), such as AlbudAb. In some embodiments, the ABD comprises albumin-binding Fab, such as dsFv CA645.

[0053] In some embodiments, the ABD comprises a heavy-chain-only (VHH) antibody, such as a nanobody, that binds to albumin. In some embodiments, the VHH antibody comprises a VHH domain having complementarity-determining region (CDR) sequences HCDR1, HCDR2, and / or HCDR3 as shown in SEQ ID NO: 191 (GFTLDYYA), SEQ ID NO: 192 (IASSGGST), and / or SEQ ID NO: 193 (AAAVLECRTVVRGYDY), respectively. In some embodiments, the VHH antibody comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO: 154. In some embodiments, the VHH antibody comprises an amino acid sequence as shown in SEQ ID NO: 154. In some embodiments, the incretin agent comprises a VHH antibody bound to an albumin fused to a unit, the unit comprising, from the N-terminus to the C-terminus: (i) a GLP1-linker peptide (e.g., SEQ ID NO: 99); (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 101); or (iii) a GLP1 receptor agonist-linker peptide-furin protease-GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 103 or 105). In some embodiments, the incretin agent comprises an amino acid sequence as shown in any of SEQ ID NO: 99, 101, 103, 105.

[0054] In some embodiments, the extended half-life portion does not contain an Fc domain, such as that derived from human IgG, optionally from human IgG1, IgG2, IgG3, or IgG4. In some embodiments, the extended half-life portion contains an Fc domain, such as that derived from human IgG, optionally from human IgG1, IgG2, IgG3, or IgG4. In some embodiments, the human IgG is human IgG4. In some embodiments, the incretin agent contains an IgG4 Fc domain fused to a unit, the unit comprising, from the N-terminus to the C-terminus: i) a GLP1 receptor agonist-linker peptide (e.g., SEQ ID NO: 10, 89, 90, 91); (ii) a GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 92, 93); or (iii) a GLP1 receptor agonist-linker peptide-furin protease-GIP receptor agonist-linker peptide (e.g., SEQ ID NO: 94, 95, 96, 97). In some embodiments, the IgG4 Fc domain contains an amino acid sequence that is at least 90%, 95%, or 99% identical to SEQ ID NO: 155. In some embodiments, the IgG4 Fc domain comprises an amino acid sequence as shown in SEQ ID NO: 155. In some embodiments, the incretin agent comprises an amino acid sequence as shown in any of SEQ ID NO: 10, 89-97.

[0055] In some embodiments, the Fc domain contains one or more mutations in one or both constant Fc domains that increase the half-life of incretins and / or induce dimerization. In some embodiments, the one or more mutations comprise one or more mutations in the CH3 domain. In some embodiments, the one or more dimerizing mutations comprise: (i) Y349C, T366S, L368A and / or Y407V (according to EU designations); or (ii) S354C and / or T366W (according to EU designations).

[0056] In some embodiments, one or more mutations comprise Y349C, T366S, L368A, and Y407V (“FcKIH-b”, according to EU designations); or S354C and T366W (“FcKIH-a”, according to EU designations). In some embodiments, the incretin agent comprises a first polypeptide chain and a second polypeptide chain, wherein the first polypeptide chain comprises an incretin peptide fused to a first Fc domain, wherein the first Fc domain comprises the mutants Y349C, T366S, L368A, and Y407V (“FcKIH-b”, according to EU designations), and wherein the second polypeptide chain comprises an incretin peptide fused to a second Fc domain, wherein the second Fc domain comprises the mutants S354C and T366W (“FcKIH-a”, according to EU designations).

[0057] In some embodiments, one or more mutations that increase the half-life of the incretin include M428L and N434S (“LS”, according to EU designations). In some embodiments, the incretin includes an Fc domain having an FcKIH-a mutation on a first polypeptide chain and an Fc domain having an FcKIH-b mutation on a second polypeptide chain, wherein the Fc domain on each polypeptide chain is independently fused to one or more units comprising, from the N-terminus to the C-terminus: (i) a GLP1 receptor agonist-linking peptide (e.g., SEQ ID NO: 84, 85, 86, 87); or (ii) a GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 88). In some embodiments, the incretin includes an amino acid sequence as shown in any of SEQ ID NO: 84-88.

[0058] In some embodiments, the Fc domain comprises one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q). In some embodiments, one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234S, L235T, and G236R (“STR”, according to EU designations). In some embodiments, one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234A and L235A (“LALA”, according to EU designations). In some embodiments, one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutation: L234A / L235A / P329G (“LALAPG”, according to EU designations).

[0059] In some embodiments, the half-life extension portion comprises VNAR that binds to albumin.

[0060] In some embodiments, the half-life extension portion comprises an XTEN sequence.

[0061] In another aspect, this disclosure provides an incretin agent comprising: a husec signal peptide; an incretin peptide comprising a GLP1 incretin peptide or a fragment or mutant thereof; wherein the GLP1 incretin peptide comprises an amino acid sequence having an A8G substitution mutation compared to the wild-type GLP1 amino acid sequence.

[0062] In another aspect, this disclosure provides a polynucleotide encoding an incretin agent comprising a husec signal peptide; an incretin peptide comprising a GLP1 incretin peptide or a fragment or mutant thereof; wherein the GLP1 incretin peptide comprises an amino acid sequence having an A8G substitution mutation compared to the wild-type GLP1 amino acid sequence.

[0063] In another aspect, this disclosure provides an incretin agent comprising: a husec signal peptide; an incretin peptide comprising a GIP incretin peptide or a fragment or mutant thereof; wherein the GIP incretin peptide comprises an amino acid sequence having an A2G substitution mutation compared to the wild-type GIP amino acid sequence.

[0064] In another aspect, this disclosure provides a polynucleotide encoding an incretin agent comprising: a husec signal peptide; an incretin peptide comprising a GIP incretin peptide or a fragment or mutant thereof; wherein the GIP incretin peptide comprises an amino acid sequence having an A2G substitution mutation compared to the wild-type GIP amino acid sequence. Attached Figure Description

[0065] Figure 1 This paper demonstrates an exemplary therapeutic strategy that utilizes polynucleotides as described herein for the delivery and in vivo expression of incretin agents.

[0066] Figure 2 A schematic diagram of an exemplary polynucleotide encoding an incretin is shown.

[0067] Figure 3 An exemplary design of the incretin agent described herein is illustrated. Specifically, Figure 3 A schematic diagram (top) of the polynucleotide encoding the signal peptide (“SP”) and a single incretin peptide (this configuration is referred to as “I:1x” in this document) and a schematic diagram (bottom) of the translated incretin protein are shown.

[0068] Figure 4 An exemplary design of the incretin agent described herein is illustrated. Specifically, Figure 4 A schematic diagram (top) of the polynucleotide encoding the signal peptide (“SP”) and the translated protein is shown (bottom). The polynucleotides are separated by a linker peptide (“L1”) and a furin cleavage site (“F”).

[0069] Figure 5 An exemplary design of the incretin agent described herein is illustrated. Specifically, Figure 5A schematic diagram (top) of the polynucleotides encoding the signal peptide (“SP”) and the four incretin peptides (this configuration is referred to herein as “I:4x”) separated by the furin cleavage site (“F”) of each free linking peptide (“L1”) is shown, along with a schematic diagram of the translated protein (bottom).

[0070] Figure 6 A-6B illustrates exemplary incretin agents, including incretin agents having a signal peptide (“SP”), a GLP1 incretin peptide, and a (GGGGS)2 linker peptide. Figure 6 A), and incretins having a signal peptide (“SP”), GLP1 incretin peptide, (GGGGS)2 linker peptide, furin cleavage site, and GIP incretin peptide. Figure 6 B). Signal peptide cleavage site at Figure 6 A and Figure 6 Indicated in B. Additionally, the furin cleavage site is located in... Figure 6 In instruction B, the GIP incretin peptide is cleaved from the GLP1 incretin peptide after expression.

[0071] Figure 7 An exemplary design of the incretin agent described herein is illustrated. Specifically, Figure 7 A schematic diagram (top) of a polynucleotide encoding a signal peptide (“SP”) and an incretin (e.g., an I:1x, I:2x, or I:4x incretin as described herein) and a schematic diagram (bottom) of the translated protein, the incretin being fused to a half-life extension (“HLE”) domain, such as human serum albumin (“HSA”) or an albumin-binding domain (“ABD”), via a linker peptide (“L2”).

[0072] Figure 8 A-8B illustrates an exemplary incretin agent that can be encoded by the polynucleotides described herein, comprising more than one incretin peptide and a half-life extension (HLE) domain. Figure 8 A demonstrates an incretin agent having a signal peptide (“SP”), a first GLP1 incretin peptide, a linker peptide, a second GLP1 incretin peptide, a second linker peptide (GGGGS)3, and an extended half-life (HLE) domain for human serum albumin (HSA). Figure 8B illustrates an incretin agent comprising a signal peptide (“SP”), a first GLP1 incretin peptide, a linker peptide, a second GLP1 incretin peptide, a second linker peptide (GGGGS)3, and a half-life extension (HLE) domain bound to the VHH domain of HSA. The furin protease and SP cleavage sites within the incretin agent are indicated by arrows, such that upon expression, the signal peptide cleaves and the first GLP1 incretin peptide cleaves from the second GLP1 incretin peptide, while the second GLP1 incretin peptide remains fused to the HLE domain (HSA or anti-HSA VHH). This design yields two incretin peptides with different half-lives and activities.

[0073] Figure 9 Exemplary incretin agents, which may be encoded by one or more polynucleotides described herein, are shown, comprising more than one incretin peptide and a half-life extension (HLE) domain. Specifically, Figure 9 The incretin in this formulation contains a signal peptide (“SP”), a first GLP1 incretin peptide, a linker peptide (GGGGS)2, a first GIP incretin peptide, a second linker peptide (GGGGS)2, a second GLP1 incretin peptide, a third linker peptide (GGGGS)2, a second GIP incretin peptide, a fourth linker peptide (GGGGS)3, and a half-life extension (HLE) domain for human serum albumin (HSA). The furin protease and SP cleavage sites within the incretin are indicated by arrows. This design produces four separate incretin peptides, with the second GIP incretin peptide remaining fused to the HLE domain.

[0074] Figure 10 An exemplary design of a polynucleotide encoding an incretin agent is shown, the incretin agent comprising an incretin peptide fused to an Fc domain, wherein the incretin peptide may be one (I:1x), two (I:2x), or four (I:4x) incretin peptides (top). When expressing two of the two separate chains of the polynucleotide, the two polypeptide chains bind and form a dimer (e.g., homodimer) structure (bottom). Each polypeptide chain also includes a signal peptide (SP) and a linker peptide (L2). In some embodiments, the Fc domain includes mutations that eliminate effector function (e.g., mutations in STR, LALA, LALAPG, etc.) and / or mutations that prolong half-life (e.g., mutations in YTE, LS, etc.).

[0075] Figure 11 Exemplary incretin agents, which may be encoded by one or more polynucleotides as described herein, are shown, comprising more than one incretin peptide on more than one polypeptide chain. Specifically, Figure 11The polypeptide chains of the incretins in the formulation contain a signal peptide (“SP”), a GLP1 incretin peptide, a linker peptide (GGGGS)3, and an Fc domain. One or both Fc domains contain an “LS” mutation (according to EU designation, M428L / N434S) to prolong the half-life of the incretin. When two polypeptide chains are expressed, they bind to form a structure as shown in the diagram. Figure 11 The homodimer structure is shown. SP cleavage sites within incretin agents are indicated by arrows.

[0076] Figure 12 Exemplary incretin agents, encoded by one or more polynucleotides described herein, are demonstrated, comprising more than one incretin peptide on more than one polypeptide chain. Specifically, each polypeptide chain of the incretin agent has a signal peptide (SP), a GLP1 incretin peptide, a linker peptide, a GIP peptide, a second linker peptide (GGGGS)3, and an Fc domain. One or two Fc domains contain an LS mutation. When two polypeptide chains are expressed, they bind to form a structure as described herein. Figure 12 The homodimer structure is shown. The furin and SP cleavage sites within the incretin are indicated by arrows.

[0077] Figure 13 An exemplary design of two polynucleotides is shown, each encoding a polypeptide chain (top) comprising an incretin peptide fused to an Fc domain. In each polypeptide chain (incretin-Fc fusion), a signal peptide (SP) and one, two, or four incretin peptides (I:1x, I:2x, or I:4x) fused to the Fc domain via a linker peptide (L2) are present, and each Fc domain has a modification that induces heterodimerization (e.g., a knock-in-hole mutation). When both polypeptide chains are expressed, they bind together to form a heterodimeric incretin agent.

[0078] Figure 14 Exemplary incretin agents, which may be encoded by one or more polynucleotides as described herein, are shown, comprising more than one incretin peptide on more than one polypeptide chain. Specifically, Figure 14 Each polypeptide chain of an incretin incretin contains a signal peptide (SP), a GLP1 or GIP incretin peptide, a linker peptide (GGGGS)3, and an Fc domain. One or two Fc domains contain “LS” mutations (M428L / N434S), “STR” mutations (L234S, L235T, and G236R mutations according to EU designations) to silence Fc effector function, and “mortar and pestle” mutations to induce heterodimerization. When two polypeptide chains are expressed, they bind to form a heterodimeric incretin containing two polypeptide chains with different incretin peptides. The SP cleavage sites within the incretin are indicated by arrows.

[0079] Figure 15 Exemplary concentrations (pg / ml) of GLP1 (7-37) in the supernatant of HEK293t17 cells transfected with polynucleotides encoding GLP1 (7-37) at 3, 6, 24, 48, and 72 hours post-transfection are shown (GLP1 n=4 + / - SD; ns = not significant; * p<0.05, ** p<0.01).

[0080] Figure 16 Exemplary concentrations (pg / ml) of GLP1 (7-37) with the K34R mutation in the supernatant of HEK293t17 cells transfected with polynucleotides encoding GLP1 (7-37)-(K34R) are shown at 3, 6, 24, 48, and 72 hours post-transfection (GLP1 n=4 + / - SD; ns = not significant). * p<0.05, ** p<0.01).

[0081] Figure 17 Exemplary concentrations (pg / ml) of GIP (1-42) in the supernatant of HEK293t17 cells transfected with polynucleotides encoding GIP (1-42) at 3, 6, 24, 48, and 72 hours post-transfection are shown (GIP n=6 + / - SD; ns = not significant; * p<0.05, ** p<0.01).

[0082] Figure 18 The concentrations (pg / ml) of exemplary GLP1 incretins in the supernatant of HEK29t17 cells transfected with polynucleotides encoding either a viral signal peptide (“viral SP”) or a husec signal peptide (“husec”) and codon-optimized using different strategies (“opt1” vs. “optp”) are shown. Specific incretins include: viral SP-GLP1 (7-37), viral SP-GLP1 (7-37)-(K34R), husec-GLP-1 (7-37)-A8G (opt1), husec-GLP-1 (7-37)-A8G-linker peptide (opt1), husec-GLP-1 (7-37)-A8G (optp), and husec-GLP-1 (7-37)-A8G-linker peptide (optp) incretins.

[0083] Figure 19The concentrations (ng / ml) of exemplary GIP incretins in the supernatant of HEK29t17 cells transfected with polynucleotides encoding either a viral signal peptide (“viral SP”) or a husec signal peptide (“husec”) and codon-optimized using different strategies (“opt1” vs. “optp”) are shown. Specific incretins include: viral SP-GIP (1-42), husec-GIP (1-42)-A2G (opt1), and GIP (1-42)-A2G (optp).

[0084] Figure 20 This diagram illustrates the theoretical cleavage sites of various signal peptides located within the amino acid sequences of incretins. Figure 20 It also indicates that the A8G mutation promotes proper N-terminal processing of GLP1 incretins with the husec signal peptide.

[0085] Figure 21 This diagram illustrates the theoretical cleavage sites of various signal peptides located within the amino acid sequences of incretins. Figure 21 It also indicates that the A2G mutation promotes proper N-terminal processing of GIP incretins with the husec signal peptide.

[0086] Figure 22 A-22B illustrates a schematic diagram of a HEK293 reporter cell line overexpressing GLP1R (A) and GIPR (B) intended for use in an assay to determine the biological activity of the exemplary GLP1 and GIP incretin agents in Example 7.

[0087] Figure 23 Results of bioactivity assays from an exemplary GLP1 incretin are presented. Specifically, results are expressed as fold induction relative to a control sample.

[0088] Figure 24 Results of bioactivity assays from exemplary GIP incretin agents are presented. Specifically, results are expressed as fold induction relative to control samples.

[0089] Figure 25 The in vitro activity (GIP expression) of some exemplary incretins tested is demonstrated.

[0090] Figure 26 The GIP bioactivity of some exemplary GIP-containing incretins tested is demonstrated.

[0091] Figure 27 The in vitro activity (GLP1 expression) of some exemplary incretins tested is demonstrated.

[0092] Figure 28 The GLP1 bioactivity of some exemplary GLP1-containing incretins tested is demonstrated.

[0093] Figure 29 The comparison of GIP expression (A) and GIP bioactivity (B) is shown in exemplary candidates with different signal peptides (husec vs. gD1).

[0094] Figure 30 The comparison of GLP1 expression (A) and GLP1 bioactivity (B) is shown in exemplary candidates with different signal peptides (husec vs. gD1).

[0095] Figure 31 The comparison of GIP expression (A) and GIP bioactivity (B) with and without various half-life extension (HLE) moieties is presented.

[0096] Figure 32 The comparison of GLP1 expression (A) and GLP1 bioactivity (32B) with and without various half-life extension (HLE) moieties is presented.

[0097] Figure 33 A comparison of GIP expression (A) and GIP bioactivity (B) in an exemplary incretin agent containing both GIP and GLP1 is shown, wherein the order of the GIP and GLP1 peptides encoded by a single polynucleotide is altered.

[0098] Figure 34 A comparison of GLP1 expression (A) and GLP1 bioactivity (B) in an exemplary incretin containing both GIP and GLP1 is shown, wherein the order of the GIP and GLP1 peptides encoded by a single polynucleotide is altered.

[0099] definition About: When used herein to refer to a value, the term “about” means a value similar to the mentioned value in the context. Generally, those skilled in the art will understand the extent of relevant variation covered by “about” in the context once they become familiar with it. For example, in some embodiments, the term “about” may cover a range of values ​​within the range of 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the mentioned value.

[0100] Reagent: As used herein, the term "reagent" may refer to a physical entity. In some embodiments, a reagent may be characterized by specific features and / or effects. For example, as used herein, the term "therapeutic agent" refers to a physical entity that has a therapeutic effect and / or induces a desired biological and / or pharmacological effect. In some embodiments, a reagent may be a compound, molecule, or entity of any chemical class, including, for example, small molecules, peptides, nucleic acids, carbohydrates, lipids, metals, or combinations or complexes thereof.

[0101] Aliphatic: The term "aliphatic" refers to a straight-chain (i.e., unbranched) or branched, substituted or unsubstituted hydrocarbon chain that is fully saturated or contains one or more unsaturated units, or a monocyclic or bicyclic hydrocarbon (also referred to herein as "cycloaliphatic") that is fully saturated or contains one or more unsaturated units but is not aromatic, having a single or more connection points with the rest of the molecule. Unless otherwise stated, the aliphatic group contains 1-12 aliphatic carbon atoms. In some embodiments, the aliphatic group contains 1-6 aliphatic carbon atoms (e.g., C64 ... 1-6 In some embodiments, the aliphatic group contains 1-5 aliphatic carbon atoms (e.g., C15, C25, C35, C45, C55, C65, C7 ... 1-5 In other embodiments, the aliphatic group contains 1-4 aliphatic carbon atoms (e.g., C46 ... 1-4 In other embodiments, the aliphatic group contains 1-3 aliphatic carbon atoms (e.g., C10, C20, C30, C40, C50, C60, C70, C80, C9 ... 1-3 In other embodiments, the aliphatic group contains 1-2 aliphatic carbon atoms (e.g., C10, C20, C30, C40, C50, C60, C70, C80, C9 ... 1-2 Suitable aliphatic groups include, but are not limited to, straight-chain or branched, substituted or unsubstituted alkyl, alkenyl or ynyl groups and their hybrids. Preferred aliphatic groups are C10 and C20. 1-6 alkyl.

[0102] Alkyl: The term "alkyl" used alone or as part of a larger part refers to having 1-12, 1-10, 1-8, 1-6, 1-4, 1-3 or 1-2 carbon atoms (e.g., C12, C23, C33, C43, C53, C63, C7 ... 1-12 C 1-10 C 1-8 C 1-6 C 1-4 C 1-3 Or C 1-2 The alkyl group is a saturated, optionally substituted, straight-chain or branched hydrocarbon group. Exemplary alkyl groups include methyl, ethyl, propyl, butyl, pentyl, hexyl, and heptyl.

[0103] Alkylene: The term "alkylene" refers to a divalent alkyl group. In some embodiments, "alkylene" is a divalent straight-chain or branched alkyl group. In some embodiments, the "alkylene chain" is a polymethylene group, i.e., -(CH2). n- where n is a positive integer, such as 1 to 6, 1 to 4, 1 to 3, 1 to 2, or 2 to 3. The optionally substituted alkylene chain is a polymethylene group in which one or more methylene hydrogen atoms are optionally substituted. Suitable substituents include those described below for substituted aliphatic groups, and also those described in this specification. It will be understood that two substituents of an alkylene group can together form a ring system. In some embodiments, two substituents can together form a 3- to 7-membered ring. Substituents can be on the same or different atoms. The suffix "-alkyl" or "-alkylene" when attached to certain groups herein is intended to refer to the bifunctional portion of that group. For example, "-alkyl" or "-alkylene" when attached to "cyclopropyl" becomes "cyclopropylene" or "cyclopropylenyl" and is intended to refer to a bifunctional cyclopropyl group, for example, .

[0104] Alkenyl: The term "alkenyl" used alone or as part of a larger part refers to a group having at least one double bond and having (unless otherwise specified) 2-12, 2-10, 2-8, 2-6, 2-4, or 2-3 carbon atoms (e.g., C12, C23, C14, C23, C2 ... 2-12 C 2-10 C 2-8 C 2-6 C 2-4 Or C 2-3 Alkenyl groups are optionally substituted linear, branched, or cyclic hydrocarbon groups. Exemplary alkenyl groups include vinyl, propenyl, butenyl, pentenyl, hexenyl, and heptenyl. The term "cycloalkenyl" refers to an optionally substituted non-aromatic monocyclic or polycyclic system containing at least one carbon-carbon double bond and having about 3 to about 10 carbon atoms. Exemplary monocyclic cycloalkenyl rings include cyclopentenyl, cyclohexenyl, and cycloheptenyl.

[0105] Alkynyl: The term "alkynyl" used alone or as part of a larger part refers to a group having at least one triple bond and having (unless otherwise specified) 2–12, 2–10, 2–8, 2–6, 2–4, or 2–3 carbon atoms (e.g., C12, C23, C14, C23, C2 ... 2-12 C 2-10 C 2-8 C 2-6 C 2-4 Or C 2-3 The optionally substituted straight-chain or branched hydrocarbon group of . Exemplary alkynyl groups include ethynyl, propynyl, butynyl, pentynyl, hexynyl, and heptynyl.

[0106] Amino acid: In its broadest sense, as used herein, the term "amino acid" refers to a compound and / or substance that can be incorporated into, incorporated into, or has been incorporated into a polypeptide chain, for example, by forming one or more peptide bonds. In some embodiments, an amino acid has the general structure H₂N-C(H)(R)-COOH. In some embodiments, an amino acid is a naturally occurring amino acid. In some embodiments, an amino acid is a non-natural amino acid; in some embodiments, an amino acid is a D-amino acid; in some embodiments, an amino acid is an L-amino acid. "Standard amino acid" refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. "Non-standard amino acid" refers to any amino acid other than a standard amino acid, whether it is synthesized or obtained from a natural source. In some embodiments, amino acids (including carboxyl-terminal and / or amino-terminal amino acids in polypeptides) may contain structural modifications compared to the general structure described above. For example, in some embodiments, amino acids may be modified by methylation, amidation, acetylation, PEGylation, glycosylation, phosphorylation, and / or substitution (e.g., amino groups, carboxylic acid groups, one or more protons and / or hydroxyl groups) compared to the general structure. In some embodiments, such modification may, for example, alter the cyclic half-life of a peptide containing modified amino acids compared to a peptide containing otherwise identical, unmodified amino acids. In some embodiments, such modification does not significantly alter the relevant activity of a peptide containing modified amino acids compared to a peptide containing otherwise identical, unmodified amino acids. As will be clear from the context, in some embodiments, the term "amino acid" may be used to refer to a free amino acid; in some embodiments, the term "amino acid" may be used to refer to an amino acid residue of a peptide.

[0107] Aromatic group: The term "aromatic group" refers to a monocyclic and bicyclic system having a total of six to fourteen ring members (e.g., C6-C14), wherein at least one ring in the system is aromatic and each ring in the system contains three to seven ring members. In some embodiments, "aromatic group" contains a total of six to twelve ring members (e.g., C6-C12). The term "aromatic group" may be used interchangeably with the term "aromatic ring". In some embodiments, "aromatic group" refers to an aromatic ring system that may have one or more substituents, including but not limited to phenyl, biphenyl, naphthyl, anthracene, and similar groups. Unless otherwise stated, "aromatic group" is a hydrocarbon. In some embodiments, an "aromatic group" ring system is an aromatic ring (e.g., phenyl) fused to a non-aromatic ring (e.g., cycloalkyl). Examples of aromatic rings include fused aromatic rings, including... , and .

[0108] Related: When used herein, the term means that two events or entities are “related” to each other if the presence, level, extent, type, and / or form of one event or entity is related to the presence, level, extent, type, and / or form of another event or entity. For example, an entity is considered related to a disease, condition, or disorder if the presence, level, and / or form of a particular entity (e.g., a polypeptide, genetic imprint, metabolite, microorganism, etc.) is related to the incidence, susceptibility, severity, stage, etc., of that disease, condition, or disorder (e.g., in a relevant population). In some embodiments, two or more entities are physically “bonded” to each other if they interact directly or indirectly such that the entities are physically close to each other and / or remain close to each other. In some embodiments, two or more physically bonded entities are covalently connected to each other; in some embodiments, two or more physically bonded entities are not covalently connected to each other, but are non-covalently bonded, for example, by hydrogen bonds, van der Waals interactions, hydrophobic interactions, magnetism, and combinations thereof.

[0109] Co-administration: As used herein, the term "co-administration" refers to the use of a composition described herein (e.g., a pharmaceutical composition) and one or more additional therapeutic agents. In some embodiments, the one or more additional therapeutic agents comprise at least one polynucleotide encoding another therapeutic agent (e.g., an incretin). The combined use of the composition described herein (e.g., a pharmaceutical composition) and the additional therapeutic agent may be performed simultaneously or separately (e.g., sequentially in any order). In some embodiments, the composition described herein (e.g., a pharmaceutical composition) and the additional therapeutic agent may be combined in a pharmaceutically acceptable excipient, or may be placed in a separate excipient and delivered to target cells or administered to an individual at different times. It is contemplated that each of these situations falls within the meaning of "co-administration" or "combination" provided that the composition described herein (e.g., a pharmaceutical composition) and the additional therapeutic agent are delivered or administered sufficiently close in time such that there is at least some temporal overlap in the biological effects produced by each on the target cells or the treated individual.

[0110] Combination therapy: As used herein, the term "combination therapy" refers to those situations where an individual is simultaneously exposed to two or more treatment regimens (e.g., two or more therapeutic agents (e.g., two or more incretin agents)). In some embodiments, the two or more regimens may be administered simultaneously; in some embodiments, such regimens may be administered sequentially (e.g., all "dose" of the first regimen are administered before any dose of the second regimen); in some embodiments, such agents are administered in an overlapping dosing regimen. In some embodiments, administration of combination therapy may involve administering one or more agents or methods to an individual receiving other agents or methods in the combination. For clarity, combination therapy does not require the individual agents to be administered together in a single composition (or even necessarily at the same time), although in some embodiments, two or more agents or their active portions may be administered together in a combined composition. In some embodiments, combination therapy comprises a polynucleotide encoding two or more incretin agents.

[0111] Comparable: As used herein, the term "comparable" means two or more sets of reagents, entities, situations, conditions, etc., that may not be consistent with each other, but are similar enough to allow comparisons between them, such that those skilled in the art will understand that reasonable conclusions can be drawn based on the observed differences or similarities. In some embodiments, a comparable set of conditions, environment, individual, or group is characterized by a plurality of substantially consistent features and one or more features that vary slightly. In the context, those skilled in the art will understand the degree of consistency required in any given environment for two or more such sets of reagents, entities, situations, conditions, etc., to be considered comparable. For example, those skilled in the art will understand that environments, individuals, or groups are comparable to each other when characterized by a sufficient number and type of substantially consistent features to ensure that differences in results or observed phenomena in different groups of environments, individuals, or groups are caused by variations in those features or to indicate reasonable conclusions about such variations.

[0112] Corresponding to: As used herein, the term “corresponding to” refers to a relationship between two or more entities. For example, the term “corresponding to” can be used to specify the position / identity of a structural component in a compound or composition relative to another compound or composition (e.g., relative to a suitable reference compound or composition). For example, in some embodiments, monomeric residues in a polymer (e.g., amino acid residues in a polypeptide or nucleic acid residues in a polynucleotide) can be identified as “corresponding to” residues in a suitable reference polymer. For example, those skilled in the art will appreciate that, for simplicity, residues in polypeptides are typically designated using a canonical numbering system based on a reference-related polypeptide, such that the amino acid “corresponding to” the residue at position 190, for example, does not actually need to be the 190th amino acid in a particular amino acid chain, but rather corresponds to the residue found at position 190 in the reference polypeptide; those skilled in the art will readily understand the manner in which “corresponding” amino acids are identified. For example, those skilled in the art will recognize various sequence alignment strategies, including software programs that can be used, for example, to identify “corresponding” residues in peptides and / or nucleic acids as described in this disclosure, such as BLAST, CS-BLAST, CUSASW++, DIAMOND, FASTA, GGSEARCH / GLSEARCH, Genoogle, HMMER, HHpred / HHsearch, IDF, Infernal, KLAST, USEARCH, parasail, PSI-BLAST, PSI-Search, ScalaBLAST, Sequilab, SAM, SSEARCH, SWAPHI, SWAPHI-LS, SWIMM, or SWIPE. Those skilled in the art will also understand that, in some cases, the term “corresponding” can be used to describe an event or entity that shares a relevant similarity with another event or entity (e.g., a suitable reference event or entity). To cite only one example, a gene or protein in one organism may be described as “corresponding” to a gene or protein from another organism in order to indicate, in some embodiments, that it plays a similar role or performs a similar function and / or exhibits a particular degree of sequence identity or homology, or shares specific characteristic sequence components.

[0113] Cycloaliphatic: As used herein, the term “cycloaliphatic” refers to a monocyclic C3-8 hydrocarbon or a bicyclic C6-10 hydrocarbon that is fully saturated or contains one or more unsaturated units but is not aromatic, and has a single or more connection points with the rest of the molecule.

[0114] Cycloalkyl: As used herein, the term "cycloalkyl" refers to a saturated cyclic monocyclic or polycyclic system of about 3 to about 10 ring carbon atoms, which may be optionally substituted. Exemplary monocyclic cycloalkyl rings include cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, and cycloheptyl.

[0115] Derivative: In the context of an amino acid sequence (peptide or polypeptide) “derived from” a specified amino acid sequence (peptide or polypeptide), it refers to a structural analog of the specified amino acid sequence. In some embodiments, the amino acid sequence derived from a particular amino acid sequence has an amino acid sequence that is identical, substantially identical, or homologous to the Peter sequence or a fragment thereof. The amino acid sequence derived from a particular amino acid sequence may be a mutant of the Peter sequence or a fragment thereof. For example, incretin agents as used herein may include amino acid sequences derived from two or more incretin agents (e.g., two or more naturally occurring incretin agents).

[0116] Detection: The term “detection” is used broadly herein to include any appropriate manner of determining the presence or absence of a target entity in a sample, or any form of measurement of the target entity. Thus, “detection” can include determining, measuring, evaluating, or detecting the presence or absence, level, quantity, and / or location of a target entity. This includes both quantitative and qualitative determination, measurement, or evaluation, including semi-quantitative methods. Such determination, measurement, or evaluation can be relative, such as when detecting a target entity relative to a control reference, or absolute. Therefore, the term “quantification” when used in the context of quantifying a target entity can refer to absolute quantification or relative quantification. Absolute quantification can be accomplished by relating the detected level of a target entity to a known control standard (e.g., by generating a standard curve). Alternatively, relative quantification can be accomplished by comparing the detected levels or quantities of two or more different target entities to provide a relative quantification of each of the two or more different target entities (i.e., relative to each other).

[0117] Dosing regimen: Those skilled in the art will understand that the term "dosing regimen" (or "treatment regimen") can be used to refer to a set of unit doses (usually more than one unit dose) that are typically administered to an individual at regular intervals. In some embodiments, a given therapeutic agent has a recommended dosing regimen, which may involve one or more doses.

[0118] Encoding: As used herein, the term "encode" or "encoding" refers to the sequence information that directs the production of a first molecule from a second molecule having a defined nucleotide sequence (e.g., a polynucleotide) or a defined amino acid sequence. For example, a DNA molecule can encode an RNA molecule (e.g., through transcription involving a DNA-dependent RNA polymerase). An RNA molecule can encode a polypeptide (e.g., through translation). Thus, if transcription and translation of RNA corresponding to a gene produces a polypeptide in a cell or other biological system, the gene, cDNA, or RNA molecule encodes that polypeptide. In some embodiments, the coding region of a polynucleotide encoding a target antigen refers to a coding strand whose nucleotide sequence is consistent with the polynucleotide sequence of such target antigen. In some embodiments, the coding region of a polynucleotide encoding a target antigen refers to a non-coding strand of such target antigen, which can be used as a template for gene or cDNA transcription.

[0119] Engineered: Generally speaking, the term "engineered" refers to aspects that are manipulated manually. For example, a polynucleotide is considered "engineered" when two or more sequences that are not linked together in nature are manually manipulated to directly link with each other in an engineered polynucleotide, and / or when a particular residue in a polynucleotide is not naturally occurring and / or is linked to an entity or part of it that is not linked to in nature through manual manipulation.

[0120] Expression: As used herein, the term “expression” of a nucleic acid sequence means the production of a gene product from a nucleic acid sequence. In some embodiments, the gene product may be a transcript, such as a polynucleotide as provided herein. In some embodiments, the gene product may be a polypeptide. In some embodiments, the expression of a nucleic acid sequence involves one or more of the following: (1) the production of an RNA template from a DNA sequence (e.g., by transcription); (2) the processing of the RNA transcript (e.g., by splicing, editing, etc.); (3) the translation of RNA into a polypeptide or protein; and / or (4) post-translational modifications of the polypeptide or protein.

[0121] Heteroaliphatic: As used herein, the term "heteroaliphatic" or "heteroaliphatic group" refers to an optionally substituted hydrocarbon moiety having one to five heteroatoms in addition to carbon atoms. It may be straight-chain (i.e., unbranched), branched, or cyclic ("heterocyclic") and may be fully saturated or contain one or more unsaturated units, but it is not aromatic. The term "heteroatom" means nitrogen, oxygen, or sulfur, and includes any oxidized form of nitrogen or sulfur, and any quaternary ammoniation of basic nitrogen. The term "nitrogen" also includes substituted nitrogen. Unless otherwise stated, a heteroaliphatic group contains 1 to 10 carbon atoms, wherein 1 to 3 carbon atoms are optionally and independently substituted with heteroatoms selected from oxygen, nitrogen, and sulfur. In some embodiments, a heteroaliphatic group contains 1 to 4 carbon atoms, wherein 1 to 2 carbon atoms are optionally and independently substituted with heteroatoms selected from oxygen, nitrogen, and sulfur. In some embodiments, the heteroaliphatic group contains 1-3 carbon atoms, wherein one carbon atom is optionally and independently replaced by a heteroatom selected from oxygen, nitrogen, and sulfur. Suitable heteroaliphatic groups include, but are not limited to, straight-chain or branched heteroalkyl, heteroalkenyl, and heteroynyl groups. For example, heteroaliphatic groups of 1 to 10 atoms include the following exemplary groups: -O-CH3, -CH2-O-CH2, -O-CH2-CH2-O-CH2-CH2-O-CH2 and similar groups.

[0122] Heteroaromatic group: The terms “heteroaromatic group” and “heteroaromatic-” used alone or as part of a larger portion (e.g., “heteroaromatic alkyl” or “heteroaromatic alkoxy”) refer to a monocyclic or bicyclic group having 5 to 10 ring atoms (e.g., a 5 to 6-membered monocyclic heteroaromatic group or a 9 to 10-membered bicyclic heteroaromatic group); having 6, 10, or 14 π electrons shared in the cyclic array; and having one to five heteroatoms in addition to carbon atoms. The heteroaromatic groups include, but are not limited to, thiophene, furanyl, pyrrole, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, oxazolyl, isoxazolyl, oxadiazolyl, thiazolyl, isothiazolyl, thiazolyl, thiazolyl, pyridinyl, pyrazinyl, indazinyl, purine, naphridinyl, pteridinyl, imidazo[1,2-a]pyrimidinyl, imidazo[1,2-a]pyridinyl, imidazo[4,5-b]pyridinyl, imidazo[4,5-c]pyridinyl, pyrrolopyridinyl, pyrrolopyrazinyl, thiophenolopyrimidinyl, triazolopyridinyl, and benzoisoxazoleyl. As used herein, the terms “heteroaromatic” and “heteroaromatic-” also include groups in which a heteroaromatic ring is fused to one or more aromatic, cycloaliphatic or heterocyclic rings, wherein the linking group or linking point is on the heteroaromatic ring (i.e., a bicyclic heteroaromatic ring having 1 to 3 heteroatoms). Non-limiting examples include indolyl, isoindolyl, benzothiophenyl, benzofuranyl, dibenzofuranyl, inzazoleyl, benzimidazolyl, benzotriazolyl, benzothiazolyl, benzothiadiazolyl, benzoxazolyl, quinolinyl, isoquinolinyl, phenolinyl, phthalazinyl, quinazolinyl, quinoxolinyl, 4H-quinazinyl, carbazoleyl, acridineyl, phenazinyl, phenthiazolyl, phenoxazinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, pyrido[2,3-b]-1,4-oxazin-3(4H)-one, 4H-thieno[3,2-b]pyrrole, and benzoisooxazolyl. The term "heteroaromatic group" may be used interchangeably with the terms "heteroaromatic ring," "heteroaromatic group," or "heteroaromatic group," any of which includes optionally substituted rings.

[0123] Heteroatoms: As used herein, the term “heteroatoms” refers to nitrogen, oxygen, or sulfur, and includes any oxidized form of nitrogen or sulfur, and any quaternary ammoniation of basic nitrogen.

[0124] Heterocycle: As used herein, the terms “heterocycle,” “heterocyclic group,” “heterocyclic ring,” and “heterocyclic ring” are used interchangeably and refer to a stable 3- to 8-membered monocyclic, 6- to 10-membered bicyclic, or 10- to 16-membered polycyclic heterocyclic portion that is saturated or partially unsaturated and has one or more heteroatoms as defined above, such as one to four heteroatoms, in addition to a carbon atom. When used to describe the ring atom of a heterocycle, the term “nitrogen” includes substituted nitrogen. As an example, in a saturated or partially unsaturated ring having 0-3 heteroatoms selected from oxygen, sulfur, or nitrogen, nitrogen may be N (as in 3,4-dihydro-2H-pyrrole), NH (as in pyrrolidinyl), or NR. + (e.g., in N-substituted pyrrolidinyl groups). The heterocycle may be attached to its side group at any heteroatom or carbon atom that produces a stable structure, and any ring atom may optionally be substituted. Examples of such saturated or partially unsaturated heterocyclic groups include, but are not limited to, azirrobutyl, oxacyclobutyl, tetrahydrofuranyl, tetrahydrothiophenyl, pyrrolidinyl, piperidinyl, decahydroquinolinyl, oxazolidinyl, pyrazinyl, dioxalyl, dioxopentyl, diazinonyl, oxonitrilenonyl, thionitrilenonyl, morpholinyl, and thiomorpholinyl. The heterocyclic group may be monocyclic, bicyclic, tricyclic, or polycyclic, preferably monocyclic, bicyclic, or tricyclic, and more preferably monocyclic or bicyclic. Bicyclic heterocycles also include groups in which the heterocycle is fused to one or more aromatic rings. Exemplary bicyclic heterocyclic groups include indololinyl, isoidololinyl, benzodioxanepentenyl, 1,3-dihydroisobenzofuranyl, 2,3-dihydrobenzofuranyl, and tetrahydroquinolinyl. Bicyclic heterocycles can also be spirocyclic systems (e.g., 7- to 11-membered spirocyclic fused heterocycles having one or more heteroatoms (e.g., one, two, three, or four heteroatoms) as defined above, in addition to a carbon atom). Bicyclic heterocycles can also be bridging ring systems (e.g., 7- to 11-membered bridging heterocycles having one, two, or three bridging atoms).

[0125] Homology: As used herein, the term "homology" or "homogeneity" refers to the overall correlation between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered "homological" if their sequences are at least 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered “homologous” to each other if their sequences are at least 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% similar (e.g., containing residues with relevant chemical properties at corresponding positions). For example, as is well known to those skilled in the art, certain amino acids are generally classified as “hydrophobic” or “hydrophilic” amino acids that are similar to each other, and / or classified as having “polar” or “nonpolar” side chains. The substitution of one amino acid for another amino acid of the same type can generally be considered a “homologous” substitution.

[0126] Identity: As used herein, the term "identity" refers to the overall correlation between polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules. In some embodiments, polynucleotide molecules (e.g., DNA molecules and / or RNA molecules) and / or polypeptide molecules are considered "substantially identical" to each other if their sequences are at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical. The percentage of identity between two nucleic acid or polypeptide sequences can be calculated, for example, by comparing the two sequences for optimal comparison purposes (e.g., for optimal comparison, gaps can be introduced in one or both of the first and second sequences, and non-identical sequences can be ignored for comparison purposes). In some embodiments, the sequence length compared for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or substantially 100% of the length of the reference sequence. Nucleotides at corresponding positions are then compared. When a position in the first sequence is occupied by the same residue (e.g., a nucleotide or amino acid) as the corresponding position in the second sequence, the molecule is considered consistent at that position. Taking into account the number of vacancies that need to be introduced for optimal alignment of the two sequences and the length of each vacancy, the percentage identity between the two sequences is a function of the number of consistent positions shared by the sequences. Sequence comparison and determination of the percentage of identity between two sequences can be accomplished using mathematical algorithms. For example, the percentage of identity between two nucleotide sequences can be determined using the algorithm of Meyers and Miller, 1989, which has been incorporated into the ALIGN program (version 2.0). In some exemplary embodiments, nucleic acid sequence comparisons performed using the ALIGN program employ a PAM120 weighted residue table, a vacancy length penalty of 12, and a vacancy penalty of 4. Alternatively, the GAP program in the GCG software package, which utilizes the NWSgapdna.CMP matrix, can be used to determine the percentage of identity between two nucleotide sequences.

[0127] Increase, induce, or decrease: As used herein, such terms, or grammatically comparable comparative terms, indicate a value relative to a comparable reference measurement. For example, in some embodiments, an assessment value achieved with the provided composition (e.g., a pharmaceutical composition) may be “increased” relative to an assessment value obtained with a comparable reference composition. Alternatively or additionally, in some embodiments, an assessment value achieved in an individual may be “increased” relative to an assessment value obtained in the same individual under different conditions (e.g., before or after an event; or with or without an event, such as the administration of a composition (e.g., a pharmaceutical composition)) or in different comparable individuals (e.g., in comparable individuals different from target individuals previously exposed to a disease (e.g., without the administration of a composition (e.g., a pharmaceutical composition)). In some embodiments, comparative terms refer to statistically relevant differences (e.g., those having a generality and / or magnitude sufficient to achieve statistical significance). Those skilled in the art will recognize or will be able to readily determine, in a given context, the degree and / or generality of difference required or sufficient to achieve such statistical significance. In some embodiments, the term "reduction" or equivalent term, compared to a comparable reference, means a reduction in the level of the assessed value by at least 5%, at least 10%, at least 20%, at least 50%, at least 75%, or higher. In some embodiments, the term "reduction" or equivalent term means complete or substantially complete suppression, i.e., a reduction to zero or substantially a reduction to zero. In some embodiments, the term "increase" or "induction," compared to a comparable reference, means an increase in the level of the assessed value by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 80%, at least 100%, at least 200%, at least 500%, or higher.

[0128] In sequence: As used herein, with respect to polynucleotides or polynucleotides, “in sequence” means the order of the features along the polynucleotide or polynucleotide from 5' to 3'. As used herein, with respect to polypeptides, “in sequence” means the order of the features along the polypeptide from the N-terminus—most features move to the C-terminus—most features. “In sequence” does not imply that additional features cannot be present among the listed features. For example, if features A, B, and C of a polynucleotide are described herein as “feature A, feature B, and feature C in sequence,” this description does not exclude, for example, feature D located between features A and B.

[0129] Ionizable: The term "ionizable" refers to a compound, group, or atom that carries a charge at a certain pH. In the context of ionizable amino lipids, such lipids or their functional groups or atoms carry a positive charge at a certain pH. In some embodiments, ionizable amino lipids carry a positive charge at acidic pH. In some embodiments, ionizable amino lipids are primarily neutral at physiological pH values ​​(e.g., about 7.0-7.4 in some examples), but become positively charged at lower pH values. In some embodiments, ionizable amino lipids may have a pKa in the range of about 5 to about 7.

[0130] Isolated: The term "isolated" means altered or removed from its natural state. For example, nucleic acids or peptides naturally present in living animals are not "isolated," but the same nucleic acid or peptide partially or completely isolated from its native coexisting material is "isolated." Isolated nucleic acids or proteins may exist in substantially purified forms or may exist in non-natural environments (such as, for example, host cells).

[0131] Lipids: As used herein, the terms “lipid” and “lipid-like material” are broadly defined as molecules comprising one or more hydrophobic moieties or groups and optionally also comprising one or more hydrophilic moieties or groups. Molecules comprising both hydrophobic and hydrophilic moieties are also commonly referred to as amphiphiles.

[0132] RNA lipid nanoparticles: As used herein, the term "RNA lipid nanoparticle" refers to a nanoparticle comprising at least one lipid and an RNA molecule, such as one or more polynucleotides as provided herein. In some embodiments, the RNA lipid nanoparticle comprises at least one cationic amino lipid. In some embodiments, the RNA lipid nanoparticle comprises at least one cationic amino lipid, at least one accessory lipid, and at least one polymer-coupled lipid (e.g., PEG-coupled lipid). In various embodiments, the RNA lipid nanoparticles as described herein may have an average size (e.g., Z-average) of about 100 nm to 1000 nm, or about 200 nm to 900 nm, or about 200 nm to 800 nm, or about 250 nm to about 700 nm. In some embodiments of this disclosure, the RNA lipid nanoparticles may have a particle size (e.g., Z-average) of about 30 nm to about 200 nm, or about 30 nm to about 150 nm, about 40 nm to about 150 nm, about 50 nm to about 150 nm, about 60 nm to about 130 nm, about 70 nm to about 110 nm, about 70 nm to about 100 nm, about 80 nm to about 100 nm, about 90 nm to about 100 nm, about 70 nm to about 90 nm, about 80 nm to about 90 nm, or about 70 nm to about 80 nm. In some embodiments, the average size of the lipid nanoparticles is determined by measuring the average particle size. In some embodiments, the RNA lipid nanoparticles may be prepared by mixing lipids with RNA molecules described herein.

[0133] Neutralization: As used herein, the term "neutralization" refers to an event in which a binding agent (such as an antibody) binds to a biologically active site of the virus (such as a receptor-binding protein), thereby inhibiting parasitic infection of the cell. In some embodiments, the term "neutralization" refers to an event in which the binding agent eliminates or significantly reduces the ability of the cell to infect it.

[0134] Nucleic Acids / Polynucleotides: As used herein, the term "nucleic acid" refers to a polymer of at least 10 nucleotides or more. In some embodiments, nucleic acids are or comprise DNA. In some embodiments, nucleic acids are or comprise RNA. In some embodiments, nucleic acids are or comprise peptide nucleic acids (PNAs). In some embodiments, nucleic acids are or comprise single-stranded nucleic acids. In some embodiments, nucleic acids are or comprise double-stranded nucleic acids. In some embodiments, nucleic acids comprise both single-stranded and double-stranded portions. In some embodiments, nucleic acids comprise a backbone comprising one or more phosphodiester bonds. In some embodiments, nucleic acids comprise a backbone comprising both phosphodiester bonds and non-phosphodiester bonds. For example, in some embodiments, nucleic acids may comprise a backbone comprising one or more thiophosphate bonds or 5'-N-phosphoamide bonds and / or one or more peptide bonds, as in "peptide nucleic acids". In some embodiments, nucleic acids comprise one or more or all of the natural residues (e.g., adenine, cytosine, deoxyadenosine, deoxycytidine, deoxyguanosine, deoxythymidine, guanine, thymine, uracil). In some embodiments, nucleic acids comprise one or more or all of the non-natural residues. In some embodiments, the non-natural residues comprise nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, C-5-propynyl-cytidine, C-5-propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazoadenosine, 7-deazoguanosine, 8-sideoxyadenosine, 8-sideoxyguanosine, 6-O-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, the non-natural residues comprise one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to the natural residues. In some embodiments, the nucleic acid has a nucleotide sequence encoding a functional gene product (such as RNA or a polypeptide). In some embodiments, the nucleic acid has a nucleotide sequence comprising one or more introns. In some embodiments, the nucleic acid can be prepared by isolating from a natural source, enzymatically synthesizing (e.g., by polymerization based on a complementary template, such as in vivo or in vitro), replicating in a recombinant cell or system, or by chemical synthesis.In some embodiments, the length of the nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 350. 0, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, 19,500, or 20,000 or more residues or nucleotides.

[0135] Pharmaceutically Effective Amount: The term "pharmaceutically effective amount" or "therapeuticly effective amount" refers to the amount, alone or in combination with other doses, that achieves the desired response or desired effect. In the treatment of a specific disease (e.g., obesity), in some embodiments, the desired response involves the inhibition of the course of the disease (e.g., obesity). In some embodiments, such inhibition may include slowing the progression of the disease (e.g., obesity) and / or interrupting or reversing the progression of the disease (e.g., obesity). In some embodiments, the desired response in the treatment of a disease (e.g., obesity) may be or include delaying or preventing the onset of the disease (e.g., obesity) or disorder (e.g., obesity-related disorder). The effective amount of the composition (e.g., pharmaceutical composition) described herein will depend on factors such as the disease (e.g., obesity) or disorder to be treated (e.g., obesity-related disorder), the severity of such disease (e.g., obesity) or disorder (e.g., obesity-related disorder), individual patient parameters (including, for example, age, physiological condition, body type, and weight), duration of treatment, type of concomitant therapy (if present), specific route of administration, and similar factors. Therefore, the dosage of the composition (e.g., pharmaceutical composition) described herein may depend on a variety of such parameters. If the patient’s response is insufficient at the initial dose, a higher dose may be used (or an effective higher dose achieved through a different, more localized route of administration).

[0136] Polypeptide: As used herein, the term "polypeptide" refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has an amino acid sequence that is naturally occurring. In some embodiments, a polypeptide has an amino acid sequence that is not naturally occurring. In some embodiments, a polypeptide has an engineered amino acid sequence because it is designed and / or generated by human manual action. In some embodiments, a polypeptide may comprise natural amino acids, non-natural amino acids, or both, or consist of them. In some embodiments, a polypeptide may comprise only natural amino acids or only non-natural amino acids, or consist of them. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may include one or more side groups or other modifications, such as modifying one or more amino acid side chains or attaching them to the N-terminus of the polypeptide, the C-terminus of the polypeptide, or any combination thereof. In some embodiments, such side groups or modifications include acetylation, amidation, esterification, methylation, polyethylene glycolation, etc., including combinations thereof. In some embodiments, a polypeptide may be cyclic and / or may contain a cyclic moiety. In some embodiments, a polypeptide is not cyclic and / or does not contain any cyclic moiety. In some embodiments, the polypeptide is linear. In some embodiments, the polypeptide may be or comprise a pinned polypeptide. In some embodiments, the term "polypeptide" may be appended to the name of a reference polypeptide, activity, or structure; in such cases, it is used herein to refer to a polypeptide that shares a relevant activity or structure and is therefore considered a member of the same polypeptide class or family. For each such class, this specification provides and / or those skilled in the art will recognize exemplary polypeptides within that class whose amino acid sequences and / or functions are known; in some embodiments, such exemplary polypeptides are reference polypeptides of a polypeptide class or family. In some embodiments, members of a polypeptide class or family exhibit significant sequence homology or identity with the reference polypeptide of that class (and in some embodiments with all polypeptides within that class), share common sequence motifs (e.g., characteristic sequence components), and / or share common activities (in some embodiments, at comparable levels or within specified ranges). For example, in some embodiments, the member polypeptide has an overall sequence homology or identity of at least about 30-40% compared to the reference polypeptide, and is typically greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher; and / or the member polypeptide contains at least one region (e.g., a conserved region, which in some embodiments may be or contains a characteristic sequence element) that exhibits extremely high sequence identity, typically greater than 90%, or even greater than 95%, 96%, 97%, 98% or 99%.Such conserved regions typically contain at least 3-4 and usually at most 35 or more amino acids; in some embodiments, the conserved region contains at least one segment consisting of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more consecutive amino acids. In some embodiments, the associated polypeptide may comprise or consist of a fragment of the parent polypeptide.

[0137] Prevention: As used herein, the terms “prevent” or “prevention” when used in connection with the occurrence of a disease, condition, and / or disorder mean reducing the risk of developing a disease, condition, and / or disorder, and / or delaying the onset of one or more features or symptoms of a disease, condition, or disorder. Prevention is considered complete when the onset of a disease, condition, or disorder has been delayed for a predetermined period of time.

[0138] Reference: As used herein, the term "reference" describes a standard or control relative to which comparisons are made. For example, in some embodiments, a reagent, animal, individual, population, sample, sequence, or value of the target is compared with a reference or control reagent, animal, individual, population, sample, sequence, or value. In some embodiments, the testing and / or determination of the reference or control is performed substantially simultaneously with the testing and / or determination of the target. In some embodiments, the reference or control is a historical reference or control optionally embodied in tangible media. Generally, as those skilled in the art will understand, the reference or control is determined or characterized under conditions or settings comparable to those being evaluated. Those skilled in the art will understand that when sufficient similarity exists, it can be justified to rely on a particular possible reference or control and / or to which comparisons are made.

[0139] Ribonucleic acid (RNA) or polynucleotide: As used herein, the terms “ribonucleic acid,” “RNA,” or “polynucleotide” refer to a polymer of ribonucleotides. In some embodiments, RNA is single-stranded. In some embodiments, RNA is double-stranded. In some embodiments, RNA comprises both single-stranded and double-stranded portions. In some embodiments, RNA may comprise a backbone structure as defined above in the definition of “nucleic acid / polynucleotide.” RNA may be regulatory RNA (e.g., siRNA, microRNA, etc.) or messenger RNA (mRNA). In some embodiments, RNA is mRNA. In some embodiments, where RNA is mRNA, RNA typically includes a poly(A) region at its 3' end. In some embodiments, where RNA is mRNA, RNA typically includes a technically recognized cap structure at its 5' end, for example, for recognizing mRNA and linking it to a ribosome to initiate translation. In some embodiments, RNA is synthetic RNA. Synthetic RNA includes RNA synthesized in vitro (e.g., by enzymatic synthesis and / or by chemical synthesis).

[0140] Ribonucleotides: As used herein, the term "ribonucleotide" encompasses both unmodified and modified ribonucleotides. For example, unmodified ribonucleotides include purine bases adenine (A) and guanine (G), and pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications, including but not limited to, for example, (a) terminal modifications, such as 5' modifications (e.g., phosphorylation, dephosphorylation, coupling, reverse bond, etc.), 3' modifications (e.g., coupling, reverse bond, etc.), (b) base modifications, such as substitution with a modified base, a stable base, a destabilized base, or a base or coupled base paired with an amplified chaperone base, (c) sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions, and (d) internucleotide bond modifications, including modifications or substitutions of phosphodiester bonds. The term "ribonucleotide" also covers ribonucleotide triphosphates, including both modified and unmodified ribonucleotide triphosphates.

[0141] Risk: As will be understood from the context, “risk” for a disease, condition, and / or disorder refers to the likelihood that a particular individual will develop a disease, condition, and / or disorder. In some embodiments, risk is expressed as a percentage. In some embodiments, risk is expressed as a risk relative to the risk associated with a reference sample or reference sample group. In some embodiments, the reference sample or reference sample group has a known risk of a disease, condition, disorder, and / or event. In some embodiments, the reference sample or reference sample group is derived from individuals comparable to the particular individual. In some embodiments, risk may reflect one or more genetic attributes, such as those that predispose an individual to (or prevent) a particular disease, condition, and / or disorder. In some embodiments, risk may reflect one or more epigenetic events or attributes and / or one or more lifestyle or environmental events or attributes.

[0142] Specificity: When used herein to refer to an active reagent, the term "specificity" should be understood by those skilled in the art to mean that the reagent distinguishes potential target entities, states, or cells. For example, in some embodiments, a reagent is said to bind "specifically" to its target if it preferentially binds to its target in the presence of one or more competing alternative targets. In many embodiments, specific interactions depend on the presence of specific structural features of the target entity (e.g., antigenic determinants, gaps, binding sites). It should be understood that specificity need not be absolute. In some embodiments, specificity may be evaluated relative to the specificity of the target-binding portion to one or more other potential target entities (e.g., competitors). In some embodiments, specificity is evaluated relative to the specificity of a reference specific binding portion. In some embodiments, specificity is evaluated relative to the specificity of a reference non-specific binding portion.

[0143] Substituted or optionally substituted: As described herein, the compounds of the present invention may contain an "optionally substituted" moiety. Generally, the term "substituted" (whether preceded by the term "optionally") means that one or more hydrogens of the specified moiety are substituted by suitable substituents. "Substituted" applies to self-structures (e.g., It means at least ;and It means at least , , ,or The group is explicitly or implicitly represented by one or more hydrogen atoms. Unless otherwise indicated, an "optionally substituted" group may have suitable substituents at each substituted position of the group, and the substituents may be the same or different at each position when more than one position in any given structure is substituted by more than one substituent selected from the specified group. The substituent combinations contemplated in this invention are preferably those that result in the formation of stable or chemically viable compounds. As used herein, the term "stable" means a compound that remains substantially unchanged when subjected to permissible production, detection, and, in some embodiments, permissible recovery, purification, and use for one or more purposes provided herein. A group described as "substituted" preferably has 1 to 4 substituents, more preferably 1 or 2 substituents. A group described as "optionally substituted" may be unsubstituted or "substituted" as described above.

[0144] The suitable monovalent substituent on the substituted carbon atom of the "optionally substituted" group is independently a halogen; -(CH2) 0-4 R 0 ;-(CH2) 0-4 OR 0 ;-O(CH2) 0-4 R 0 -O-(CH2) 0-4 C(O)OR 0 ;-(CH2) 0-4 CH(OR 0 )2;-(CH2) 0- 4SR 0 ;-(CH2) 0-4 Ph, which can be transmitted via R 0 Substitution; -(CH2) 0-4 O(CH2) 0-1 Ph, which can be transmitted via R 0 Substitution; -CH=CHPh, which can be obtained via R 0 Substitution; -(CH2) 0-4 O(CH2) 0-1 -pyridyl group, which can be transmitted via R 0 Substitution; -NO2; -CN; -N3; ​​-(CH2) 0-4 N(R 0 )2;-(CH2) 0-4 N(R 0 )C(O)R 0 ;-N(R 0 )C(S)R 0 ;-(CH2) 0-4 N(R 0 )C(O)NR 0 2; -N(R) 0 )C(S)NR 02;-(CH2) 0-4 N(R 0 )C(O)OR 0 ;-N(R 0 )N(R 0 )C(O)R 0 ;-N(R 0 )N(R 0 )C(O)NR 0 2;-N(R 0 )N(R 0 )C(O)OR 0 ;-(CH2) 0-4 C(O)R 0 ;C(S)R 0 ;-(CH2) 0-4 C(O)OR 0 ;-(CH2) 0-4 C(O)SR 0 ;-(CH2) 0-4 C(O)OSiR 0 3;-(CH2) 0-4 OC(O)R 0 ;-OC(O)(CH2) 0-4 SR 0 ;-(CH2) 0-4 SC(O)R 0 ;-(CH2) 0-4 C(O)NR 0 2;-C(S)NR 0 2;-C(S)SR 0 ;-SC(S)SR 0 、-(CH2) 0- 4OC(O)NR 0 2;-C(O)N(OR 0 )R 0 ;-C(O)C(O)R 0 ;-C(O)CH2C(O)R 0 ;-C(NOR 0 )R 0 ;-(CH2) 0-4 SSR 0 ;-(CH2) 0-4 S(O)2R 0 ;-(CH2) 0-4 S(O)2OR 0 ;-(CH2) 0-4 OS(O)2R 0 ;-S(O)2NR 0 2;-(CH2) 0-4 S(O)R0 ;-N(R 0 )S(O)2NR 0 2; -N(R) 0 )S(O)2R 0 ;-N(OR) 0 )R 0 ;-C(NH)NR 0 2; -P(O)2R 0 ;-P(O)R 0 2; -OP(O)R 0 2; -OP(O)(OR 0 )2; SiR 0 3; -(C 1-4 (linear or branched alkylene)ON(R) 0 )2; or -(C 1-4 (straight-chain or branched alkylene)C(O)ON(R) 0 )2, where each R 0 It can be substituted and independently formed as hydrogen or C as defined below. 1-6 Aliphatic, -CH2Ph, -O(CH2) 0-1 Ph, -CH2- (5- to 6-membered heteroaryl ring), or 3- to 6-membered saturated, partially unsaturated, or aromatic ring having 0-4 independent heteroatoms selected from nitrogen, oxygen, or sulfur, or, despite the above definition, two independently occurring R... 0 Together with intercalated atoms, they form 3 to 12 saturated, partially unsaturated or aromatic monocyclic or bicyclic rings with 0 to 4 independent heteroatoms selected from nitrogen, oxygen or sulfur, which may be substituted as defined below.

[0145] R 0 (or through two independently occurring R) 0 Suitable monovalent substituents on the ring formed together with the intercalated atoms are independently halogens, -(CH2). 0-2 R l -(halogenated R) l -(CH2) 0-2 OH, -(CH2) 0-2 OR l -(CH2) 0-2 CH(OR l 2. -O(halogenated R) l -CN, -N3, -(CH2) 0-2 C(O)R l -(CH2) 0-2 C(O)OH, -(CH2) 0-2 C(O)OR l -(CH2) 0-2 SR l-(CH2) 0-2 SH, -(CH2) 0-2 NH2、-(CH2) 0-2 NHR l -(CH2) 0-2 NR l 2, -NO2, -SiR l 3. -OSiR l 3. -C(O)SR l -(C 1-4 (straight-chain or branched alkylene)C(O)OR l or -SSR l , where each R l Unsubstituted or, in the case of a preceding "halogen group," substituted with only one or more halogens, and independently selected from C 1-4 Aliphatic, -CH2Ph, -O(CH2) 0-1 Ph or a 3- to 6-membered saturated, partially unsaturated, or aromatic ring having 0-4 independent heteroatoms selected from nitrogen, oxygen, or sulfur. R 0 Suitable divalent substituents on saturated carbon atoms include =O and =S.

[0146] Suitable divalent substituents on the saturated carbon atom of the "optionally substituted" group include the following: =O ("side oxygen"), =S, =NNR * 2、=NNHC(O)R * =NNHC(O)OR * =NNHS(O)2R * =NR * =NOR * -O(C(R) * 2)) 2-3 O- or -S(C(R) * 2)) 2-3 S-, where each R appears independently * Selected from hydrogen, and substituted C as defined below. 1-6 Aliphatic, or an unsubstituted 5- to 6-membered saturated, partially unsaturated, or aromatic ring having 0-4 independent heteroatoms selected from nitrogen, oxygen, or sulfur. Suitable divalent substituents attached to the ortho-substituted carbon of the "optionally substituted" group include: -O(CR * 2) 2-3 O-, where each R appears independently * Selected from hydrogen, and substituted C as defined below. 1-6 Aliphatic, or having 0-4 unsubstituted 5-6 member saturated, partially unsaturated or aromatic rings independently selected from nitrogen, oxygen or sulfur.

[0147] R *Suitable substituents on aliphatic groups include halogens, -R l -(halogenated R) l -OH, -OR l -O(halogenated R) l -CN, -C(O)OH, -C(O)OR l -NH2, -NHR l -NR l 2 or -NO2, where each R l It is unsubstituted or, in the case of a preceding "halogen group", substituted with only one or more halogens, and is independently C. 1-4 Aliphatic, -CH2Ph, -O(CH2) 0-1 Ph or a 3- to 6-membered saturated, partially unsaturated, or aromatic ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.

[0148] Suitable substituents on the substituted nitrogen of the "optionally substituted" group include -R † -NR † 2. -C(O)R † -C(O)OR † -C(O)C(O)R † -C(O)CH2C(O)R † -S(O)2R † -S(O)2NR † 2. -C(S)NR † 2. -C(NH)NR † 2, or -N(R) † )S(O)2R † ; where each R † Independently, hydrogen, or substituted C as defined below 1-6 Aliphatic, unsubstituted -OPh, or unsubstituted 3- to 6-membered saturated, partially unsaturated, or aromatic rings having 0-4 independently selected heteroatoms chosen from nitrogen, oxygen, or sulfur, or, despite the above definition, two independently occurring R... † Together with intercalated atoms, they form unsubstituted 3 to 12-membered saturated, partially unsaturated, or aromatic monocyclic or bicyclic rings with 0 to 4 independent heteroatoms selected from nitrogen, oxygen, or sulfur.

[0149] R † Suitable substituents on the aliphatic group can be halogens, -R l -(halogenated R) l -OH, -OR l -O(halogenated R) l -CN, -C(O)OH, -C(O)OR l -NH2, -NHRl -NR l 2 or -NO2, where each R l It is unsubstituted or, in the case of a preceding "halogen group", substituted with only one or more halogens, and is independently C. 1-4 Aliphatic, -CH2Ph, -O(CH2) 0-1 Ph or a 3- to 6-membered saturated, partially unsaturated, or aromatic ring having 0 to 4 heteroatoms independently selected from nitrogen, oxygen, or sulfur.

[0150] Individual: As used herein, the term "individual" refers to an organism to which the compositions described herein are intended to be administered, for example, for experimental, diagnostic, preventative, and / or therapeutic purposes. Typical individuals include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, domestic pets, etc.) and humans. In some embodiments, the individual is a human individual. In some embodiments, the individual suffers from a disease, condition, or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, the individual is susceptible to a disease, condition, or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, the individual exhibits one or more symptoms or features of a disease, condition, or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, the individual exhibits one or more nonspecific symptoms of a disease, condition, or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, the individual does not exhibit any symptoms or features of a disease, condition, or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, the individual is a person having one or more characteristics that characterize susceptibility or risk to a disease, condition, or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, the individual is a patient. In some embodiments, the individual is an individual who has received and / or has received diagnostic and / or therapeutic treatments.

[0151] "Having": An individual who has been diagnosed with and / or exhibits one or more symptoms of a disease, condition and / or disorder (e.g., obesity, obesity-related disorders, etc.) is considered to have a disease, condition and / or disorder.

[0152] Susceptible: An individual "susceptible" to a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.) is someone who has a higher risk of developing that disease, condition, and / or disorder than the general population. In some embodiments, an individual susceptible to a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.) may not be diagnosed with that disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, an individual susceptible to a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.) may express symptoms of that disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, an individual susceptible to a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.) may not express symptoms of that disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, an individual susceptible to a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.) will develop that disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, an individual susceptible to a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.) will not develop that disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.).

[0153] Therapy: The term "therapy" refers to the administration or delivery of an agent or intervention that has a therapeutic effect and / or induces a desired biological and / or pharmacological effect (e.g., an effect that has been shown to be statistically likely when administered to a relevant population). In some embodiments, a therapeutic agent or therapy is any substance that can be used to reduce, improve, alleviate, suppress, prevent, or delay the onset, severity, and / or incidence of one or more symptoms or features of a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, a therapeutic agent or therapy is a medical intervention (e.g., surgery, radiation, phototherapy) that can be performed to reduce, alleviate, suppress, or induce one or more symptoms or features of a disease, condition, and / or disorder, delay its onset, reduce its severity, and / or decrease its incidence.

[0154] Treatment: As used herein, the terms “treat,” “treatment,” or “treating” refer to any method used to partially or completely reduce, improve, alleviate, suppress, prevent, delay the onset, reduce the severity, and / or decrease the incidence of one or more symptoms or features of a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). Treatment may be administered to individuals who do not express signs of a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.). In some embodiments, treatment may be administered to individuals who express only early signs of a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.), for example, for the purpose of reducing the risk of developing a pathology associated with that disease, condition, and / or disorder. In some embodiments, treatment may be administered to individuals in later stages of a disease, condition, and / or disorder (e.g., obesity, obesity-related disorders, etc.).

[0155] The compounds disclosed herein include those generally described above, and are further described by the classes, subclasses and species disclosed herein. Unless otherwise indicated, the following definitions should apply as used herein. For the purposes of this disclosure, chemical elements are identified according to the periodic table, CAS version, Handbook of Chemistry and Physics, 75th edition. Furthermore, the general principles of organic chemistry are described in “Organic Chemistry”, Thomas Sorrell, University Science Books, Sausalito: 1999 and “March's Advanced Organic Chemistry”, 5th edition, edited by Smith, MB and March, J., John Wiley and Sons, New York: 2001, the entire contents of which are hereby incorporated by reference.

[0156] Unless otherwise stated, the structures described herein are intended to include all stereoisomers (e.g., mirror-image or non-mirror-image isomers) of the structure, as well as all geometric or configurational isomers of the structure. For example, the R and S configurations of each stereocenter are considered as part of this disclosure. Therefore, single stereochemical isomers of the provided compounds, as well as mirror-image, non-mirror-image, and geometric (or configurational) mixtures, are within the scope of this disclosure. For example, in some cases, the provided compounds exhibit one or more stereoisomers of the compound, and unless otherwise indicated, each stereoisomer represents individually and / or as a mixture. Unless otherwise stated, all tautomeristic forms of the provided compounds are within the scope of this disclosure.

[0157] Unless otherwise indicated, the structures described herein are intended to include compounds that differ only in the presence of one or more isotopically enriched atoms. For example, compounds having this structure (including those with hydrogen replaced by deuterium or tritium, or carbon replaced by carbon enriched by 13C or 14C) are within the scope of this disclosure. Detailed Implementation

[0158] Incretins and their use in disease treatment Incretins are peptide hormones released in the gastrointestinal (GI) tract in response to glucose consumption. They stimulate the pancreas to secrete insulin and reduce glucagon production, thereby lowering blood glucose levels. Incretins exert their effects by binding to their corresponding receptors on pancreatic β-cells, leading to insulin release. Glucagon-like peptide-1 (GLP1) and glucose-dependent insulinotropic peptide (GIP) are two incretins discovered for their role in postprandial insulin secretion. GIP is primarily responsible for releasing insulin in response to glucose intake. GLP1 stimulates satiety, slows gastric emptying, reduces glucagon secretion, and decreases food intake, leading to weight loss. It has been shown that activation of GIP and GLP1 receptors during insulin secretion has an additive effect. See Chim, US Pharm. 2022; 47(10):18-22, which is incorporated herein by reference in its entirety.

[0159] Due to their roles in controlling blood sugar and satiety, incretins and incretin mimics have the potential to treat a variety of diseases, including obesity, prediabetes, type 2 diabetes (T2D, and its complications), early type 1 diabetes (e.g., within 3 months of T1D diagnosis), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), cardiovascular (CV) diseases (e.g., characterized by major cardiovascular events (MACE), including CV death, non-fatal myocardial infarction, non-fatal stroke, or heart failure with preserved ejection fraction (HFpEF)), kidney disease, and an increased risk of premature death. These chronic diseases are prevalent worldwide and often coexist as comorbidities.

[0160] obesity Obesity is the most prevalent chronic disease worldwide, affecting approximately 650 million adults. Obesity is considered a starting point and key contributing factor to prediabetes, type 2 diabetes (T2D, and its complications), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), cardiovascular disease and kidney disease, and premature death. Obesity imposes a substantial economic burden, including additional direct healthcare costs, productivity costs (absence, attendance, disability support, premature death), transportation costs (including an increased CO2 footprint), and human capital accumulation costs (school absences, highest level of education achieved). It is estimated that by 2030, obesity (BMI > 30 kg / m²) will become a major health concern. 2 The number of people with this condition will exceed one billion, of whom approximately 10% will suffer from severe type III obesity (BMI > 40 kg / m²). 2 Half of all obese men live in just nine countries: the United States, China, India, Brazil, Mexico, Russia, Egypt, Germany, and Turkey. Childhood obesity is also rising sharply worldwide.

[0161] Obesity was only declared a disease by the American Association of Clinical Endocrinologists (AACE) in 2011 and is managed based on severity, starting with lifestyle / behavioral interventions and increasing physical activity, followed by drug therapy, and finally bariatric surgery.

[0162] Based on short-term studies, it is recommended that overweight or obese individuals with type 2 diabetes mellitus (T2D) lose weight. The Look AHEAD study, which investigated the long-term effects of weight loss on cardiovascular disease in over 5,100 people with T2D, was discontinued in 2012 after nearly 10 years because it showed that intensive lifestyle interventions focused on weight loss did not reduce the rate of cardiovascular events in overweight or obese adults with T2D.

[0163] In summary, drug therapy for obesity has had limited success to date, with only modest placebo-corrected weight loss, such as 3% for Xenical® (orlistat) and 4-5% for Belviq® (lorcaserin) and Contrave® (naltrexone SR / bupropion SR), while also carrying the social restriction side effects of Xenical® and adverse central nervous system effects of its centrally acting agents. The FDA rejected Acomplia® (rimonabant) due to concerns that its use might increase suicidal ideation and depression.

[0164] The most effective intervention for obesity remains bariatric surgery; however, due to perceived serious complications (including mortality), only 1% of eligible patients undergo the procedure. Bariatric surgery has been shown to significantly alter the release of endocrine hormones, which has sparked research interest in endogenous nutrient-stimulating hormone pathways, aiming to essentially mimic the effects of bariatric surgery through chemical agents called "incretin mimics."

[0165] Existing treatments using incretin mimics Current therapies using incretin mimics include glucagon-like peptide-1 (GLP1) receptor agonists, such as Trulicity® (dulaglutide), Byetta® (exenatide), Ozempic® / Rybelsus® (semaglutide, injectable / oral), Victoza® (liraglutide), and Suliqua® (lixisenatide, in combination with insulin glargine only). These receptor agonists are approved for lowering blood glucose in people with type 2 diabetes without requiring continuous monitoring of blood glucose levels. Additional benefits include weight loss (2-4%) and positive effects on cardiovascular and renal parameters. Recent studies have combined the activity of GLP1 receptor agonists with GIP receptor agonists and / or glucagon (GCG) receptor agonists (dual / triple agonists) to achieve better glycemic control and greater weight loss. The dual GLP1 / GCG receptor agonist SAR425899 demonstrated glycemic control and weight loss; however, the program was discontinued in 2019 due to unacceptable gastrointestinal side effects. Recently, the dual GLP1 / GIP receptor agonist tirzepatide (now marketed as Mounjaro®) was approved as an injectable medication for adults with type 2 diabetes (T2D), used in conjunction with diet and exercise to improve glycemic control. GIP functions by regulating energy balance through cell surface receptor signaling in the brain and adipose tissue. The SURPASS-2 study demonstrated tirzepatide's non-inferiority and superiority over smegglutide in lowering glycemic control. However, weight loss was only a secondary endpoint.

[0166] Phase 3 trials with weight loss as the primary endpoint have been conducted in non-diabetic, overweight / obese patient populations – see the SCALE (liraglutide), STEP-1 (semaglutide), and SURMOUNT-1 (telpoglide) trials. The SCALE trial demonstrated a 5.4% placebo-corrected weight loss and a lower progression to prediabetes with a dose of 3 mg liraglutide once daily plus a lifestyle intervention. The STEP-1 trial produced a 12.4% weight loss in overweight or obese participants receiving 2.4 mg semaglutide once weekly plus a lifestyle intervention. The SURMOUNT-1 trial showed a 17.8% placebo-corrected weight loss at the highest weekly dose of 15 mg telpoglide. Cardiometabolic measures were improved in all studies. Gastrointestinal side effects were as expected, especially at the start of treatment, and were manageable. Approved obesity products now include Saxenda® (liraglutide 3 mg daily injection) and Wegovy® (semaglutide once weekly injection). In October 2022, the FDA granted semaglutide Fast Track designation for the treatment of adults with obesity or overweight and weight-related comorbidities, based on rolling submissions of further data from the SURMOUNT study series, which could lead to approval for the indication in early 2024. Topline results recently published from the SELECT trial, testing Wegovy® in patients with established cardiovascular disease and overweight / obese but without type 2 diabetes (T2D), found that 2.4 mg semaglutide reduced the risk of major adverse cardiovascular events by 20% in overweight or obese adults. Another study using an incretin mimic was OASIS-1, which investigated oral administration of 25 or 50 mg semaglutide in non-diabetic overweight / obese patients. A 14 mg maintenance dose of oral semaglutide has been approved as Rybelsus® for T2D. Recent Phase 2 results for the triple receptor agonist (GIP / GLP1 / GCG receptor) retratutide LY3437943 have been published, showing a placebo-corrected weight loss in 22.1% of obese individuals after 48 weeks of treatment. Recent studies have also shown benefit for semaglutide in early-stage T1D (within 3 months of T1D diagnosis). In a small study, all participants no longer required mealtime insulin, and most participants no longer required basal insulin.

[0167] Exemplary treatments for obesity / T2D using incretin mimics are shown in Table 1 below.

[0168] Table 1: Exemplary incretin mimics for the treatment of obesity / T2D This disclosure particularly recognizes current problems in the market for incretin analogues for the treatment of obesity, prediabetes, type 2 diabetes, early type 1 diabetes, NAFLD, NASH, cardiovascular disease, kidney disease, and increased risk of premature death, including but not limited to limited supply, high prices, lack of health insurance coverage for such treatments, frequent injections (e.g., once a week), high injection volumes, and gastrointestinal side effects.

[0169] This disclosure provides, in particular, a more effective and cost-efficient method for treating obesity, prediabetes, type 2 diabetes (T2D), early type 1 diabetes (T1D), NAFLD, NASH, cardiovascular disease, kidney disease, and increased risk of premature death by delivering incretins and incretin mimics (collectively covered by the term "incretin agent") encoded by one or more polynucleotides. In some embodiments, the polynucleotide encoding the incretin agent is used for therapeutic treatment of obesity and / or prediabetes, T2D, early type 1 diabetes, NAFLD, NASH, cardiovascular disease, kidney disease, and increased risk of premature death (e.g., obesity-related diseases), which has improved properties compared to known incretin mimic therapies, including requiring fewer injections (no more than once a week), smaller injection volumes (e.g., no more than 0.5 ml), and fewer or less severe side effects. Additionally, the polynucleotide used to deliver the incretin agent provides expression of the incretin agent in cells at therapeutically relevant levels, comparable to the dosage of current peptide-based therapies.

[0170] Polynucleotides for delivering incretins This disclosure utilizes RNA technology, in particular, as a modality for the direct expression in an individual of incretins as a novel class of therapeutic agents that activate GLP1, GIP, and / or GCG receptors to effectively treat disease states such as obesity, prediabetes, type 2 diabetes, early type 1 diabetes, NAFLD, NASH, cardiovascular disease, kidney disease, and / or an increased risk of premature death. In some embodiments, the polynucleotides described herein encode incretins found in nature, or fragments or mutants thereof. In some embodiments, the polynucleotides described herein encode incretins modified from their natural form, or fragments or mutants thereof.

[0171] As used herein, the term "incretin agent" refers to a reagent comprising incretin or an incretin mimic (whereincretin and incretin mimic are collectively referred to herein as "incretin peptides"). Exemplary incretins include GLP1, GIP, and GCG. Exemplary incretin mimics are shown, for example, in Table 1. In some embodiments, the incretin agent is a biologically active portion or fragment of incretin or an incretin mimic. In some embodiments, the incretin agent comprises an incretin peptide as part of a fusion. For example, in some embodiments, the incretin agent is an incretin peptide fused to another peptide portion (e.g., a half-life extension (HLE) domain).

[0172] In some embodiments, the incretin agent comprises a GLP1 receptor agonist, such as GLP1. In some embodiments, the incretin agent comprises a GIP receptor agonist, such as GIP. In some embodiments, the incretin agent comprises dual GIP and GLP1 receptor agonists. In some embodiments, the incretin agent comprises triple GIP, GLP1, and GCG receptor agonists.

[0173] Exemplary incretin peptide In some embodiments, the incretin agent comprises a wild-type (i.e., unmutated) incretin peptide sequence or a fragment thereof. For example, in some embodiments, the incretin agent comprises any of the incretin peptides shown in SEQ ID NO: 5-15 and 62-64.

[0174] In some embodiments, the incretin agent comprises an incretin peptide or fragment thereof having at least one mutated amino acid residue compared to a wild-type reference sequence. In some embodiments, the mutated amino acid residue comprises a natural amino acid residue substituted by another natural amino acid residue. In some embodiments, the incretin agent comprises any of the incretin peptides shown in SEQ ID NO: 5-10. In some embodiments, the mutated amino acid residue may confer dual or triple activation properties to the incretin agent. For example, in some embodiments, one or more amino acid substitutions are introduced into the GLP1, GIP, or GCG peptide sequence (e.g., as shown in SEQ ID NO: 12-15 and 62-64) to confer binding properties to two or more of the GLP1, GIP, and / or GCG receptors.

[0175] In some embodiments, the incretin agent has an amino acid sequence that is at least 85%, at least 90%, at least 95%, or 100% identical to any of the incretin peptides detailed in Table 2. In some embodiments, the incretin agent comprises any of the incretin peptides detailed in Table 2 below, or combinations thereof, or mutants thereof.

[0176] Table 2: Exemplary incretin peptides (mutations are shown in bold) Linking peptides In some embodiments, the incretin agents described herein comprise a single incretin peptide (this configuration is referred to herein as "I:1x") (see, for example) Figure 3 In some embodiments, the incretin peptide is fused to another peptide (e.g., a half-life extension (HLE) domain) via a linker peptide (see example...). Figures 4-14 In some embodiments, the linker peptide contains at least one Gly (G) amino acid residue. Suitable linkers can be readily selected and can have various lengths, such as 1 amino acid (e.g., Gly) to 25 amino acids, 2 amino acids to 15 amino acids, 3 amino acids to 12 amino acids, including 4 amino acids to 10 amino acids, 5 amino acids to 9 amino acids, 6 amino acids to 8 amino acids, or 7 amino acids to 8 amino acids (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids). Exemplary linkers include glycine polymers (G). n Glycine-serine polymers (including, for example, (GS)) n (GGGGS: SEQ ID NO: 1) n and (GGGS: SEQ ID NO: 2) n(where n is an integer of at least one), glycine-alanine polymers, alanine-serine polymers, and other flexible linker peptides known in the art. Glycine and glycine-serine polymers are relatively unstructured and therefore capable of acting as neutral linkers between components. Glycine is even closer to a significantly larger phi-psi space than alanine and is much less restricted than residues with longer side chains (see Scheraga, Rev. Computational Chem. 11:173-142 (1992)). In some embodiments, the linker peptide comprises the amino acid sequence of SEQ ID NO: 3 (GGGGSGGGGSGGGGSGGGS or "(G4S)4" linker peptide) or SEQ ID NO: 4 (GGGGSGGGGSGGGGSGGGSGGGS or "(G4S)5" linker peptide). In some embodiments, the linker peptide comprises the amino acid sequence of SEQ ID NO: 68 (GGGGSGGGGS or "(G4S)2" linker peptide), SEQ ID NO: 156 (GGGSGGGS or "(G3S)2" linker peptide), SEQ ID NO: 157 (GGGGSGGGGSGGGGS or "(G4S)3" linker peptide) or GGGGSGGGS (SEQ ID NO: 186).

[0177] In some embodiments, an incretin agent comprises an incretin peptide linked to another peptide (e.g., the HLE domain described herein) using a (G4S)3 linker peptide. In some embodiments, an incretin agent described herein comprises an incretin peptide linked to another incretin peptide using a (G4S)2 linker peptide.

[0178] cleavage site In some embodiments, the incretin peptide is fused to another peptide (e.g., another incretin peptide and / or a half-life extension (HLE) domain) via a protease cleavage site, such as a furin protease cleavage site (e.g., a peptide including the motif RXK / RR SEQ ID NO: 158, such as SEQ ID NO: 160 RRKR or SEQ ID NO: 153 NVRRKR). In some embodiments, the protease cleavage site (e.g., a furin protease cleavage site) comprises any one of SEQ ID NO: 160 RRKR, SEQ ID NO: 153 NVRRKR, SEQ ID NO: 189 RKKR, SEQ ID NO: 190 RMQR, or SEQ ID NO: 191 VFRR. The terms “furin protease cleavage site” and “furin protease recognition site” are used interchangeably herein and refer to a sequence that facilitates furin protease cleavage.

[0179] In some embodiments, the furin cleavage site is operatively linked to a linker peptide (e.g., a glycine linker peptide, such as a (G4S)2 linker peptide) (e.g., on its C-terminal side). In some embodiments, the incretin agent comprises a plurality of (e.g., 2, 3, 4 or more) incretin peptides, each separated from the furin cleavage site and optionally a linker peptide (e.g., a (G4S)2 linker peptide located on the N-terminal side of the furin cleavage site).

[0180] The incretin agents described herein include furin cleavage sites, particularly in cases where the incretin peptide is fused to another peptide (e.g., another incretin peptide and / or a half-life extension (HLE) domain), to facilitate the proper cleavage of the incretin peptide from the other peptide and allow the incretin peptide to be fully processed and functional after its expression.

[0181] Without being bound by any theory, in the context of polynucleotides encoding incretins, the type of cleavage sites and recognition sites is important for ensuring proper processing of the N-terminus of the incretin peptide. Certain furin cleavage and recognition sites, and the placement of those sites relative to the incretin peptide within the incretin agent, can create alternative processing or cleavage sites, ultimately altering the final amino acid sequence of the mature incretin peptide. In such relatively small peptides, such as GLP1 or GIP (or their mutants, and other peptides of similar size / properties), any variation in amino acid residues can affect the peptide's biological activity. In some embodiments, furin recognition / cleavage sites are selected and located within the incretin agent to promote proper cleavage of the N-terminus of the incretin peptide, or in other words, to produce a "scarless" N-terminus of the incretin peptide to maintain its biological activity. Figure 20 and Figure 21 A schematic diagram showing the location of the theoretical cleavage sites of some exemplary signal peptides. Figure 20 The A8G mutation indicates that it promotes the proper N-terminal processing of the GLP1 incretin peptide with the husec signal peptide. Figure 21 The A2G mutation indicates that it promotes the proper N-terminal processing of the GIP incretin peptide with the husec signal peptide.

[0182] This concept and utilization of furin cleavage sites at the N-terminus of incretin peptides linked to another peptide can also be applied to other intestinal peptides (e.g., glucagon) and / or other peptides of comparable size / properties to GLP1 and GIP as described herein. This is particularly important in the context of delivering incretin agents (or other similar peptides) as one or more polynucleotides encoding incretin agents. Such delivery requires proper translation of the protein within the cell, in addition to post-translational processing (including proper cleavage of one incretin peptide from another). In some embodiments, incretin agents comprising one or more incretin peptides fused to another peptide as described herein have been engineered and generated to include protease cleavage sites within the incretin agent, such that cleavage of the incretin peptide from the other peptide is accurate and does not affect the amino acid sequence of the mature peptide (i.e., producing scarless N-termini). A “scarless” N-terminus, as referred to herein, includes a peptide that has been cleaved from another peptide via a cleavage site, wherein the cleavage occurs in such a manner that no remaining amino acids are not part of the mature peptide and all amino acids of the mature peptide are retained at the N-terminus of the peptide. The scarless N-terminus of incretin peptides (and other similar peptides, such as other intestinal peptides, such as glucagon) allows the peptides to retain their proper function after processing into mature peptides.

[0183] In some embodiments, the furin cleavage site is positioned immediately adjacent to the 5' of the second incretin peptide in the incretin agent to ensure that cleavage on the second incretin peptide produces a scarless N-terminus. Such cleavage can be important for maintaining the function of mature incretin peptides.

[0184] In some embodiments, the furin cleavage site is selected to be compatible with the N-terminal sequence of an incretin peptide (e.g., a wild-type or mutant incretin peptide, such as GIP with an A2G mutation or GLP1 with an H1Y mutation and / or an A8G mutation). Mutations as described herein can be introduced into the incretin peptide to promote efficient cleavage and maintain a scar-free N-terminus of the incretin peptide.

[0185] As disclosed herein, various furin cleavage sequences can be utilized to facilitate appropriate cleavage. For example, in some embodiments, the furin cleavage site is, for instance, NVRRKR (SEQ ID NO: 153), which is derived from the human MT-MMP 1 protein. This furin cleavage site is derived from human proteins and is compatible with the N-terminal sequences of GIP and GLP1 incretin peptides (including wild-type and mutant GIP and GLP1 incretin peptides). Other human furin cleavage sequences can be utilized because different human furin cleavage sequences can exhibit different cleavage efficiencies based on adjacent amino acid sequences (see Izidoro et al., Archives of biochemistry and biophysics 2009, 487.2, 105-114, which are incorporated herein by reference in their entirety).

[0186] In some embodiments, mutations are introduced into the incretin peptides described herein to promote signal peptide cleavage and produce a mature incretin peptide with a scar-free N-terminus. In some embodiments, such mutations include the A2G mutation in GIP incretin peptides (e.g., GIP (1-42) incretin peptides). In some embodiments, such mutations include the A8G mutation in GLP1 (7-37) incretin peptides. In some embodiments, such mutations may also prolong the half-life of the incretin agent, for example, by preventing proteolysis of amino acids at a second position of the incretin peptide (thus avoiding the production of a truncated mature incretin peptide with one or two amino acids missing from the N-terminus). In some embodiments, the mutation is selected to increase the probability of correct cleavage (i.e., cleavage at the N-terminus that prevents the mature incretin peptide from being truncated). In some embodiments, the compatibility of the signal peptide and cleavage site sequences used in the incretin agents described herein depends on the specific amino acid sequences of adjacent incretin peptides, particularly the amino acid residues at the N-terminus.

[0187] Incretins containing multiple incretin peptides In some embodiments, an incretin agent comprises a single incretin peptide. In some embodiments, an incretin agent comprises one or more incretin peptides. In some embodiments, the two or more incretin peptides included in the incretin agent described herein are the same or derived from the same incretin peptide (e.g., a GLP1 receptor agonist). In some embodiments, the two or more incretin peptides included in the incretin agent described herein are different or derived from different incretin peptides.

[0188] In some embodiments, the incretin agent comprises a combination of incretin peptides, such as fused to a single polypeptide chain. For example, in some embodiments, the incretin agent comprises a GLP1 receptor agonist (e.g., a GLP1 peptide or a fragment or mutant thereof) and a GIP receptor agonist (e.g., a GIP peptide or a fragment or mutant thereof). In some embodiments, the incretin agent comprises one or more incretin peptides selected from SEQ ID NO: 5-15 and 62-64 and one or more incretin peptides selected from SEQ ID NO: 5-15 and 62-64. In some embodiments, the incretin agent comprises a GLP1 receptor agonist (e.g., a GLP1 peptide or a fragment or mutant thereof), a GIP receptor agonist (e.g., a GIP peptide or a fragment or mutant thereof), and a GCG receptor agonist (e.g., a fragment of a GCG peptide or a mutant thereof). In some embodiments, the incretin agent comprises more than one of the same or different GLP1 receptor agonists (e.g., a GLP1 peptide or a fragment or mutant thereof). In some embodiments, the incretin agent comprises more than one of the same or different GIP receptor agonists (e.g., a GIP peptide or a fragment or mutant thereof). In some embodiments, the incretin agent comprises more than one copy of the same incretin peptide and / or a combination of different incretin peptides.

[0189] In some embodiments, an incretin peptide is fused to another incretin peptide via a linker peptide. In some embodiments, the linker peptide contains at least one Gly (G) amino acid residue. Suitable linker peptides can be readily selected and can have various lengths, such as 1 amino acid (e.g., Gly) to 25 amino acids, 2 amino acids to 15 amino acids, 3 amino acids to 12 amino acids, including 4 amino acids to 10 amino acids, 5 amino acids to 9 amino acids, 6 amino acids to 8 amino acids, or 7 amino acids to 8 amino acids (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 amino acids). Exemplary linker peptides include glycine polymers (G). nGlycine-serine polymers (including, for example, (GS)) n (GGGGS: SEQ ID NO: 1) n and (GGGS: SEQ ID NO: 2) n (where n is an integer of at least one), glycine-alanine polymers, alanine-serine polymers, and other flexible linker peptides known in the art. Glycine and glycine-serine polymers are relatively unstructured and therefore capable of acting as neutral linkers between components. Glycine is even closer to a significantly larger phi-psi space than alanine and is much less restricted than residues with longer side chains (see Scheraga, Rev. Computational Chem. 11:173-142 (1992)). In some embodiments, the linker peptide comprises the amino acid sequence of SEQ ID NO: 3 (GGGGSGGGGSGGGGSGGGS or "(G4S)4" linker peptide) or SEQ ID NO: 4 (GGGGSGGGGSGGGGSGGGSGGGS or "(G4S)5" linker peptide). In some embodiments, the linker peptide comprises the amino acid sequence of SEQ ID NO: 68 (GGGGSGGGGS or "(G4S)2" linker peptide), SEQ ID NO: 156 (GGGSGGGS or "(G3S)2" linker peptide), SEQ ID NO: 157 (GGGGSGGGGSGGGGS or "(G4S)3" linker peptide) or GGGGSGGGS (SEQ ID NO: 186).

[0190] In some embodiments, the incretin peptide is fused to another incretin peptide via a protease cleavage site, such as a furin protease cleavage site (e.g., a peptide comprising the motif RXK / RR SEQ ID NO: 158, such as RRKR SEQ ID NO: 160 or SEQ ID NO: 153 NVRRKR) and optionally any of the aforementioned linker peptides. In some embodiments, the protease cleavage site (e.g., a furin protease cleavage site) comprises any of SEQ ID NO: 160 RRKR, SEQ ID NO: 153 NVRRKR, SEQ ID NO: 189 RKKR, SEQ ID NO: 190 RMQR, or SEQ ID NO: 191 VFRR. In some embodiments, the furin protease cleavage site is operatively linked to a linker peptide (e.g., a glycine linker peptide, such as a (G4S)2 linker peptide) (e.g., its 3'). In some embodiments, the incretin agent comprises a plurality of (e.g., 2, 3, 4 or more) incretin peptides, each separated from a furin cleavage site and optionally a linker peptide (e.g., a (G4S)2 linker peptide located at the 5' of the furin cleavage site). As described herein, in some embodiments, the sequence and placement of the cleavage site (e.g., the furin cleavage site) within the incretin agent are important to facilitate proper cleavage of the peptide and to produce a scarless N-terminus of the incretin peptide.

[0191] In some embodiments, the polynucleotide encodes an incretin agent comprising a signal peptide and a single incretin peptide (this configuration is referred to herein as "I:1x") (see, for example...). Figure 3 In some embodiments, the polynucleotide encodes an incretin comprising a signal peptide and two incretins separated by a linker peptide and a cleavage site (e.g., a furin cleavage site) (this configuration is referred to herein as "I:2x"). (See example...) Figure 4 This design allows for the cleavage of the first incretin peptide from the second incretin peptide after translation from the polynucleotide. In some embodiments, the polynucleotide encodes an incretin agent comprising a signal peptide and each of its own linker peptides and cleavage sites (e.g., furin cleavage sites) separated by a cleavage site (this configuration is referred to herein as "I:4x"). Figure 5 This design allows for the cleavage of four incretin peptides after translation from a polynucleotide. Those skilled in the art will understand that while this document specifically describes incretin agents containing 1, 2, and 4 incretin peptides, incretin agents with varying numbers of incretin peptides separated by cleavage sites and linker peptides can be used. Figures 3-5In any of the designs shown, the incretin peptide may be any incretin peptide described herein (e.g., GLP1 or GIP incretin peptide, such as any of the incretin peptides shown in Table 2).

[0192] In some embodiments, the polynucleotides described herein encode a GIP incretin peptide (e.g., any of the GIP incretin peptides described in Table 2) located upstream or at the 5' end of a GLP1 incretin peptide (e.g., any of the GLP1 incretin peptides described in Table 2). In some embodiments, the polynucleotides described herein encode a GLP1 incretin peptide (e.g., any of the GLP1 incretin peptides described in Table 2) located upstream or at the 5' end of a GIP incretin peptide (e.g., any of the GIP incretin peptides described in Table 2). In some embodiments, the sequence of the incretin peptide (N-terminal to C-terminal direction) is determined by the intended manner of incretin peptide cleavage, such that the incretin peptide maintains its amino acid sequence and a scarless N-terminus.

[0193] In some embodiments, the polynucleotides described herein encode an incretin having an amino acid sequence having at least 85%, at least 90%, at least 95%, or 100% identity with any of the incretins detailed in Table 3. In some embodiments, the polynucleotides described herein encode an incretin having an amino acid sequence as shown in any of the incretins detailed in Table 3.

[0194] Table 3: Exemplary incretin agents including more than one incretin peptide (mutations are shown in bold, linking peptides are shown underlined, and furin cleavage sites are shown in italics), where examples x2 and x4 include linking peptides and furin cleavage sites between the repeating units. In some embodiments, this disclosure provides one or more polynucleotides encoding an incretin agent comprising a combination of incretin peptides. In such embodiments, one or more polynucleotides may encode an incretin agent. In some embodiments, a first polynucleotide may encode a first incretin peptide of the incretin agent, and a second polynucleotide may encode a second incretin peptide of the incretin agent.

[0195] Extended Half-Life (HLE) Domain In some embodiments, the polynucleotides described herein encode incretin agents comprising one or more incretin peptides fused to a half-life extension (HLE) domain (see, for example...). Figures 7-14 (The incretin agent shown herein). In some embodiments, where the incretin agent comprises more than one incretin peptide, the HLE domain may be included in the incretin agent described herein to increase the half-life of one or each of the incretin peptides.

[0196] Human serum albumin (HSA) In some embodiments, the incretin agent comprises one or more incretin peptides fused to an HLE domain comprising albumin (e.g., human serum albumin (HSA)). In some embodiments, the half-life extension portion comprises albumin, e.g., human serum albumin. In some embodiments, the human serum albumin (HSA) sequence has at least 90%, 95%, or 99% identity with the amino acid sequence shown in SEQ ID NO: 159, or a fragment thereof or a mutant thereof. In some embodiments, the HSA sequence comprises or is composed of the amino acid sequence shown in SEQ ID NO: 159, or a fragment thereof or a mutant thereof. In some embodiments, the HSA sequence comprises or is composed of an amino acid sequence that is a mutant of wild-type HSA (i.e., SEQ ID NO: 159) containing one or more amino acid mutations. In some embodiments, one or more mutations comprise the mutation at position 573 in SEQ ID NO: 159.In some embodiments, the K residue at position 573 of SEQ ID NO: 159 is substituted with any of the following amino acid residues: A, C, D, F, G, H, I, L, M, N, P, Q, R, S, V, W, and Y (SEQ ID NO: 187). In some embodiments, the K residue at position 573 of SEQ ID NO: 159 is substituted with a P residue (SEQ ID NO: 188). In some embodiments, the HSA mutant comprises any HSA mutant disclosed in U.S. Patent No. 8,748,380, which is hereby incorporated herein by reference in its entirety.

[0197] In some embodiments, the polynucleotides described herein encode fusions to the HLE domain (e.g., HSA or HSA mutants as described herein) Figure 7 The incretin agents shown include 1:1x, 1:2x, or 1:4x configurations (i.e., 1, 2, or 4 incretin peptides). In some embodiments, the incretin peptides may be the GLP1 or GIP incretin peptides described herein or mutants thereof. In some embodiments, where the incretin agent comprises more than one incretin peptide, the incretin peptide adjacent to the HLE domain will remain fused to the HLE domain after post-translational processing, and the incretin peptide not adjacent to the HLE domain will cleave from the adjacent incretin peptides and the HLE domain. Such a design may be used when administration of multiple incretin peptides having various half-lives is desirable. Such a design may also be desirable when one of the incretin peptides is intended to cross the blood-brain barrier (i.e., where the HLE domain is undesirable) and one of the incretin peptides is intended to remain in circulation for a longer period of time (i.e., where the HLE remains connected).

[0198] In some embodiments, the incretin agent comprises more than one incretin peptide separated by a linker peptide and a protease cleavage site (e.g., a furin protease cleavage site), wherein one of the incretin peptides is the GLP1 peptide described herein adjacent to an HLE domain (e.g., HSA or an HSA mutant), and another of the incretin peptides is a GIP peptide not adjacent to an HLE domain (e.g., at the N-terminus of the polypeptide chain). In this embodiment, when the incretin agent is expressed from the polynucleotides described herein, the GIP peptide will cleave from the GLP1 peptide linked to the HLE domain, such that the GLP1 incretin peptide will have a longer half-life than the GIP incretin peptide. In some embodiments, the incretin agent comprises more than one incretin peptide separated by a linker peptide and a protease cleavage site (e.g., furin cleavage site), wherein one of the incretin peptides is the GIP incretin peptide described herein adjacent to an HLE domain (e.g., HSA or an HSA mutant), and another of the incretin peptides is a GLP1 peptide not adjacent to an HLE domain (e.g., at the N-terminus of the polypeptide chain). In this embodiment, when the incretin agent is expressed from the polynucleotides described herein, the GLP1 incretin peptide will cleave from the GIP peptide linked to the HLE domain, such that the GIP incretin peptide will have a longer half-life than the GLP1 incretin peptide (see, for example...). Figure 9 ).

[0199] In some embodiments, the polynucleotides described herein encode, for example, Figure 8 The incretin agent shown in A comprises a signal peptide (“SP”), a GLP1 incretin peptide, a linker peptide, a second GLP1 incretin peptide, a second linker peptide (GGGS)3, and a half-life extension (HLE) domain for human serum albumin (HSA) or an HSA mutant. Additionally, the incretin agent may include a protease cleavage site (e.g., a furin protease cleavage site) between the GLP1 incretin peptides. In this embodiment, once the incretin agent is expressed, the first (N-terminal) GLP1 incretin peptide is cleaved from the second incretin GLP1 incretin peptide adjacent to the HLE domain. The resulting GLP1 incretin peptide will have two distinct half-lives (i.e., the GLP1 incretin peptide remaining linked to the HLE domain will have a longer half-life than the cleaved GLP1 incretin peptide).

[0200] In some embodiments, the polynucleotides described herein encode, for example, Figure 9The incretin agent shown comprises a signal peptide (SP), a first GLP1 incretin peptide, a linker peptide (GGGS)2, a first GIP incretin peptide, a second linker peptide (GGGS)2, a second GLP1 incretin peptide, a third linker peptide (GGGS)2, a second GIP incretin peptide, a fourth linker peptide (GGGS)3, and a half-life extension (HLE) domain of human serum albumin (HSA) or an HSA mutant. The furin protease and SP cleavage sites within the incretin agent are indicated by arrows. In this embodiment, when the incretin agent is expressed, the first GLP1 incretin peptide, the first GIP incretin peptide, and the second GLP1 incretin peptide are cleaved, and the second GIP incretin peptide remains fused to the HLE domain. The resulting second GIP-HLE fusion will have a longer half-life than other incretin peptides.

[0201] In some embodiments, the polynucleotides described herein encode incretins, which comprise an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the incretin sequences in Table 4. In some embodiments, the incretins comprise any of the incretins detailed in Table 4 below, or combinations thereof, or mutants thereof.

[0202] Table 4: Exemplary incretin agents including hAlbumin (mutations are shown in bold, linker peptides are shown underlined, and furin cleavage sites are shown in italics), where examples x2 and x4 include linker peptides and furin cleavage sites between the repeating units. Albumin-binding domain In some embodiments, the incretin peptide is fused to an albumin-binding half-life extension portion. Various albumin-binding portions (i.e., albumin-binding protein domains) can be used as the half-life extension portion of the incretin agents described herein (see, for example, Zorzi et al., MedChemComm, 2019, 10.7, 1068-1081, which are incorporated herein by reference in their entirety). In some embodiments, the albumin-binding protein domain comprises an albumin-binding domain (ABD) of protein G derived from Streptococcus strain GI48 and / or protein PAB derived from *Gynostemma pentaphyllum*, such as ABD035 and SA21 (described in Levy et al., PLoS One, 2014, 9(2), e87704, which is incorporated herein by reference in its entirety) and ABD094 (NCT02690142) (described in Frejd and Kim, Exp. Mol. Med., 2017, 49(3), e306, which is incorporated herein by reference in its entirety). In some embodiments, the ABD binds to domain II of human serum albumin without overlapping with or interfering with binding to the FcRn binding site on albumin.

[0203] In some embodiments, ABD comprises ABDCon (a triple-helix bundle albumin-binding domain), as described in Jacobs et al., Protein Eng., Des. Sel., 2015, 28(10), 385-393, which is incorporated herein by reference in its entirety. In some embodiments, ABD is derived from the bacterial protein Sso7d from the hyperthermophilic archaea *Sulphozoa*, such as M11.12 and M18.2.5 (as described in Gao et al., Nat. Struct. Biol., 1998, 5(9), 782-786 and Traxlmayr et al., J. Biol. Chem., 2016, 291(43), 22496–22508, which are incorporated herein by reference in their entirety). In some embodiments, ABD includes DARPin, as described in Pluckthun, Annu. Rev.Pharmacol. Toxicol., 2015, 55, 489-511, which is incorporated herein by reference in its entirety.

[0204] In some embodiments, the ABD comprises an immunoglobulin domain or a fragment thereof. In some embodiments, the ABD comprises a fully human domain antibody (dAb). For example, in some embodiments, the ABD comprises AlbudAb, as described in Holt et al., Protein Eng., Des. Sel., 2008, 21(5), 283-288, which is incorporated herein by reference in its entirety. In some embodiments, the ABD comprises Fab, such as dsFv CA645, as described in Adams et al., mAbs, 2016, 8(7), 1336-1346, which is incorporated herein by reference in its entirety.

[0205] In some embodiments, the ABD comprises a heavy-chain-only (VHH) antibody, i.e., a nanobody, as described in Steeland et al., DrugDiscovery Today, 2016, 21(7), 1076-1113, which is incorporated herein by reference in its entirety. In some embodiments, the ABD comprises a VHH domain containing one or more of the complementarity-determining region (CDR) sequences HCDR1, HCDR2, and / or HCDR3 as shown in SEQ ID NO: 191 (GFTLDYYA), SEQ ID NO: 192 (IASSGGST), and / or SEQ ID NO: 193 (AAAVLECRTVVRGYDY). In some embodiments, the ABD comprises a VHH domain containing the CDR sequences HCDR1, HCDR2, and HCDR3 as shown in SEQ ID NO: 191 (GFTLDYYA), SEQ ID NO: 192 (IASSGGST), and SEQ ID NO: 193 (AAAVLECRTVVRGYDY), respectively. In some embodiments, the ABD comprises a VHH domain having at least 90%, 95%, or 99% identity with the amino acid sequence shown in SEQ ID NO: 154 (EVQLLESGGGLVQPGGSLRLSCAASGFTLDYYAIGWFRQAPGKEREGVSCIASSGGSTNYADSVKGRFTISRDNSKNTVYLQMNSLKPEDTAVYYCAAAVLECRTVVRGYDYWGQGTQVTVSS or “aHSA-VHH”). In some embodiments, the VHH domain comprises the amino acid sequence shown in SEQ ID NO: 154. In some embodiments, ABD includes VNAR, as described in Muller et al., mAbs, 2012, 4(6), 673-685, which is incorporated herein by reference in its entirety.

[0206] In some embodiments, the polynucleotides described herein encode, for example, Figure 8 The incretin agent shown in B is an incretin agent having a signal peptide (SP), a first GLP1 incretin peptide, a linker peptide, a second GLP1 incretin peptide, a second linker peptide (GGGS)3, and an extended half-life (HLE) domain for binding to the VHH domain of HSA. The furin protease and SP cleavage sites within the incretin agent are indicated by arrows. In this embodiment, when the incretin agent is expressed, the first GLP1 incretin peptide is cleaved from the second GLP1 incretin peptide, and the second GLP1 incretin peptide remains fused to the HLE domain (anti-HSA VHH domain). Therefore, the second GLP1 incretin peptide will have a longer half-life than the first GLP1 incretin peptide. In some embodiments, the incretin agent is... Figure 8 The incretins shown in B, wherein one or both of the GLP1 incretin peptides may instead be different incretins (e.g., the GIP incretin peptide described herein). In some embodiments, a single incretin peptide is fused to a VHH domain that binds to HSA.

[0207] Other albumin-binding domains are known in the art; see, for example, Zorzi et al., MedChemComm, 2019, 10.7, 1068-1081, which are incorporated herein by reference in their entirety.

[0208] In some embodiments, the polynucleotides described herein encode incretins, which comprise an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the incretin sequences in Table 5. In some embodiments, the incretins comprise any of the incretins detailed in Table 5 below, or combinations thereof, or mutants thereof.

[0209] Table 5: Exemplary incretin agents including the aHSA-VHH domain (mutations are shown in bold, linking peptides are shown underlined) XTEN In some embodiments, the half-life extension portion is an XTEN sequence as described in U.S. Patent No. 8,673,860 and Podust et al., Journal of Controlled Release, 2016, 240, 52-66, which are incorporated herein by reference in their entirety.

[0210] In some embodiments, the XTEN sequence comprises about 100 to about 3000 amino acid residues, preferably 400 to about 3000 residues, wherein at least about 80%, or at least about 90%, or about 91%, or about 92%, or about 93%, or about 94%, or about 95%, or about 96%, or about 97%, or about 98%, or about 99% to about 100% of the sequence consists of multiple units of two or more non-overlapping sequence motifs selected from the amino acid sequences in Table 6 or Table 7. In some cases, the XTEN comprises non-overlapping sequence motifs wherein about 80%, or at least about 90%, or about 91%, or about 92%, or about 93%, or about 94%, or about 95%, or about 96%, or about 97%, or about 98%, or about 99% to about 100% of the sequence consists of two or more non-overlapping sequences selected from a single motif family in Table 6 or Table 7, thereby producing a “family” sequence in which the overall sequence remains substantially non-repetitive. Therefore, in this type of embodiment, the XTEN sequence may comprise multiple units of non-overlapping sequence motifs from the AD, AE, AF, AG, AM, AQ, BC, or BD families of the sequences in Table 6. In other cases, the XTEN comprises motif sequences from two or more motif families in Table 6. In still other cases, the XTEN comprises motif sequences from one or more motif families in Table 7.

[0211] Table 6: XTEN sequence motifs and motif families of 12 amino acids Table 7: XTEN sequence motifs and motif families of 12 amino acids Fc domain with multiple polypeptide chains and incretin In some embodiments, the half-life extension portion is or includes, for example, an Fc domain of human IgG (e.g., human IgG1, IgG2, IgG3, or IgG4). In some embodiments, the half-life extension portion does not include, for example, an Fc domain of human IgG (e.g., human IgG1, IgG2, IgG3, or IgG4). In some embodiments, the half-life extension portion includes an Fc domain of human IgG4 or a mutant thereof (e.g., as included in dulaglutide). In some embodiments, the Fc domain of the IgG4 sequence has at least 90, 95, 96, 97, 97, or 99% identity with SEQ ID NO: 155 (AESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQEEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSRLTVDKSRWQEGNVFSCSVMHEALHNHYTQKSLSLSLG). In some embodiments, the Fc domain of the IgG4 sequence is or contains the amino acid sequence shown in SEQ ID NO: 155.

[0212] Homodimeric incretin In some embodiments where the incretin agent comprises more than one incretin peptide, the incretin agent comprises two or more incretin peptides on a single polypeptide chain. In some embodiments where the incretin agent comprises more than one incretin peptide, the incretin agent comprises one or more incretin peptides on a single polypeptide chain.

[0213] In some embodiments, the individual polypeptide chains are polymerized (e.g., dimerized). In some embodiments where the incretin agent comprises one or more incretin peptides on individual polypeptide chains, the individual polypeptide chain comprises two polypeptide chains, each containing an immunoglobulin constant domain, and the two polypeptide chains are dimerized by two constant domains, which are combined to form an Fc domain.

[0214] In some embodiments, the Fc domain comprises an IgG4 Fc domain (e.g., as included in dulaglutide). In some embodiments, the Fc domain comprises an IgG1 Fc domain. Exemplary designs of incretin agents comprising multiple polypeptide chains are shown, for example... Figure 10 In this process, the polypeptide chain includes an Fc domain, which causes the polypeptide chain to dimerize. Figure 10The design may include incretin peptides of I:1x, I:2x, or I:4x configuration (or other numbers of incretin peptides may be used), and each may be a GLP1 or GIP incretin peptide (or any of the mutants described herein). Figure 10 In this model, the two Fc domains are identical. The incretins on each polypeptide chain may be the same or different (or, in the case of multiple incretins on a single chain, may contain different combinations of incretin peptides). An exemplary incretin agent comprising two polypeptide chains forming a homodimer through dimerization of the Fc domains on each polypeptide chain is shown in [the diagram / illustration]. Figure 11 In. Figure 11 In this formulation, the incretin peptide comprises two polypeptide chains, each including a signal peptide (SP), a GLP1 incretin peptide, a linker peptide (GGGGS)3, and an Fc domain. In some embodiments, each polypeptide chain may contain two or more incretin peptides (see, for example...). Figure 12 In some embodiments, two or more incretin peptides may be the same incretin peptide. In some embodiments, two or more incretin peptides may be different incretin peptides (see, for example...). Figure 12 In cases where each polypeptide chain comprises two or more incretin peptides, cleavage sites can be introduced to cleave the incretin peptides, thereby leaving one incretin peptide linked to the Fc domain. Those skilled in the art will understand that the incretin peptide remaining linked to the Fc domain will have a longer half-life and different activity than other cleaved incretin peptides.

[0215] The Fc domain included in the incretin agents described in this article not only allows for the dimerization of the two polypeptide chains but also increases the half-life of the incretin peptide. Other mutations can also be introduced into the Fc domain to increase the half-life of the incretin agents.

[0216] Fc mutation in HLE In some embodiments, the Fc domain within the incretin agent contains one or more mutations to increase the half-life of the incretin agent. For example, in some embodiments, the Fc domain may include an LS mutation in the CH3 region (for enhanced FcRn binding) (see Zalevsky et al., Nature biotechnology, 2010, 28.2: 157-159, which is incorporated herein by reference). Such mutations are designated M428L and N434S according to EU designations (i.e., M88L and N94S within the CH3 domain) and are referred to herein as “LS”. Exemplary incretin agents having such mutations are shown in, for example... Figure 11 , Figure 12 and Figure 14 middle.

[0217] In some embodiments, the Fc domain of the IgG4 (LS) sequence is consistent with SEQ ID NO: 299 (AESKYGPPCPPCPAPEAAGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSQEDPEVQFNWYVDGVEVHNAKTKPREEQFNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKGLPSSIEKTISKAKGQPREPQVYTLPPSQEEMTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSRLTVDKSRWQEGNVFSCSVLHEALHSHYTQKSLSLSLG).

[0218] In some embodiments, the Fc domain of the incretin includes one or more mutated amino acid residues that increase the half-life. In some embodiments, the Fc domain includes one of the following mutated amino acid residues according to the EU numbering scheme: M252Y, S254T, and T256E (“YTE”) to increase the half-life. In some embodiments, the Fc domain includes a combination of the following mutated amino acid residues according to the EU numbering scheme: M252Y, S254T, and T256E to increase the half-life.

[0219] In some embodiments, the Fc domain comprises one of the following mutated amino acid residues: T250Q and M428L (“QL”) according to the EU numbering scheme, to increase the half-life. In some embodiments, the Fc domain comprises one of the following mutated amino acid residues: H433K and N434F (“KF”) according to the EU numbering scheme, to increase the half-life. In some embodiments, the second Fc domain comprises one of the following mutated amino acid residues: T307A, E380A, and N434A (“AAA”) according to the EU numbering scheme, to increase the half-life. In some embodiments, the Fc domain comprises one of the following mutated amino acid residues: V308P according to the EU numbering scheme, to increase the half-life. In some embodiments, the Fc domain comprises one of the following mutated amino acid residues: M252Y, V308P, and N434Y (“YPY”) according to the EU numbering scheme, to increase the half-life. In some embodiments, the Fc domain comprises one of the following mutated amino acid residues: H285D, T307Q, and A378V (“DQV”) according to the EU numbering scheme, to increase the half-life. In some embodiments, the Fc domain comprises one of the following mutated amino acid residues: L309D, Q311H, and N434S (“DHS”) according to the EU numbering scheme, to increase the half-life. Exemplary Fc mutations are described, for example, in Liu et al., Antibodies 9.4: 64 (2020), which are hereby incorporated by reference in their entirety.

[0220] In some embodiments, the polynucleotides described herein encode incretins, which comprise an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the incretin sequences in Table 8. In some embodiments, the incretins comprise any of the incretins detailed in Table 8 below, or combinations thereof, or mutants thereof.

[0221] Table 8: Exemplary incretin agents including the IgG4 Fc domain (mutations are shown in bold, linker peptides are shown underlined, and furin cleavage sites are shown in italics), where x4 examples include linker peptides between repeat units and furin cleavage sites. Mutations that eliminate Fc effector function In some embodiments, the Fc domain includes one or more mutations that eliminate the effector function of the Fc domain included in the incretin. By eliminating the Fc effector function of the incretin, the delivery of the incretin is less likely to cause undesirable immune responses due to immune cell-triggered cytotoxicity and other effector activities. In some embodiments, the Fc domain of the molecule includes one or more mutations that silence the effector function.

[0222] In some embodiments, one or more mutations in the Fc domain include “STR” modifications, or, according to the EU numbering scheme, combinations of mutations of L234S, L235T, and G236R. When such mutations are introduced into the Fc domain, the Fc domain exhibits minimal or undetectable binding to the Fcγ receptor or C1q and does not promote inflammatory interleukin responses (see, for example, Wilkinson et al., (2021) PLoS One 16.12: e0260954, which is incorporated herein by reference in its entirety). Exemplary incretin agents containing STR modifications are shown in, for example... Figure 14 middle.

[0223] In some embodiments, modifications that reduce or silence the effector function of the Fc domain included in the incretin agents described herein comprise one or more of the following mutations: L234A, L235A, P329G, P329A, N297A, or N297D according to the EU numbering scheme. In some embodiments, modifications that silence the effector function of the Fc domain comprise the following mutated amino acid residues: L234A and L235A (“LALA”) according to the EU numbering scheme. In some embodiments, mutations for eliminating the effector function of the Fc domain comprise the following: L234A / L235A / P329G (“LALAPG”) according to the EU numbering scheme. In some embodiments, modifications that silence the effector function of the Fc domain comprise the following mutated amino acid residues: L234A, L235A, and P329A (“LALAPA”) according to the EU numbering scheme. In some embodiments, modifications that silence the effector function of the Fc domain further comprise N297A or N297D according to the EU numbering scheme. In some embodiments, the modification of the effector function of the silencing Fc domain includes the following Fc mutations: according to EU numbers, L234A / L235A / P329G and N297A. In some embodiments, the modification of the effector function of the silencing Fc domain includes the following Fc mutations: according to EU numbers, L234A / L235A / P329G and N297D. In some embodiments, the modification of the effector function of the silencing Fc domain includes the following Fc mutations: according to EU numbers, L234A, L235A and N297A. In some embodiments, the modification of the effector function of the silencing Fc domain includes the following Fc mutations: according to EU numbers, L234A, L235A and N297D. In some embodiments, the modification of the effector function of the silencing Fc domain includes the following Fc mutations: according to EU numbers, L234A, L235A, P329A and N297A. In some embodiments, the modification of the effector function of the silencing Fc domain includes the following Fc mutations: according to EU numbers, L234A, L235A, P329A, and N297D.

[0224] In some embodiments, modifications that reduce or silence the effector function of the Fc domain included in the incretin agents described herein comprise the following mutation: L234F / L235E / P331S (“FES”) according to the EU numbering scheme. In some embodiments, modifications that reduce or silence the effector function of the Fc domain included in the incretin agents described herein comprise the following mutation: L234F / L235Q / K322Q (“FQQ”) according to the EU numbering scheme. In some embodiments, modifications that reduce or silence the effector function of the Fc domain included in the incretin agents described herein comprise the following mutation: A330S / P331S according to the EU numbering scheme.

[0225] Those skilled in the art will understand that other modifications known in the art can be used to eliminate the effect function.

[0226] Heterodimeric incretin In some embodiments, the one or more polynucleotides described herein encode an incretin agent comprising more than one incretin peptide on a single polypeptide chain. In some embodiments, the single polypeptide chain is polymerized (e.g., dimerized). In some embodiments where the incretin agent comprises one or more incretin peptides on a single polypeptide chain, the single polypeptide chain comprises two polypeptide chains, each comprising an immunoglobulin constant domain, and the two polypeptide chains are dimerized by two constant domains, which combine to form an Fc domain.

[0227] In some embodiments, the Fc domain within the incretin agent contains one or more mutations that induce dimerization. For example, in some embodiments, the incretin agent comprises a first polypeptide chain containing an incretin peptide fused to an immunoglobulin, wherein the constant domain contains one or more mutations that induce dimerization with a second polypeptide chain containing an incretin peptide fused to an immunoglobulin. In some embodiments, the constant domains of both the first and second polypeptides contain one or more mutations that induce dimerization. In some embodiments, one or more incretin peptides in the first and second polypeptides are different.

[0228] One method for inducing dimerization is called the "Groove-and-Pepper Structure Technology" (KIH), which aims to force two distinct constant domains to pair by introducing mutations into the CH3 domain to modify the contact interface. In one CH3 domain, a bulky amino acid is replaced by an amino acid with a short side chain to produce a "groove," and an amino acid with a large side chain is introduced into the other CH3 domain to produce a "pestle." Co-expression of these two constant domains induces dimerization. In some embodiments, the Fc domain described herein utilizes, for example, the KIH technology described in WO1998 / 050431, which is incorporated herein by reference in its entirety. As described herein, the Fc domain of an incretin agent may contain certain mutations utilizing the KIH technology, including but not limited to CH3 modification. In some embodiments, the Fc domain of an incretin agent comprises a CH3 domain containing one or more of the following mutations: Y349C, T366S, L368A, and Y407V (according to EU numbers). In some embodiments, the Fc domain of the incretin agent comprises a CH3 domain, wherein the CH3 domain comprises one or more of the following mutations: Y349C, T366S, L368A, and Y407V (according to EU designations). This combination of mutations is referred to herein as “FcKIH-b”. In some embodiments, the Fc domain of the incretin agent comprises a CH3 domain, which comprises one or more mutations selected from: S354C and T366W (according to EU designations). In some embodiments, the Fc domain of the incretin agent comprises a CH3 domain, which comprises one or more of the following mutations: S354C and T366W (according to EU designations). This combination of CH3 mutations is referred to herein as “FcKIH-a”. In some embodiments, the incretin agent comprises an Fc domain including both the FcKIH-a and FcKIH-b sequences.

[0229] In some embodiments, the Fc domain within the incretin contains one or more "KiH" mutations and LS mutations.

[0230] Therefore, in some embodiments, the incretin agent encoded by one or more polynucleotides as described herein comprises one or more incretin peptides fused to the Fc domain, wherein the CH3 domain of the Fc domain comprises one or more mutations selected from the following: Y349C, T366S, L368A, and Y407V (according to EU designations). In some embodiments, the incretin agent encoded by one or more polynucleotides as described herein comprises one or more incretin peptides fused to the Fc domain, wherein the Fc domain comprises a CH3 domain comprising one or more mutations selected from the following: S354C and T366W (according to EU designations).

[0231] In some embodiments, the incretin agent comprises a heterodimer, such as... Figure 13 or Figure 14 As shown in the image. Figure 13 An exemplary design of two polypeptide chains comprising an incretin peptide fused to an Fc domain is shown. In each polypeptide chain (incretin-Fc fusion), a signal peptide (SP) and one, two, or four incretin peptides (I:1x, I:2x, or I:4x) fused to the Fc domain are present, and each incretin peptide includes an Fc mutation that induces heterodimerization (e.g., a club-and-mortar structure mutation). The Fc domain may also include modifications that eliminate effector function and / or increase half-life as described herein. When expressing... Figure 13 When one or more polynucleotides are present in the two polypeptide chains (top), the two polypeptide chains bind and form a heterodimer incretin (bottom). According to... Figure 13 An exemplary incretin agent of the design shown is illustrated in Figure 14 In China. Specifically, Figure 14 Each polypeptide chain of the incretin in the formulation has a signal peptide (SP), a GLP1 or GIP incretin peptide, a linker peptide (GGGS)3, and an Fc domain. One or both of the Fc domains contain an “LS” mutation (M428L / N434S) that prolongs the half-life of the incretin, a “STR” mutation that silences the Fc effector function, and a “mortar and pestle” mutation that promotes heterodimerization. In some embodiments, instead of LS and / or STR mutations, or in addition to LS and / or STR mutations, any mutation described herein may be included. When two polypeptide chains are expressed, they bind to form a heterodimeric structure containing two polypeptide chains with different incretin peptides. The SP cleavage site within the incretin is indicated by an arrow.

[0232] In some embodiments, the incretin agent comprises two polypeptide chains that bind to and form a heterodimeric incretin agent, one polypeptide chain comprising the sequence shown in SEQ ID NO: 300 (DKTHTCPPCPAPESTRGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVEVHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPIEKTISKAKGQPREPQVYTLPPCREEMTKNQVSLWCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVLHEALHSHYTQKSLSLSPGK) (“FcKIH-a (LS and STR)”), and the other polypeptide chain comprising the sequence shown in SEQ ID NO: 301(DKTHTCPPCPAPESTRGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVEVHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPI EKTISKAKGQPREPQVCTLPPSREEMTKNQVSLSCAVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLVSKLTVDKSRWQQGNVFSCSVLHEALHSHYTQKSLSLSPGK) ("FcKIH-b (LS and STR)").

[0233] In some embodiments, the polynucleotides described herein encode incretins, which comprise an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the incretin sequences in Table 9. In some embodiments, the incretins comprise any of the incretins detailed in Table 9 below, or combinations thereof, or mutants thereof.

[0234] Table 9: Exemplary incretin agents containing FcKIH-a or FcKIH-b domains that form heterodimers (mutations are shown in bold, linking peptides are shown underlined). In some embodiments, any of the exemplary incretin agents including the FcKIH-a domain may be combined with an exemplary incretin agent including the FcKIH-b domain. For example, in some embodiments, the incretin agent of SEQ ID NO: 84, 85, 86, 87, 173 or 174 may be combined with the incretin agent of SEQ ID NO: 88.

[0235] signal peptide According to some embodiments, the signal peptide is fused directly or via a linker peptide to the encoded incretin peptide described herein. In some embodiments, the open reading frame of the polynucleotide described herein encodes an incretin having, for example, a signal peptide that functions in mammalian cells.

[0236] In some embodiments, the signal peptide is typically characterized by a sequence of about 15 to 30 amino acids in length. In many embodiments, the signal peptide is located at the N-terminus of the incretin, but is not limited thereto. In some embodiments, the signal peptide preferably allows the transport of the incretin encoded by the polynucleotide of this disclosure bound thereto to a defined cellular compartment, preferably on the cell surface, in the endoplasmic reticulum (ER), or in an endosome-lysosome compartment.

[0237] In some embodiments, the ribonucleic acid sequence encoding the signal peptide allows an incretin encoded by a polynucleotide to be secreted post-translation by cells, for example, present in an individual, thus producing a plasma concentration of a biologically active incretin.

[0238] In some embodiments, the ribonucleic acid sequence encoding a signal peptide included in the polynucleotide comprises or contains a nucleotide sequence encoding a human signal peptide. In some embodiments, the ribonucleic acid sequence encoding a secretion signal included in the polynucleotide comprises or contains a nucleotide sequence encoding a non-human secretion signal. In some embodiments, the signal peptide may be or contains a viral signal peptide. In some embodiments, such a signal peptide may be or contains the amino acid sequence of MRVLVLLACLAAASNA (SP1-2; SEQ ID NO: 17). In some embodiments, the signal peptide may be or contains the amino acid sequence of MRVMAPRTLILLLSGALALTETWA (husec signal peptide δ GS; SEQ ID NO: 65).

[0239] In some embodiments, the signal peptide sequence is selected from those sequences included in Table 10 below, or fragments thereof or mutants thereof: Table 10: Exemplary signal peptides This disclosure particularly recognizes that the selection of a signal peptide is important for predicting the cleavage site between the signal peptide and the incretin peptide. To enable the delivery and expression of the incretin peptide via polynucleotides, wherein the expressed incretin peptide maintains appropriate function and biological activity, in some embodiments, the signal peptide is selected and included in the incretin agent to achieve appropriate cleavage of the incretin into its mature form.

[0240] Without being bound by any theory, in the context of the polynucleotides encoding incretins described herein, the cleavage site and type or sequence of the signal peptide are important for ensuring the proper processing of the N-terminus of the incretin peptide. The signal peptide may contain specific sequences or structures that lead to alternative processing or cleavage sites, ultimately altering the final amino acid sequence of the mature incretin peptide. In such relatively small peptides, such as GLP1 or GIP (or their mutants, and other peptides of similar size / properties), variations in amino acid residues can affect the peptide's biological activity. In some embodiments, the signal peptide is selected for inclusion in the incretins described herein to promote proper cleavage of the N-terminus of the incretin peptide, or in other words, to produce a "scarless" N-terminus of the incretin peptide to maintain its biological activity. Figure 20 and Figure 21 A schematic diagram showing the location of the theoretical cleavage sites of certain exemplary signal peptides and incretin agents. Figure 20 The A8G mutation indicates that it promotes proper N-terminal processing of GLP1 incretins with the husec signal peptide. Figure 21 The A2G mutation indicates that it promotes proper N-terminal processing of GIP incretins with the husec signal peptide.

[0241] The concept and utilization of a specific signal peptide promoting cleavage to produce a mature peptide with a scarless N-terminus can also be applied to other intestinal peptides (e.g., glucagon) and / or other peptides of comparable size / properties to GLP1 and GIP as described herein. This is important in the context of delivering incretin agents (or other similar peptides) as one or more polynucleotides encoding incretin agents. Such delivery requires proper translation of the protein within the cell, in addition to post-translational processing (including proper cleavage of the post-translational peptide). Incretin agents comprising one or more incretin peptides fused to another peptide as described herein can be engineered and generated such that the signal peptide cleavage is accurate and does not affect the amino acid sequence of the mature peptide (i.e., producing a “scarless” N-terminus). The scarless N-terminus of the incretin peptide (and other similar peptides, such as other intestinal peptides, such as glucagon) allows the peptide to have appropriate function after processing into a mature peptide.

[0242] In some embodiments, the polynucleotide comprises a signal peptide coding sequence and one or more incretin peptide coding sequences in the 5' to 3' direction. In some embodiments, the signal peptide coding sequence and one or more incretin peptide coding sequences encode any of the sequences shown in Table 11 below, or sequences having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the sequences shown in Table 11 below.

[0243] Table 11: Exemplary incretin agents, wherein examples x2 and x4 include linker peptides between the repeating units and furin cleavage sites. In some embodiments, the polynucleotide includes, in the 5' to 3' direction, a signal peptide coding sequence; an incretin peptide coding sequence; a linker peptide coding sequence; and a half-life extension coding sequence.

[0244] In some embodiments, the polynucleotide comprises a signal peptide coding sequence and one or more incretin peptide coding sequences in the 5' to 3' direction, each being independently separated from the linker peptide coding sequence and the protease cleavage site coding sequence, such as the furin protease cleavage site coding sequence. In some such embodiments, one or more of the incretin peptide coding sequences are preceded or followed by the linker peptide coding sequence and the half-life extension coding sequence.

[0245] Exemplary polynucleotides encoding incretins In some embodiments, the polynucleotide comprises a ribonucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the sequences shown in Table 12 below.

[0246] Table 12: Exemplary polynucleotides encoding incretin agents, wherein examples x2 and x4 include linker peptides between repeating units and furin cleavage sites. Exemplary polynucleotide features The polynucleotides described herein encode incretins as described herein. Additionally, in some embodiments, the polynucleotides described herein include those encoding other components, such as signal peptides. In some embodiments, the polynucleotides described herein may comprise nucleotide sequences encoding a 5' UTR and / or a 3' UTR. In some embodiments, the polynucleotides described herein may comprise nucleotide sequences encoding a polyA tail. In some embodiments, the polynucleotides described herein may include a 5' cap, which may be incorporated during transcription or conjugated to the polynucleotide post-transcriptionally.

[0247] 5' cap A structural feature of mRNA is the 5' cap structure. Native eukaryotic mRNA contains a 7-methylguanosine cap linked to the mRNA via a 5'-to-5'-triphosphate bridge, resulting in the cap 0 structure (m7GpppN). In most eukaryotic mRNAs and some viral mRNAs, further modifications can occur at the 2'-hydroxyl group (2'-OH) of the first and subsequent nucleotides (e.g., the 2'-hydroxyl group can be methylated to form 2'-O-Me), producing the "cap 1" and "cap 2" pentatonic ends, respectively. Diamond et al., (2014) Cytokine & Growth Factor Reviews, 25:543-550, reported that cap 0 mRNA cannot be translated as efficiently as cap 1 mRNA, with the 2'-O-Me playing a decisive role at the penultimate position of the 5' end of the mRNA. The absence of 2'-O-Me has been shown to trigger innate immunity and activate the IFN response. Daffis et al., (2010) Nature, 468:452-456; and Züst et al., (2011) Nature Immunology, 12:137-143.

[0248] RNA capping has been well studied and described, for example, in Decroly et al., (2012) Nature Reviews 10:51-65; and Ramanathan et al., (2016) Nucleic Acids Res; 44(16): 7511-7526, the entire contents of which are hereby incorporated by reference. For example, in some embodiments, the 5' cap structure suitable for the context of this invention may be cap 0 (methylation of the first nucleobase, e.g., m7GpppN), cap 1 (additional methylation of the ribose of the adjacent nucleotide of m7GpppN), cap 2 (additional methylation of the ribose of the second nucleotide downstream of m7GpppN), cap 3 (additional methylation of the ribose of the third nucleotide downstream of m7GpppN), cap 4 (additional methylation of the ribose of the fourth nucleotide downstream of m7GpppN), ARCA (“anti-reverse cap analog”), modified ARCA (e.g., phosphate thioester modified ARCA), inosine, N1-methyl-guanosine, 2'-fluoro-guanosine, 7-deazon-guanosine, 8-sideoxy-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.

[0249] As used herein, the term "5'-cap" refers to a structure found at the 5' end of RNA (e.g., mRNA) and generally comprises a guanosine nucleotide linked to RNA (e.g., mRNA) via a 5'-to-5'-triphosphate bond (also referred to as Gppp or G(5')ppp(5')). In some embodiments, the guanosine nucleotide included in the 5'-cap may be modified, for example, by methylation at one or more positions (e.g., at the 7 position) of the base (guanine) and / or by methylation at one or more positions of the ribose. In some embodiments, the guanosine nucleotide included in the 5'-cap comprises a 3'O methylation at the ribose (3'OmeG). In some embodiments, the guanosine nucleotide included in the 5'-cap comprises a methylation at the 7 position of guanine (m7G). In some embodiments, the guanosine nucleotide included in the 5'-cap comprises a methylation at the 7 position of guanine and a 3'O methylation at the ribose (m7(3'OmeG)). It will be understood that the symbols used in the preceding paragraph, such as "(m2 7,3’-O ")G" or "m7(3'OmeG)" applies to other structures described herein.

[0250] In some embodiments, providing RNA with the 5'-cap disclosed herein can be achieved through in vitro transcription, wherein the 5'-cap is co-transcribed into the RNA strand, or it can be ligated to the RNA post-transcriptionally using a capping enzyme. In some embodiments, co-transcriptional capping using the cap disclosed herein improves the capping efficiency of RNA compared to co-transcriptional capping using a suitable reference. In some embodiments, improved capping efficiency can increase RNA translation efficiency and / or translation rate, and / or increase the expression of encoded polypeptides. In some embodiments, alterations to the polynucleotide produce a non-hydrolyzable cap structure, which can, for example, prevent uncapping and increase the RNA half-life.

[0251] In some embodiments, the 5' cap used is a cap 0, cap 1, or cap 2 structure. See, for example, Ramanathan et al. Figure 1 and Decroly et al. Figure 1 The references mentioned herein are incorporated herein by reference in their entirety. In some embodiments, the RNA described herein includes a cap 1 structure. In some embodiments, the RNA described herein includes a cap 2 structure.

[0252] In some embodiments, the RNA described herein comprises a cap 0 structure. In some embodiments, the cap 0 structure comprises a guanosine nucleotide ((m7)G) methylated at the 7' position of guanine. In some embodiments, this cap 0 structure is linked to the RNA via a 5'-to-5'-triphosphate bond and is also referred to herein as (m7)Gppp. In some embodiments, the cap 0 structure comprises a guanosine nucleotide methylated at the 2' position of the guanosine ribose. In some embodiments, the cap 0 structure comprises a guanosine nucleotide methylated at the 3' position of the guanosine ribose. In some embodiments, the guanosine nucleotide included in the 5' cap comprises methylation ((m7)G) at the 7' position of guanine and at the 2' position of the ribose. 27,2’-O In some embodiments, the guanosine nucleotide included in the 5' cap comprises methylation at the 7' position of guanine and at the 2' position of ribose ((m 27,3’-O )G).

[0253] In some embodiments, the cap 1 structure comprises a guanosine nucleotide methylated at the 7' position ((m7)G) and optionally methylated at the 2' or 3' position of the ribose, and a first nucleotide methylated at the 2'O position in the RNA ((m7)G). 2’-O In some embodiments, the cap 1 structure comprises a guanosine nucleotide methylated at the 7' position ((m7)G) and at the 3' position of the ribose, and a first nucleotide methylated at the 2'O position in the RNA ((m7)G). 2’-O In some embodiments, the cap 1 structure is linked to the RNA via a 5'-to-5'-triphosphate bond, and is also referred to herein as, for example, ((m7)Gppp( 2’-O )N1) or (m 27,3’-O )Gppp( 2’-O N1), wherein N1 is as defined and described herein. In some embodiments, the cap 1 structure comprises a second nucleotide N2, which is at position 2 and selected from A, G, C, or U, for example (m7)Gppp( 2’-O )N1pN2 or (m 27,3’-O )Gppp( 2’-O )N1pN2, where each of N1 and N2 is as defined and described herein.

[0254] In some embodiments, the cap 2 structure comprises a guanosine nucleotide methylated at the guanine 7-position ((m7)G) and optionally methylated at the 2' or 3' position of the ribose, and a first nucleotide and a second nucleotide in RNA methylated at the 2'O position ((m7)G). 2 ’-O )N1p(m 2’-OIn some embodiments, the cap 2 structure comprises a guanosine nucleotide methylated at the 7' position ((m7)G) and methylated at the 3' position of the ribose, and a first and second nucleotide methylated at the 2'O position in the RNA. In some embodiments, the cap 2 structure is linked to the RNA via a 5'-to-5'-triphosphate bond and is also referred to herein as, for example, ((m7)Gppp( 2’-O )N1p( 2’-O )N2) or (m 27,3’-O )Gppp( 2’-O )N1p( 2’-O )N2), where each of N1 and N2 is as defined and described herein.

[0255] In some embodiments, the 5' cap is a dinucleotide cap structure. In some embodiments, the 5' cap is a dinucleotide cap structure containing N1, wherein N1 is as defined and described herein. In some embodiments, the 5' cap is a dinucleotide cap G. * N1, where N1 is as defined above and here, and G * The structure of inclusion formula (I): (I) or its salt, in R2 and R3 are each -OH or -OCH3; and X is either O or S.

[0256] In some embodiments, R2 is -OH. In some embodiments, R2 is -OCH3. In some embodiments, R3 is -OH. In some embodiments, R3 is -OCH3. In some embodiments, R2 is -OH and R3 is -OH. In some embodiments, R2 is -OH and R3 is -CH3. In some embodiments, R2 is -CH3 and R3 is -OH. In some embodiments, R2 is -CH3 and R3 is -CH3.

[0257] In some embodiments, X is O. In some embodiments, X is S.

[0258] In some embodiments, the 5' cap is a dinucleotide cap O structure (e.g., (m 7 )GpppN1、(m2 7,2’-O )GpppN1、(m2 7,3’-O )GpppN1、(m 7 )GppSpN1、(m2 7,2’-O )GppSpN1 or (m2 7,3’-O )GppSpN1), where N1 is as defined and described herein. In some embodiments, the 5' cap is a dinucleotide cap O structure (e.g., (m7 )GpppN1、(m2 7,2’-O )GpppN1、(m2 7,3’-O )GpppN1、(m 7 )GppSpN1、(m2 7,2’-O )GppSpN1 or (m2 7,3’-O )GppSpN1), where N1 is G. In some embodiments, the 5' cap is a dinucleotide cap O structure (e.g., (m 7 )GpppN1、(m2 7,2’-O )GpppN1、(m2 7,3’-O )GpppN1、(m 7 )GppSpN1、(m2 7,2’-O )GppSpN1 or (m2 7,3’-O )GppSpN1), where N1 is A, U, or C. In some embodiments, the 5' cap is a dinucleotide cap 1 structure (e.g., (m 7 )Gppp(m 2’-O N1, (m2) 7,2’-O )Gppp(m 2’-O N1, (m2) 7,3’-O )Gppp(m 2’-O )N1、(m 7 )GppSp(m 2’-O N1, (m2) 7,2’-O )GppSp(m 2’-O )N1 or (m2) 7,3’-O )GppSp(m 2’-O )N1), where N1 is as defined and described herein. In some embodiments, the 5' cap is selected from the group consisting of: (m 7 )GpppG (“Ecap0”), (m 7 )Gppp(m 2'-O )G(“Ecap1”), (m2 7,3’-O )GpppG (“ARCA”) and (m2 7,2'-O )GppSpG (“β-S-ARCA”). In some embodiments, the 5' cap has the following structure (m 7 )GpppG (“Ecap0”): Or its salt.

[0259] In some embodiments, the 5' cap is a (m) structure having the following characteristics. 7 )Gppp(m 2'-O G (“Ecap1”): Or its salt.

[0260] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,2'-O )GpppG (“ARCA”): Or its salt.

[0261] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,2'-O GppSpG (“β-S-ARCA”): Or its salt.

[0262] In some embodiments, the 5' cap is a trinucleotide cap structure. In some embodiments, the 5' cap is a trinucleotide cap structure comprising N1pN2, wherein N1 and N2 are as defined and described herein. In some embodiments, the 5' cap is a dinucleotide cap G. * N1pN2, where N1 and N2 are as defined above and here, and G * The structure of inclusion formula (I): (I) Or its salts, wherein R2, R3 and X are as defined and described herein.

[0263] In some embodiments, the 5' cap is a trinucleotide cap O structure (e.g., (m 7 )GpppN1pN2、(m2 7,2’-O )GpppN1pN2 or (m2 7,3’-O ()GpppN1pN2), where N1 and N2 are as defined and described herein). In some embodiments, the 5' cap is a trinucleotide cap 1 structure (e.g., (m7)Gppp(m 2’-O )N1pN2、(m2 7,2’-O )Gppp(m 2’-O )N1pN2、(m2 7,3’-O )Gppp(m 2’-O (N1pN2), where N1 and N2 are as defined and described herein. In some embodiments, the 5' cap is a trinucleotide cap 2 structure (e.g., (m7)Gppp(m 2’-O )N1p(m 2’-O N2, (m2) 7,2’-O )Gppp(m 2’-O )N1p(m 2’-O N2, (m2) 7,3’-O )Gppp(m 2’-O )N1p(m 2’-O)N2), where N1 and N2 are as defined and described herein. In some embodiments, the 5' cap is selected from the group consisting of: (m2 7,3’-O )Gppp(m 2’-O )ApG ("CleanCap AG", "CC413"), (m2 7,3’-O )Gppp(m 2’-O )GpG ("CleanCap GG"), (m7)Gppp(m 2’-O ApG, (m7)Gppp(m) 2’-O )GpG、(m2 7,3’-O )Gppp(m2 6,2’-O ApG and (m7)Gppp(m 2’-O )ApU.

[0264] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,3’-O )Gppp(m 2’-O )ApG ("CleanCapAG", "CC413"): Or its salt.

[0265] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,3’-O )Gppp(m 2’-O GpG (“CleanCapGG”): Or its salt.

[0266] In some embodiments, the 5' cap is a (m7)Gppp(m) structure having the following structure. 2’-O ApG: Or its salt.

[0267] In some embodiments, the 5' cap is a (m7)Gppp(m) structure having the following structure. 2’-O GpG: Or its salt.

[0268] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,3’-O )Gppp(m2 6,2’-O ApG: Or its salt.

[0269] In some embodiments, the 5' cap is a (m7)Gppp(m) structure having the following structure.2’-O ApU: Or its salt.

[0270] In some embodiments, the 5' cap is a tetranucleotide cap structure. In some embodiments, the 5' cap is a tetranucleotide cap structure comprising N1pN2pN3, wherein N1, N2, and N3 are as defined and described herein. In some embodiments, the 5' cap is a tetranucleotide cap G. * N1pN2pN3, where N1, N2, and N3 are as defined above and here, and G * The structure of inclusion formula (I): (I) Or its salts, wherein R2, R3 and X are as defined and described herein.

[0271] In some embodiments, the 5' cap is a tetranucleotide cap O structure (e.g., (m 7 )GpppN1pN2pN3、(m2 7,2’-O )GpppN1pN2pN3 or (m2 7,3’-O (GpppN1N2pN3), where N1, N2, and N3 are as defined and described herein). In some embodiments, the 5' cap is a tetranucleotide cap 1 structure (e.g., (m 7 )Gppp(m 2’-O )N1pN2pN3、(m2 7,2’-O )Gppp(m 2’-O )N1pN2pN3、(m2 7,3’-O )Gppp(m 2’-O )N1pN2N3), wherein N1, N2, and N3 are as defined and described herein. In some embodiments, the 5' cap is a tetranucleotide cap 2 structure (e.g., (m 7 )Gppp(m 2’-O )N1p(m 2’-O )N2pN3、(m2 7,2’-O )Gppp(m 2’-O )N1p(m 2’-O )N2pN3、(m2 7,3’-O )Gppp(m 2’-O )N1p(m 2’-O )N2pN3), where N1, N2, and N3 are as defined and described herein. In some embodiments, the 5' cap is selected from the group consisting of: (m2 7,3’-O )Gppp(m 2’-O Ap(m) 2’-O )GpG、(m2 7,3’-O )Gppp(m 2’-O)Gp(m 2’-O GpC, (m 7 )Gppp(m 2’-O Ap(m) 2’-O UpA and (m 7 )Gppp(m 2’-O Ap(m) 2’-O )GpG.

[0272] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,3’-O )Gppp(m 2’-O Ap(m) 2’-O GpG: Or its salt.

[0273] In some embodiments, the 5' cap is a (m2) structure having the following characteristics. 7,3’-O )Gppp(m 2’-O )Gp(m 2’-O GpC: Or its salt.

[0274] In some embodiments, the 5' cap is a (m) structure having the following characteristics. 7 )Gppp(m 2’-O Ap(m) 2’-O UpA: Or its salt.

[0275] In some embodiments, the 5' cap is a (m) structure having the following characteristics. 7 )Gppp(m 2’-O Ap(m) 2’-O GpG: Or its salt.

[0276] Proximal sequence of the cap In some embodiments, the 5' UTR as used herein includes a cap proximal sequence, for example, as disclosed herein. In some embodiments, the cap proximal sequence includes a sequence adjacent to the 5' cap. In some embodiments, the cap proximal sequence includes nucleotides at RNA polynucleotide positions +1, +2, +3, +4, and / or +5.

[0277] In some embodiments, the cap structure comprises one or more polynucleotides of a proximal cap sequence. In some embodiments, the cap structure comprises an m7 guanosine cap and nucleotide +1 (N1) of an RNA polynucleotide. In some embodiments, the cap structure comprises an m7 guanosine cap and nucleotide +2 (N2) of an RNA polynucleotide. In some embodiments, the cap structure comprises an m7 guanosine cap and nucleotides +1 and +2 (N1 and N2) of an RNA polynucleotide. In some embodiments, the cap structure comprises an m7 guanosine cap and nucleotides +1, +2, and +3 (N1, N2, and N3) of an RNA polynucleotide.

[0278] Those skilled in the art will appreciate upon reading this disclosure that, in some embodiments, one or more residues of the cap proximal sequence (e.g., one or more of residues +1, +2, +3, +4, and / or +5) may be included in the RNA because they are already included in the cap body (e.g., cap 1 structure or cap 2 structure, etc.); or, in some embodiments, at least some residues in the cap proximal sequence may be added enzymatically (e.g., by a polymerase such as T7 polymerase). For example, when utilizing m2 7,3’-O Gppp(m1 2’-O In some exemplary embodiments of the ApG cap, +1 (i.e., N1) and +2 (i.e., N2) are the cap's (m1) 2’-O A and G residues, and +3, +4 and +5 residues are added by polymerase (e.g., T7 polymerase).

[0279] In some embodiments, the 5' cap is a dinucleotide cap structure, wherein the proximal sequence of the cap contains N1 of the 5' cap, where N1 is any nucleotide, such as A, C, G, or U. In some embodiments, the 5' cap is a trinucleotide cap structure (e.g., the trinucleotide cap structures described above and herein), wherein the proximal sequence of the cap contains N1 and N2 of the 5' cap, where N1 and N2 are independently any nucleotide, such as A, C, G, or U. In some embodiments, the 5' cap is a tetranucleotide cap structure (e.g., the trinucleotide cap structures described above and herein), wherein the proximal sequence of the cap contains N1, N2, and N3 of the 5' cap, where N1, N2, and N3 are any nucleotide, such as A, C, G, or U.

[0280] In some embodiments, for example, when the 5' cap is a dinucleotide cap structure, the proximal cap sequence includes N1 of the 5' cap, and N2, N3, N4, and N5, wherein N1 to N5 correspond to positions +1, +2, +3, +4, and / or +5 of the RNA polynucleotide. In some embodiments, for example, when the 5' cap is a trinucleotide cap structure, the proximal cap sequence includes N1 and N2 of the 5' cap, and N3, N4, and N5, wherein N1 to N5 correspond to positions +1, +2, +3, +4, and / or +5 of the RNA polynucleotide. In some embodiments, for example, when the 5' cap is a tetranucleotide cap structure, the proximal cap sequence includes N1, N2, and N3 of the 5' cap, and N4 and N5, wherein N1 to N5 correspond to positions +1, +2, +3, +4, and / or +5 of the RNA polynucleotide.

[0281] In some embodiments, N1 is A. In some embodiments, N1 is C. In some embodiments, N1 is G. In some embodiments, N1 is U. In some embodiments, N2 is A. In some embodiments, N2 is C. In some embodiments, N2 is G. In some embodiments, N2 is U. In some embodiments, N3 is A. In some embodiments, N3 is C. In some embodiments, N3 is G. In some embodiments, N3 is U. In some embodiments, N4 is A. In some embodiments, N4 is C. In some embodiments, N4 is G. In some embodiments, N4 is U. In some embodiments, N5 is A. In some embodiments, N5 is C. In some embodiments, N5 is G. In some embodiments, N5 is U. It will be understood that the various embodiments described above and herein (e.g., for N1 to N5) may be used alone or in combination and / or may be combined with other embodiments of the variables described above and herein (e.g., 5' cap).

[0282] 5' UTR In some embodiments, nucleic acids (e.g., DNA, RNA) as used herein comprise a 5' UTR. In some embodiments, the 5' UTR may comprise a plurality of different sequence components; in some embodiments, such a plurality may be or comprise a plurality of copies of one or more assay sequence components (e.g., such as those derived from a particular source or otherwise referred to as functional or characteristic sequence components). In some embodiments, the 5' UTR comprises a plurality of different sequence components.

[0283] The term “untranslated region” or “UTR” is generally used in the art to refer to a region in a DNA molecule that is transcribed but not translated into an amino acid sequence, or to the corresponding region in an RNA polynucleotide (such as mRNA). Untranslated regions (UTRs) can be present at the 5' (upstream) (5' UTR) and / or the 3' (downstream) (3' UTR) of an open reading frame. As used herein, the term “5' UTR” refers to a polynucleotide sequence between the 5' end of a polynucleotide (e.g., a transcription start site) and the start codon of the polynucleotide coding region. In some embodiments, a “5' UTR” refers to a polynucleotide sequence that begins at the 5' end of a polynucleotide (e.g., a transcription start site) and ends one nucleotide (nt) before the start codon (typically AUG) of the polynucleotide coding region, for example, in its natural background. In some embodiments, a 5' UTR contains a Kozak sequence. A 5' UTR is downstream of a 5' cap (if present), for example, directly adjacent to the 5' cap. In some embodiments, the 5' UTR disclosed herein includes, for example, a cap proximal sequence as defined and described herein. In some embodiments, the cap proximal sequence includes a sequence adjacent to the 5' cap.

[0284] Exemplary 5' UTRs include human alpha globulin (hAg) 5' UTR or a fragment thereof, TEV 5' UTR or a fragment thereof, HSP70 5' UTR or a fragment thereof, or c-Jun 5' UTR or a fragment thereof.

[0285] In some embodiments, the RNA disclosed herein comprises hAg 5' UTR or a fragment thereof.

[0286] In some embodiments, the RNA disclosed herein comprises a 5' UTR having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a 5' UTR having the sequence shown as AACUAGUAUUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACC (SEQ ID NO: 49). In some embodiments, the RNA disclosed herein comprises the 5' UTR provided in SEQ ID NO: 49.

[0287] PolyA tail In some embodiments, the polynucleotides (e.g., DNA, RNA) disclosed herein comprise, for example, polyadenylated or poly(A) sequences as described herein. In some embodiments, the poly(A) sequence is located downstream of the 3' UTR, for example, adjacent to the 3' UTR.

[0288] As used herein, the term "poly(A) sequence" or "polyA tail" refers to a continuous or discontinuous sequence of adenosine residues typically located at the 3' end of an RNA polynucleotide. Poly(A) sequences are known to those skilled in the art and may follow the 3' UTR in the RNA described herein. A continuous poly(A) sequence is characterized by a continuous sequence of adenosine residues. Continuous poly(A) sequences are typical in nature. In some embodiments, the polynucleotides disclosed herein comprise continuous poly(A) sequences. In some embodiments, the polynucleotides disclosed herein comprise discontinuous poly(A) sequences. In some embodiments, the RNA disclosed herein may have a poly(A) sequence that is ligated to a free 3' end of the RNA post-transcriptionally by a template-independent RNA polymerase, or a poly(A) sequence encoded by DNA and transcribed by a template-dependent RNA polymerase.

[0289] It has been demonstrated that a poly(A) sequence of approximately 120 A nucleotides has a beneficial effect on RNA levels in transfected eukaryotic cells, as well as on protein levels translated from open reading frames located upstream (5') of the poly(A) sequence (Holtkamp et al., 2006, Blood, Vol. 108, pp. 4009-4017, which are incorporated herein by reference).

[0290] In some embodiments, the poly(A) sequence as described in this disclosure is not limited to a specific length; in some embodiments, the poly(A) sequence is of any length. In some embodiments, the poly(A) sequence comprises at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 A nucleotides, and particularly about 120 A nucleotides, consisting substantially of or composed of the stated number of A nucleotides. In this context, “consistently composed of” means that the majority of the nucleotides in the poly(A) sequence, typically at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% by weight, are A nucleotides, but the remaining nucleotides are permitted to be nucleotides other than A nucleotides, such as U nucleotides (uridine monophosphate), G nucleotides (guanosine monophosphate), or C nucleotides (cytidine monophosphate). In this context, "composed of" means all nucleotides in the poly(A) sequence, i.e., 100% of the nucleotides in the poly(A) sequence are A nucleotides. The term "A nucleotide" or "A" refers to adenosine monophosphate.

[0291] In some embodiments, based on a DNA template containing repeating dT nucleotides (deoxythymidines) in a strand complementary to the coding strand, a poly(A) sequence is ligated during RNA transcription, such as during the preparation of RNA transcribed in vitro. The DNA sequence (coding strand) encoding the poly(A) sequence is called a poly(A) box.

[0292] In some embodiments, the poly(A) box present in the DNA coding strand is essentially composed of dA nucleotides, but interrupted by random sequences of four nucleotides (dA, dC, dG, and dT). The length of such random sequences can be 5 to 50, 10 to 30, or 10 to 20 nucleotides. Such a box is disclosed in WO2016 / 005324, which is hereby incorporated by reference. Any poly(A) box disclosed in WO2016 / 005324 may be used according to this disclosure. Covering poly(A) boxes that are essentially composed of dA nucleotides, but interrupted by random sequences of equal distribution of the four nucleotides (dA, dC, dG, and dT) and of a length of, for example, 5 to 50 nucleotides, exhibit constant amplification of plasmid DNA in *E. coli* at the DNA level, and remain relevant to beneficial properties supporting RNA stability and translation efficiency at the RNA level. In some embodiments, the poly(A) sequence contained in the RNA polynucleotide described herein consists essentially of A nucleotides, but is interrupted by random sequences of four nucleotides (A, C, G, U). The length of such random sequences can be 5 to 50, 10 to 30, or 10 to 20 nucleotides.

[0293] In some embodiments, no nucleotides other than A are located on the 3' flanking side of the poly(A) sequence, meaning that the poly(A) sequence is not masked by nucleotides other than A or followed by nucleotides other than A at its 3' end.

[0294] In some embodiments, the poly(A) sequence may comprise at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 nucleotides. In some embodiments, the poly(A) sequence may consist substantially of at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 nucleotides. In some embodiments, the poly(A) sequence may consist of at least 20, at least 30, at least 40, at least 80, or at least 100 and at most 500, at most 400, at most 300, at most 200, or at most 150 nucleotides. In some embodiments, the poly(A) sequence comprises at least 100 nucleotides. In some embodiments, the poly(A) sequence comprises about 150 nucleotides. In some embodiments, the poly(A) sequence comprises about 120 nucleotides.

[0295] In some embodiments, the polyA tail comprises a specific number of adenosine residues, such as about 50 or more, about 60 or more, about 70 or more, about 80 or more, about 90 or more, about 100 or more, about 120, about 150, or about 200. In some embodiments, the polyA tail of the string construct may comprise 200 or fewer A residues. In some embodiments, the polyA tail of the string construct may comprise about 200 A residues. In some embodiments, the polyA tail of the string construct may comprise 180 or fewer A residues. In some embodiments, the polyA tail of the string construct may comprise about 180 A residues. In some embodiments, the polyA tail may comprise 150 or fewer residues.

[0296] In some embodiments, the RNA comprises a poly(A) sequence containing a nucleotide sequence of AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAGCAUAUGACUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 50), or a nucleotide sequence having at least 99%, 98%, 97%, 96%, 95%, 90%, 85%, or 80% identity with the nucleotide sequence of SEQ ID NO: 50. In some embodiments, the poly(A) tail comprises a nucleotide sequence as in SEQ ID NO: 50.

[0297] In some embodiments, the polyA tail comprises a plurality of A residues interrupted by the linker peptide. In some embodiments, the linker peptide comprises the nucleotide sequence GCAUAUGACU (SEQ ID NO: 40).

[0298] 3' UTR In some embodiments, the RNA, as used herein, comprises a 3' UTR. As used herein, the terms "3' untranslated region," "3' untranslated region," or "3' UTR" refer to the mRNA molecule sequence that begins after the stop codon of the open reading frame (OPG) coding region. In some embodiments, the 3' UTR begins immediately after the stop codon of the OPG coding region, for example, in its natural context. In other embodiments, the 3' UTR does not begin immediately after the stop codon of the OPG coding region, for example, in its natural context. The term "3' UTR" preferably does not include a poly(A) sequence. Thus, the 3' UTR is upstream of the poly(A) sequence (if present), for example, directly adjacent to the poly(A) sequence.

[0299] In some embodiments, the RNA disclosed herein comprises a 3' UTR containing an F component and / or an I component. In some embodiments, the 3' UTR or its proximal sequence contains a restriction site. In some embodiments, the restriction site is a BamHI site. In some embodiments, the restriction site is an XhoI site.

[0300] In some embodiments, the RNA construct includes the F component. In some embodiments, the F component sequence is the 3' UTR of an N-terminal cleavage enhancer (AES).

[0301] In some embodiments, the RNA disclosed herein comprises a 3' UTR having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with a 3' UTR having the sequence shown in SEQ ID NO: 51. In some embodiments, the RNA disclosed herein comprises the 3' UTR provided in SEQ ID NO: 51.

[0302] In some embodiments, 3'UTR is a FI component as described in WO2017 / 060314, which is incorporated herein by reference in its entirety.

[0303] RNA form At least three distinct forms have been developed for use in RNA compositions (e.g., pharmaceutical compositions): unmodified uridine-containing mRNA (uRNA), nucleoside-modified mRNA (modRNA), and self-amplified mRNA (saRNA). Each of these platforms exhibits unique characteristics. Generally, in all three forms, the RNA is capped, contains an open reading frame (ORF) flanked by an untranslated region (UTR), and has a polyA tail at the 3' end. The ORFs of uRNA and modRNA vectors encode incretin agents. saRNA has multiple ORFs.

[0304] In some embodiments, the RNA described herein may have modified nucleosides. In some embodiments, the RNA contains at least one (e.g., each) modified nucleoside that replaces uridine.

[0305] As used herein, the term "uracil" describes something that can be present in every nucleobase of RNA nucleic acid. The structure of uracil is: .

[0306] As used herein, the term "uridine" describes every nucleoside that can be found in RNA. The structure of uridine is: .

[0307] UTP (uridine 5'-triphosphate) has the following structure: .

[0308] Pseudo-UTP (pseudouridine 5'-triphosphate) has the following structure: .

[0309] "Pseudouridine" is an example of a modified nucleoside that is an isomer of uridine, in which uracil is attached to the pentose ring via a carbon-carbon bond rather than a nitrogen-carbon glycosidic bond.

[0310] Another exemplary modified nucleoside is N1-methyl-pseuuridine (m1Ψ), which has the following structure: .

[0311] N1-Methyl-pseudo-UTP has the following structure: .

[0312] Another exemplary modified nucleoside is 5-methyluridine (m5U), which has the following structure: .

[0313] In some embodiments, one or more uridines in the RNA described herein are replaced by modified nucleosides. In some embodiments, the modified nucleosides are modified uridines.

[0314] In some embodiments, the RNA comprises at least one modified nucleoside replacing uridine. In some embodiments, the RNA comprises a modified nucleoside replacing each uridine.

[0315] In some embodiments, the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U). In some embodiments, the modified nucleosides comprise pseudouridine (ψ). In some embodiments, the modified nucleosides comprise N1-methyl-pseudouridine (m1ψ). In some embodiments, the modified nucleosides comprise 5-methyl-uridine (m5U). In some embodiments, RNA may comprise more than one type of modified nucleoside, and the modified nucleosides are independently selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U). In some embodiments, the modified nucleosides comprise pseudouridine (ψ) and N1-methyl-pseudouridine (m1ψ). In some embodiments, the modified nucleosides comprise pseudouridine (ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleosides comprise N1-methyl-pseudouridine (m1ψ) and 5-methyl-uridine (m5U). In some embodiments, the modified nucleosides include pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U).

[0316] In some embodiments, the modified nucleosides replacing one or more (e.g., all) uridines in the RNA may be any one or more of the following: 3-methyluridine (m3U), 5-methoxyuridine (mo5U), 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s2U), 4-thio-uridine (s4U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine (ho5U), 5-aminoallyl-uridine, 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine), uridine 5-oxyacetic acid (cmo5U), uridine 5-oxyacetic acid methyl Ester (mcmo5U), 5-carboxymethyl-uridine (cm5U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyl-uridine (chm5U), 5-carboxyhydroxymethyl-uridine methyl ester (mchm5U), 5-methoxycarbonylmethyl-uridine (mcm5U), 5-methoxycarbonylmethyl-2-thio-uridine (mcm5s2U), 5-aminomethyl-2-thio-uridine (nm5s2U), 5-methylaminomethyl-uridine (mnm5U), 1-ethyl-pseudouridine, 5-methylaminomethyl-2-thio-uridine (mnm5s2U), 5-methylaminomethyl-2-seleno-uridine (mnm5s) e2U), 5-carboxymethylaminomethyluridine (ncm5U), 5-carboxymethylaminomethyluridine (cmnm5U), 5-carboxymethylaminomethyl-2-thiouridine (cmnm5s2U), 5-propynyluridine, 1-propynyl-pseudouridine, 5-tauronic acid methyluridine (τm5U), 1-tauronic acid methyl-pseudouridine, 5-tauronic acid methyl-2-thiouridine (τm5s2U), 1-tauronic acid methyl-4-thio-pseudouridine), 5-methyl-2-thio-uridine (m5s2U), 1-methyl-4-thio-pseudouridine (m1s4ψ), 4-thio-1-methyl-pseudouridine, 3-methyl- Pseudouridine (m3ψ), 2-thio-1-methyl-pseudouridine, 1-methyl-1-deazon-pseudouridine, 2-thio-1-methyl-1-deazon-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-dihydrouridine, 5-methyl-dihydrouridine (m5D), 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxy-uridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, N1-methyl-pseudouridine, 3-(3-amino-3-carboxypropyl)uridine (acp3U), 1-methyl-3-(3-amino-3-carboxypropyl)pseudouridine (acp3) ψ), 5-(isopentenylaminomethyl)uridine (inm5U), 5-(isopentenylaminomethyl)-2-thio-uridine (inm5s2U), α-thio-uridine, 2'-O-methyl-uridine (Um), 5,2'-O-dimethyluridine (m5Um), 2'-O-methyl-pseudouridine (ψm), 2-thio-2'-O-methyluridine (s2Um), 5-methoxycarbonylmethyl-2'-O-methyluridine (mcm5Um), 5-carbamoylmethyl-2'-O-methyluridine (ncm5Um), 5-carboxymethylaminomethyl-2'-O-methyluridine (cmnm5Um), 3,2 '-O-dimethyluridine (m3Um), 5-(isopentenylaminomethyl)-2'-O-methyluridine (inm5Um), 1-thiouridine, deoxythymidine, 2'-F-arasu-uridine, 2'-F-uridine, 2'-OH-arasu-uridine, 5-(2-carbonmethoxyvinyl)uridine, 5-[3-(1-E-propenylamino)uridine, or any other modified uridine known in the art.

[0317] In some embodiments, the RNA comprises other modified nucleosides or further modified nucleosides, such as modified cytidine. For example, in some embodiments, 5-methylcytidine partially or completely, preferably completely, replaces cytidine in the RNA. In some embodiments, the RNA comprises 5-methylcytidine and one or more selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U). In some embodiments, the RNA comprises 5-methylcytidine and N1-methyl-pseudouridine (m1ψ). In some embodiments, the RNA comprises 5-methylcytidine replacing each cytidine and N1-methyl-pseudouridine (m1ψ) replacing each uridine.

[0318] In some embodiments of this disclosure, the RNA is referred to as "replicon RNA" or simply "replicon," particularly "self-replicating RNA" or "self-amplifying RNA." In a particularly preferred embodiment, the replicon or self-replicating RNA is derived from a single-stranded (ss) RNA virus (particularly a positive-stranded ss RNA virus, such as alphavirus) or contains components derived from that virus. Alphavirus is a typical example of a positive-stranded RNA virus. Alphavirus replicates in the cytoplasm of infected cells (for a review of the alphavirus life cycle, see José et al., Future Microbiol., 2009, Vol. 4, pp. 837–856, which are incorporated herein by reference in their entirety). The total genome length of many alphaviruses is typically in the range of 11,000 to 12,000 nucleotides, and the genome RNA typically has a 5' cap and a 3' poly(A) tail. The alphavirus genome encodes non-structural proteins (involved in transcription, modification, and replication of viral RNA, as well as protein modification) and structural proteins (forming viral particles). Two open reading frames (ORFs) are typically present in the genomic body. The four non-structural proteins (nsP1-nsP4) are usually encoded together by a first ORF located near the 5' end of the genomic body, while the alphavirus structural proteins are encoded together by a second ORF located downstream of the first ORF and extending near the 3' end of the genomic body. Typically, the first ORF is larger than the second ORF, with a ratio of approximately 2:1. In alphavirus-infected cells, only the nucleic acid sequences encoding the non-structural proteins are translated from the genomic body RNA, while the genetic information encoding the structural proteins is translated from the subgenomic transcript, which is an RNA molecule similar to eukaryotic messenger RNA (mRNA; Gould et al., 2010, Antiviral Res., Vol. 87, pp. 111-124). Post-infection, i.e., in the early stages of the viral life cycle, the (+) strand of the genomic body RNA is used directly, like messenger RNA, to translate the open reading frame encoding the non-structural multi-protein (nsP1234).

[0319] Alphavirus-derived vectors have been proposed for delivering foreign genetic information to target cells or organisms. In a simplified approach, the first ORF encodes an alphavirus-derived RNA-dependent RNA polymerase (replicaase), which mediates the self-amplification of RNA post-translational. The second ORF, encoding an alphavirus structural protein, is replaced by an open reading frame encoding the target protein (e.g., an incretin). Alphavirus-based trans-replication systems rely on alphavirus nucleotide sequence components on two separate nucleic acid molecules: one molecule encodes a viral replicase, and the other is capable of trans-replication by that replicase (hence the name trans-replication system). Trans-replication requires the presence of two such nucleic acid molecules in a given host cell. The nucleic acid molecule capable of trans-replication by the replicase must contain certain alphavirus sequence components that allow the alphavirus replicase to recognize and synthesize RNA.

[0320] Unmodified uridine platforms may be characterized by, for example, one or more inherent adjuvant effects, as well as good tolerability and safety. Modified uridine (e.g., pseudouridine) platforms may be characterized by reduced adjuvant effects, deactivated innate immune sensor activation capacity, and thus good tolerability and safety. Self-amplification platforms may be characterized by, for example, long duration of protein expression, good tolerability and safety, and the potential for higher efficacy at very low vaccine doses.

[0321] This disclosure provides optimized specific RNA constructs, for example, to improve manufacturability, encapsulation, expression levels (and / or timing). Certain components are discussed below, and some preferred embodiments are illustrated herein.

[0322] Codon optimization and GC enrichment As used herein, the term "codon-optimized" refers to altering codons in the coding region of a nucleic acid molecule (e.g., a polynucleotide) to reflect typical codon usage in a host organism (e.g., an individual receiving a polynucleotide), preferably without altering the amino acid sequence encoded by the nucleic acid molecule. In the context of this disclosure, in some embodiments, the coding region is codon-optimized to achieve optimal expression in a subject treated with the RNA molecule described herein. In some embodiments, codon optimization may be performed by replacing codons with common tRNA supplies with "rare codons." In some embodiments, codon optimization may include increasing the G / C content of the RNA coding region described herein compared to the guanosine / cytosine (G / C) content of the corresponding coding sequence of wild-type RNA, wherein the amino acid sequence encoded by the RNA is preferably unmodified compared to the amino acid sequence.

[0323] In some embodiments, the coding sequence (also referred to as the “coding region”) is codon-optimized for expression in an individual (e.g., a human) to whom the composition (e.g., a pharmaceutical composition) is to be administered. Therefore, in some embodiments, the sequence in the polynucleotide (e.g., a polynucleotide) may differ from the wild-type sequence encoding the associated incretin, even when the amino acid sequence of the incretin is wild-type.

[0324] In some embodiments, a codon optimization strategy for expression in relevant individuals (e.g., humans) and even in some cases for expression in specific cells or tissues.

[0325] Different species exhibit specific preferences for certain codons of particular amino acids. Without being bound by any single theory, codon preference (differences in codon usage between organisms) is generally associated with the translation efficiency of messenger RNA (mRNA), which in turn is considered to depend particularly on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the most frequently used codons in peptide synthesis. Therefore, gene expression can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are available, for example, at the "Codon Usage Database" at www.kazusa.orjp / codon / , and these tables can be modified in various ways. Computer algorithms for codon optimization of specific sequences for expression in a particular individual or its cells are also available, such as GeneForge (Aptagen; Jacobus, PA).

[0326] In some embodiments, the polynucleotides (e.g., polynucleotides) of this disclosure are codon-optimized, wherein the codons in the polynucleotides (e.g., polynucleotides) are adapted for use with human codons (referred to herein as "human codon-optimized polynucleotides"). Codons encoding the same amino acid appear at different frequencies in individuals (e.g., humans). Therefore, in some embodiments, the coding sequences of the polynucleotides of this disclosure are modified such that the frequencies of codons encoding the same amino acid correspond to the natural frequency of codons used according to human codons, for example, as shown in Table 13. For example, in the case of amino acid Ala, the wild-type coding sequence is preferably adjusted such that codon "GCC" is used at a frequency of 0.40, codon "GCT" at a frequency of 0.28, codon "GCA" at a frequency of 0.22, and codon "GCG" at a frequency of 0.10, etc. (see Table 13). Therefore, in some embodiments, this procedure (as illustrated for Ala) is applied to each amino acid encoded by the coding sequence of the polynucleotide to obtain a sequence adapted for use with human codons.

[0327] Table 13: Table of eukaryotic codon usage, including the frequency of use for each amino acid. Certain strategies for codon optimization and / or G / C enrichment for human expression are described in WO2002 / 098443, which is incorporated herein by reference in its entirety. In some embodiments, multi-parameter optimization strategies may be used to optimize coding sequences. In some embodiments, optimization parameters may include parameters that affect protein expression, such as those affecting transcriptional, mRNA, and / or translational levels. In some embodiments, exemplary optimization parameters include, but are not limited to, transcriptional level parameters (including, for example, GC content, shared splicing sites, recessive splicing sites, SD sequences, TATA boxes, termination signals, artificial recombination sites, and combinations thereof); mRNA level parameters (including, for example, RNA instability motifs, ribosome entry sites, repetitive sequences, and combinations thereof); translational level parameters (including, for example, codon usage, premature poly(A) sites, ribosome entry sites, secondary structures, and combinations thereof); or combinations thereof. In some embodiments, the coding sequence may be optimized by the GeneOptimizer algorithm as described in the following literature: Fath et al., “Multiparameter RNA and Codon Optimization: A Standardized Tool to Assess and Enhance Autologous Mammalian Gene Expression” PloS ONE 6(3): e17596; Rabb et al., “The GeneOptimizer Algorithm: using a sliding window approach to cope with the vast sequence space in multiparameter DNA sequence optimization” Systems and Synthetic Biology (2010) 4:215-225; and Graft et al., “Codon-optimized genes that enable increased heterologous expression in mammalian cells and elicit efficient immune responses in mice after vaccination of naked DNA” Methods Mol Med (2004) 94:197-210, the entire contents of which are incorporated herein by reference for the purposes described herein.In some embodiments, the encoded sequence may be optimized by Eurofins' adaptation and optimization algorithm "GENEius" compared to other optimization algorithms, as described in Eurofins' Application Notes: Eurofins' Adaptation and the optimization software "GENEius", the entire contents of which are incorporated herein by reference for the purposes described herein.

[0328] In some embodiments, the coding sequence used as described herein has an increased G / C content compared to the wild-type coding sequence of the relevant incretin. In some embodiments, the guanosine / cytidine (G / C) content of the coding region is modified relative to the wild-type coding sequence of the relevant incretin, but the amino acid sequence encoded by the polynucleotide is not modified.

[0329] Without being bound by any particular theory, it is proposed that GC enrichment can improve the translation of rewarded sequences. Generally, sequences with increased G (guanosine) / C (cytidine) content are more stable than sequences with increased A (adenosine) / U (uridine) content. The most favorable codons for stability (so-called alternative codon use) can be determined based on the fact that several codons encode the same amino acid (so-called genetic code degeneracy). Depending on the amino acid encoded by a polynucleotide, there are various possibilities for modifying the ribonucleic acid sequence compared to its wild-type sequence. In particular, codons containing A and / or U nucleotides can be modified by replacing them with other codons encoding the same amino acid but without A and / or U, or with lower A and / or U nucleotide content.

[0330] In some embodiments, the G / C content of the polynucleotide coding region described herein is increased by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or even more compared to the G / C content of the coding region before codon optimization, such as in wild-type RNA. In some embodiments, the G / C content of the polynucleotide coding region described herein is decreased by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or even more compared to the G / C content of the coding region before codon optimization, such as in wild-type RNA.

[0331] In some embodiments, the stability and translation efficiency of the polynucleotide may be incorporated into one or more components designed to contribute to the stability and / or translation efficiency of the polynucleotide; exemplary such components are described, for example, in WO2007 / 036366, which is incorporated herein by reference. In some embodiments, to increase the expression of the polynucleotide as used in this disclosure, the polynucleotide may be modified within its coding region (i.e., the sequence encoding the expressed peptide or protein) without altering the sequence of the expressed peptide or protein, for example, by increasing GC content to increase mRNA stability and / or by codon optimization, and thus enhancing translation in the cell.

[0332] Exemplary polynucleotides encoding incretins In some embodiments, the polynucleotide comprises a ribonucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identity with any of the sequences shown in Table 14 below. In addition to sequences encoding incretin agents and signal peptides, the polynucleotide includes exemplary polynucleotide features described herein, including a proximal cap sequence (AAUA), a 5' UTR sequence AACUAGUAUUCUUCUGGUCCCCACAGACUCAGAGAGAACCCGCCACC (SEQ ID NO: 49), a polyA tail sequence AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAGCAUAUGAGACUAAA ... UTR sequence CUGGUACUGCAUGCACGCAAUGCUAGCUGCCCCUUUCCCGUCCUGGGUACCCCGAGUCUCCCCCGACCUCGGGUCCCAGGUAUGCUCCCACCUCCACCUGCCCCACACCACCUCUGCUAGUUCCAGACACCUCC CAAGCACGCAGCAAUGCAGCUCAAAACGCUUAGCCUAGCCACACCCCCACGGGAAACAGCAGUGAUUAACCUUUAGCAAUAAACGAAAGUUUAACUAAGCUAUACUAACCCCAGGGUUGGUCAAUUUCGUGCCAGCCACACC (SEQ ID NO: 51).

[0333] Table 14: Exemplary polynucleotides encoding incretin agents, wherein examples x2 and x4 include linker peptides between repeating units and furin cleavage sites. RNA delivery technology The provided polynucleotides can be delivered for the therapeutic applications described herein using any suitable method known in the art, including, for example, delivery as naked RNA, or delivery mediated by viral and / or non-viral vectors, polymer-based vectors, lipid-based vectors, nanoparticles (e.g., lipid nanoparticles, polymer nanoparticles, lipid-polymer hybrid nanoparticles, etc.) and / or peptide-based vectors. See, for example, Wadhwa et al., “Opportunities and Challenges in the Delivery of mRNA-Based Vaccines”, Pharmaceutics (2020) 102 (27 pages), the contents of which are incorporated herein by reference for information on various methods that can be used to deliver the polynucleotides described herein.

[0334] In some embodiments, one or more polynucleotides may be formulated together with lipid nanoparticles for delivery (e.g., administration).

[0335] In some embodiments, lipid nanoparticles may be designed to protect polynucleotides from extracellular RNases and / or engineered for the systematic delivery of RNA to target cells (e.g., liver, intestinal, or pancreatic cells). In some embodiments, such lipid nanoparticles may be particularly suitable for delivering polynucleotides when administered intraperitoneally, intravenously, or intramuscularly to an individual.

[0336] Particles for delivering at least one polynucleotide The polynucleotides provided herein can be delivered via particles. In the context of this disclosure, the term "particle" refers to a structured entity formed from molecules or molecular complexes. In some embodiments, the term "particle" refers to a micrometer-sized or nanometer-sized structure, such as a micrometer-sized or nanometer-sized compact structure dispersed in a medium. In some embodiments, the particle is a particle containing nucleic acids, such as a particle containing polynucleotides.

[0337] Electrostatic interactions between positively charged molecules (such as polymers and lipids) and negatively charged nucleic acids (e.g., polynucleotides) are involved in particle formation. This leads to the recombination and spontaneous formation of nucleic acid particles (e.g., ribonucleic acid particles). In some embodiments, the nucleic acid particles (e.g., ribonucleic acid particles) are nanoparticles.

[0338] A “nucleic acid particle” (e.g., a ribonucleic acid particle) is a particle that encompasses or contains nucleic acids and is used to deliver nucleic acids (e.g., polynucleotides) to a target site (e.g., a cell, tissue, organ, or similar site). Nucleic acid particles (e.g., ribonucleic acid particles) can be formed from: (i) at least one cation or cation-ionizable lipid or lipid-like material, (ii) at least one cationic polymer (such as protamine), or a mixture of (i) and (ii), and (iii) a nucleic acid (e.g., a polynucleotide). Nucleic acid particles (e.g., ribonucleic acid particles) comprise multiple lipid nanoparticles (single lipid nanoparticles) and lipid complexes (LPX).

[0339] In some embodiments, nucleic acid particles (e.g., ribonucleic acid particles) comprise more than one type of nucleic acid molecule (e.g., polynucleotides), wherein the molecular parameters of the nucleic acid molecules may be similar to or different from each other, such as with respect to molar mass or basic structural components, such as molecular architecture, capping, coding regions or other features.

[0340] In some embodiments, the provided nucleic acid particles (e.g., ribonucleic acid particles) may comprise lipid nanoparticles. As used herein, "nanoparticle" refers to a particle having an average diameter suitable for non-enteric administration. In various embodiments, lipid nanoparticles may have an average size (e.g., average diameter) of about 30 nm to about 150 nm, about 40 nm to about 150 nm, about 50 nm to about 150 nm, about 60 nm to about 130 nm, about 70 nm to about 110 nm, about 70 nm to about 100 nm, about 70 nm to about 90 nm, or about 70 nm to about 80 nm. In some embodiments, lipid nanoparticles as described herein may have an average size (e.g., average diameter) of about 50 nm to about 100 nm. In some embodiments, lipid nanoparticles may have an average size (e.g., average diameter) of about 50 nm to about 150 nm. In some embodiments, lipid nanoparticles may have an average size (e.g., average diameter) of about 60 nm to about 120 nm. In some embodiments, the lipid nanoparticles as described in this disclosure may have an average size (e.g., average diameter) of about 30 nm, 35 nm, 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, or 150 nm.

[0341] The nucleic acid particles (e.g., ribonucleic acid particles) described herein may exhibit a polydispersity index of less than about 0.5, less than about 0.4, less than about 0.3, or about 0.2 or smaller. For example, nucleic acid particles (e.g., ribonucleic acid particles) may exhibit a polydispersity index in the range of about 0.1 to about 0.3 or about 0.2 to about 0.3.

[0342] The nucleic acid particles (e.g., ribonucleic acid particles) described herein are characterized by an "N / P ratio," which is the molar ratio of cationic (nitrogen) groups ("N" in N / P) in the cationic polymer to anionic (phosphate) groups ("P" in N / P) in the RNA. It should be understood that cationic groups are groups in cationic form (e.g., N...). +( ), or groups that can ionize to become cations. Using a single number in the N / P ratio (e.g., an N / P ratio of about 5) is intended to mean that the number is greater than 1; for example, an N / P ratio of about 5 means 5:1. In some embodiments, the nucleic acid particles (e.g., ribonucleic acid particles) described herein have an N / P ratio greater than or equal to 5. In some embodiments, the nucleic acid particles (e.g., ribonucleic acid particles) described herein have an N / P ratio of about 5, 6, 7, 8, 9, or 10. In some embodiments, the N / P ratio of the nucleic acid particles (e.g., ribonucleic acid particles) described herein is about 10 to about 50. In some embodiments, the N / P ratio of the nucleic acid particles (e.g., ribonucleic acid particles) described herein is about 10 to about 70. In some embodiments, the N / P ratio of the nucleic acid particles (e.g., ribonucleic acid particles) described herein is about 10 to about 120.

[0343] The nucleic acid particles (e.g., ribonucleic acid particles) described herein can be prepared using a variety of methods, which may involve obtaining a colloid from at least one cation or cation-ionizable lipid or lipid-like material and / or at least one cationic polymer and mixing the colloid with nucleic acids to obtain nucleic acid particles.

[0344] As used herein, the term "colloid" refers to a class of homogeneous mixtures in which dispersed particles do not settle. The insoluble particles in the mixture can be microscopic, ranging in size from 1 to 1000 nanometers. The mixture may be referred to as a colloid or a colloidal suspension. Sometimes the term "colloid" refers only to the particles in the mixture rather than the entire suspension.

[0345] The term "average diameter" or "mean diameter" refers to the average hydrodynamic diameter of a particle, as measured by dynamic laser light scattering (DLS), where data analysis is performed using a so-called cumulant algorithm, the results of which provide a so-called Z-mean with a length dimension and a dimensionless polydispersity index (PI) (Koppel, D., J. Chem. Phys. 57, 1972, pp. 4814-4820, ISO 13321, which is incorporated herein by reference). In this document, the terms "average diameter," "mean diameter," "diameter," or "size" of a particle are used synonymously with the value of the Z-mean.

[0346] The "polydispersity index" is preferably calculated based on dynamic light scattering measurements using the so-called cumulative analysis mentioned in the definition of "mean diameter". Under certain prerequisites, it can be regarded as a measure of the size distribution of an aggregate of ribonucleic acid nanoparticles (e.g., ribonucleic acid nanoparticles).

[0347] Different types of nucleic acid particles have previously been described as suitable for delivering nucleic acids in particulate form (e.g., Kaczmarek et al., 2017, Genome Medicine 9, 60, which is incorporated herein by reference). For non-viral nucleic acid delivery media, the encapsulation of nucleic acids in nanoparticles physically protects the nucleic acids from degradation and, depending on specific chemical properties, can facilitate cellular uptake and endosome escape.

[0348] This disclosure describes particles comprising nucleic acids (e.g., polynucleotides), at least one cationic or cationically ionizable lipid or lipid-like material, and / or at least one cationic polymer bound to a nucleic acid (e.g., polynucleotides) to form nucleic acid particles (e.g., ribonucleic acid particles, such as ribonucleic acid nanoparticles), and compositions comprising such particles. The nucleic acid particles (e.g., ribonucleic acid particles, such as ribonucleic acid nanoparticles) may comprise nucleic acids (e.g., polynucleotides) compounded with the particles in various forms through non-covalent interactions. The particles described herein are not viral particles, and in particular, they are not infectious viral particles, i.e., they cannot virally infect cells.

[0349] Some of the embodiments described herein relate to compositions, methods, and uses that relate to more than one, such as 2, 3, 4, 5, 6, or even more nucleic acid species (e.g., polynucleotide species).

[0350] In nucleic acid particle (e.g., ribonucleic acid particles, e.g., ribonucleic acid nanoparticles) formulations, it is possible to individually formulate each nucleic acid species (e.g., polynucleotide species) into a single nucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulation. In this case, each single nucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulation will contain one nucleic acid species (e.g., polynucleotide species). The single nucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulation can exist as a separate entity, for example, in a separate container. Such formulations can be obtained by separately providing each nucleic acid species (e.g., polynucleotide species) (typically in the form of a solution containing nucleic acid) and a particle-forming reagent, thereby allowing particle formation. The corresponding particle will exclusively contain the specific nucleic acid species (e.g., polynucleotide species) provided during particle formation (single-particle formulation).

[0351] In some embodiments, the composition (such as a pharmaceutical composition) comprises more than one type of mononuclear particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulation. The corresponding pharmaceutical composition is referred to as a “mixed microparticle formulation.” The mixed microparticle formulation as described in this invention can be obtained by separately forming a mononucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulation as described above, followed by a step of mixing the mononucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulation. Through the mixing step, a formulation comprising a mixed population of nucleic acid particles is obtained. The mononucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) population can be together in a container containing the mixed population of mononucleic acid particle (e.g., ribonucleic acid particle, e.g., ribonucleic acid nanoparticle) formulations.

[0352] Alternatively, different nucleic acid species (e.g., polynucleotide species) can be formulated together as a "combined microparticle formulation." Such formulations are obtained by providing a combination of different nucleic acid species (e.g., polynucleotide species) and particle-forming agents (typically a combination solution), thereby allowing particle formation. In contrast to a "mixed microparticle formulation," a "combined microparticle formulation" will typically contain particles comprising more than one nucleic acid species (e.g., polynucleotide species). In a combined microparticle composition, different nucleic acid species (e.g., polynucleotide species) are typically present together within a single particle.

[0353] In some embodiments, when present in the provided nucleic acid particles (e.g., ribonucleic acid particles, such as lipid nanoparticles), the nucleic acid (e.g., polynucleotides) is resistant to degradation by nucleases in aqueous solution.

[0354] In some embodiments, the nucleic acid particles (e.g., ribonucleic acid particles) are lipid nanoparticles. In some embodiments, the lipid nanoparticles are liver-targeting lipid nanoparticles. In some embodiments, the lipid nanoparticles are cationic lipid nanoparticles comprising one or more cationic lipids (e.g., the lipids described herein). In some embodiments, the cationic lipid nanoparticles may comprise at least one cationic lipid, at least one polymer-coupled lipid, and at least one accessory lipid (e.g., at least one neutral lipid).

[0355] cationic polymer materials Cationic polymers have been considered for use in the development of particulate delivery media, as reported in PCT Application Publication No. WO2021 / 001417, the entire contents of which are incorporated herein by reference. As used herein, the term “polymer” refers to a composition comprising one or more molecules, said molecules comprising repeating units of one or more monomers. As used herein, “polymer,” “polymer material,” and “polymer composition” are used interchangeably and, unless otherwise stated, refer to a composition of polymer molecules. Those skilled in the art will appreciate that polymer compositions comprise polymer molecules of different lengths (e.g., comprising different amounts of monomers). Polymer compositions described herein are characterized by one or more of a normalized molecular weight (Mn), a weight-average molecular weight (Mw), and / or a polydispersity index (PDI). In some embodiments, such repeating units may be all identical (“homomer”); or, in some cases, more than one type of repeating unit may be present within the polymer material (“hybrid” or “copolymer”). In some cases, the polymer is biologically derived, such as a biopolymer, like a protein. In some cases, additional portions may also exist in polymer materials, such as targeted portions, as described herein.

[0356] In some embodiments, the polymer used as described herein may be a copolymer. The repeating units forming the copolymer may be arranged in any manner. For example, in some embodiments, the repeating units may be arranged in a random order; alternatively or additionally, in some embodiments, the repeating units may be arranged in an alternating order or as a “block” copolymer, for example, a “block” copolymer comprising one or more regions each containing a first repeating unit (e.g., a first block) and one or more regions each containing a second repeating unit (e.g., a second block). Block copolymers may have two (diblock copolymers), three (triblock copolymers), or more different blocks.

[0357] In some embodiments, the polymeric materials used in this disclosure are biocompatible. In some embodiments, the biocompatible materials are biodegradable, for example, capable of chemical and / or biological degradation within physiological environments (e.g., in vivo).

[0358] In some embodiments, the polymer material may be or contain protamine or polyalkylene imide.

[0359] As those skilled in the art will recognize, the term "protamine" is generally used to refer to any of a variety of strongly basic proteins of relatively low molecular weight, rich in arginine, and found to be particularly associated with DNA, replacing somatic histones in sperm cells of various animals (such as fish). Specifically, the term "protamine" is generally used to refer to a protein found in fish sperm that is strongly basic, water-soluble, does not coagulate upon heat, and primarily produces arginine upon hydrolysis. In purified form, protamine is used in long-acting formulations of insulin and to neutralize the anticoagulant effect of heparin.

[0360] In some embodiments, as used herein, the term "protamine" refers to a protamine amino acid sequence obtained from or derived from a natural or biological source, including fragments thereof and / or a multimeric form of such amino acid sequence or fragments thereof, as well as (synthetic) polypeptides that are artificial and specifically designed for a particular purpose and cannot be isolated from a natural or biological source.

[0361] In some embodiments, the polyalkylene imide comprises polyethyleneimine (PEI) and / or polypropylene imide. In some embodiments, the preferred polyalkylene imide is polyethyleneimine (PEI). In some embodiments, the average molecular weight of PEI is preferably 0.75∙10² to 10⁷ Da, more preferably 1000 to 10⁵ Da, more preferably 10000 to 40000 Da, more preferably 15000 to 30000 Da, and even more preferably 20000 to 25000 Da.

[0362] The cationic materials (e.g., polymeric materials, including polycationic polymers) considered herein include those capable of electrostatically binding nucleic acids. In some embodiments, the cationic polymeric materials considered herein include any cationic polymeric material to which nucleic acids can bind (e.g., by forming a complex with the nucleic acid or forming vesicles therein that encapsulate or encapsulate the nucleic acid).

[0363] In some embodiments, the particles described herein may comprise polymers other than cationic polymers, such as non-cationic polymer materials and / or anionic polymer materials. In summary, anionic and neutral polymer materials are referred to herein as non-cationic polymer materials.

[0364] lipid compositions Lipids and lipid-like materials Lipids and lipid-like materials have also been considered for use in the development of particulate delivery media. The terms “lipid” and “lipid-like material” are broadly defined herein as molecules comprising one or more hydrophobic moieties or groups and optionally one or more hydrophilic moieties or groups. Molecules comprising both hydrophobic and hydrophilic moieties are also commonly referred to as amphiphiles. Lipids are generally poorly soluble in water. In an aqueous environment, the amphiphilic nature allows the molecules to self-assemble into organized structures and different phases. One of these phases consists of a lipid bilayer, as it exists in vesicles, multilayer / monolayer lipid particles, or membranes in an aqueous environment. Hydrophobicity can be imparted by including nonpolar groups, including but not limited to long-chain saturated and unsaturated aliphatic hydrocarbon groups and such groups substituted with one or more aromatic, cyclic aliphatic, or heterocyclic groups. Hydrophilic groups may include polar and / or charged groups and include carbohydrate, phosphate, carboxyl, sulfate, amino, thiosulfate, nitro, hydroxyl, and other similar groups.

[0365] Typically, amphiphilic compounds have a polar head connected to a long hydrophobic tail. In some embodiments, the polar portion is soluble in water, while the nonpolar portion is insoluble in water. Additionally, the polar portion may carry a positive or negative charge. Alternatively, the polar portion may carry both a positive and negative charge and be an amphoteric ion or an internal salt. For the purposes of this disclosure, the amphiphilic compound may be, but is not limited to, one or more natural or non-natural lipids and lipid-like compounds.

[0366] "Lipid-like materials" are substances that are structurally and / or functionally related to lipids but may not be considered lipids in the strict sense. For example, the term includes compounds capable of forming amphiphilic layers when present in vesicles, multilayer / monolayer lipid particles, or membranes in an aqueous environment, and includes surfactants or synthetic compounds having both hydrophilic and hydrophobic portions. Generally, the term refers to molecules containing hydrophilic and hydrophobic portions with different structural organization, which may be similar to or dissimilar to the structural organization of lipids.

[0367] Specific examples of amphiphilic compounds that may be included in the amphiphilic layer include, but are not limited to, phospholipids, aminolipids, and sphingolipids.

[0368] Generally, lipids can be classified into eight categories: fatty acids, glycerolipids, glycerophospholipids, sphingolipids, glycolipids, polyketides (derived from the condensation of ketoacyl subunits), sterols, and isoprenolol lipids (derived from the condensation of isoprenoid subunits). Although the term "lipid" is sometimes used synonymously with fat, fat is a subgroup of lipids called triglycerides. Lipids also encompass molecules such as fatty acids and their derivatives (including triglycerides, diglycerides, monoglycerides, and phospholipids), as well as sterol-containing metabolites such as cholesterol.

[0369] Fatty acids are a group of distinct molecules composed of hydrocarbon chains terminated by carboxylic acid groups; this arrangement gives the molecule a polar hydrophilic end and a water-insoluble, nonpolar hydrophobic end. The carbon chain, typically between four and 24 carbons in length, can be saturated or unsaturated and can be linked to functional groups containing oxygen, halogens, nitrogen, and sulfur. If a fatty acid contains double bonds, it may exhibit cis or trans geometric isomerism, which significantly affects the molecular configuration. Cis double bonds cause the fatty acid chain to bend, an effect of complexation with more double bonds in the chain. Other major lipid classes within the fatty acid category are fatty esters and fatty amides.

[0370] Glycerol lipids consist of mono-, di-, and tri-substituted glycerols, the most well-known being fatty acid triesters of glycerol, called triglycerides. The term "triacylglycerol" is sometimes used synonymously with "triglyceride." In this class of compounds, the three hydroxyl groups of glycerol are typically esterified from different fatty acids. Other subclasses of glycerol lipids are represented by glycosylglycerols, characterized by the presence of one or more sugar residues linked to glycerol via glycosidic bonds.

[0371] Glycerophospholipids are amphiphilic molecules (containing both hydrophobic and hydrophilic regions) with a glycerol core consisting of two fatty acid-derived "tails" linked by ester bonds and a "head" group linked by phosphate ester bonds. Examples of glycerophospholipids commonly referred to as phospholipids (although sphingomyelins are also classified as phospholipids) include phosphatidylcholine (also known as PC, GPCho, or lecithin), phosphatidylethanolamine (PE or GPEtn), and phosphatidylserine (PS or GPSer).

[0372] Sphingolipids are members of a family of complex compounds that share a common structural feature: a sphingosine-like base backbone. The major sphingosine-like base in mammals is generally referred to as sphingosine. Ceramides (N-acyl-sphingosine-like bases) are the major subclass of sphingosine-like base derivatives of fatty acids with amide linkages. These fatty acids are typically saturated or monounsaturated, with chain lengths of 16 to 26 carbon atoms. The major sphingomyelin in mammals is sphingomyelin (ceramide phosphocholine), while insects primarily contain ceramide phosphoethanolamine, and fungi possess phytoceramide phosphoinositol and a mannose-containing head group. Glycosphingolipids are a distinct family of molecules composed of one or more sugar residues linked to a sphingosine-like base via glycosidic bonds. Examples of this class include simple and complex glycosphingolipids such as cerebrosides and gangliosides.

[0373] Sterols, such as cholesterol and its derivatives, or tocopherol and its derivatives, along with glycerophospholipids and sphingomyelin, are important components of membrane lipids.

[0374] Glycolipids are compounds in which fatty acids are directly linked to the sugar backbone, forming a structure compatible with the membrane bilayer. In glycolipids, monosaccharides replace the glycerol backbone present in glycerol lipids and glycerophospholipids. The most familiar glycolipid is the acylated glucosamine precursor of the lipid A component of lipopolysaccharides in Gram-negative bacteria. Typical lipid A molecules are disaccharides of glucosamine, derived from up to seven fatty acyl chains. The smallest lipopolysaccharide required for growth in *E. coli* is Kdo2-lipid A, a hexaacylated disaccharide of glucosamine formed by glycosylation of two 3-deoxy-D-manno-octulose (Kdo) residues.

[0375] Polyketides are synthesized via the polymerization of acetyl and propionyl subunits using classical enzymes and iterative and modular enzymes that share mechanical features with fatty acid synthases. Polyketides comprise a vast array of secondary metabolites and natural products from animal, plant, bacterial, fungal, and marine sources, exhibiting immense structural diversity. Many polyketides are cyclic molecules, and their backbones are typically further modified through glycosylation, methylation, hydroxylation, oxidation, or other processes.

[0376] Lipids and lipid-like materials can be cationic, anionic, or neutral. Neutral lipids or lipid-like materials exist as uncharged or neutral zwitterionic forms at a selected pH.

[0377] In some embodiments, suitable lipids or lipid-like materials as used in this disclosure include those described in WO2020 / 128031 and US2020 / 0163878, the entire contents of which are incorporated herein by reference for the purposes described herein.

[0378] Cations or cation-ionizable lipids or lipid-like materials In some embodiments, the cationic or cationically ionizable lipid or lipid-like material contemplated herein includes any cationic or cationically ionizable lipid or lipid-like material capable of electrostatically binding nucleic acids. In one embodiment, the cationic or cationically ionizable lipid or lipid-like material contemplated herein may bind to nucleic acids, for example, by forming a complex with nucleic acids or forming vesicles therein that encapsulate or seal nucleic acids.

[0379] Cationic lipids or lipid-like materials are characterized by having a net positive charge (e.g., at the relevant pH). Cationic lipids or lipid-like materials bind negatively charged nucleic acids through electrostatic interactions. Generally, cationic lipids have lipophilic moieties, such as sterols, acyl chains, diacyl groups, or multiple acyl chains, and the head group of the lipid typically carries a positive charge.

[0380] In some embodiments, cationic lipids or lipid-like materials possess a net positive charge only at certain pH levels, particularly acidic pH levels, and preferably no net positive charge at different, preferably higher pH levels, such as physiological pH levels, i.e., neutral. This ionization behavior is thought to enhance efficacy by facilitating endosome escape and reducing toxicity compared to particles that remain cationic at physiological pH levels.

[0381] In some embodiments, the cation or cation-ionizable lipid or lipid-like material comprises a head group that includes at least one nitrogen atom (N) that is positively charged or capable of being protonated.

[0382] Examples of cationic lipids include, but are not limited to, 1,2-dioleoyl-3-trimethylammonium propane (DOTAP); N,N-dimethyl-2,3-dioleoyloxypropylamine (DODMA), 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA), 3-(N-(N′,N′-dimethylaminoethane)-carbamoyl)cholesterol (DC-Chol), dimethylbis(octadecylammonium) (DDAB); 1,2-dioleoyl-3-dimethylammonium-propane (DODAP); 1,2-diacyloxy-3-dimethylammonium propane; 1,2-dialkyloxy-3-dimethylammonium propane; bis(octadecyl)dimethylammonium chloride (DODAC); 1,2-distearate-N,N-dimethylammonium chloride (DODAC); and 1,2-distearate-N,N-dimethylammonium chloride (DODAC). -3-aminopropane (DSDMA), 2,3-di(tetradecoxy)propyl-(2-hydroxyethyl)-dimethylazonium (DMRIE), 1,2-dimyristoyl-sn-glycerol-3-ethylphosphocholine (DMEPC), 1,2-dimyristoyl-3-trimethylammonium propane (DMTAP), 1,2-dioleoyloxypropyl-3-dimethyl-hydroxyethylammonium bromide (DORIE), and trifluoroacetic acid 2,3-dioleoyloxy-N-[2-(sperminecarbamate)ethyl]-N,N-dimethyl-1-propanium (DOSPA), 1,2-dilinoleoyloxy-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinyloxy-N,N-dimethylaminopropane ( DLenDMA), bis(octadecylaminoglycyl)sperylamine (DOGS), 3-dimethylamino-2-(cholesterol-5-en-3-β-oxybut-4-oxy)-1-(cis,cis-9,12-octadecadienoxy)propane (CLinDMA), 2-[5′-(cholesterol-5-en-3-β-oxy)-3′-oxaproloxy)-3-dimethyl-1-(cis,cis-9′,12′-octadecadienoxy)propane (CpLinDMA), N,N-dimethyl-3,4-dioleoyloxybenzylamine (DMOBA), 1,2-N,N′-dioleoylaminoformyl-3-dimethylaminopropane (DOcarbDAP), 2,3-dilinoleoyloxy- N,N-Dimethylpropylamine (DLinDAP), 1,2-N,N'-Dilinoleylaminoformyl-3-dimethylaminopropane (DLincarbDAP), 1,2-Dilinoleylaminoformyl-3-dimethylaminopropane (DLinCDAP), 2,2-Dilinoleyl-4-dimethylaminomethyl-[1,3]-dioxacyclopentane (DLin-K-DMA), 2,2-Dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxacyclopentane (DLin-K-XTC2-DMA), 2,2-Dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxacyclopentane (DLin-KC2-DMA), heptadecane-6,9,28,31-Tetraen-19-yl-4-(dimethylamino)butyrate (DLin-MC3-DMA), N-(2-hydroxyethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)-1-propaneammonium bromide (DMRIE), (±)-N-(3-aminopropyl)-N,N-dimethyl-2,3-bis(cis-9-tetradecenyloxy)-1-propaneammonium bromide (GAP-DMORIE), (±)-N-(3-aminopropyl)-N,N-dimethyl-2,3-bis(dodecyloxy)-1-propaneammonium bromide (GAP-DLR) IE), (±)-N-(3-aminopropyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)-1-propaneammonium bromide (GAP-DMRIE), N-(2-aminoethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)-1-propaneammonium bromide (βAE-DMRIE), N-(4-carboxybenzyl)-N,N-dimethyl-2,3-bis(oleoyloxy)prop-1-ammonium (DOBAQ), 2-({8-[(3β)-cholesterol-5-en-3-yloxy]octyl}oxy)-N,N-dimethyl-3-[(9 [Z,12Z)-octadec-9,12-dien-1-yloxy]prop-1-amine (octyl-CLinDMA), 1,2-dimyristoyl-3-dimethylammonium-propane (DMDAP), 1,2-dipalmitoyl-3-dimethylammonium-propane (DPDAP), N1-[2-((1S)-1-[(3-aminopropyl)amino]-4-[di(3-aminopropyl)amino]butylcarbamoylamino)ethyl]-3,4-di[oleoyloxy]-benzamide (MVL5), 1,2-dioleoyl-sn-glycerol-3-ethylphosphocholine (DO) EPC), 2,3-bis(dodecyloxy)-N-(2-hydroxyethyl)-N,N-dimethylprop-1-ammonium bromide (DLRIE), N-(2-aminoethyl)-N,N-dimethyl-2,3-bis(tetradecyloxy)prop-1-ammonium bromide (DMORIE), 8,8'-((((2(dimethylamino)ethyl)thio)carbonyl)azanyl)dioctanoic acid di((Z)-non-2-en-1-yl) ester (ATX), N,N-dimethyl-2,3-bis(dodecyloxy)prop-1-amine (DLDMA), N,N-dimethyl-2,3-Bis(tetradecyloxy)prop-1-amine (DMDMA), di((Z)-non-2-en-1-yl)-9-((4-(dimethylaminobutyryl)oxy)heptadecanoic acid ester (L319), N-dodecyl-3-((2-dodecylaminoformyl-ethyl)-{2-[(2-dodecylaminoformyl-ethyl)-2-{(2-dodecylaminoformyl-ethyl)-[2-(2-dodecylaminoformyl-ethylamino)-ethyl]-amino}-ethylamino)propionamide (lipid 98N12-5), 1-[2-[bis(2-hydroxydodecyl)amino]ethyl-[2-[4-[2-[bis(2-hydroxydodecyl)amino]ethyl]pyrazin-1-yl]ethyl]amino]dodecane-2-ol (lipid C12-200), LIPOFECTIN® (Commercially available cationic liposomes containing DOTMA and 1,2-dioleoyl-sn-3-phosphate ethanolamine (DOPE), from GIBCO / BRL, Grand Island, NY); LIPOFECTAMINE® (commercially available cationic liposomes containing N-(1-(2,3-dioleoyloxy)propyl)-N-(2-(spermineformylamino)ethyl)-N,N-dimethyltrifluoroacetate ammonium (DOSPA) and (DOPE), from GIBCO / BRL); and TRANSFECTAM® (commercially available cationic liposomes containing ethanol-containing bis(octadecylamino)glycylcarboxyspermine (DOGS), from PromegaCorp., Madison, NY); Wis.), or any combination thereof. Other suitable cationic lipids used in this disclosure include those described in WO2020 / 128031 and US2020 / 0163878, the entire contents of which are incorporated herein by reference for the purposes described herein. Other suitable cationic lipids used in this disclosure include those described in WO2010 / 053572 (including C12-200 as described in paragraph

[00225] ) and WO2012 / 170930, which are incorporated herein by reference for the purposes described herein. Additional suitable cationic lipids used in this disclosure include HGT4003, HGT5000, HGTS001, HGT5001, and HGT5002 (see US2015 / 0140070, which is incorporated herein by reference in its entirety).

[0383] In some embodiments, formulations that can be used in pharmaceutical compositions (e.g., immunogenic compositions, such as vaccines) as described herein may comprise at least one cationic lipid. Representative cationic lipids include, but are not limited to, 1,2-dilinoleoyloxy-3-(dimethylamino)acetoxypropane (DLin-DAC), 1,2-dilinoleoyloxy-3-morpholinopropane (DLin-MA), 1,2-dilinoleoyl-3-dimethylaminopropane (DLinDAP), 1,2-dilinoleothio-3-dimethylaminopropane (DLin-S-DMA), 1-linoleoyl-2-linoleoyloxy-3-dimethylaminopropane (DLin-2-DMAP), 1,2-dilinoleoyloxy-3-trimethylaminopropane chloride (DLin-TMA.CI), 1,2-dilinoleoyl-3-trimethylaminopropane chloride (DLin-TAP.CI), 1,2-dilinoleoyl... 3-(N-methylpyrazinyl)propane (DLin-MPZ), 3-(N,N-dioleinylamino)-1,2-propanediol (DLinAP), 3-(N,N-dioleinylamino)-1,2-propanediol (DOAP), 1,2-dioleinyl-3-(2-N,N-dimethylamino)ethoxypropane (DLin-EG-DMA), and 2,2-dioleinyl-4-dimethylaminomethyl-[1,3]-dioxane (DLin-K-DMA), 2,2-dioleinyl-4-(2-dimethylaminoethyl)-[1,3]-dioxane (DLin-KC2-DMA); dioleinyl-methyl-4-dimethylaminobutyrate (DLin-MC3-DMA); MC3 (US2010 / 0324120, which is incorporated herein by reference in its entirety).

[0384] In some embodiments, the amino or cationic lipids that can be used as described in this disclosure have at least one protonable or deprotonable group, such that the lipid is positively charged at a pH equal to or below physiological pH (e.g., pH 7.4) and neutral at a second pH, preferably equal to or above physiological pH. It will be understood that the addition or removal of protons varies with pH as an equilibrium process, and the reference to charged or neutral lipids refers to the properties of the dominant species and does not require all lipids to be present in a charged or neutral form. Lipids having more than one protonable or deprotonable group or being zwitterionic are not excluded and are equally applicable in the context of this invention.

[0385] In some embodiments, the protonable lipid has a pKa of a protonable group in the range of about 4 to about 11, for example, about 5 to about 7.

[0386] In some embodiments, in the lipid composition as used in this disclosure, cationic lipids may comprise about 10 mol% to about 100 mol%, about 20 mol% to about 100 mol%, about 30 mol% to about 100 mol%, about 40 mol% to about 100 mol%, or about 50 mol% to about 100 mol% of total lipids.

[0387] Additional lipids or lipid-like materials In some embodiments, formulations as used herein may comprise lipids or lipid-like materials other than cationic or cationic ionizable lipids or lipid-like materials, i.e., non-cationic lipids or lipid-like materials (including non-cationic ionizable lipids or lipid-like materials). In summary, anionic and neutral lipids or lipid-like materials are referred to herein as non-cationic lipids or lipid-like materials. In some embodiments, optimizing the formulation of nucleic acid particles by adding other hydrophobic components (such as cholesterol and lipids) in addition to ionizable / cationic lipids or lipid-like substances can, for example, enhance particle stability and nucleic acid delivery efficiency.

[0388] In some embodiments, lipids or lipid-like materials may be incorporated that may or may not affect the overall charge of the particles. In some embodiments, such lipids or lipid-like materials are non-cationic lipids or lipid-like materials.

[0389] In some embodiments, non-cationic lipids may comprise, for example, one or more anionic lipids and / or neutral lipids. “Anionic lipids” carry a negative charge (e.g., at a selected pH).

[0390] "Assistant lipids" are present in a neutral zwitterionic form (e.g., at a selected pH) or, in some embodiments, in a cationic or positively charged form at physiological pH. In some embodiments, the formulation comprises one of the following assistant lipid components: (1) phospholipids, (2) cholesterol or a derivative thereof; or (3) a mixture of phospholipids and cholesterol or a derivative thereof. Examples of cholesterol derivatives include, but are not limited to, cholesterol, cholesterone, cholesterol, coprostinol, cholesteryl-2'-hydroxyethyl ether, cholesteryl-4'-hydroxybutyl ether, tocopherol and its derivatives and mixtures thereof.

[0391] Specific exemplary phospholipids that may be used include, but are not limited to, phosphatidylcholine, phosphatidylethanolamine, phosphatidylglycerol, phosphatidic acid, phosphatidylserine, or sphingomyelin. Specifically, these phospholipids include diacylphosphatidylcholine, such as distearylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), dimyristoylphosphatidylcholine (DMPC), octacoylphosphatidylcholine, dilauroylphosphatidylcholine, dipalmitoylphosphatidylcholine (DPPC), arachidoylphosphatidylcholine (DAPC), disambacylphosphatidylcholine (DBPC), triacylphosphatidylcholine (DTPC), bis(tetracoylphosphatidylcholine) (DLPC), palmitoyloleyl-phosphatidylcholine (POPC), 1,2-di-O-octadecenyl-sn-glycerol-3-phosphocholine (18:0 diether PC), 1-oleoyl-2-cholesterolyl-semisuccinyl-sn-glycerol-3-phosphocholine (OChemsPC), and 1-hexadecyl-sn-glycerol-3-phosphocholine (C16 Lyso). PC) and phosphatidylethanolamines, particularly diacylphosphatidylethanolamines such as dioleoylphosphatidylethanolamine (DOPE), distearate-phosphatidylethanolamine (DSPE), dipalmitoylphosphatidylethanolamine (DPPE), dimyristoylphosphatidylethanolamine (DMPE), dilauroylphosphatidylethanolamine (DLPE), diphyranoylphosphatidylethanolamine (DPyPE), and other phosphatidylethanolamine lipids with different hydrophobic chains.

[0392] In some embodiments, the formulations used as described in this disclosure include DSPC or DSPC and cholesterol.

[0393] In some embodiments, the formulations used in this disclosure comprise ionizable or cationic auxiliary lipids.

[0394] In some embodiments, the formulations used in this disclosure include both cationic lipids and additional (non-cationic) lipids.

[0395] In some embodiments, the formulations described herein comprise polymer-coupled lipids, such as PEGylated lipids. "PEGylated lipids" or "PEG-coupled lipids" comprise both a lipid portion and a polyethylene glycol portion. PEGylated lipids are known in the art.

[0396] To avoid being bound by theory, the amount of (total) cationic lipids, compared to the amount of other lipids in the formulation, can affect important characteristics such as the charge, particle size, stability, tissue selectivity, and bioactivity of nucleic acids. In some embodiments, the molar ratio of at least one cationic lipid to at least one additional lipid is about 10:0 to about 1:9, about 4:1 to about 1:2, or about 3:1 to about 1:1.

[0397] lipid complex particles In some embodiments of this disclosure, the RNA described herein may be present in RNA-lipid complex particles.

[0398] "RNA-lipid complex particles" contain lipids (particularly cationic lipids) and RNA. Electrostatic interactions between positively charged liposomes and negatively charged RNA lead to the recombination and spontaneous formation of RNA-lipid complex particles. Positively charged liposomes are typically synthesized using cationic lipids (such as DOTMA) and additional lipids (such as DOPE). In one embodiment, the RNA-lipid complex particle is a nanoparticle.

[0399] In some embodiments, the RNA-lipid complex particles comprise both cationic lipids and additional lipids. In one exemplary embodiment, the cationic lipid is DOTMA and the additional lipid is DOPE.

[0400] In some embodiments, the molar ratio of at least one cationic lipid to at least one additional lipid is about 10:0 to about 1:9, about 4:1 to about 1:2, or about 3:1 to about 1:1. In specific embodiments, the molar ratio may be about 3:1, about 2.75:1, about 2.5:1, about 2.25:1, about 2:1, about 1.75:1, about 1.5:1, about 1.25:1, or about 1:1. In one exemplary embodiment, the molar ratio of at least one cationic lipid to at least one additional lipid is about 2:1.

[0401] In some embodiments, the RNA lipid complex particles have an average diameter in one embodiment ranging from about 200 nm to about 1000 nm, from about 200 nm to about 800 nm, from about 250 nm to about 700 nm, from about 400 nm to about 600 nm, from about 300 nm to about 500 nm, or from about 350 nm to about 400 nm. In a specific embodiment, the RNA-lipid complex particles have an average diameter of about 200 nm, about 225 nm, about 250 nm, about 275 nm, about 300 nm, about 325 nm, about 350 nm, about 375 nm, about 400 nm, about 425 nm, about 450 nm, about 475 nm, about 500 nm, about 525 nm, about 550 nm, about 575 nm, about 600 nm, about 625 nm, about 650 nm, about 700 nm, about 725 nm, about 750 nm, about 775 nm, about 800 nm, about 825 nm, about 850 nm, about 875 nm, about 900 nm, about 925 nm, about 950 nm, about 975 nm, or about 1000 nm. In one embodiment, the RNA-lipid complex particles have an average diameter in the range of about 250 nm to about 700 nm. In another embodiment, the RNA-lipid complex particles have an average diameter in the range of about 300 nm to about 500 nm. In an exemplary embodiment, the RNA-lipid complex particles have an average diameter of about 400 nm.

[0402] The RNA-lipid complex particles and compositions comprising RNA-lipid complex particles described herein can be used to deliver RNA to target tissues after non-enteral administration, particularly after intravenous administration. The RNA-lipid complex particles can be prepared using liposomes, which can be obtained by injecting an ethanolic solution of lipids into water or a suitable aqueous phase. In one embodiment, the aqueous phase has an acidic pH. In one embodiment, the aqueous phase contains, for example, about 5 mM of acetic acid. Liposomes can be used to prepare RNA-lipid complex particles by mixing liposomes with RNA. In one embodiment, the liposomes and RNA-lipid complex particles comprise at least one cationic lipid and at least one additional lipid. In one embodiment, the at least one cationic lipid comprises 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA) and / or 1,2-dioleoyl-3-trimethylammonium propane (DOTAP). In one embodiment, the at least one additional lipid comprises 1,2-di-(9Z-octadecenoyl)-sn-glycerol-3-phosphate ethanolamine (DOPE), cholesterol (Chol), and / or 1,2-dioleoyl-sn-glycerol-3-phosphate choline (DOPC). In one embodiment, at least one cationic lipid comprises 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA) and at least one additional lipid comprises 1,2-di-(9Z-octadecenyl)-sn-glycerol-3-phosphate ethanolamine (DOPE). In one embodiment, the lipid particle and RNA lipid complex particle comprises 1,2-di-O-octadecenyl-3-trimethylammonium propane (DOTMA) and 1,2-di-(9Z-octadecenyl)-sn-glycerol-3-phosphate ethanolamine (DOPE).

[0403] RNA-lipid complex particles targeting the spleen are described in WO2013 / 143683, which is incorporated herein by reference. It has been found that RNA-lipid complex particles with a net negative charge can be used to preferentially target spleen tissue or spleen cells, such as antigen-presenting cells, particularly dendritic cells. Therefore, RNA accumulation and / or RNA expression occur in the spleen after administration of the RNA-lipid complex particles. Therefore, the RNA-lipid complex particles of this disclosure can be used to express RNA in the spleen. In one embodiment, no or substantially no RNA accumulation and / or RNA expression occurs in the lungs and / or liver after administration of the RNA-lipid complex particles. In one embodiment, RNA accumulation and / or RNA expression occur in antigen-presenting cells (such as professional antigen-presenting cells in the spleen) after administration of the RNA-lipid complex particles. Therefore, the RNA-lipid complex particles of this disclosure can be used to express RNA in such antigen-presenting cells. In one embodiment, the antigen-presenting cells are dendritic cells and / or macrophages.

[0404] Lipid nanoparticles (LNP) In some embodiments, the nucleic acids (such as RNA) described herein are administered in the form of lipid nanoparticles (LNPs). In some embodiments, the LNP may comprise any lipid capable of forming a particle, one or more nucleic acid molecules being linked to the particle, or one or more nucleic acid molecules being encapsulated therein.

[0405] In some embodiments, the LNP comprises one or more cationic lipids and one or more stable lipids. Stable lipids include neutral lipids, accessory lipids, and polyethylene glycol-modified lipids.

[0406] In some embodiments, the LNP comprises cationic lipids, auxiliary lipids, sterols, polymer-coupled lipids; and RNA encapsulated within or bound to lipid nanoparticles.

[0407] In some embodiments, the cofactor lipid is a phosphatidylcholine, such as 1,2-distearyl-sn-glycerol-3-phosphate choline (DSPC), 1,2-dipalmitoyl-sn-glycerol-3-phosphate choline (DPPC), 1,2-dimyristoyl-sn-glycerol-3-phosphate choline (DMPC), 1-palmitoyl-2-oleoyl-sn-glycerol-3-phosphate choline (POPC), 1,2-dioleoyl-sn-glycerol-3-phosphate choline (DOPC), phosphatidylethanolamines such as 1,2-dioleoyl-sn-glycerol-3-phosphate ethanolamine (DOPE), or sphingomyelin (SM). In some embodiments, the cofactor lipid is 1,2-dioleoyl-3-trimethylammonium propane (DOTAP), 1-α-phosphatidylserine (PS), or DOPE. In some embodiments, the cofactor lipid is selected from the group consisting of DSPC, DPPC, DMPC, DOPC, POPC, DOPE, DOPG, DPPG, POPE, DPPE, DMPE, DSPE, DOTAP, PS, and SM. In some embodiments, the cofactor lipid is DSPC. In some embodiments, the cofactor lipid is DOTAP. In some embodiments, the cofactor lipid is DOPE. In some embodiments, the cofactor lipid is PS.

[0408] In some embodiments, the sterol is cholesterol.

[0409] In some embodiments, the polymer-coupled lipid is a polyethylene glycol-modified lipid (PEG lipid). In some embodiments, the PEG lipid is selected from polyethylene glycol-modified diacylglycerols (PEG-DAG), such as 1-(monomethoxy-polyethylene glycol)-2,3-dimyristoylglycerol (PEG-DMG) (e.g., 1,2-dimyristoyl-rac-glycerol-3-methoxypolyethylene glycol-2000 (PEG2000-DMG)); polyethylene glycol-modified phosphatidylethanolamine (PEG-PE); PEG-S-DAG, such as 4-O-(2',3'-di(tetradecanoyloxy)propyl-1-O-(ω-methoxy(polyethoxy)ethyl)succinate (PEG-S-DMG); polyethylene glycol-modified ceramides. (PEG-cer); or PEG dialkoxypropyl carbamate, such as ω-methoxy(polyethoxy)ethyl-N-(2,3-di(tetradecoxy)propyl)carbamate and 2,3-di(tetradecoxy)propyl-1-N-(ω-methoxy(polyethoxy)ethyl)carbamate. In some embodiments, the PEGylated lipid is 1,2-dimyristoyl-sn-glycerol-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (C14-PEG2000). In some embodiments, the PEGylated lipid has the following structure: Or its pharmaceutically acceptable salts, tautomers or stereoisomers, wherein: R 12 and R 13 Each is independently a straight-chain or branched, saturated or unsaturated alkyl chain containing 10 to 30 carbon atoms, wherein the alkyl chain is optionally interrupted by one or more ester bonds; and w has an average value in the range of 30 to 60. In some embodiments, R 12 and R 13 Each is independently a straight-chain saturated alkyl chain containing 12 to 16 carbon atoms. In some embodiments, w has an average value in the range of 40 to 55. In some embodiments, the average w is about 45. In some embodiments, R 12 and R 13 Each is an independent straight-chain saturated alkyl chain containing about 14 carbon atoms, and w has an average value of about 45.

[0410] In some embodiments, the polyethylene glycolated lipoprotein is DMG-PEG 2000, for example having the following structure: .

[0411] In some embodiments, the PEGylated lipid is or comprises 2-[(PEG)-2000]-N,N-bistetradecylacetamide having the following chemical structure: Or a pharmaceutically acceptable salt thereof, wherein n' is an integer from 45 to 50. In some embodiments, the cationic lipid component of LNP has the structure of formula (III): (III) Or its pharmaceutically acceptable salt, tautomer, prodrug, or stereoisomer, wherein: L 1 or L 2 One of is -O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O)x-, -SS-, -C(=O)S-, SC(=O)-, -NRaC(=O)-, -C(=O)NRa-, NRaC(=O)NRa-, -OC(=O)NRa-, or -NRaC(=O)O-, and L 1 or L 2 The other one is -O(C=O)-, -(C=O)O-, -C(=O)-, -O-, -S(O)x-, -SS-, -C(=O)S-, SC(=O)-, -NRaC(=O)NRa-, -C(=O)NRa-, NRaC(=O)NRa-, -OC(=O)NRa- or -NRaC(=O)O- or a direct bond; G 1 and G 2 Each independently constitutes unsubstituted C1-C 12 Alkylene or C1-C 12 alkenyl; G 3 For C1-C 24 Alkylene, C1-C 24 alkenyl, C3-C8 cycloalkylene, C3-C8 cycloalkylene; R a For H or C1-C 12 alkyl; R 1 and R 2 Each independently is C6-C 24 Alkyl or C6-C 24 alkenyl; R 3 For H, OR 5 CN, -C(=O)OR 4 -OC(=O)R 4 or -NR 5 C(=O)R 4 ; R 4For C1-C 12 alkyl; R 5 It is H or C1-C6 alkyl; and x is 0, 1, or 2.

[0412] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures (IIIA) or (IIIB): or in: A is a 3- to 8-membered cycloalkyl or cycloidene cycloalkyl ring; R 6 Each time it appears, it is independently H, OH, or Cl-C. 24 alkyl; n is an integer in the range of 1 to 15.

[0413] In some of the foregoing embodiments of formula (III), the lipid has structure (IIIA), and in other embodiments, the lipid has structure (IIIB).

[0414] In other embodiments of formula (III), the lipid has one of the following structures (IIIC) or (IIID): or Where y and z are each an independent integer in the range of 1 to 12.

[0415] In any of the foregoing embodiments of formula (III), L 1 or L 2 One of them is -O (C=O)-. For example, in some embodiments, L 1 and L 2 Each of these is -O (C=O)-. In some different embodiments shown in any of the foregoing, L 1 and L 2 Each is independently -(C=O)O- or -O(C=O)-. For example, in some embodiments, L 1 and L 2 The components in the equation are -(C=O)O-.

[0416] In some different embodiments of formula (III), the lipid has one of the following structures (IIIE) or (IIIF): or .

[0417] In some of the foregoing embodiments of formula (III), the lipid has one of the following structures: (IIIG), (IIIH), (IIII), or (IIIJ): ; ; or .

[0418] In some of the foregoing embodiments of formula (III), n is an integer in the range of 2 to 12, for example, 2 to 8 or 2 to 4. For example, in some embodiments, n is 3, 4, 5 or 6. In some embodiments, n is 3. In some embodiments, n is 4. In some embodiments, n is 5. In some embodiments, n is 6.

[0419] In some other of the foregoing embodiments of equation (III), y and z are each independently an integer in the range of 2 to 10. For example, in some embodiments, y and z are each independently an integer in the range of 4 to 9 or 4 to 6.

[0420] In some of the foregoing embodiments of formula (III), R 6 For H. In other aforementioned embodiments, R 6 For C1-C 24 Alkyl group. In other embodiments, R 6 It is OH.

[0421] In some embodiments of formula (III), G 3 No replacement. In other embodiments, G 3 Replaced. In various different embodiments, G 3 For straight chain C1-C 24 Alkylene or straight-chain C1-C 24 Alkenyl group.

[0422] In some other of the foregoing embodiments of equation (III), R 1 or R 2 Or both are C6-C 24 Alkenyl. For example, in some embodiments, R 1 and R 2 Each of them independently has the following structure: , in: R 7a and R 7b Each occurrence is independently H or Cl-C. 12 Alkyl; and a is an integer from 2 to 12. Where R 7a R 7bEach of a and a is chosen such that R 1 and R 2 Each contains 6 to 20 carbon atoms independently. For example, in some embodiments, a is an integer in the range of 5 to 9 or 8 to 12.

[0423] In some of the foregoing embodiments of equation (III), R appears at least once. 7a For example, in some embodiments, R is H. 7a H occurs each time it appears. In the other different embodiments described above, R occurs at least once. 7b It is a C1-C8 alkyl group. For example, in some embodiments, the C1-C8 alkyl group is methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, tributyl, n-hexyl, or n-octyl.

[0424] In different embodiments of formula (III), R 1 or R 2 Or both have one of the following structures: ; ; ; ; ; ; ; ; ; .

[0425] In some of the foregoing embodiments of formula (III), R 3 For OH, CN, -C(=O)OR 4 -OC(=O) R 4 or -NHC(=O)R 4 In some embodiments, R 4 It can be methyl or ethyl.

[0426] In various embodiments, the cationic lipid of formula (III) has one of the structures listed in Table 15 below.

[0427] Table 15: Exemplary cationic lipid structures of formula (III) In various embodiments, the cationic lipids have one of the structures listed in Table 16 below.

[0428] Table 16: Exemplary cationic lipid structures of formula AF In some embodiments, the LNP comprises cationic lipids as ionizable lipid-like materials (lipids). In some embodiments, the cationic lipids have one of the following structures as described in Melamed et al., Science Advances, 2023, 9, eade1444 (the entire contents of which are incorporated herein by reference): X-1 X-2 X-3 X-4.

[0429] In some embodiments, lipid nanoparticles may have an average size (e.g., average diameter) of about 30 nm to about 150 nm, about 40 nm to about 150 nm, about 50 nm to about 150 nm, about 60 nm to about 130 nm, about 70 nm to about 110 nm, about 70 nm to about 100 nm, about 70 nm to about 90 nm, or about 70 nm to about 80 nm. In some embodiments, lipid nanoparticles as described in this disclosure may have an average size (e.g., average diameter) of about 50 nm to about 100 nm. In some embodiments, lipid nanoparticles may have an average size (e.g., average diameter) of about 50 nm to about 150 nm. In some embodiments, lipid nanoparticles may have an average size (e.g., average diameter) of about 60 nm to about 120 nm. In some embodiments, the lipid nanoparticles as described in this disclosure may have an average size (e.g., average diameter) of about 30 nm, 35 nm, 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, or 150 nm. The term "average diameter" or "mean diameter" refers to the average hydrodynamic diameter of a particle, as measured by dynamic laser light scattering (DLS), where data analysis is performed using a so-called cumulant algorithm, the results of which provide a so-called Z-mean with a length dimension and a dimensionless polydispersity index (PI) (Koppel, Chem. Phys. 57, 1972, pp 4814-4820, ISO 13321, which is incorporated herein by reference). Here, the terms "average diameter," "mean diameter," "diameter," or "size" of a particle are used synonymously with the value of the Z-mean.

[0430] In some embodiments, the lipid nanoparticles described herein may exhibit a polydispersity index of less than about 0.5, less than about 0.4, less than about 0.3, or about 0.2 or smaller. For example, lipid nanoparticles may exhibit a polydispersity index in the range of about 0.1 to about 0.3 or about 0.2 to about 0.3. The “polydispersity index” is preferably calculated based on dynamic light scattering measurements using so-called cumulative analysis as described in the definition of “mean diameter.” Under certain prerequisites, it can be considered a measure of the size distribution of an aggregate of ribonucleic acid nanoparticles (e.g., ribonucleic acid nanoparticles).

[0431] The lipid nanoparticles described herein are characterized by an "N / P ratio," which is the molar ratio of cationic (nitrogen) groups ("N" in N / P) in the cationic polymer to anionic (phosphate) groups ("P" in N / P) in the RNA. It should be understood that cationic groups are groups in cationic form (e.g., N...). + ( ), or groups that can ionize to become cations. Using a single number in the N / P ratio (e.g., an N / P ratio of about 5) is intended to mean that the number is greater than 1; for example, an N / P ratio of about 5 is intended to mean 5:1. In some embodiments, the lipid nanoparticles described herein have an N / P ratio greater than or equal to 5. In some embodiments, the lipid nanoparticles described herein have an N / P ratio of about 5, 6, 7, 8, 9, or 10. In some embodiments, the N / P ratio of the lipid nanoparticles described herein is about 10 to about 50. In some embodiments, the N / P ratio of the lipid nanoparticles described herein is about 10 to about 70. In some embodiments, the N / P ratio of the lipid nanoparticles described herein is about 10 to about 120.

[0432] In some embodiments, the formulation described herein comprising lipid nanoparticles is formulated for intramuscular (im) or intravenous (iv) delivery, and the lipid nanoparticles comprise i) about 30 to about 50 mol% cationic lipids; ii) about 1 to about 5 mol% PEG-coupled lipids; iii) about 5 to about 15 mol% accessory lipids; and iv) about 30 to about 50 mol% steroids.

[0433] In some embodiments, the formulations described herein comprising lipid nanoparticles are formulated for intraperitoneal (ip) delivery, and the lipid nanoparticles comprise: i) about 30 mol% to about 50 mol% cationic lipids; ii) about 1 mol% to 5 mol% PEG-conjugated lipids; iii) about 30 mol% to about 50 mol% accessory lipids; and iv) about 20 mol% to about 40 mol% cholesterol. Without wishing to be limited by any theory, it is anticipated that intraperitoneal (ip) delivery of such formulations will result in enhanced delivery of the lipid nanoparticles to pancreatic β-cells and enhanced in vivo expression of the encoded incretin in pancreatic β-cells.

[0434] In some embodiments, the formulation for IP delivery comprises lipid nanoparticles, wherein the lipid nanoparticles comprise about 35 mol% cationic lipids; about 40 mol% cofactor lipids; about 22.5 mol% cholesterol; and about 2.5 mol% PEG-coupled lipids. In some embodiments, the lipid nanoparticles comprise about 35 mol% cationic lipids X-2, X-3, or X-4; about 40 mol% DOTAP, DOPE, or PS; about 22.5 mol% cholesterol; and about 2.5 mol% C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% cationic lipid X-2; about 40 mol% DOTAP; about 22.5 mol% cholesterol; and about 2.5 mol% C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000. In some embodiments, the lipid nanoparticl...

Claims

1. A composition comprising a polynucleotide encoding an incretin.

2. The composition of claim 1, wherein the incretin is a GLP1 receptor agonist.

3. The composition of claim 1, wherein the incretin is a GIP receptor agonist.

4. The composition of claim 1, wherein the incretin is a GLP1 / GIP dual receptor agonist.

5. The composition of claim 1, wherein the incretin is a GLP1 / GCG dual receptor agonist.

6. The composition of claim 1, wherein the incretin is a GLP1 / GIP / GCG triple receptor agonist.

7. The composition of claim 2, wherein the incretin agent comprises an incretin peptide having an amino acid sequence as shown in any one of SEQ ID NO: 5-7, 63-64, 69-70, and 74-75.

8. The composition of claim 3, wherein the incretin agent comprises an incretin peptide having an amino acid sequence as shown in any one of SEQ ID NO: 8-9, 62 and 72.

9. The composition of claim 2, wherein the incretin agent comprises an incretin peptide having an amino acid sequence as shown in SEQ ID NO:

11.

10. The composition of claim 4, wherein the incretin agent comprises an incretin peptide having an amino acid sequence as shown in any of SEQ ID NO: 12-14.

11. The composition of claim 6, wherein the incretin agent comprises an incretin peptide having an amino acid sequence as shown in SEQ ID NO:

15.

12. The composition of any one of claims 7 to 11, wherein the incretin peptide is optionally fused to a signal peptide via the N-terminus of the incretin peptide, optionally via a linker peptide.

13. The composition of claim 12, wherein the signal peptide has an amino acid sequence as shown in any one of SEQ ID NO: 16-39 and 65-67.

14. The composition of claim 12, wherein the signal peptide has an amino acid sequence as shown in any one of SEQ ID NO: 16-21 and 65-67.

15. The composition of claim 12, wherein the signal peptide has an amino acid sequence as shown in SEQ ID NO:

17.

16. The composition of claim 12, wherein the signal peptide has an amino acid sequence as shown in SEQ ID NO:

65.

17. The composition of claim 12, wherein the signal peptide has an amino acid sequence as shown in SEQ ID NO:

66.

18. The composition of claim 1, wherein the incretin comprises an amino acid sequence as shown in any one of SEQ ID NO: 41-45, 52-61 and 108-152.

19. The composition of any one of claims 1 to 18, wherein the incretin agent comprises an incretin peptide optionally fused to one or more additional incretin peptides via one or more linking peptides.

20. The composition of claim 19, wherein the one or more linker peptides comprise an amino acid sequence as shown in any one of SEQ ID NO: 1-5, 68 or 156.

21. The composition of claim 19 or 20, wherein the incretin agent comprises an incretin peptide fused to two or more incretin peptides.

22. The composition of any one of claims 19 to 21, wherein the incretin comprises at least one GLP1 receptor agonist and at least one GIP receptor agonist.

23. The composition of any one of claims 19 to 22, wherein the incretin comprises at least two GLP1 receptor agonists.

24. The composition of any one of claims 19 to 23, wherein the incretin comprises at least two GIP receptor agonists.

25. The composition of any one of claims 19 to 24, wherein the incretin contains one or more furin cleavage sites.

26. The composition of claim 25, wherein the one or more furin cleavage sites are located between adjacent incretin peptides.

27. The composition of claim 27 or 28, wherein the one or more furin cleavage sites comprise an amino acid sequence as shown in SEQ ID NO:

153.

28. The composition of any one of claims 19 to 27, wherein the incretin agent comprises one or more units, each of the one or more units comprising, from the N-terminus to the C-terminus: a GLP1 receptor agonist-linking peptide-furin protease cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 76, 77, 78, 79, 80, 81), two units (e.g., SEQ ID NO: 82); or four units (e.g., SEQ ID NO: 83).

29. The composition of any one of claims 19 to 30, wherein the incretin comprises an amino acid sequence as shown in any one of SEQ ID NO: 76-83, 94-97, 102-107.

30. The composition of any one of claims 1 to 29, wherein the incretin comprises a half-life extended portion.

31. The composition of claim 30, wherein the extended half-life portion comprises albumin (e.g., human serum albumin).

32. The composition of claim 31, wherein the human serum albumin comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO:

159.

33. The composition of claim 31 or 32, wherein the human serum albumin comprises an amino acid sequence as shown in SEQ ID NO:

159.

34. The composition of any one of claims 31 to 33, wherein the incretin comprises albumin (e.g., human serum albumin) fused to one or more units, each of the one or more units comprising, from the N-terminus to the C-terminus: (i) GLP1 receptor agonist-linking peptide (e.g., SEQ ID NO: 98); (ii) GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 100); (iii) GLP1 receptor agonist-linked peptide-furin protease cleavage site (e.g., SEQ ID NO: 102); or (iv) GLP1 receptor agonist-linking peptide-furin protease cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 104), two units (e.g., SEQ ID NO: 106), or four units (e.g., SEQ ID NO: 107).

35. The composition of any one of claims 31 to 34, wherein the incretin comprises an amino acid sequence as shown in any one of SEQ ID NO: 98, 100, 102, 104, 106, 107 or any combination thereof.

36. The composition of claim 30, wherein the extended half-life portion comprises an albumin-binding domain (ABD).

37. The composition of claim 36, wherein the ABD is derived from protein G of Streptococcus strain GI48 and / or protein PAB of Finegoldia magna, such as ABD035 and SA21.

38. The composition of claim 36, wherein the extended half-life portion comprises ABD that binds to domain II of human serum albumin without overlapping or interfering with binding to the FcRn binding site on albumin.

39. The composition of claim 36, wherein the half-life extension portion comprises ABDCon.

40. The composition of claim 36, wherein the extended half-life portion comprises a derivative derived from hyperthermophilic archaea ( hyperthermophilic archaeon Sulfur-oxidizing leaf fungi ( Sulfolobus solfataricus The albumin-binding domain (ABD) of bacterial proteins Sso7d, such as M11.12 and M18.2.

5.

41. The composition of claim 36, wherein the extended half-life portion comprises DARPin bound to albumin.

42. The composition of claim 36, wherein the ABD comprises an immunoglobulin domain or a fragment thereof that binds albumin.

43. The composition of claim 36 or 42, wherein the ABD comprises a fully human domain antibody (dAb) that binds to albumin, such as AlbudAb.

44. The composition of claims 36 or 42 to 43, wherein the ABD comprises albumin-binding Fab, such as dsFvCA645.

45. The composition of any one of claims 36 or 42 to 44, wherein the ABD comprises a heavy-chain-only (VHH) antibody, such as a nanobody, that binds to albumin.

46. ​​The composition of claim 45, wherein the VHH antibody comprises a VHH domain having complementarity-determining region (CDR) sequences HCDR1, HCDR2 and / or HCDR3 as shown in SEQ ID NO: 191 (GFTLDYYA), SEQ ID NO: 192 (IASSGGST) and / or SEQ ID NO: 193 (AAAVLECRTVVRGYDY), respectively.

47. The composition of claim 46, wherein the VHH antibody comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO:

154.

48. The composition of claim 46 or 47, wherein the VHH antibody comprises the amino acid sequence shown in SEQ ID NO:

154.

49. The composition of any one of claims 45 to 47, wherein the incretin comprises a VHH antibody that binds to an albumin fused to a unit, the unit comprising, from the N-terminus to the C-terminus: (i) GLP1-linked peptide (e.g., SEQ ID NO: 99); (ii) GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 101); or (iii) GLP1 receptor agonist-linked peptide-furin protease-GIP receptor agonist-linked peptide (e.g., SEQ ID NO: 103 or 105).

50. The composition of any one of claims 45 to 49, wherein the incretin comprises an amino acid sequence as shown in any one of SEQ ID NO: 99, 101, 103, 105.

51. The composition of claim 30, wherein the extended half-life portion does not contain an Fc domain, such as that derived from human IgG, optionally derived from human IgG1, IgG2, IgG3 or IgG4.

52. The composition of claim 30, wherein the extended half-life portion comprises an Fc domain, such as that derived from human IgG, optionally derived from human IgG1, IgG2, IgG3 or IgG4.

53. The composition of claim 52, wherein the human IgG is human IgG4.

54. The composition of claim 52 or 53, wherein the incretin comprises an IgG4Fc domain fused to a unit, the unit comprising, from the N-terminus to the C-terminus: (i) GLP1 receptor agonist-linking peptides (e.g., SEQ ID NO: 10, 89, 90, 91); (ii) GIP receptor agonist-linking peptides (e.g., SEQ ID NO: 92, 93); or (iii) GLP1 receptor agonist-linked peptide-furin protease-GIP receptor agonist-linked peptide (e.g., SEQ ID NO: 94, 95, 96, 97).

55. The composition of claim 53 or 54, wherein the IgG4 Fc domain comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO:

155.

56. The composition of claim 55, wherein the IgG4 Fc domain comprises the amino acid sequence shown in SEQ ID NO:

155.

57. The composition of any one of claims 53 to 56, wherein the incretin comprises an amino acid sequence as shown in any one of SEQ ID NO: 10 and 89-97.

58. The composition of any one of claims 52 to 57, wherein the Fc domain contains one or more mutations in one or both constant Fc domains that increase the half-life of the incretin and / or induce dimerization.

59. The composition of claim 58, wherein the one or more mutations comprise one or more mutations in the CH3 domain.

60. The composition of claim 58 or 59, wherein the one or more mutations inducing dimerization comprise: (i) Y349C, T366S, L368A and / or Y407V (according to EU designation); or (ii) S354C and / or T366W (according to EU number).

61. The composition of any one of claims 58 to 60, wherein the one or more mutations comprise Y349C, T366S, L368A and Y407V ("FcKIH-b", according to EU designation); or S354C and T366W ("FcKIH-a", according to EU designation).

62. The composition of any one of claims 58 to 61, wherein the incretin agent comprises a first polypeptide chain and a second polypeptide chain, wherein the first polypeptide chain comprises an incretin peptide fused to a first Fc domain, wherein the first Fc domain comprises mutants Y349C, T366S, L368A, and Y407V ("FcKIH-b", according to EU designation), and wherein the second polypeptide chain comprises an incretin peptide fused to a second Fc domain, wherein the second Fc domain comprises mutants S354C and T366W ("FcKIH-a", according to EU designation).

63. The composition of any one of claims 58 to 62, wherein one or more mutations that increase the half-life of the incretin agent comprise M428L and N434S ("LS", according to EU designation).

64. The composition of any one of claims 58 to 63, wherein the incretin comprises an Fc domain having an FcKIH-a mutation on a first polypeptide chain and an Fc domain having an FcKIH-b mutation on a second polypeptide chain, wherein each Fc domain on each polypeptide chain is independently fused to one or more units, the one or more units comprising, from the N-terminus to the C-terminus: (i) GLP1 receptor agonist-linked peptides (e.g., SEQ ID NO: 84, 85, 86, 87); or (ii) GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 88).

65. The composition of any one of claims 58 to 64, wherein the incretin comprises an amino acid sequence as shown in any one of SEQ ID NO: 84-88.

66. The composition of any one of claims 58 to 64, wherein the Fc domain comprises one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q).

67. The composition of claim 66, wherein one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234S, L235T, and G236R ("STR", according to EU designation).

68. The composition of claim 66, wherein one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234A and L235A ("LALA", according to EU designation).

69. The composition of claim 66, wherein the one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234A / L235A / P329G ("LALAPG", according to EU designation).

70. The composition of claim 30, wherein the extended half-life portion comprises albumin-bound VNAR.

71. The composition of claim 30, wherein the half-life extension portion comprises an XTEN sequence.

72. The composition of any one of claims 1 to 71, wherein the polynucleotide has a ribonucleic acid sequence having at least 90% identity with any one of SEQ ID NO: 177-185 and 224-256.

73. The composition of claim 72, wherein the polynucleotide has a ribonucleic acid sequence as shown in any one of SEQ ID No: 177-185 and 224-256.

74. The composition of any one of claims 1 to 73, wherein the polynucleotide comprises at least one non-coding sequence component that enhances RNA stability and / or translation efficiency.

75. The composition of claim 74, wherein the at least one non-coding sequence component comprises a 5' cap structure, a 5' UTR, a 3' UTR, and / or a polyA tail.

76. The composition of claim 75, wherein the polynucleotide comprises, in the 5' to 3' direction: a. 5' UTR; b. Signal peptide coding sequence; c. Encoding sequence of incretin peptide; d. 3' UTR; and e. polyA tail.

77. The composition of claim 75, wherein the polynucleotide comprises, in the 5' to 3' direction: (1) a. 5' UTR; b. Signal peptide coding sequence; c. Encoding sequence of incretin peptide; d. Linker peptide coding sequence; e. The coding sequence for the extended half-life; f. 3' UTR; and g. polyA tail; or (2) a. 5' UTR; b. Signal peptide coding sequence; c. The coding sequence for the extended half-life; d. Linker peptide coding sequence; e. The encoding sequence of incretin peptide; f. 3' UTR; and g. polyA tail.

78. The composition of any one of claims 1 to 77, wherein the incretin peptide is encoded by a coding sequence that has been codon-optimized and / or has an increased G / C content compared to the wild-type coding sequence, wherein the codon optimization and / or the increase in G / C content does not alter the sequence of the encoded amino acid sequence.

79. The composition of any one of claims 1 to 78, wherein the polynucleotide comprises at least one modified ribonucleotide.

80. The composition of claim 79, wherein the polynucleotide comprises a modified nucleoside replacing uridine.

81. The composition of claim 80, wherein the polynucleotide comprises a modified nucleoside replacing each uridine.

82. The composition of claim 81, wherein the modified nucleoside is selected from pseudouridine (ψ), N1-methyl-pseudouridine (m1ψ), and 5-methyl-uridine (m5U).

83. The composition according to any one of claims 80 to 83, wherein the modified nucleoside is N1-methyl-pseudouridine (m1ψ).

84. The composition of any one of claims 1 to 83, wherein the polynucleotide comprises a 5' cap structure.

85. The composition of any one of claims 1 to 84, wherein the polynucleotide comprises 5' UTR.

86. The composition of any one of claims 1 to 85, wherein the polynucleotide comprises 3' UTR.

87. The composition of any one of claims 1 to 86, wherein the polynucleotide comprises a polyA tail.

88. The composition of claim 87, wherein the polyA tail comprises at least 100 nucleotides.

89. The composition of any one of claims 1 to 88, wherein the polynucleotide is mRNA.

90. The composition of any one of claims 1 to 89, wherein the polynucleotide is formulated as a liquid, formulated as a solid, or a combination thereof.

91. The composition of any one of claims 1 to 90, wherein the polynucleotide is formulated for injection.

92. The composition of any one of claims 1 to 91, wherein the polynucleotide is formulated for intraperitoneal or intravenous administration.

93. The composition of any one of claims 1 to 92, wherein the polynucleotide is formulated or is intended to be formulated into lipid particles.

94. The composition of claim 93, wherein the polynucleotide is formulated or is intended to be formulated as lipid nanoparticles.

95. The composition of claim 94, wherein the polynucleotide is encapsulated within isolipid nanoparticles.

96. The composition of claim 94 or 95, wherein the isolipid nanoparticles are lipid nanoparticles targeting the pancreas and / or the intestine.

97. The composition of any one of claims 94 to 96, wherein the isolipid nanoparticles are cationic lipid nanoparticles.

98. The composition of claim 97, wherein the lipids forming the isolipid nanoparticles comprise: a. Polymer-coupled lipids; b. Cationic lipids; and c. Neutral lipids.

99. The composition of claim 98, wherein the polymer-coupled lipid is a PEG-coupled lipid.

100. The composition of claim 98 or 99, wherein the cationic lipid is an ionizable lipid-like material (lipid-like substance).

101. The composition of claim 100, wherein the cationic lipid has one of the following structures: X-1 X-2 X-3 X-4。 102. The composition of any one of claims 99 to 101, wherein the neutral lipid comprises an accessory lipid, such as 1,2-distearyl-sn-glycerol-3-phosphocholine (DPSC) and / or cholesterol.

103. The composition of any one of claims 99 to 102, wherein the cationic lipid is selected from cationic lipids X-2, X-3 or X-4, and the neutral lipid comprises accessory lipids such as DOTAP, DOPE or PS and cholesterol.

104. The composition of claim 103, wherein the polymer-coupled lipid is C14-PEG2000.

105. The composition of any one of claims 99 to 104, wherein the isolipid nanoparticles comprise: i) about 30 mol% to about 50 mol% of cationic lipids; ii) about 1 mol% to 5 mol% of PEG-coupled lipids; iii) about 30 mol% to about 50 mol% of auxiliary lipids; and iv) about 20 mol% to about 40 mol% of cholesterol.

106. The composition of any one of claims 99 to 104, wherein the isolipid nanoparticles comprise about 35 mol% cationic lipids; about 40 mol% auxiliary lipids; about 22.5 mol% cholesterol; and about 2.5 mol% PEG-coupled lipids.

107. The composition of claim 106, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipids X-2, X-3 or X-4, about 40 mol% of DOTAP, DOPE or PS, about 22.5 mol% of cholesterol, and about 2.5 mol% of C14-PEG2000.

108. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

109. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

110. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of DOTAP; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

111. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

112. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

113. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of DOPE; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

114. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-2; about 40 mol% of PS; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

115. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-3; about 40 mol% of PS; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

116. The composition of claim 107, wherein the isolipid nanoparticles comprise about 35 mol% of cationic lipid X-4; about 40 mol% of PS; about 22.5 mol% of cholesterol; and about 2.5 mol% of C14-PEG2000.

117. The composition of any one of claims 94 to 116, wherein the isolipid nanoparticles are formulated for intraperitoneal (ip) delivery.

118. The composition of any one of claims 94 to 117, wherein the isolipid nanoparticles have an average size of about 50-150 nm.

119. The composition of any one of claims 1 to 118, further comprising one or more pharmaceutically acceptable carriers, diluents and / or excipients.

120. The composition of claim 119, further comprising a cryoprotectant.

121. The composition of claim 120, wherein the cryoprotectant is sucrose.

122. The composition of any one of claims 119 to 121, further comprising an aqueous buffer solution.

123. The composition of claim 122, wherein the aqueous buffer solution comprises sodium ions.

124. A method of treating a disease state in an individual in need, comprising administering to the individual a therapeutically effective amount of the composition as described in any one of claims 1 to 123.

125. The method of claim 124, further comprising administering one or more DPP-4 inhibitors.

126. The method of claim 125, wherein the one or more DPP-4 inhibitors and the composition are administered simultaneously.

127. The method of claim 125, wherein the one or more DPP-4 inhibitors and the composition are administered sequentially.

128. The method of claim 127, wherein the one or more DPP-4 inhibitors are administered prior to the composition.

129. The method of claim 127, wherein the one or more DPP-4 inhibitors are administered after the composition.

130. The method of any one of claims 125 to 129, wherein the one or more DPP-4 inhibitors comprise sitagliptin, vildagliptin, saxagliptin, linagliptin, gemigliptin, anagliptin, teneligliptin, alogliptin, trelagliptin, omarigliptin, evogliptin, gosogliptin, dutogliptin, neogliptin, retagliptin, denagliptin, cofrogliptin, fotagliptin, prusogliptin, berberine, or any combination thereof.

131. The method of any one of claims 125 to 130, wherein the one or more DPP-4 inhibitors are administered orally.

132. The method of claim 124, wherein the disease state is obesity or an obesity-related condition.

133. The method of claim 132, wherein the obesity-related condition is prediabetes, type 2 diabetes (T2D), early type 1 diabetes (T1D), non-alcoholic fatty liver disease (NAFLD), non-alcoholic steatohepatitis (NASH), cardiovascular (CV) disease, kidney disease, or an increased risk of premature death.

134. The method of claim 133, wherein the cardiovascular (CV) disease includes major cardiovascular events (MACE), including CV death, nonfatal myocardial infarction, nonfatal stroke and / or heart failure with preserved ejection fraction (HFpEF).

135. The method of claim 132, wherein the method improves the individual's weight management.

136. The method of claim 132, wherein the method reduces the individual's weight gain or induces weight loss.

137. The method of claim 124, wherein the disease state is diabetes.

138. The method of claim 137, wherein the method improves the individual's blood glucose control.

139. The method of claim 137, wherein the method reduces the individual's HbA1c.

140. The method of claim 137, wherein the diabetes is prediabetes, type 2 diabetes (T2D), or early type 1 diabetes (T1D).

141. The method of claim 124, wherein the disease state is cardiovascular (CV) disease.

142. The method of claim 141, wherein the cardiovascular disease includes major cardiovascular events (MACE), including CV death, nonfatal myocardial infarction, nonfatal stroke and / or heart failure with preserved ejection fraction (HfpEF).

143. The method of claim 141, wherein the method improves the individual's blood pressure and / or blood lipids.

144. The method of claim 124, wherein the disease state is kidney disease.

145. The method of claim 124, wherein the disease state is non-alcoholic fatty liver disease (NAFLD).

146. The method of claim 124, wherein the disease state is non-alcoholic steatohepatitis (NASH) and optionally its sequelae liver fibrosis and cirrhosis.

147. The method of any one of claims 124 to 146, wherein administering the composition to the individual comprises administering one or more doses of the composition to the individual.

148. The method of claim 147, wherein one or more doses of the composition are administered to the individual daily, every other day, or weekly.

149. The method of claim 147, wherein one or more doses of the composition are administered to the individual at a frequency of less than once a week.

150. The method of claim 147, wherein one or more doses of the composition are administered to the individual every 2, 3 or 4 weeks.

151. The method of any one of claims 124 to 150, wherein the composition is administered by injection.

152. The method of claim 151, wherein the composition is administered subcutaneously, intravenously, intramuscularly, or intraperitoneally.

153. The method of claim 152, wherein the composition is administered intraperitoneally.

154. The method of any one of claims 124 to 150, wherein the composition is applied non-invasively (e.g., orally or nasally).

155. The method of any one of claims 124 to 154, wherein administration of the composition results in the expression of the incretin in the individual.

156. The method of any one of claims 124 to 155, wherein the composition is administered in a volume of less than 0.5 mL.

157. Use of a composition as described in any one of claims 1 to 123 for treating a disease state in an individual in need.

158. A method of producing an incretin, comprising administering to cells the composition of any one of claims 1 to 123, such that the cells express and secrete the incretin.

159. An incretin agent comprising an incretin peptide fused to a signal peptide.

160. The incretin agent of claim 159, wherein the incretin peptide is fused to the signal peptide via a linker peptide at its N-terminus.

161. The incretin agent of claim 159 or 160, wherein the signal peptide has an amino acid sequence as shown in any one of SEQ ID NO: 16-39 and 65-67.

162. The incretin agent of claim 161, wherein the signal peptide has an amino acid sequence as shown in any one of SEQ ID NO: 16-21 and 65-67.

163. The incretin agent of claim 161, wherein the signal peptide has an amino acid sequence as shown in SEQ ID NO:

17.

164. The incretin agent of claim 161, wherein the signal peptide has an amino acid sequence as shown in SEQ ID NO:

65.

165. The incretin agent of claim 161, wherein the signal peptide has an amino acid sequence as shown in SEQ ID NO:

66.

166. An incretin agent according to any one of claims 159 to 165, wherein the incretin agent comprises an incretin peptide fused to a signal peptide, comprising an amino acid sequence as shown in any one of SEQ ID NO: 41-45, 52-61, and 108-152.

167. An incretin agent of any one of claims 159 to 166, wherein the incretin agent comprises an incretin peptide optionally fused to one or more additional incretin peptides via one or more linking peptides.

168. The incretin agent of claim 167, wherein the one or more linker peptides comprise an amino acid sequence as shown in any one of SEQ ID NO: 1-5, 68 or 156.

169. The incretin agent of claim 167 or 168, wherein the incretin agent comprises an incretin peptide fused to two or more incretin peptides.

170. An incretin agent according to any one of claims 167 to 169, wherein the incretin agent comprises at least one GLP1 receptor agonist and at least one GIP receptor agonist.

171. An incretin agent according to any one of claims 167 to 170, wherein the incretin agent comprises at least two GLP1 receptor agonists.

172. An incretin agent according to any one of claims 167 to 171, wherein the incretin agent comprises at least two GIP receptor agonists.

173. The composition of any one of claims 167 to 172, wherein the incretin contains one or more furin cleavage sites.

174. The incretin agent of claim 173, wherein the one or more furin cleavage sites are located between adjacent incretin peptides.

175. The incretin agent of claim 173 or 174, wherein the one or more furin cleavage sites comprise an amino acid sequence as shown in SEQ ID NO:

153.

176. The incretin agent of any one of claims 167 to 175, wherein the incretin agent comprises one or more units, each of the one or more units comprising, from the N-terminus to the C-terminus: a GLP1 receptor agonist-linking peptide-furin protease cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 76, 77, 78, 79, 80, 81), two units (e.g., SEQ ID NO: 82); or four units (e.g., SEQ ID NO: 83).

177. The incretin agent of any one of claims 167 to 176, wherein the incretin agent comprises an amino acid sequence as shown in any one of SEQ ID NO: 76-83, 94-97, 102-107.

178. The incretin agent of any one of claims 159 to 177, wherein the incretin agent comprises a prolonged half-life portion.

179. The incretin agent of claim 178, wherein the extended half-life portion comprises albumin (e.g., human serum albumin).

180. The incretin agent of claim 180, wherein the human serum albumin comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO:

159.

181. The incretin agent of claim 179 or 180, wherein the human serum albumin comprises the amino acid sequence shown in SEQ ID NO:

159.

182. The incretin agent of any one of claims 172 to 174, wherein the incretin agent comprises albumin (e.g., human serum albumin) fused to one or more units, each of the one or more units comprising, from the N-terminus to the C-terminus: (i) GLP1 receptor agonist-linking peptide (e.g., SEQ ID NO: 98); (ii) GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 100); (iii) GLP1 receptor agonist-linked peptide-furin protease cleavage site (e.g., SEQ ID NO: 102); or (iv) GLP1 receptor agonist-linking peptide-furin protease cleavage site-GIP receptor agonist, for example, wherein the incretin agent comprises one unit (e.g., SEQ ID NO: 104), two units (e.g., SEQ ID NO: 106), or four units (e.g., SEQ ID NO: 107).

183. The incretin agent of any one of claims 179 to 182, wherein the incretin agent comprises an amino acid sequence as shown in any one of SEQ ID NO: 98, 100, 102, 104, 106, 107 or any combination thereof.

184. The incretin agent of claim 178, wherein the extended half-life portion comprises an albumin-binding domain (ABD).

185. The incretin agent of claim 184, wherein the ABD is derived from protein G of Streptococcus strain GI48 and / or protein PAB of *Gynostemma pentaphyllum*, such as ABD035 and SA21.

186. The incretin agent of claim 184, wherein the extended half-life portion comprises ABD, which binds to domain II of human serum albumin without overlapping or interfering with binding to the FcRn binding site on albumin.

187. The incretin agent of claim 184, wherein the extended half-life portion comprises ABDCon.

188. The incretin agent of claim 184, wherein the extended half-life portion comprises an ABD derived from bacterial proteins Sso7d, such as M11.12 and M18.2.5, from the hyperthermophilic archaea *Sulphozoa sulfophylla*.

189. The incretin agent of claim 178, wherein the extended half-life portion comprises DARPin bound to albumin.

190. The incretin agent of claim 184, wherein the ABD comprises an immunoglobulin domain or a fragment thereof that binds albumin.

191. The incretin agent of claim 184 or 190, wherein the ABD comprises a fully human domain antibody (dAb) that binds to albumin, such as AlbudAb.

192. The incretin agent of any one of claims 184 or 190 to 191, wherein the ABD comprises albumin-binding Fab, such as dsFv CA645.

193. The incretin agent of claims 184 or 190 to 192, wherein the ABD comprises a heavy-chain-only (VHH) antibody, such as a nanobody, that binds to albumin.

194. The incretin agent of claim 193, wherein the VHH antibody comprises a VHH domain having complementarity-determining region (CDR) sequences HCDR1, HCDR2 and / or HCDR3 as shown in SEQ ID NO: 191 (GFTLDYYA), SEQ ID NO: 192 (IASSGGST) and / or SEQ ID NO: 193 (AAAVLECRTVVRGYDY), respectively.

195. The incretin agent of claim 193 or 194, wherein the VHH antibody comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO:

154.

196. The incretin agent of claim 195, wherein the VHH antibody comprises the amino acid sequence shown in SEQ ID NO:

154.

197. The incretin agent of any one of claims 193 to 196, wherein the incretin agent comprises a VHH antibody that binds to an albumin fused to a unit, the unit comprising, from the N-terminus to the C-terminus: (i) GLP1-linked peptide (e.g., SEQ ID NO: 99); (ii) GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 101); or (iii) GLP1 receptor agonist-linked peptide-furin protease-GIP receptor agonist-linked peptide (e.g., SEQ ID NO: 103 or 105).

198. The incretin agent of any one of claims 193 to 197, wherein the incretin agent comprises an amino acid sequence as shown in any one of SEQ ID NO: 99, 101, 103, 105.

199. The incretin agent of claim 199, wherein the extended half-life portion does not contain an Fc domain, such as that derived from human IgG, optionally derived from human IgG1, IgG2, IgG3 or IgG4.

200. The incretin agent of claim 178, wherein the extended half-life portion comprises an Fc domain, such as that derived from human IgG, optionally derived from human IgG1, IgG2, IgG3 or IgG4.

201. The incretin agent of claim 200, wherein the human IgG is human IgG4.

202. The incretin agent of claim 200 or 201, wherein the incretin agent comprises an IgG4 Fc domain fused to a unit, the unit comprising, from the N-terminus to the C-terminus: (i) GLP1 receptor agonist-linking peptides (e.g., SEQ ID NO: 10, 89, 90, 91); (ii) GIP receptor agonist-linking peptides (e.g., SEQ ID NO: 92, 93); or (iii) GLP1 receptor agonist-linked peptide-furin protease-GIP receptor agonist-linked peptide (e.g., SEQ ID NO: 94, 95, 96, 97).

203. The incretin agent of any one of claims 200 to 202, wherein the IgG4 Fc domain comprises an amino acid sequence having at least 90%, 95%, or 99% identity with SEQ ID NO:

155.

204. The incretin agent of claim 203, wherein the IgG4 Fc domain comprises the amino acid sequence shown in SEQ ID NO:

155.

205. The incretin agent of any one of claims 200 to 204, wherein the incretin agent comprises an amino acid sequence as shown in any one of SEQ ID NO: 10, 89-97.

206. The incretin agent of any one of claims 200 to 205, wherein the Fc domain contains one or more mutations in one or both constant Fc domains that increase the half-life of the incretin agent and / or induce dimerization.

207. The incretin agent of claim 206, wherein the one or more mutations comprise one or more mutations in the CH3 domain.

208. The incretin agent of claim 206 or 207, wherein the one or more mutations inducing dimerization comprise: (i) Y349C, T366S, L368A and / or Y407V (according to EU designation); or (ii) S354C and / or T366W (according to EU number).

209. The incretin agent of any one of claims 206 to 208, wherein the one or more mutations comprise Y349C, T366S, L368A and Y407V ("FcKIH-b", according to EU designation); or S354C and T366W ("FcKIH-a", according to EU designation).

210. The incretin agent of any one of claims 206 to 209, wherein the incretin agent comprises a first polypeptide chain and a second polypeptide chain, wherein the first polypeptide chain comprises an incretin peptide fused to a first Fc domain, wherein the first Fc domain comprises mutants Y349C, T366S, L368A, and Y407V ("FcKIH-b", according to EU designation), and wherein the second polypeptide chain comprises an incretin peptide fused to a second Fc domain, wherein the second Fc domain comprises mutants S354C and T366W ("FcKIH-a", according to EU designation).

211. The incretin agent of any one of claims 206 to 210, wherein the one or more mutations that increase the half-life of the incretin agent comprise M428L and N434S ("LS", according to EU designation).

212. The incretin agent of any one of claims 206 to 211, wherein the incretin agent comprises an Fc domain having an FcKIH-a mutation on a first polypeptide chain and an Fc domain having an FcKIH-b mutation on a second polypeptide chain, wherein the Fc domain on each polypeptide chain is independently fused to one or more units, the one or more units comprising, from the N-terminus to the C-terminus: (i) GLP1 receptor agonist-linked peptides (e.g., SEQ ID NO: 84, 85, 86, 87); or (ii) GIP receptor agonist-linking peptide (e.g., SEQ ID NO: 88).

213. The incretin agent of any one of claims 206 to 212, wherein the incretin agent comprises an amino acid sequence as shown in any one of SEQ ID NO: 84-88.

214. The incretin agent of any one of claims 206 to 213, wherein the Fc domain comprises one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q).

215. The composition of claim 214, wherein one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234S, L235T, and G236R ("STR", according to EU designation).

216. The composition of claim 215, wherein one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234A and L235A ("LALA", according to EU designation).

217. The composition of claim 216, wherein the one or more mutations that eliminate the effector activity of the Fc domain (e.g., binding to the Fcγ receptor or C1q) comprise the following mutations: L234A / L235A / P329G ("LALAPG", according to EU designation).

218. The incretin agent of claim 178, wherein the extended half-life portion comprises albumin-bound VNAR.

219. The incretin agent of claim 178, wherein the extended half-life portion comprises an XTEN sequence.

220. An incretin agent comprising: husec signal peptide; Incretin peptides containing GLP1 incretin peptide or fragments or mutants thereof; The GLP1 incretin peptide described therein comprises an amino acid sequence with an A8G substitution mutation compared to the wild-type GLP1 amino acid sequence.

221. A polynucleotide encoding an incretin as described in claim 220.

222. An incretin agent comprising: husec signal peptide; Incretin peptides containing GIP incretin peptide or fragments or mutants thereof; The GIP incretin peptide comprises an amino acid sequence with an A2G substitution mutation compared to the wild-type GIP amino acid sequence.

223. A polynucleotide encoding an incretin as described in claim 222.

Citation Information

Patent Citations

  • Liposomal apparatus and manufacturing methods

    US20040142025A1

  • Flyback transformer wire attach method to printed circuit board

    US20050017054A1

  • Lipid encapsulated interfering RNA

    US20050064595A1

  • Systemic delivery of serum stable plasmid lipid particles for cancer therapy

    US20050118253A1

  • Polyethyleneglycol-modified lipid compounds and uses thereof

    US20050175682A1