Compositions and methods for optimizing expression of proteins in subcutaneous tissues
Optimized genetic cassettes with secretion signals enhance protein expression and secretion in subcutaneous adipocytes, addressing the limitations of viral vectors and immune responses in gene therapies.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- REMEDIUM BIO INC
- Filing Date
- 2026-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
Current gene therapies predominantly use viral vectors, which can cause liver damage and elicit immune responses, and are not optimized for non-viral delivery to subcutaneous adipose tissues for systemic protein expression.
Optimized genetic cassettes with natural and synthetic secretion signals for non-viral delivery to subcutaneous adipocytes, enhancing protein expression and secretion for systemic biodistribution.
Enables repeat dosing and optimal protein expression without immune response, leveraging adipocytes' secretory nature for systemic biodistribution.
Smart Images

Figure US2026012218_30072026_PF_FP_ABST
Abstract
Description
[0001] Attorney Docket No.: R0872.70006WO00
[0002] COMPOSITIONS AND METHODS FOR OPTIMIZING EXPRESSION OF PROTEINS IN SUBCUTANEOUS TISSUES
[0003] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit under 35 U. S. C. § 119(e) of U. S. Provisional Application No. 63 / 748,112, filed January 22, 2025, entitled “METHODS FOR OPTIMIZING EXPRESSION OF PROTEINS IN SUBCUTANEOUS TISSUES”, the entire disclosure of which is hereby incorporated by reference in its entirety.
[0004] FIELD
[0005] The present invention generally pertains, at least in part, to compositions and methods for optimizing gene therapies, such as for optimizing non- viral gene therapies, including non-viral gene therapies targeting the subcutaneous adipose tissues. More specifically, the present invention provides compositions and methods for optimizing expression of therapeutic proteins in the subcutaneous space, such as following administration of DNA using non- viral vectors.
[0006] BACKGROUND
[0007] Current generation gene therapies are delivered predominantly using viral vectors and administered locally to the target cells or systemically via the intravenous route of administration. Local administration works optimally when the target cell type is confined to a given, easily accessible site, such as for example the joint or intra-articular space. Intravenous administration generally results in the delivery of genetic cargo to predominantly the liver and can cause hepatitis or liver damage. Moreover, current generation gene therapies are delivered predominantly by viral vectors or virus-like particles produced by recombinant processes and containing potentially antigenic or immunogenic peptides and proteins.
[0008] SUMMARY
[0009] Subcutaneous adipose tissues, and specifically subcutaneous adipocytes are locally a non-function- and non-life-sustaining organ, tissue, or cell type. They can be an optimal target for gene therapy to provide de novo or augmented expression of proteins that can biodistribute systemically. Moreover, delivering DNA cargo to adipocytes using non-viral or non-antigenic vectors can enable repeat dosing and post-treatment dose augmentation without eliciting a neutralizing or memory immune response. Such an approach is especially promising sincemany biologies (including insulin, GLP-1 RAs, monoclonal antibodies, etc.) are currently administered subcutaneously and biodistribute optimally to target tissues. However, to enable optimal protein expression and export from adipocytes, a cell type that is secretory but may not normally produce and export a given therapeutic protein, genetic cassette parameters may need to be specifically optimized for each individual therapeutic protein of interest. The present invention provides compositions and methods for optimizing genetic cassettes, and specifically secretion signals in some embodiments, and can provide optimal production, expression and / or secretion of proteins in subcutaneous tissues, such as by subcutaneous adipocytes for systemic biodistribution.
[0010] In one aspect, a composition for optimized expression from DNA or RNA sequences delivered by non-viral vectors of therapeutic proteins in subcutaneous tissue is provided. The composition may comprise one or more or a set of natural secretion signals from secreted proteins expressed by human adipocytes or adipocytes of non-human primates, or one ore more or a set of secretion signals optimized from secretion signals of proteins expressed by human adipocytes or adipocytes of non-human primates in silico encoded in the DNA or RNA cassettes. Optionally, in one embodiment, the one or more or set of secretion signals are or were prioritized in silico, screened and prioritized in vitro, and / or screened and prioritized in vivo.
[0011] In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) are derived at least in part or comprise at least 4 consecutive and up to all consecutive amino acids of the following MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG, MALWMRLLPLLALLALWGPDPAAA, MDLWQLLLTLALAGSSDA, MDYLLMIFSLLFVACQG, MELWGAYLLLCLFSLLTQVTT, METPAQLLFLLLLWLPDTTG, MGPTSGPSLLLLLLTHLPLALG, MGSRGQGLLLAYCLLLAFASGLVLS, MHLLAILFCALWSAVLA, MHSSALLCCLVLLTGVRA, MHSWERLAVLVLLGAAACAA, MHWGTLCGFLWLWPYLFYVQA, MKALCLLLLPVLGLLVSS, MKASSLAFSLLSAAFYLLWTPSTG, MKCLLYLAFLFIGVNC, MKFVPCLLLVTLSCLGTLG, MKGLRSLAATTLALFLVFVFLGNSSC, MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA, MKLVSVALMYLGSLAFLGADT, MKSIYFVAGLFVMLVQGSWQ, MKSLPILLLLCVAVCSA, MKTLLLLAVIMIFGLLQAHG,MKWVTFISLLFLFSSAYSRGVFRRD, MKWVTFISLLFSSAYS, MLAATVLTLALLGNAHA, MLLLGAVLLLLALPGHDQ, MLLLLLL LLGLRL QLSLG, MNQLSFLLFLIATTRGWS, MNSFSTSAFGPVAFSLGLLLVLPAAFPAP, MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG, MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG, MSGAPTAGAALMLCAATAVLLSAQG, MTSKLAVALLAAFLISAAL C, MVFTPQILGLMLFWISASRG, MYSAPSACTCLCLHFLLLCFQVQVLVA, MHWGTLCGFLWLWPYLFYVQA.
[0012] In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are derived from secretion signal(s) obtained from proteins produced and secreted by adipocytes. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are derived from secretion signals obtained from proteins produced and secreted by adipocytes under a given disease state. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are appended to therapeutic proteins such that the first amino acid of the therapeutic protein is identical to the first post-secretion signal amino acid of the original secreted protein. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are appended to therapeutic proteins such that the first two amino acids of the therapeutic protein are sequentially identical to the first two post-secretion signal amino acid of the original secreted protein. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are appended to therapeutic proteins such that the first three amino acids of the therapeutic protein are sequentially identical to the first three post-secretion signal amino acid of the original secreted protein.
[0013] In one embodiment of any one of the compositions or methods provided herein, the in silico optimized secretion signal(s) include natural secretion signals in combination with synthetic secretion signals, or fully synthetic signals that are optimized completely in silico with or without initial derivation from natural sequences or motifs.
[0014] In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are from proteins produced and secreted by adipocytes, preadipocytes, or adipocyte progenitors. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are from a pool of known eukaryotic secretion signals. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are from a pool of known mammalian secretion signals.In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are are derived from secretion signals obtained at least in part from a pool of known human secretion signals. In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) include matched amino acids such that at least one of the three amino acids in the pre-cleavage site sequence and at least one of the three amino acids in the post-cleavage site sequence, preferentially the n=-l and n=0 position relative to the secretion signal cleavage site of 0, which indicates the future N-terminal of the protein.
[0015] In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are from the following: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), and spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC). In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) are from the following: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1(MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), and interleukin 10 (MHSSALLCCLVLLTGVRA).
[0016] In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) have an identically matching first post cleavage site amino acid of the therapeutic protein of interest or the minimal active sequence of the therapeutic protein of interest to the first post cleavage site of the protein that the secretion signal was obtained from.
[0017] In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) are synthetic, semi-synthetic, or natural, or a combination of synthetic and natural secretion signals.
[0018] In one embodiment of any one of the compositions or methods provided herein, the natural secretion signal(s) additionally include an identically matching n=l amino acid of the secretion signal to the secretion signal of the therapeutic protein of interest.
[0019] In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) are synthetic or composite synthetic-natural amino acid sequences.
[0020] In one embodiment of any one of the compositions or methods provided herein, the matching of amino acids, or identity of amino acids includes identical matches, or matches based on charge, hydrophilicity, hydrophobicity, pKa, or other physico-chemical parameter.
[0021] In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) include one or more parts of a natural secretion signal. In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) include a natural secretion signal followed by a stretch of modified amino acids derived from natural or in silico designed sequences and the therapeutic protein to be secreted. In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) are derived from secretion signals of predicted proteins or proteins or nucleic acids in bioinformatics databases. In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) are codon optimized.
[0022] In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) include at least a part from at least one of the following: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1(MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(l-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
[0023] In one embodiment of any one of the compositions or methods provided herein, the signal(s) contain motifs designed by modeling, simulation, artificial intelligence, regression, bioinformatics analysis, or combination thereof. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include sequence homology before and / or after the cleavage site, within up to the first 4 amino acids in each direction inclusive of the therapeutic protein to be secreted and the original protein that the parent signal is derived from. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include sequences or patterns of sequences common to adipocyte or preadipocyte secretion signals. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include patterns of nucleic acid or amino acid sequences used by adipocytes to enable secretion, partitioning, cleavage, or export of proteins. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include pairwisematched natural secretion signals with the secreted sequence of the protein of interest, or a functional fragment of the protein of interest, or a functional fragment or full protein of interest coupled with one or more spacer sequences and identifying sequence sets, which have a cleavage site and predictable n-, h-, and c-regions quantified as likely secreted and cleavable by one or more bioinformatics tools. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include a secretion signal sequence and one or more cleavage sites ahead of the therapeutic protein of interest. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include a part or a consensus sequence cleaved by an enzyme from a set that includes: serine protease 1 (PRSS1), transmembrane protease serine 2 (TMPRSS2), cathepsin D (CTSD), cathepsin E (CTSE), renin (REN), caspase 3 (CASP3), caspase 9 (CASP9), sentrin-specific protease 1 (SENP1), sentrin- specific protease 6 (SENP6), interstitial collagenase (MMP1), 72 kDa type IV collagenase (MMP2), stromelysin 1 (MMP3), or furin. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include a cleavage site from the following set: P-P-T-I-F-F-R-L, K-P-I-E-F-F-R-L, R-[S / K]-[R / S / K]-[R / K]-[]-[]-[]-[G / E], []-[]-[]-[E / F]-[V]-[]-[]-[], P-F-H-E-[E / V / K]-[V / I / Y]-[Y / H / G]-[S / N], D-E-V-D-[G / S]-[]-[]-[], []-[E / D]-[]-D-[]-[]-[]-[], Q-T-G-G-K-[]-E-[], P-Q-G-I-A-G-Q, D-[]-[]-D, []-R-[]-[K]-R-R-[].
[0024] In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) include a flexible amino acid linker between the secretion signal and the therapeutic protein. In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) include a flexible amino acid linker which contains one or more endopeptidase cleavage sites. In one embodiment of any one of the compositions or methods provided herein, the secretion signal(s) include a flexible linker region or spacer region from a set that includes: QPEEQKPFKYTTVTKRSRRIRPTHPA, GGGGS, SGSG, APSVAPEPDGC, AAAAA, PAAAA, GEAAEGPAAA, AAGVGGERSS, GGPSGAGAGDE, VRTHGTEESVNGPKA, DQKVRPNEENNKDADE, GVKDTD, EPVQNGCPESAMEMN.
[0025] In one embodiment of any one of the compositions or methods provided herein, the signal(s) include two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein ofinterest with one or more amino acid spacers based on the cleavage site score. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position and score. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal and a homology match to a secretion signal from a naturally adipocyte-secreted protein. In one embodiment of any one of the compositions or methods provided herein, the signal(s) include two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal, and a homology match to a secretion signal from a naturally adipocyte-secreted protein, and at least one amino acid homology or similarity match to the naturally secreted protein.
[0026] In one embodiment of any one of the compositions or methods provided herein, the in silico designed or optimized signal(s) include at least a portion of the following secretion signals: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA),pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(l-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
[0027] In one embodiment of any one of the compositions or methods provided herein, the signal(s) contain homology of the secretion signal sequence, and the first three amino acid sequence post cleavage site to the naturally secreted protein section signal sequence and the first three amino acid sequence of the naturally secreted protein.
[0028] In one embodiment of any one of the compositions or methods provided herein, the DNA cassette(s) are linear double stranded DNA, linear single stranded DNA, circular double stranded DNA, circular single stranded DNA with or without chemical modifications or substitutions of one or more bases. In one embodiment of any one of the compositions or methods provided herein, the DNA cassette(s) include at least one synthetic DNA containing at least one synthetic DNA encoding a therapeutic protein of interest and a secretion signal.
[0029] In one embodiment of any one of the compositions or methods provided herein, the RNA cassette(s) include at least one mRNA containing at least one mRNA encoding a therapeutic protein of interest and a secretion signal.
[0030] In one embodiment of any one of the compositions or methods provided herein, the DNA or RNA cassette(s) include synthetic, semi-synthetic, enzymatic, or recombinant DNA or RNA. In one embodiment of any one of the compositions or methods provided herein, the DNA or RNA cassette(s) include at least one DNA or RNA strand encoding at least a part of theamino acid sequence from the following proteins: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCEEEVTESCEGTEG), Clq and TNF related 1 (MGSRGQGEEEAYCEEEAFASGEVES), intelectin 1 (MNQESFEEFEIATTRGWS), leptin (MHWGTECGFEWEWPYEFYVQA), angiopoietin like 4 (MSGAPTAGAAEMECAATAVEESAQG), spexin hormone (MKGERSEAATTEAEFEVFVFEGNSSC), matrix metallopeptidase 3 (MKSEPIEEEECVAVCSA), growth hormone receptor (MDEWQEEETEAEAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTEEEEAVIMIFGEEQAHG), interleukin 20 (MKASSEAFSEESAAFYEEWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATESAAPSNPREERVAEEEEEEVAASRRAAG), endothelin 1 (MDYEEMIFSEEFVACQG), interleukin 6 (MNSFSTSAFGPVAFSEGEEEVEPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSERATIECEEAWAGEAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MEEEEEEEGEREQESEG), exendin-4 (MKIIEWECVFGEFEATEFPISWQMPVESGESSEDSASSESFA), immunoglobulin kappa light chain (METPAQEEFEEEEWEPDTTG), proglucagon (MKSIYFVAGEFVMEVQGSWQ), insulin (MAEWMREEPEEAEEAEWGPDPAAA), albumin(l-25) (MKWVTFISEEFEFSSAYSRGVFRRD), resistin (MKAECEEEEPVEGEEVSS), complement C3 (MGPTSGPSEEEEEETHEPEAEG), vesicular stomatitis virus G protein (MKCEEYEAFEFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPERTVPGAEGAWEEGGEWAWTECGECSEGAVG), or an amino acid sequence of a natural secretion signal.
[0031] In one aspect, the one or more or set of in vitro, in vivo, or ex vivo screened secretion signal(s) of any one of the compositions or methods provided herein is provided and, optionally, include the use of fluorescent or bioluminescent reporter proteins expressed in tandem or as a fusion with the therapeutic protein of interest to quantify expression.
[0032] In one aspect, any one of the compositions provided herein, or a substantially similar composition, is provided, where, optionally, the composition is designed with the purpose ofincreasing expression of therapeutic proteins from subcutaneous adipocytes treated with a non-viral vector for the delivery of DNA. In one embodiment of any one of the compositions provided herein, the composition or substantially similar composition is for the attainment of sufficient efficacy or the reduction of a therapeutic dose of a non-viral vector gene therapy targeting or predominantly targeting subcutaneous adipocytes or preadipocytes.
[0033] In one embodiment of any one of the compositions or methods provided herein, the DNA or RNA cassette(s) are comprised of natural chemical structures of DNA or RNA, hybrids, or chemically modified structures of DNA or RNA.
[0034] In one embodiment of any one of the compositions or methods provided herein, the DNA cassette(s) are plasmids, double stranded DNA, single stranded DNA, linear or circular DNA, and may be used in vitro or in vivo to select optimal secretion signals or cleavage sites.
[0035] In one embodiment of any one of the compositions or methods provided herein, the RNA cassette(s) are mRNA, circular RNA, linear RNA, double stranded RNA, single stranded RNA or a combination of more than one RNA type.
[0036] In one aspect, is a composition derived from any one of the methods provided herein, including any one of the methods described in the Examples or a method that is a combination of one or more parts thereof.
[0037] In one aspect, is an in vitro screening derived composition and / or secretion signal as provided in any one of the compositions or methods provided herein.
[0038] In one embodiment of any one of the compositions or methods provided herein, the composition or secretion signal is derived from or utilized in the process of introducing plasmids, mRNA, circRNA, self-amplifying RNA, nanoplasmids, DNA minicircles, double stranded linear DNA, single stranded linear or circular DNA, covalently closed single stranded DNA or other genetic cargo into primary human adipocytes via one or more means such as electroporation, metal particle bombardment, chemical transfection, viral transduction, or combination of one or more approaches. In one embodiment of any one of the compositions or methods provided herein, the composition or secretion signal is obtained via quantification of the secreted therapeutic protein of interest, a reporter protein such as a fluorescent or bioluminescent reporter, or a fusion protein between the therapeutic protein and a reporter protein.
[0039] In one embodiment of any one of the compositions or methods provided herein, the therapeutic protein is a natural protein, human protein, hybrid protein, fusion protein, or part of a natural protein with or without additional amino acid sequences, motifs, or domains.In one aspect is a method for optimizing expression of therapeutic proteins in the subcutaneous tissues from DNA or RNA sequences delivered by non-viral vectors. Such a methods may include selection of a set of natural secretion signals or secretion signals optimized in silico; optionally, in silico prioritization of secretion signals using signal prediction tools; production of DNA or RNA cassettes including the secretion signals or a subset thereof, and the therapeutic protein of interest; screening and prioritization of secretion signals in vitro using primary adipocytes or an adipocyte cell line, via quantification of secreted therapeutic protein; and optionally, confirmation of optimally performing secretion signals ex vivo or in vivo via quantification of secreted therapeutic protein.
[0040] In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals involves selecting from a set of natural secretion signals for proteins produced and secreted by adipocytes. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals involves selecting from a set of natural secretion signals for proteins produced and secreted by adipocytes under a given disease state. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals involves selecting from a set of natural secretion signals for proteins produced and secreted by adipocytes and matching of the post-secretion signal amino acid sequence of the therapeutic protein to be delivered to the natural protein’s sequence that follows the secretion signal cleavage site.
[0041] In one embodiment of any one of the compositions or methods provided herein, the selection from a set of in silico optimized secretion signals includes natural secretion signals in combination with synthetic secretion signals, or fully synthetic signals that are optimized completely in silico with or without initial derivation from natural sequences or motifs.
[0042] In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals involves selecting from a set of natural secretion signals for proteins produced and secreted by adipocytes and matching at least one of the first three amino acids of the post-secretion signal amino acid sequence of the therapeutic protein to be delivered to the natural protein’s sequence that follows the secretion signal cleavage site. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes selection of secretion signals from proteins produced and secreted by adipocytes, pre-adipocytes, or adipocyte progenitors. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes the selection of secretion signals from a pool of known eukaryotic secretion signals.In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes the selection of secretion signals from a pool of known mammalian secretion signals. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes the selection of secretion signals from a pool of known human secretion signals. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes matching of at least one of the three amino acids in the pre-cleavage site sequence and at least one of the three amino acids in the post-cleavage site sequence, preferentially the n=-l and n=0 position relative to the secretion signal cleavage site of 0, which indicates the future N-terminal of the protein.
[0043] In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes selection of secretion signals from the following pool of proteins that are secreted by adipocytes: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), and spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC). In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes selection from at least a portion (2 or greater) of secretion signals from the following pool of proteins that are secreted by adipocytes: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA),pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), and interleukin 10 (MHSSALLCCLVLLTGVRA).
[0044] In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes the matching of the first post cleavage site amino acid of the therapeutic protein of interest or the minimal active sequence of the therapeutic protein of interest to the first post cleavage site of the protein that the secretion signal was obtained from.
[0045] In one embodiment of any one of the compositions or methods provided herein, the selection of secretion signals are synthetic, semi-synthetic, or natural, or a combination of synthetic and natural secretion signals.
[0046] In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals additionally includes matching of the n=l amino acid of the secretion signal to the secretion signal of the therapeutic protein of interest.
[0047] In one embodiment of any one of the compositions or methods provided herein, the selection of secretion signals include selection of elements of synthetic sequences.
[0048] In one embodiment of any one of the compositions or methods provided herein, the matching of amino acids includes identical matches, or matches based on charge, hydrophilicity, hydrophobicity, pKa, or other physico-chemical parameter.
[0049] In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes a partial selection of the natural secretion signal. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes a selection of a natural secretion signal followed by modification of the secretion signal and or the therapeutic protein to be secreted. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes selection of secretion signals from predicted proteins in protein or nucleic acid databases. In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes selection at either a protein or nucleic acid level followed by codon optimization.In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes selection at least in part from at least a part of the following: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(l-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
[0050] In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes computer aided design of amino acid or nucleic acid sequences with the goal of optimizing at least a part of the amino acid sequence to promote increase production and / or secretion of a therapeutic protein by adipocytes or preadipocytes. In one embodiment of any one of the compositions or methods provided herein, the in silico design ofsignals includes the use of simulation, artificial intelligence, regression, bioinformatics analysis, or combination thereof. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes the identification of a set of amino acid sequences, sequence types, or sequence patterns in natural secretion signals utilized by adipocytes and their application to the development of an optimal secretion signal sequence for a given therapeutic protein of interest. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes the use of sequence homology before and / or after the cleavage site, within the first 4 amino acids in each direction. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes the use of sequence homology before and / or after the cleavage site, within the first 3 amino acids in each direction. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes the use of sequence homology before and / or after the cleavage site, within the first 2 amino acids in each direction. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes identification of sequences or patterns of sequences common to adipocyte or preadipocyte secretion signals or nucleic acids encoding secretion signals and the use of one or more said sequences or patterns to enable secretion of a given therapeutic protein. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes design and / or verification of sequences using one or more of the following tools: Phobius, PrediSi, SIGCLEAVE, ANTHEPROT, SOSUIsignal, SPD – Secreted Protein Database, SignalP, TargetP, DeepLoc, NetGPI, DeepSig, TSignal, UniProt Knowledgebase, NCBI, PubMed, signalpeptide.de, or tools with similar function or purpose. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes the use of bioinformatics tools to identify patters of nucleic acid or amino acid sequences used in adipocytes to enable secretion, partitioning, cleavage, or export of proteins and application of said sequences to the therapeutic protein of interest in whole or in part. In one embodiment of any one of the compositions or methods provided herein, the in silico design of signals includes pairwise matching of natural secretion signals with the secreted sequence of the protein of interest, or a functional fragment of the protein of interest, or a functional fragment or full protein of interest coupled with one or more spacer sequences and identifying sequence sets, which have a cleavage site and predictable n-, h-, and c-regions quantified as likely secreted and cleavable by one or more bioinformatics tools.In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes the design of a secretion signal sequence and the design of one or more cleavage site to ensure optimal processing and exposure of the appropriate N-terminal sequence. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes the design of a secretion signal sequence and the design of one or more cleavage site that includes the use of databases algorithms or bioinformatics tools designed to select or optimized cleavage site sequences. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes the design of a secretion signal sequence and the design of one or more cleavage sites that includes the use of one or more of the following tools: PeptideCutter, PoPS, SitePrediction, CASVM, Cascleave, Cascleave 2.0, Pripper, PROSPER, PROSPEROUS, iProt-Sub, Procleave, CAT3, LabCaS, ScreenCap3, Proteasix, MEROPS database, CutDB, or tools and databases with similar function or purpose. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes selection of a consensus sequence cleaved by an enzyme from a set that includes: serine protease 1 (PRSS1), transmembrane protease serine 2 (TMPRSS2), cathepsin D (CTSD), cathepsin E (CTSE), renin (REN), caspase 3 (CASP3), caspase 9 (CASP9), sentrin-specific protease 1 (SENP1), sentrin-specific protease 6 (SENP6), interstitial collagenase (MMP1), 72 kDa type IV collagenase (MMP2), stromelysin 1 (MMP3), or furin. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes selection of a cleavage site from the following pool: P-P-T-I-F-F-R-L, K-P-I-E-F-F-R-L, R-[S / K]-[R / S / K]-[R / K]-[]-[]-[]-[G / E], []-[]-[] -[L / F]-[V]-[]-[]-[], P-F-H-L-[L / V / K]-[V / I / Y]-[Y / H / G]-[S / N], D-E-V-D-[G / S] -[]-[]-[], []-[E / D]-[]-D-[]-[]-[]-[], Q-T-G-G-K-[]-E-[], P-Q-G-I-A-G-Q, D-[]-[]-D, []-R-[]-[K]-R-R-[]. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes the insertion of a flexible amino acid linker between the secretion signal and the therapeutic protein that is identifiable by bioinformatics or structural analysis techniques as flexible and or unstructured. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes the selection and insertion of a flexible amino acid linker between the secretion signal and the therapeutic protein that is identifiable by bioinformatics or structural analysis techniques as flexible and or unstructured when coupled with the secretion signal and the therapeutic protein of interest. In one embodiment of any one of the compositions or methods provided herein, theprocess of designing secretion signals includes the selection and insertion of a flexible amino acid linker which contains one or more endopeptidase cleavage sites that are specific to extracellular endopeptidases. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes the selection and insertion of a flexible amino acid linker which contains one or more endopeptidase cleavage sites that are specific to adipocyte intracellular or extracellular endopeptidases. In one embodiment of any one of the compositions or methods provided herein, the process of designing secretion signals includes at least in part selection of a flexible linker region or spacer region from a set that includes: QPELQKPFKYTTVTKRSRRIRPTHPA, GGGGS, SGSG, APSVAPEPDGC, AAAAA, PAAAA, GEAAEGPAAA, AAGVGGERSS, GGPSGAGAGDE, VRTHGTLESVNGPKA, DQKVRPNEENNKDADL, GVKDTD, LPVQNGCPESAMEMN.
[0051] In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals includes prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals includes prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site score. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals includes prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position and score. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals includes prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals includes prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequencehaving a secretion signal and a homology match to a secretion signal from a naturally adipocyte-secreted protein. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals includes prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal, and a homology match to a secretion signal from a naturally adipocyte- secreted protein, and at least one amino acid homology or similarity match to the naturally secreted protein.
[0052] In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals is from a set that includes at least one of the following peptides: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(l-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin(MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
[0053] In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals is based on homology of the secretion signal sequence, and the first three amino acid sequence post cleavage site to the naturally secreted protein section signal sequence and the first three amino acid sequence of the naturally secreted protein. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals is by using one or more of the parameters reported by one or more of the following tools: Phobius, PrediSi, SIGCLEAVE, ANTHEPROT, SOSUIsignal, SPD – Secreted Protein Database, SignalP, TargetP, DeepLoc, NetGPI, DeepSig, TSignal, or tools with similar function or purpose. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals relies at least in part on artificial intelligence. In one embodiment of any one of the compositions or methods provided herein, the in silico prioritization of signals relies at least in part on comparing amino acid sequences, motifs, homology, or similarity to natural sequences utilized by adipocytes.
[0054] In one embodiment of any one of the compositions or methods provided herein, the production of DNA cassettes includes production of at least one DNA plasmid library containing at least one DNA plasmid encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. In one embodiment of any one of the compositions or methods provided herein, the production of DNA cassettes includes production of at least one DNA nanoplasmid library containing at least one DNA nanoplasmid encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. In one embodiment of any one of the compositions or methods provided herein, the production of DNA cassettes includes production of at least one DNA minicircle library containing at least one DNA minicircle encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. In one embodiment of any one of the compositions or methods provided herein, the production of DNA cassettes includes production of at least one DNA cassette library containing at least one DNA cassette encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion.In one embodiment of any one of the compositions or methods provided herein, the production of DNA cassettes includes production of at least one DNA virus library containing at least one DNA virus encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. In one embodiment of any one of the compositions or methods provided herein, the production of DNA cassettes includes production of at least one synthetic DNA library containing at least one synthetic DNA encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion.
[0055] In one embodiment of any one of the compositions or methods provided herein, the production of RNA cassettes includes includes production of at least one mRNA library containing at least one mRNA encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. In one embodiment of any one of the compositions or methods provided herein, the production of RNA cassettes includes production of at least one RNA virus library containing at least one RNA virus encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion.
[0056] In one embodiment of any one of the compositions or methods provided herein, the production of DNA or RNA cassettes includes synthetic, semi- synthetic, enzymatic, or recombinant production of DNA or RNA.
[0057] In one embodiment of any one of the compositions or methods provided herein, the production of DNA or RNA cassettes includes the production of at least one DNA or RNA strand encoding at least a part of the amino acid sequence from the following proteins: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like
[0058]
[0059] (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA(MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(1-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG), or an amino acid sequence of a natural secretion signal.
[0060] In one embodiment of any one of the compositions or methods provided herein, the in vitro screening on primary adipocytes includes transfection of at least one mRNA or DNA into cells comprised of at least 10% primary adipocytes or preadipocytes and measurement of expression and / or secretion of the therapeutic protein of interest encoded in the said mRNA or DNA. In one embodiment of any one of the compositions or methods provided herein, the in vitro screening on primary adipocytes includes quantitative or semi-quantitative secretion efficiency comparison between at least one or more secretion signals and at least one or more therapeutic proteins of interest, which may be comprised of a protein, fusion protein, functional fragment of a protein or a protein, functional fragment of a protein, or fusion protein fused to at least one amino acid spacer. In one embodiment of any one of the compositions or methods provided herein, the in vitro screening on primary adipocytes includes one or more screens on a culture comprised of at least 10% primary adipocytes or preadipocytes using one or more of the following assays: fluorescent protein secretion assay, bioluminescent protein secretion assay, or enzyme-linked immunosorbent assay (ELISA) to measure a protein, fragment, or fusion thereof. In one embodiment of any one of the compositions or methods provided herein, the in vitro screening on primary adipocytes includes one or more screens on a culture comprised of an adipocyte cell line using one or more of the following assays: fluorescentprotein secretion assay, bioluminescent protein secretion assay, or enzyme-linked immunosorbent assay (ELISA) to measure a protein, fragment, or fusion thereof.
[0061] In one embodiment of any one of the compositions or methods provided herein, the in vitro screening using cell lines is derived from embryonic stem cells, induced pluripotent stem cells, immortalized from primary cells, or derived from tumors.
[0062] In one embodiment of any one of the compositions or methods provided herein, the in vitro screening using mesenchymal stem cell or adipocyte cell lines includes one or more of the following cell lines, either differentiated to adipocytes or not: Simpson-Golabi-Behmel syndrome cells (SGBS), hMADS (human multipotent adipose-derived stem cells), Chub-S7, LiSa-2, PAZ6, hTERT A41hBAT-SVF, hTERT A41hWAT-SVF, hAMSC76telo, 3T3-L1, 3T3-F442A, C3H / 10T1 / 2, Obl7, BFC-1, OP9, ICP1, ICP2, ISP4.
[0063] In one embodiment of any one of the compositions or methods provided herein, the ex vivo confirmation of performance includes screening for therapeutic protein secretion efficiency by comparing at least two secretion signals using tissue explants containing at least 20% subcutaneous adipose tissues. In one embodiment of any one of the compositions or methods provided herein, the ex vivo confirmation of performance includes screening for therapeutic protein secretion efficiency by comparing at least two secretion signals using spheroids containing at least 10% adipocytes by total cell count. In one embodiment of any one of the compositions or methods provided herein, the ex vivo confirmation of performance includes screening for therapeutic protein secretion efficiency by comparing at least two secretion signals using 3D culture, or lab-on-a-chip tissue culture with at least 5% total cell count comprising adipocytes or preadipocytes. In one embodiment of any one of the compositions or methods provided herein, the ex vivo confirmation of performance is performed in lieu of in vitro screening. In one embodiment of any one of the compositions or methods provided herein, the ex vivo confirmation of performance ranks secretion signals for a given therapeutic protein of interest, functional fragment of a therapeutic protein of interest, or a fusion of a functional fragment of a therapeutic protein of interest or a full length therapeutic protein of interest and a spacer sequence by quantifying secretion efficiency of the therapeutic protein by primary adipocytes or a mixture of cells comprising at least 10% primary adipocytes.
[0064] In one embodiment of any one of the compositions or methods provided herein, the in vivo confirmation of performance includes screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the saidprotein, fragment, or fusion levels in circulation. In one embodiment of any one of the compositions or methods provided herein, the in vivo confirmation of performance includes screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels in local adipose tissues at the site of injection. In one embodiment of any one of the compositions or methods provided herein, the in vivo confirmation of performance includes screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels via quantification of bioluminescence of a conjugated bioluminescent protein at the site of injection relative to the areas distal from the injection site. In one embodiment of any one of the compositions or methods provided herein, the in vivo confirmation of performance includes screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by biomolecular assay. In one embodiment of any one of the compositions or methods provided herein, the in vivo confirmation of performance includes screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels in blood, serum, plasma, or tissue by ELISA, mass spectrometry, untargeted proteomics, targeted proteomics, Western blot, or other related analytical or bioanalytical technique. In one embodiment of any one of the compositions or methods provided herein, the in vivo confirmation of performance includes screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by quantitative polymerase chain reaction, mass spectrometry, untargeted proteomics, targeted proteomics, or Western blot of the cells expressing the protein, fragment, or fusion.
[0065] In one embodiment of any one of the compositions or methods provided herein, the in vitro, in vivo, or ex vivo method includes the use of fluorescent or bioluminescent reporter proteins expressed in tandem or as a fusion with the therapeutic protein of interest to quantify expression.
[0066] In one embodiment of any one of the compositions or methods provided herein, the performing of the process or a substantially similar process is with the purpose of increasing expression of therapeutic proteins from subcutaneous adipocytes treated with a non-viral vector for the delivery of DNA. In one embodiment of any one of the compositions or methods provided herein, the performing of the process or a substantially similar process is for theattainment of sufficient efficacy or the reduction of a therapeutic dose of a non-viral vector gene therapy targeting or predominantly targeting subcutaneous adipocytes or preadipocytes.
[0067] In one embodiment of any one of the compositions or methods provided herein, the use of DNA cassette(s), which may be plasmids, double stranded DNA, single stranded DNA, linear or circular DNA, is for in vitro or in vivo selection of optimal secretion signals or cleavage sites.
[0068] In one embodiment of any one of the compositions or methods provided herein, the RNA cassette(s) are mRNA, circular RNA, linear RNA, double stranded RNA, single stranded RNA or a combination of more than one RNA type.
[0069] In one embodiment of any one of the compositions or methods provided herein, the in vitro screening includes the process of introducing plasmids, mRNA, circRNA, self-amplifying RNA, nanoplasmids, DNA minicircles, double stranded linear DNA, single stranded linear or circular DNA, covalently closed single stranded DNA or other genetic cargo into primary human adipocytes via one or more means such as electroporation, metal particle bombardment, chemical transfection, viral transduction, or combination of one or more approaches.
[0070] In one embodiment of any one of the compositions or methods provided herein, the in vitro or in vivo screening is performed via quantification of the secreted therapeutic protein of interest, a reporter protein such as a fluorescent or bioluminescent reporter, or a fusion protein between the therapeutic protein and a reporter protein.
[0071] In one aspect is a method for optimizing expression of a therapeutic protein in the subcutaneous tissues from DNA or RNA sequences delivered by non-viral vectors. Such a method may include selection of a set of natural secretion signals or secretion signals optimized in silico; optionally, in silico prioritization of secretion signals using signal prediction tools; production of DNA or RNA cassettes including the aforementioned secretion signals or a subset thereof, and the therapeutic protein of interest; screening and prioritization of secretion signals in vitro using primary adipocytes or an adipocyte cell line, via quantification of secreted therapeutic protein; and optionally confirmation of optimally performing secretion signals ex vivo or in vivo via quantification of secreted therapeutic protein.
[0072] In one aspect is any one of the methods provided herein, such as in the Examples, or a method that comprises any combination of one or more parts thereof.
[0073] In one aspect is a composition used in or produced by any one of the methods provided herein, such as in the Examples, or a methods that comprises any combination of one or more parts thereof.In one aspect is any one of the one or more or sets of signals provided herein.
[0074] In one aspect is a nucleic acid encoding any one of the one or more or sets of signals provided herein.
[0075] In one aspect is any one of the cassettes described herein. In one aspect is a DNA, RNA or DNA and RNA cassette that encodes any one of the one or more or sets of signals provided herein. In one aspect is a composition comprising any one of the foregoing.
[0076] In one embodiment, any one of the nucleic acids or cassettes provided herein may further enode a therapeutic peptide or protein. In one embodiment, any one of the nucleic acids or cassettes provided herein may further encode any one or more of the features as described herein for expression of the therapeutic peptide or protein and / or for peptide or protein expression in adipocytes.
[0077] Any one of the compositions provided herein may be used in any one of the methods provided herein. Any one of the methods or compositions provided herein may be for dose adjustment and / or may be included in a package insert for any one of the compositions provided herein for therapy.
[0078] BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 provides the flow diagram for Example 1.
[0079] FIG. 2 provides the flow diagram for Example 2.
[0080] FIG. 3 provides the flow diagram for Example 3.
[0081] FIG. 4 provides the flow diagram for Example 4.
[0082] FIG. 5 provides the flow diagram for Example 5.
[0083] FIG. 6 provides the flow diagram for Example 6.
[0084] FIG. 7 provides the flow diagram for Example 7.
[0085] DETAILED DESCRIPTION
[0086] The present invention generally discloses methods for optimizing secretion signals to enable efficient production and secretion of therapeutic proteins in subcutaneous adipocytes as well as the compositions that result from such methods as well as those that encode or comprise such secretion signals. The secretion signals may be encoded by RNA or DNA delivered by a viral or non- viral vector delivery system to subcutaneous adipocytes for local production and systemic biodistribution of therapeutic proteins. In some general methodological approaches, the adipocytes are targeted for DNA delivery via non- viral vectors and said DNA encodes atleast one therapeutic protein that is secreted by the adipocyte following protein expression. In the aforementioned compositions and methods, protein expression is relatively durable following the delivery of a DNA cassette, which is optionally optimized for transport to the nucleus and episomal formation. The expression from the DNA cassette may be transient as the DNA cassette may be a simple plasmid, which may be shut down after a short period of transcription. In yet other embodiments the delivered genetic cargo may be RNA or RNA and DNA, and may have transient or permanent expression. Regardless of the iterations, the optimization strategies for the secretion signal nevertheless may apply at least in some form to any one of the compositions or methods provided herein.
[0087] The present invention provides methods for optimizing expression of therapeutic proteins in the subcutaneous tissues from DNA or RNA sequences delivered by non-viral vectors and specifically within adipocytes, which may or may not serve as the target of the vector, but are at least preferentially transfected or transduced due to the local nature of the treatment, and related compositions. The methods provided within the present invention generally include an initial selection of a set of natural, synthetic, or semi-synthetic secretion signals, optional additional optimization in silico, optional prioritization of secretion signals using signal prediction tools, followed by production of DNA or RNA libraries encoding said secretion signals in cis with the therapeutic gene of interest, or in cis with the therapeutic gene of interest fused with a reporter protein or simply in cis with a reporter protein. The methods additionally may include screening and prioritization of secretion signals in vitro using specifically primary adipocyts or an adipocyte cell line of the species intended to be the target of the eventual therapy or a related species. Ranking of the secretion signal / therapeutic protein pairs is performed by quantifying the secreted therapeutic protein, reporter protein, or their fusions in the supernatant following a given time post transfection or transduction in vitro (such as from 1 hour to approximately 8 days). Finally, the prioritized secretion signal therapeutic protein sets or a subset thereof may be advanced for confirmation of functionality and prioritization ex vivo or in vivo via quantification of secreted therapeutic proteins in for example systemic biodistribution, blood, or plasma, and may be included in any one of the compositions provided herein.
[0088] Some more specific compositions and methods provided in the present invention include the selection of a set of specifically natural secretion signals from proteins expressed and secreted by adipocytes. This may involve the selection of said sequences after obtaining a set of genes expressed in adipocytes from a gene expression database, cross-referencing thisset with secretome databases or with published literature disclosing adipocyte secreted proteins and subsequently advancing this set for quantification of expression and secretion in vitro in primary human adipocytes or adipocyte cell lines. In an alternative method, the secretome databases or gene expression databases may be searched for adipocyte- secreted proteins that are upregulated in a given disease state and their secretion signals initially selected.
[0089] Alternatively, initial selection may be from a set of natural secretion signals for proteins produced and secreted by adipocytes and matching of the post-secretion signal amino acid sequence of the therapeutic protein to be delivered to the natural protein’s sequence that follows the secretion signal cleavage site. Such cleavage site matching can be performed in Signal P or other similar bioinformatics databases, alternatively or optionally cleavage site functionality can be verified in vitro by techniques such as mass spectroscopy, western blot, or HPLC on the secreted protein of interest in the supernatant.
[0090] High-throughput combinatorial approaches that involve the combination of natura secretion signal sequences with synthetic secretion signal sequences or derivation of fully synthetic signals from natural secretion signal motifs by identifying high-use consensus sequences in adipocytes and employing those motifs in the design of the final sequence, with eventual verification of said sequence in vitro using primary human adipocytes, may be used in an embodiment of any one of the compositions or methods provided herein.
[0091] In an embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals may be by using a selection from a set of natural secretion signals for proteins produced and secreted by adipocytes and matching at least one of the first three amino acids of the post-secretion signal amino acid sequence of the therapeutic protein to be delivered to the natural protein’s sequence that follows the secretion signal cleavage site. Other sequence components may be combined with the pre- and post-cleavage site matched with alternative secretion signal sequences, followed by optional confirmation in silico and testing in vitro using primary human adipocytes or adipocyte cell lines.
[0092] An initial step of the selection of natural secretion signals can include selection of secretion signals from proteins produced and secreted by adipocytes, pre-adipocytes, or adipocyte progenitors. Alternatively, an initial step can include the selection of secretion signals from a pool of known eukaryotic secretion signals. Alternatively, the selection of natural secretion signals can include the selection of secretion signals from a pool of known mammalian secretion signals or specifically human secretion signals or secretion signals in the target species, for which the therapy is eventually intended, or a related, or a model species.In one embodiment of any one of the compositions or methods provided herein, the selection of natural secretion signals includes matching of at least one of the three amino acids in the pre-cleavage site sequence and at least one of the three amino acids in the post-cleavage site sequence, preferentially the n=-l and n=0 position relative to the secretion signal cleavage site of 0, which indicates the future N-terminal of the protein. In some specific approaches, the selection of natural secretion signals step can include selection of secretion signals from the following pool of proteins that are secreted by adipocytes: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), and spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), or related secreted proteins, or secreted proteins upregulation of which is correlated following some disease state or biologic activity.
[0093] One embodiment of any one of the compositions or methods provided herein, involves the use of multiple secretion signals in tandem or at least one secretion signal with at least a portion of another, initial sets for the first step of the selection process can be selected from the following pool of proteins that are secreted by adipocytes (secretion signal shown in round brackets following the protein): adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3(MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), and interleukin 10 (MHSSALLCCLVLLTGVRA), or a related secreted protein, or a secreted protein the expression of which is upregulated or downregulated with one or more aforementioned proteins in a diseased or given biological state.
[0094] The selection of natural, synthetic, or semi- synthetic secretion signals can include the matching of the first post cleavage site amino acid of the therapeutic protein of interest or the minimal active sequence of the therapeutic protein of interest to the first post cleavage site of the protein that the secretion signal was obtained from. The signal sequences can be derived from a set that includes matching of the n=l amino acid of the secretion signal to the secretion signal of the therapeutic protein of interest. This approach involves fist identifying the therapeutic protein of interest, selecting the minimal N-terminal sequence that enables functionality and identifying the leading amino acids that can serve as the n=l position amino acid (first amino acid after the cleavage site). Matching of a secretion signal from a protein that has the same n=l amino acid may be an extra step undertaken in the method to improve processibility and enhance the process of sequence selection.
[0095] Optionally, adipocyte specific motifs in the N, H, or C region or across regions that can enhance processability in adipocytes may be identified, followed by the introduction of said motifs into sequences screened in vitro, and such may be a feature of any one of the compositions or methods provided herein. The secretion signal design can include matching of amino acids in given regions pre-cleavage position and post-cleavage position on the target protein or via the introduction of a spacer sequence. Up to 3 amino acids can be matched in each direction between the natural secretion signal of the therapeutic protein or peptide and the selected secretion signal for optimal processing in adipocytes, between the natural sequence of the therapeutic protein or peptide and the original protein or peptide from which a given secretion signal sequence is derived. The process of said matching during sequence design and selection can be based on direct matching (100% identity) or by matching charge, hydrophobicity, structure of the side chain, steric hinderance, size of the side chain, or electron density of the side chain, or the functional group of the side chain, pKa, hydrophilicity, or other physico-chemical properties of the amino acid, or one or more elements across one or more amino acids preceding or following the cleavage site.
[0096] One embodiment of any one of the compositions or methods provided herein, may involve the selection of natural secretion signals, which includes a partial selection of thenatural secretion signal or optionally a selection of a natural secretion signal followed by modification of the secretion signal and or the therapeutic protein to be secreted. Optionally the selection step can be performed from a set of secretion signals from predicted proteins in protein or nucleic acid databases. Following said selection and prior to production of DNA or RNA libraries for in vitro screening codon optimization may be performed to ensure high expression.
[0097] Selection of secretion signals can be performed from subsets obtained from databases, some such subset can include a list of at least the following proteins or related proteins in the adipocyte secretome that are related by pathway, disease, or metabolic state: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(1-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin(MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG),vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
[0098] Methods to select optimal peptides can involve in silico design of secretion signals, which include computer aided design of amino acid or nucleic acid sequences with the goal of optimizing at least a part of the amino acid sequence to promote increase production and / or secretion of a therapeutic protein by adipocytes or preadipocytes. Alternatively or optionally said approaches can include the use of simulation, artificial intelligence, regression, bioinformatics analysis, or combination thereof. In some approaches the in silico design methods can include the identification of a set of amino acid sequences, sequence types, or sequence patterns in natural secretion signals utilized by adipocytes and their application to the development of an optimal secretion signal sequence for a given therapeutic protein of interest. Alternative approaches to the in silico design step can include the use of sequence homology before and / or after the cleavage site, within the first 4 amino acids in each direction, or the use of sequence homology before and / or after the cleavage site, within the first 3 amino acids in each direction, or the use of sequence homology before and / or after the cleavage site, within the first 2 amino acids in each direction.
[0099] In silico design of signals can include identification of sequences or patterns of sequences common to adipocyte or preadipocyte secretion signals or nucleic acids encoding secretion signals and the use of one or more said sequences or patterns to enable secretion of a given therapeutic protein. Alternatively, in silico design of signals can include design and / or verification of sequences using one or more of the following tools: Phobius, PrediSi, SIGCLEAVE, ANTHEPROT, SOSUIsignal, SPD – Secreted Protein Database, SignalP, TargetP, DeepLoc, NetGPI, DeepSig, TSignal, UniProt Knowledgebase, NCBI, PubMed, signalpeptide.de, or tools with similar function or purpose.
[0100] Other approaches to in silico design of signals can include the use of bioinformatics tools to identify patters of nucleic acid or amino acid sequences used in adipocytes to enable secretion, partitioning, cleavage, or export of proteins and application of said sequences to the therapeutic protein of interest in whole or in part. And yet other in silico design approaches can include pairwise matching of natural secretion signals with the secreted sequence of the protein of interest, or a functional fragment of the protein of interest, or a functional fragment or full protein of interest coupled with one or more spacer sequences and identifying sequence sets, which have a cleavage site and predictable n-, h-, and c-regions quantified as likely secreted and cleavable by one or more bioinformatics tools. The process of designing secretion signalscan additionally include the design of a secretion signal sequence and the design of one or more cleavage site to ensure optimal processing and exposure of the appropriate N-terminal sequence such that the cleavage site first releases the secretion signal and thereafter inside or outside of the cell releases the functional protein by cleaving a spacer that precedes the N-terminal of the functional protein. Intracellular processing of the cleavage site may be performed by intracellular endopeptidases and extracellular cleavage may be performed by surface bound or extracellular endopeptidases.
[0101] The process of designing secretion signals can include the design of a secretion signal sequence and the design of one or more cleavage site that includes the use of databases algorithms or bioinformatics tools designed to select or optimized cleavage site sequences. Cleavage sites can be at the end of the secretion signal or also following a spacer sequence to enable optimal cleavage site to expose an active N-terminal of the therapeutic protein or peptide. The process of designing secretion signals can include the design of a secretion signal sequence and the design of one or more cleavage sites that includes the use of one or more of the following tools: PeptideCutter, PoPS, SitePrediction, CASVM, Cascleave, Cascleave 2.0, Pripper, PROSPER, PROSPEROUS, iProt-Sub, Procleave, CAT3, LabCaS, ScreenCap3, Proteasix, MEROPS database, CutDB, or tools and databases with similar function or purpose. Alternatively, the process can include selection of a consensus sequence cleaved by an enzyme from a set that includes: serine protease 1 (PRSS1), transmembrane protease serine 2 (TMPRSS2), cathepsin D (CTSD), cathepsin E (CTSE), renin (REN), caspase 3 (CASP3), caspase 9 (CASP9), sentrin-specific protease 1 (SENP1), sentrin-specific protease 6 (SENP6), interstitial collagenase (MMP1), 72 kDa type IV collagenase (MMP2), stromelysin 1 (MMP3), or furin or proteins related or co-expressed with one or more of the aforementioned proteins.
[0102] The process of designing secretion signals can include the addition of one or more cleavage sites from the following pool of motifs, where a slash indicates alternative amino acids and empty square brackets indicate wildcard amino acids: P-P-T-I-F-F-R-L, K-P-I-E-F-F-R-L, R-[S / K]-[R / S / K]-[R / K]-[]-[]-[]-[G / E], []-[]-[]-[L / F]-[V]-[]-[]-[], P-F-H-L-[L / V / K]-[V / I / Y]-[Y / H / G]-[S / N], D-E-V-D-[G / S]-[]-[]-[], []-[E / D]-[]-D-[]-[]-[]-[], Q-T-G-G-K-[]-E-[], P-Q-G-I-A-G-Q, D-[]-[]-D, []-R-[]-[K]-R-R-[],
[0103] The process of designing secretion signals can optionally include the insertion of a flexible amino acid linker between the secretion signal and the therapeutic protein that is identifiable by bioinformatics or structural analysis techniques as flexible and or unstructured. Said process can include the selection and insertion of a flexible amino acid linker between thesecretion signal and the therapeutic protein that is identifiable by bioinformatics or structural analysis techniques as flexible and or unstructured when coupled with the secretion signal and the therapeutic protein of interest. A specific process can include the selection and insertion of a flexible amino acid linker which contains one or more endopeptidase cleavage sites that are specific to extracellular endopeptidases. Other specific processes can include the selection and insertion of a flexible amino acid linker which contains one or more endopeptidase cleavage sites that are specific to adipocyte intracellular or extracellular endopeptidases. Yet other specific processes can include the selection of at least a fragment of a flexible amino acid sequence that is flexible or amorphous or generally does not form a rigid structure, or is derived from a set of: QPELQKPFKYTTVTKRSRRIRPTHPA, GGGGS, SGSG, APSVAPEPDGC, AAAAA, PAAAA, GEAAEGPAAA, AAGVGGERSS, GGPSGAGAGDE, VRTHGTLESVNGPKA, DQKVRPNEENNKDADL, GVKDTD, LPVQNGCPESAMEMN.
[0104] In addition to the design or selection of the secretion signal, the signal sequences can be ranked or prioritized. Said processes can include prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position.
[0105] The process of in silico prioritization of signals can include prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site score. Or alternatively, prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position and score. Optionally, the process of in silico prioritization of signals can include prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal. Alternatively the prioritization process can include prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal and a homology match to a secretion signal from a naturally adipocyte- secreted protein. And yet alternatively the in silicoprioritization process can include prioritization of two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal, and a homology match to a secretion signal from a naturally adipocyte- secreted protein, and at least one amino acid homology or similarity match to the naturally secreted protein. In some specific approaches, the in silico prioritization process can be performed from a set of at least one of the following peptides: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(1-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG),vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG). Otherapproaches to in silico prioritization of secretion signals can be based on homology of the secretion signal sequence, and the first three amino acid sequence post cleavage site to the naturally secreted protein section signal sequence and the first three amino acid sequence of the naturally secreted protein. Yet other approaches can use one or more of the parameters reported by one or more of the following tools: Phobius, PrediSi, SIGCLEAVE, ANTHEPROT, SOSUIsignal, SPD – Secreted Protein Database, SignalP, TargetP, DeepLoc, NetGPI, DeepSig, TSignal, or tools with similar function or purpose or rely at least in part on artificial intelligence. Prioritization can rely in sequence matching or at least in part on comparing amino acid sequences, motifs, homology, or similarity to natural sequences utilized by adipocytes.
[0106] Following design and optimization of secretion signal and protein pairings, a full set or a subset of said pairs can be advanced to in vitro screening. To enable in vitro screening the prioritized or all sequences must be advanced to production in library form, which may be DNA or RNA. Production of DNA cassettes can include production of at least one DNA plasmid library containing at least one DNA plasmid encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. Alternatively, the process of producing the DNA cassettes can include production of at least one DNA nanoplasmid library containing at least one DNA nanoplasmid encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. The step of producing DNA cassettes can include the production of at least one of DNA minicircle library containing at least one DNA minicircle encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion, alternatively to a DNA minicircle, a plasmid, nanoplasmid, doggy bone DNA, single stranded DNA, double stranded DNA, linear DNA, circular DNA, linear covalently closed DNA, linear single stranded DNA, linear double stranded DNA, circular single stranded DNA, or circular double stranded DNA with or without supercoiling may be used. The said DNA library, or an alternative RNA library, is eventually introduced at least one plasmid per well or per cell or per set of cells to measure efficiency of the protein secretion.
[0107] Production of DNA cassettes can include production of at least one DNA virus library containing at least one DNA virus encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. Alternatively, the DNA library can be comprised of synthetic DNA encoding atherapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion.
[0108] In other approaches, production of RNA libraries or RNA cassettes can be used in lieu or in conjunction with DNA libraries or DNA cassettes and include production of at least one mRNA library containing at least one mRNA encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. The RNA library may be comprised of at least one RNA virus encoding a therapeutic protein of interest and a secretion signal for subsequent introduction into adipocytes to measure efficiency of the protein secretion. The DNA or RNA cassettes used in the method disclosed in the present invention may includes synthetic, semi-synthetic, enzymatic, or recombinant production of DNA or RNA.
[0109] In some specific methods, the process of producing DNA or RNA cassettes may include the production of at least one DNA or RNA strand encoding at least a part of the amino acid sequence from the following proteins: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGEEEAYCEEEAFASGEVES), intelectin 1 (MNQESFEEFEIATTRGWS), leptin (MHWGTECGFEWEWPYEFYVQA), angiopoietin like 4 (MSGAPTAGAAEMECAATAVEESAQG), spexin hormone (MKGERSEAATTEAEFEVFVFEGNSSC), matrix metallopeptidase 3 (MKSEPIEEEECVAVCSA), growth hormone receptor (MDEWQEEETEAEAGSSDA), pentraxin 3 (MHEEAIEFCAEWSAVEA), phospholipase A2 group IIA (MKTEEEEAVIMIFGEEQAHG), interleukin 20 (MKASSEAFSEESAAFYEEWTPSTG), adrenomedullin (MKEVSVAEMYEGSEAFEGADT), C-X-C motif chemokine ligand 3 (MAHATESAAPSNPREERVAEEEEEEVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MEEEEEEEGEREQESEG), exendin-4 (MKIIEWECVFGEFEATEFPISWQMPVESGESSEDSASSESFA), immunoglobulin kappalight chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(1-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG), or an amino acid sequence of a natural secretion signal.
[0110] Following production of DNA or RNA libraries or at least two DNA or RNA sets, in vitro screening on primary adipocytes can be performed to select the top expression and top secreting constructs. The process of said in vitro screening can include transfection of at least one mRNA or DNA into cells comprised of at least 10% primary adipocytes or preadipocytes and measurement of expression and / or secretion of the therapeutic protein of interest encoded in the said mRNA or DNA. In other approaches, the process of in vitro screening can include quantitative or semi-quantitative secretion efficiency comparison between at least one or more secretion signals and at least one or more therapeutic proteins of interest, which may be comprised of a protein, fusion protein, functional fragment of a protein or a protein, functional fragment of a protein, or fusion protein fused to at least one amino acid spacer. In yet other process of in vitro screening the screening is performed on primary adipocytes and includes one or more screens on a culture comprised of at least 10% primary adipocytes or preadipocytes using one or more of the following assays: fluorescent protein secretion assay, bioluminescent protein secretion assay, or enzyme-linked immunosorbent assay (ELISA) to measure a protein, fragment, or fusion thereof. In yet other process of in vitro screening the screening is performed on an adipocyte cell line or cell lines and includes one or more screens on a culture comprised of an adipocyte cell line using one or more of the following assays: fluorescent protein secretion assay, bioluminescent protein secretion assay, or enzyme-linked immunosorbent assay (ELISA) to measure a protein, fragment, or fusion thereof. Said in vitro screening using cell lines may rely on cell lines derived from embryonic stem cells, induced pluripotent stem cells, immortalized from primary cells, or derived from tumors. Similarly, in vitro screening may utilize mesenchymal stem cell or adipocyte cell lines in and include one or more of the following cell lines, either differentiated to adipocytes or not: Simpson-Golabi-Behmel syndrome cells (SGBS), hMADS (human multipotent adipose-derived stem cells), Chub-S7, LiSa-2, PAZ6, hTERT A41hBAT-SVF, hTERT A41hWAT-SVF, hAMSC76telo, 3T3-L1, 3T3-F442A, C3H / 10T1 / 2, Ob17, BFC-1, OP9, ICP1, ICP2, ISP4, with the ultimate goal ofreplicating or substantially mimicking the production and secretion of proteins in human adipocytes in vivo. An alternative approach includes screening, prioritizing, ranking or confirming performance of constructs ex vivo and includes screening for therapeutic protein secretion efficiency by comparing at least two secretion signals using tissue explants containing at least 5% and ideally at least 20% subcutaneous adipose tissues. Methods of ex vivo confirmation, ranking, qualification, or prioritization can include screening for therapeutic protein secretion efficiency by comparing at least two secretion signals using spheroids containing at least 5% and ideally at least 10% adipocytes by total cell count. Other methods of ex vivo confirmation of performance can include screening for therapeutic protein secretion efficiency by comparing at least two secretion signals using 3D culture, or lab-on-a-chip tissue culture with at least 5% total cell count comprising adipocytes or preadipocytes. Yet other ex vivo confirmation of performance methods can be performed in parallel, in conjunction, or in place of an in vitro screen. Yet other ex vivo confirmation of performance may be utilized to rank or prioritize secretion signals for a given therapeutic protein of interest, functional fragment of a therapeutic protein of interest, or a fusion of a functional fragment of a therapeutic protein of interest or a full length therapeutic protein of interest and a spacer sequence by quantifying secretion efficiency of the therapeutic protein by primary adipocytes or a mixture of cells comprising at least 5% and preferably at least 10% primary adipocytes.
[0111] All in vitro or ex vivo methods may be quantitative and functional, to ensure that the protein is not only secreted in the maximum quantity and of optimal length, but also in a functional conformation to ensure eventual therapeutic activity. Some examples of functional assays include cell proliferation assays, signaling assays, receptor binding and activation assays, epitope binding assays, or other biological assays used in vitro, ex vivo, or in vivo.
[0112] The selected constructs or a subset of selected constructs may be advanced to in vivo prioritization, ranking, screening, or final confirmation of performance. Said in vivo tests in general can include screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels in circulation. In other approaches said in vivo tests can include screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels in local adipose tissues at the site of injection. Alternatively, said in vivo tests can include screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels via quantification of bioluminescenceof a conjugated bioluminescent protein at the site of injection relative to the areas distal from the injection site. Yet alternatively, in vivo tests can include screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by biomolecular assay. Finally, in vivo testing can include screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by measuring the said protein, fragment, or fusion levels in blood, serum, plasma, or tissue by ELISA, mass spectrometry, untargeted proteomics, targeted proteomics, Western blot, or other related analytical or bioanalytical technique. In vivo confirmation of performance or testing in general in the method of developing an optimal secretion signal - therapeutic protein or peptide pair can include screening of one or more combinations of secretion signal and therapeutic protein of interest, or fragment, fusion, thereof by quantitative polymerase chain reaction, mass spectrometry, untargeted proteomics, targeted proteomics, or Western blot of the cells expressing the protein, fragment, or fusion. Optionally methods can include the use of fluorescent or bioluminescent reporter proteins expressed in tandem or as a fusion with the therapeutic protein of interest to quantify expression. All aforementioned in vivo tests or methods can be quantitative, semi-quantitative, and optionally in conjunction to quantification functional measuring a given biological function or efficacy in a disease model.
[0113] Confirmation of expression in vitro, ex vivo, or in vivo can be performed by transducing or transfection adipocytes, adipocyte progenitors, or adipocyte cell lines in culture or in situ with viral or non- viral vectors, using lipid nanoparticle carriers, polymeric carriers, polymeric transfection, gold particle bombardment or other methods with or without targeting of the adipose tissues and adipocytes specifically and can deliver DNA or RNA in various forms.
[0114] DNA molecules or cassettes used in screening can include plasmids, double stranded DNA, single stranded DNA, linear or circular DNA, and are used in vitro or in vivo to select optimal secretion signals or cleavage sites. RNA molecules or cassettes used in screening or confirmation of expression and secretion can include mRNA, circular RNA, linear RNA, double stranded RNA, single stranded RNA or a combination of more than one RNA type.
[0115] An important step of the disclosed methods is the confirmation of expression and secretion in primary adipocytes, preadipocytes, adipocyte cell lines, or progenitors in vitro. Said in vitro screening method steps can include the process of introducing plasmids, mRNA, circRNA, self-amplifying RNA, nanoplasmids, DNA minicircles, double stranded linear DNA, single stranded linear or circular DNA, covalently closed single stranded DNA or other genetic cargo into primary human adipocytes via one or more means such as electroporation, metalparticle bombardment, chemical transfection, viral transduction, or combination of one or more approaches.
[0116] Other in vitro screening method steps can include quantification of the secreted therapeutic protein of interest, a reporter protein such as a fluorescent or bioluminescent reporter, or a fusion protein between the therapeutic protein and a reporter protein. The reporter protein can include any protein that produces a quantifiable measure that is linked to its concentration in vitro or specifically within the supernatant, and can be a protein, tag, motif, or element of a peptide, or a glycosylation, or other post-translational modification or a co-translational modification such as for example the incorporation of a specific amino acid, natural or synthetic, which may be quantified by an analytical technique following in vitro secretion.
[0117] Some specific methods within the scope of this invention are outlined in the examples of the current disclosure and serve as explenary approaches that may be effective on their own, in modified form, or when used in tandem or with elements of other steps or methods, combining one or more part of the methods outlided in the examples from methods or parts of the methods disclosed elsewhere within the invention.
[0118] The methods or method components described in this invention can optionally work in concert to enable any functionality of the overall method or approach or more elements of the method, but may also optionally have additional functionality such as methods or approaches intended to increase transcription, decrease immunogenicity, increase mRNA stability, or improve cytocompatibility. Methods or approaches to improve one or more elements described in the presence invention may be combined with methods or approaches to improve one or more other elements of the derived system.
[0119] It should be understood that within the scope of this invention the methods or derived sequences, genetic constructs, formulations, methods for confirming functionality in vitro, in silico, or in vivo may be varied by one skilled in the art, to the extent that the methods described herewithin perform the desired function and remain within the scope of the present invention. Various parts, components or characteristics of one or more methods described in the present invention may be used in combination, with or without modification by someone skilled in the art to achieve the desired functionality of the aforedescribed approach.
[0120] Moreover, all individual features and methods of use described herein, and each and every combination of two or more of such features and methods of use, are included within the scope of the present invention provided that these features and methods of use in such acombination are not mutually inconsistent. It is understood that certain portions or combinations of such portions can be varied by someone trained in the art while still achieving the main goal of the invention.
[0121] Finally, it is understood that the specific ranges, tools, and systems provided in the current invention are not restrictive and are for example purposes only, values outside of the specified ranges or alternative bioinformatics tools or systems may be used to achieve the goal of the invention without modification to the proposed mechanistic principles.
[0122] EXAMPLES
[0123] Example 1
[0124] One example of using said method applies the sequence of steps shown in FIG. 1 to identify an optimal initial set in silico. The full set is screened in vitro using nLuc reporter in the supernatant (quantification performed in the presence of F-furimazine on supernatant collected from 24- well plates transfected with plasmid library hits).
[0125] Following selection of top in silico hits, 6 combinations of secretion signal / therapeutic peptide are screened in vitro using primary human adipocytes. Confirmation of secretion into supernatant is optionally confirmed by targeted proteomics. The 3 highest expressing combinations are advanced into in vivo efficacy model verification in C57BL / 6 high-fat diet induced obesity and type 2 diabetes mice.
[0126] Example 2
[0127] Another example using a number of optional steps in the selection of optimal secretion signal / therapeutic protein pairs includes the elements of the selection process shown in FIG.
[0128] 2. The screen produces the following priority list prior to final rank-ordering:
[0129]
[0130] PSSGAP
[0131] PPS
[0132] Complex,
[0133] multiple
[0134] alpha / YGEGTF
[0135] MRKRA
[0136] beta, TSDYSI
[0137] PQSEMA
[0138] globlular AMDKIA
[0139] Angiotensinogen PAGVSL DRVYIH
[0140] head, QKAFVQ Bad
[0141] sinogen RATILCL PFHL 0.547 G-Y multiple WLIAGG
[0142] LAWAGL
[0143] peptides PSSGAP
[0144] AAG
[0145] cut after PPS
[0146] secretion
[0147] signal
[0148] YGEGTF
[0149] 253AA
[0150] TSDLSK
[0151] multiple MHSWE
[0152] W-d QMEEEA
[0153] beta RLAVLV PPRGRIL A-C C-A Adipsin hydrophoVRLFIE Ok
[0154] LLG i Bad strands AAA GGR Ok A-AA-Y bic WLMNT
[0155] and CAA
[0156] KRNRN
[0157] helicies
[0158] NIA
[0159] 99AA two YGEGTF
[0160] helicies TSDYSI
[0161] MTSKLA
[0162] small ALDKIA
[0163] VALLAA EGAVLP G-E and ILS globlular QKAFVQ
[0164] FLISAAL RSAK
[0165] Bad
[0166] head with WLIAGG
[0167] C PSSGAP short beta
[0168] sheet PPS
[0169] 212 AA YGEGTF
[0170] multiple MNSFST TSDYSI
[0171] alpha SAFGPV AMDKIA
[0172] VPPGED
[0173] 11.6 helecies, AFSLGL QKAFVQ Good P-A SKDV 0.9997
[0174] 2 LLVLPA Good WLIAGG
[0175] disulfide AFPAP PSSGAP
[0176] bridges PPS
[0177] Large,
[0178] YGEGTF
[0179] globular
[0180] TSDLSK
[0181] head,
[0182] MLLLGA QMEEEA
[0183] short ETTTQGdiponec VLLLLA VRLFIE Bad 0.9973 G-E tin secretion PGVL iiiiiii
[0184] LPGHDQ WLMNT
[0185] signal,
[0186] KRNRN
[0187] two long
[0188] NIA
[0189] helicies
[0190] IDS hgegtftsdl
[0191] MHWGT
[0192] bond, skqmeeea
[0193] LCGFLW VPIQKV Leptin 167AA, 4 vrlfiewlkn Good 0.9992 A-H LWPYLF QDDT Bad ggpssrhylntral a- YVQA
[0194] helix nlvtrqry
[0195] HGEGTF
[0196] MLLLLL IIPVEEE Alkaline P05187.2, Hydropho TSDVSS
[0197] LLGLRL 0.9996 483AA NPD YLEGQA Phospha I bic
[0198] QLSLG
[0199] asc AKEFIA
[0200]
[0201] WLVKG
[0202] RG
[0203] MKIILW LCVFGL HGEGTF
[0204] Ex4 + FLATLFP TSDVSS
[0205] furin P26349, ISWQMP HGEGTF YLEGQA
[0206] Positive Positive High
[0207] cleavage 87AA VESGLS TSDL AKEFIA Good 0.3848 site SEDSAS WLVKG
[0208] SESFA + RG
[0209] KRIKR HAEGTF METPAQ TSDVSS
[0210] A0A0C4 LLFLLL EIVLTQS YLEGQA
[0211] IgK Negative Positive 0.9996 DH25 LWLPDT PAT AKEFIA iiiiiii Hii TG WLVKG RG HGEGTF TSDLSK METPAQ QMEEEA
[0212] A0A0C4 LLFLLL EIVLTQS
[0213] IgK Negative VRLFIE Positive 0.9997 G-H DH25 LWLPDT PAT
[0214] WLKNG TG GPSSGA PPPSG HDEFER HAEGTF MKSIYF TSDVSS
[0215] I’rogluca P01275.3, VAGLFV RSLQDT
[0216] Positive YLEGQA Positive
[0217] gon 180AA MLVQGS EEKS Good iiiiii AKEFIA WQ WLVKG RG MALWM HAEGTF
[0218] Insulin (+ RLLPLL TSDVSS
[0219] furin P01308, ALLALW FVNQHL Hydropho YLEGQA
[0220] Positive
[0221] cleavage 110AA GPDPAA CGSH bic AKEFIA iiiiii BW
[0222] site) A + WLVKG
[0223] RGRR RG HAEGTF MKWVT TSDVSS FISLLFL
[0224] Albumin P02768, AHKSEV Hydropho YLEGQA
[0225] FSSAYS Positive 0.9995 ■ liiii 609AA AHRF bic AKEFIA
[0226] RGVFRR WLVKG D RG HAEGTF MKALC TSDVSS LLLLPV KTLCSM YLEGQA
[0227] Resistin 108AA Positive Positive Good Medium
[0228] LGLLVS EEAI AKEFIA 0.9996 iiiiii S WLVKG RG
[0229]
[0230] YGEGTF
[0231] TSDYSI MGPTSG ALDKIA PSLLLLL SPMYSII Uncharge Hydropho
[0232] 03 1663AA QKAFVQ
[0233] LTHLPL TP iiiiiii 0.9997
[0234] N d bic G-Y WLIAGG ALG PSSGAP PPS YGEGTF MYSAPS TSDYSI ACTCLC ALDKIA
[0235] 076093, EENVDF Hydropho
[0236] FGF18 LHFLLL Negative QKAFVQ iiiiii 0.9994
[0237] 207AA RIHV bic Iiiiii CFQVQV WLIAGG LVA PSSGAP PPS YPYDVP DYA+
[0238] IgK+HA MVFTPQ RGRR+H
[0239] +FCS+G A0A125T ILGLML DVVLTQ GEGTFT Hydropho
[0240] Negative i 0.9997 G-Y 904 FWISAS SPAT SDVSSY bic iiiii
[0241] 37) ARG RG LEGQAA
[0242] KEFIAW LVKGRG YPYDVP DYA+
[0243] Leptin+HA IDS
[0244] MHWGT RGRR+H
[0245] bond,
[0246] A+FCS+ LCGFLW VPIQKV GEGTFT Hydropho
[0247] 167AA, 4 Good
[0248] LWPYLF QDDT SDVSSY 0.9997 A-Y GLP1(7- central a- lllibiii bic
[0249] l|Mi YVQA LEGQAA
[0250] helix
[0251] KEFIAW LVKGRG YPYDVP
[0252] 253AA DYA+
[0253] Adipsin+ multiple MHSWE RGRR+H
[0254] HA+FCS beta RLAVLV PPRGRIL Rigid GEGTFT Hydropho
[0255] Ok 0.9996 A-Y +GLP1(7 strands LLGAAA GGR Hydropho SDVSSY bic
[0256] -37) ARG and CAA bic LEGQAA
[0257] helicies KEFIAW
[0258] LVKGRG
[0259] Large, YPYDVP
[0260] globular DYA+
[0261] Adiponec
[0262] head, RGRR+H
[0263] tin+HA+ MLLLGA
[0264] short ETTTQG GEGTFT Hydropho
[0265] FCS+GL VLLLLA
[0266] secretion PGVL SDVSSY 0.9993
[0267] 1’1(7-37) bic
[0268] LPGHDQ
[0269] signal, LEGQAA
[0270] ASG
[0271] two long KEFIAW
[0272] helicies LVKGRG
[0273] HAEGTF TSDVSS MKWVT
[0274] Albumin P02768, RGVFRR YLEGQA
[0275] FISLLFS Positive Positive 0.9993
[0276] <1-16) 609AA DAHK AKEFIA Good
[0277] SAYS WLVKG RG
[0278]
[0279] HAEGTF
[0280] MKCLLY TSDVSS
[0281] VSV-G P04884, KFTIVFP YLEGQA
[0282] LAFLFIG Positive Positive Good Low
[0283] 511AA HNQ AKEFIA QW iiiiii VNC WLVKG RG YPYDVP
[0284] 224 AA MPGRAP
[0285] DYA+
[0286] protein, LRTVPG
[0287] KPDRI+ RGRR+H
[0288] HA+ two ALGAW FCS APRPCQ Hydropho GEGTFT Hydropho
[0289] glycosylat LLGGLW Good
[0290] +GLPK7 APQQ bic SDVSSY bic iiiiii ion sites AWTLCG iiiiii.37 > ARG LEGQAA
[0291] several LCSLGA
[0292] KEFIAW
[0293] disulfides VG
[0294] LVKGRG
[0295]
[0296] The priority list is then translated into genetic sequences with codon optimization as follows:
[0297]
[0298] MHWGTL ATGCACTGGGGAACACTG CGFLWL TGCGGCTTCCTGTGGCTGT YGEGT WPYLFY GGCCCTACCTGTTCTACGT FTSDYS MHWGTL VQAYGE GCAGGCCTATGGCGAGGG IALDKI CGFLWL GTFTSDY 62.22 CACCTTCACCAGCGACTAC Leptin AQKAF i 60aa 180bp WPYLFY SIALDKIA AGCATCGCTCTGGATAAGA %
[0299] VQWLI VQA TCGCCCAGAAGGCCTTTGT QKAFVQ AGGPSS WLIAGGP GCAGTGGCTGATCGCCGG GAPPPS SSGAPPP CGGACCTTCTAGCGGCGC S CCCTCCACCTTCC ATGAGAAAGCGGGCCCCC MRKRAP CAGAGCGAGATGGCCCCC QSEMAPA GCCGGAGTGTCCCTGAGA YGEGT GVSLRAT MRKRAP GCTACAATCCTGTGCCTGC FTSDYS ILCLLAW QSEMAP TGGCCTGGGCCGGCCTGG IAMDKI AGLAAG
[0300] Angiotensin AGVSLR 65.74 CTGCTGGCTACGGCGAAG
[0301] AQKAF YGEGTFT 72aa 216bp $ ATILCLL GCACCTTCACCAGCGACT % ogen VQWLI SDYSIAM
[0302] AWAGLA ACAGCATCGCCATGGATAA DKIAQKA AGGPSS GATCGCCCAGAAAGCCTT AG FVQWLIA GAPPPS CGTGCAGTGGCTGATCGCC GGPSSGA GGCGGACCTAGCAGCGGC PPPS GCCCCTCCTCCATCT
[0303]
[0304] ATGCACTCTTGGGAGAGA MHSWER CTGGCCGTGCTGGTGCTGC YGEGT LAVLVLL TGGGCGCCGCTGCTTGCG FTSDLS GAAACA MHSWER CCGCCTACGGCGAGGGAA KQMEE AYGEGTF LAVLVLL CATTCACCAGCGACCTGA 59.06 Adipsin EAVRLF TSDLSKQ 57aa 171bp GAAACA GCAAGCAGATGGAAGAGG %
[0305] IEWLM MEEEAVR A AAGCCGTGCGGCTGTTCAT NTKRN LFIEWLM CGAGTGGCTGATGAACAC RNNIA NTKRNR CAAGAGAAATAGAAACAA NNIA CATCGCC ATGACAAGCAAGCTGGCC MTSKLAV GTGGCCCTGCTGGCTGCTT YGEGT ALLAAFL TTCTGATCAGCGCCGCCCT FTSDYS ISAALCY MTSKLA GTGCTACGGCGAGGGCAC IALDKI GEGTFTS VALLAA CTTCACCAGCGACTACAG 64.41 ILS AQKAF DYSIALD 59aa 177bp FLISAAL CATCGCCCTGGATAAGATC %
[0306] VQWLI KIAQKAF C GCCCAGAAAGCCTTCGTG AGGPSS VQWLIA CAGTGGCTGATCGCCGGA GAPPPS GGPSSGA GGCCCCTCCAGCGGCGCC PPPS CCTCCACCTTCT ATGAACAGCTTCAGCACAT MNSFSTS CCGCCTTCGGCCCCGTGGC AFGPVAF YGEGT CTTTAGCCTGGGCCTGCTG SLGLLLV MNSFSTS FTSDYS CTGGTGCTGCCTGCTGCTT LPAAFPA AFGPVAF IAMDKI TCCCCGCCCCTTACGGCGA PYGEGTF 64.22 IL6 SLGLLLV AQKAF 68aa GGGCACCTTCACCTCTGAT 204bp TSDYSIA % LPAAFPA VQWLI TACAGCATCGCCATGGACA MDKIAQ P AGGPSS AGATCGCCCAGAAGGCCT KAFVQW GAPPPS TCGTGCAGTGGCTGATCGC LIAGGPS CGGCGGACCTAGCAGCGG SGAPPPS CGCCCCTCCTCCATCT ATGCTGCTGCTCGGCGCTG MLLLGAV YGEGT TGCTGCTGCTGCTGGCCCT LLLLALP FTSDLS GCCTGGCCACGATCAGTAC GHDQYG MLLLGA KQMEE GGCGAGGGAACATTCACC EGTFTSD 57.58 Adiponectin VLLLLAL EAVRLF 55aa AGCGACCTGTCTAAGCAG 165bp LSKQMEE % PGHDQ IEWLM ATGGAAGAGGAAGCCGTG EAVRLFIE NTKRN CGGCTGTTCATCGAGTGGC WLMNTK RNNIA TGATGAACACCAAGAGAA RNRNNIA ATAGAAACAACATCGCC ATGCACTGGGGCACCCTGT GCGGCTTCCTGTGGCTGTG MHWGTL GCCCTATCTGTTCTACGTG
[0307] hgegtftsd CGFLWL
[0308] CAGGCCCACGGCGAGGGC MHWGTL Iskqmeee WPYLFY
[0309] ACATTCACCAGCGACCTGT CGFLWL avrlfiewl VQAhgegtf 59.09 Leptin 66aa CCAAGCAGATGGAAGAGG 198bp WPYLFY knggpssr tsdlskqmee %
[0310] AAGCCGTGCGGCTGTTTAT VQA hylnlvtrq eavrlfiewlk
[0311] CGAGTGGCTCAAGAACGG
[0312] ry nggpssrhyl
[0313] CGGACCTTCTAGCAGACA
[0314] nlvtrqry
[0315] CTACCTGAACCTGGTGACC AGACAGAGATAC
[0316]
[0317] ATGCTGCTGCTGCTGCTGC MLLLLLL HGEGT TCCTGGGACTGCGGCTGC LGLRLQL
[0318] Secreated FTSDVS AGCTGTCTCTGGGCCACG
[0319] MLLLLL SLGHGEG
[0320] Alkaline SYLEG GCGAGGGAACATTCACCA 64.58
[0321] LLGLRL TFTSDVS 48aa 144bp I’hosplialiis QAAKE GCGACGTGTCCAGCTACCT %
[0322] QLSLG SYLEGQA
[0323] e FIAWLV GGAAGGCCAGGCCGCTAA AKEFIAW KGRG GGAGTTCATCGCCTGGCTG LVKGRG GTGAAGGGCAGAGGC ATGAAGATCATCCTGTGGC MKIILWL TGTGCGTGTTCGGCCTCTT CVFGLFL TCTGGCCACCCTGTTCCCC MKIILWL ATLFPIS ATCAGCTGGCAGATGCCTG CVFGLFL HGEGT WQMPVE TGGAAAGCGGCCTGAGCT ATLFPIS FTSDVS SGLSSED
[0324] Ex4 + furin CTGAGGATAGCGCCTCTAG WQMPVE SYLEG SASSESFA 59.40 cleavage 78aa CGAGAGCTTCGCCAAGAG 234bp SGLSSED QAAKE KRIKRHG %
[0325] site AATCAAGCGGCACGGCGA
[0326] SASSESF FIAWLV EGTFTSD GGGAACATTCACCAGCGA
[0327] A + KGRG VSSYLEG
[0328] CGTGTCCTCCTACCTGGAA KRIKR QAAKEFI GGCCAGGCCGCTAAAGAG AWLVKG TTCATCGCCTGGCTGGTCA RG AGGGCAGAGGA ATGGAAACACCTGCTCAG METPAQL CTGCTGTTTCTGCTGCTGC HAEGT LFLLLLW TGTGGCTCCCCGATACCAC METPAQ FTSDVS LPDTTGH CGGCCACGCCGAGGGAAC LLFLLLL SYLEG AEGTFTS 62.09 IgK 51aa ATTCACCAGCGACGTGTCT 153bp WLPDTT QAAKE DVSSYLE %
[0329] AGCTACCTGGAAGGCCAG G FIAWLV GQAAKE GCCGCCAAGGAGTTCATC KGRG FIAWLVK GCCTGGCTGGTGAAGGGC GRG AGAGGC ATGGAGACACCCGCTCAG METPAQL HGEGT CTGCTGTTTCTGCTGCTGC LFLLLLW FTSDLS TGTGGCTCCCCGATACAAC LPDTTGH METPAQ KQMEE CGGCCACGGCGAGGGCAC GEGTFTS LLFLLLL EAVRLF CTTCACCAGCGACCTGAG 62.78 IgK DLSKQM 60aa 180bp WLPDTT IEWLK CAAGCAGATGGAAGAGGA %
[0330] EEEAVRL G NGGPSS AGCCGTGAGACTGTTCATC FIEWLKN GAPPPS GAGTGGCTGAAGAACGGC GGPSSGA G GGACCTTCCAGCGGCGCC PPPSG CCTCCTCCATCTGGC ATGAAAAGCATCTACTTCG MKSIYFV TGGCCGGCCTGTTCGTGAT HDEFER AGLFVM GCTGGTGCAGGGCTCTTG HAEGT LVQGSW MKSIYFV GCAGCACGATGAGTTCGA FTSDVS QHDEFER
[0331] l’r< tglucago AGLFVM GCGGCACGCCGAGGGAAC 59.65
[0332] SYLEG HAEGTFT 57aa 171bp n LVQGSW ATTCACCAGCGACGTGTCC %
[0333] QAAKE SDVSSYL Q AGCTACCTGGAAGGCCAG FIAWLV EGQAAK GCCGCTAAGGAATTTATCG KGRG EFIAWLV CCTGGCTGGTCAAGGGCA KGRG GAGGC
[0334]
[0335] MALWMR ATGGCCCTGTGGATGCGGC TGCTGCCTCTGCTCGCCCT LLPLLAL HAEGT GCTGGCCCTGTGGGGCCC MALWM LALWGP
[0336] Insulin {+ FTSDVS CGATCCTGCCGCTGCTAGA RLLPLLA DPAAARG
[0337] furin SYLEG GGAAGACGGCACGCCGAG 67.23 LLALWG RRHAEGT 59aa 177bp cleavage QAAKE GGCACCTTCACAAGCGAC %
[0338] PDPAAA FTSDVS S
[0339] site) FIAWLV GTGTCTAGCTACCTGGAAG
[0340] + RGRR YLEGQA
[0341] KGRG GCCAGGCCGCCAAGGAGT AKEFIAW TCATCGCCTGGCTGGTGAA LVKGRG GGGCAGAGGC ATGAAGTGGGTCACCTTCA MKWVTF TCAGCCTGCTGTTTCTGTT HAEGT ISLLFLFS CAGCAGCGCCTACAGCAG SAYSRGV MKWVTF FTSDVS AGGCGTGTTCAGACGGGA FRRDHAE
[0342] Albumin (1- ISLLFLFS SYLEG CCACGCCGAGGGAACATT 58.33
[0343] GTFTSDV 56aa 168bp SAYSRG QAAKE CACCTCTGATGTGTCCAGC %
[0344] SSYLEGQ VFRRD FIAWLV TACCTGGAAGGCCAGGCC AAKEFIA KGRG GCTAAAGAGTTCATCGCCT WLVKGR GGCTGGTGAAGGGCAGAG G GC MKALCL ATGAAGGCCCTGTGCCTGC HAEGT LLLPVLG TCCTGCTGCCTGTGCTGGG FTSDVS LLVSSHA CCTGCTGGTGTCCAGCCAC MKALCL
[0345] Resistin SYLEG EGTFTSD GCCGAGGGAACATTCACC 63.27
[0346] LLLPVLG 49aa 147bp QAAKE VSSYLEG AGCGACGTGTCTAGCTACC % LLVSS FIAWLV QAAKEFI TGGAAGGCCAGGCCGCTA KGRG AWLVKG AAGAGTTCATCGCCTGGCT RG GGTCAAGGGCAGAGGC ATGGGCCCCACCTCCGGCC MGPTSGP CTAGCCTGCTGCTGCTGCT YGEGT SLLLLLLT GCTCACCCACCTGCCTCTG FTSDYS HLPLALG MGPTSG IALDKI YGEGTFT GCCCTGGGCTACGGCGAG PSLLLLL GGAACATTCACCAGCGAC 66.67
[0347] (3 AQKAF SDYSIAL 61aa 183bp LTHLPLA TACAGCATCGCTCTGGATA %
[0348] VQWLI DKIAQKA LG AGATCGCCCAGAAGGCCT AGGPSS FVQWLIA TCGTGCAGTGGCTGATCGC GAPPPS GGPSSGA CGGCGGACCCAGCAGCGG PPPS CGCCCCTCCTCCATCT ATGTACAGCGCTCCCAGCG MYSAPSA CCTGCACATGCCTCTGTCT CTCLCLH YGEGT GCACTTCCTGCTGCTGTGC FLLLCFQ MYSAPS FTSDYS TTCCAGGTGCAAGTGCTG VQVLVAY ACTCLC IALDKI GTCGCCTACGGCGAGGGC GEGTFTS 62.63 FGF18 LHFLLLC AQKAF 66aa ACCTTCACCAGCGACTACT 198bp DYSIALD % FQVQVL VQWLI CCATCGCTCTGGATAAGAT KIAQKAF VA AGGPSS CGCCCAGAAGGCCTTTGT VQWLIA GAPPPS GGPSSGA GCAGTGGCTGATCGCCGG CGGACCTAGCAGCGGCGC PPPS CCCTCCTCCATCT
[0349]
[0350] ATGGTCTTTACACCTCAGA MVFTPQI YPYDV TCCTGGGCCTGATGCTGTT LGLMLF PDYA+ CTGGATCAGCGCCTCTAGA WISASRG RGRR+ GGCTACCCCTACGACGTGC YPYDVPD
[0351] IgK+HA+F MVFTPQI HGEGT CTGATTACGCCAGAGGAA
[0352] YARGRR 60.94 CS+GLP1(7 LGLMLF FTSDVS 64aa GACGGCACGGCGAGGGCA 192bp HGEGTFT
[0353] -37) ASG %
[0354] WISASRG SYLEG CCTTCACCAGCGACGTGTC SDVSSYL QAAKE CAGCTATCTGGAAGGCCA EGQAAK FIAWLV GGCCGCTAAGGAGTTCATC EFIAWLV KGRG GCCTGGCTGGTGAAGGGC KGRG CGGGGA MHWGTL ATGCACTGGGGAACACTG YPYDV CGFLWL TGCGGCTTTCTGTGGCTGT PDYA+ WPYLFY GGCCCTACCTGTTCTACGT
[0355] RGRR+ VQAYPY GCAGGCTTATCCTTACGAC MHWGTL
[0356] Leptin+HA HGEGT DVPDYAR
[0357] CGFLWL GTGCCTGATTACGCCAGAG
[0358] 61.03 +FCS+GLP FTSDVS GRRHGE 65aa GCCGGAGACACGGCGAGG 195bp WPYLFY %
[0359] 1(7-37) A8G SYLEG GTFTSDV GCACCTTCACCAGCGACG
[0360] VQA QAAKE SSYLEGQ TGTCTAGCTACCTGGAAGG FIAWLV AAKEFIA CCAGGCCGCCAAGGAGTT KGRG WLVKGR CATCGCCTGGCTGGTCAAG G GGCAGAGGA ATGCACAGCTGGGAGCGG MHSWER YPYDV LAVLVLL CTGGCCGTGCTGGTGCTGC PDYA+ TGGGCGCCGCTGCCTGCG
[0361] GAAACA RGRR+ CCGCTTATCCCTACGACGT Adipsin+I 1 MHSWER AYPYDVP
[0362] HGEGT GCCTGACTACGCCAGAGG
[0363] A+FCS+GL LAVLVLL DYARGR 65.62
[0364] FTSDVS 64aa AAGACGGCACGGCGAGGG 192bp Pl (7-37) GAAACA RHGEGTF %
[0365] SYLEG CACCTTCACATCTGATGTG
[0366] A8G A TSDVSSY
[0367] QAAKE TCCAGCTACCTGGAAGGC LEGQAA FIAWLV KEFIAWL CAGGCCGCCAAGGAATTC KGRG ATCGCCTGGCTGGTCAAG VKGRG GGCAGAGGC ATGCTGCTGCTCGGCGCCG MLLLGAV YPYDV TGCTGCTGCTGCTGGCCCT LLLLALP PDYA+ GCCTGGCCACGACCAGTA GHDQYP RGRR+ CCCCTACGACGTGCCTGAT Adiponectin YDVPDY
[0368] MLLLGA HGEGT TACGCCAGAGGAAGACGG +I1A+FCS* ARGRRH 63.98
[0369] VLLLLAL GLPK7-37) FTSDVS 62aa CACGGCGAGGGCACCTTC 186bp GEGTFTS % PGHDQ SYLEG ACAAGCGACGTGTCTAGC
[0370] A8G DVSSYLE
[0371] QAAKE TATCTGGAAGGCCAGGCC GQAAKE FIAWLV GCTAAGGAGTTCATCGCCT FIAWLVK KGRG GGCTGGTCAAGGGCAGAG GRG GA ATGAAGTGGGTGACCTTTA MKWVTF HAEGT TCAGCCTGCTGTTCAGCTC ISLLFSSA TGCCTACAGCCACGCCGA MKWVTF FTSDVS YSHAEGT
[0372] Albumin (1- SYLEG GGGAACATTCACCAGCGA 58.87
[0373] ISLLFSSA FTSDVS S 47aa 141bp 16) QAAKE CGTGTCCAGCTACCTGGA %
[0374] YS YLEGQA FIAWLV AGGCCAGGCCGCTAAAGA AKEFIAW KGRG GTTCATCGCCTGGCTGGTG LVKGRG AAGGGCAGAGGC
[0375]
[0376] MKCLLYL ATGAAGTGTCTGCTGTACC
[0377] HAEGT TGGCCTTTCTGTTCATCGG AFLFIGV FTSDVS CGTGAACTGCCACGCCGA MKCLLY NCHAEG SYLEG GGGAACATTCACCAGCGA 58.16 23 LAFLFIG QAAKE TFTSDVS 47aa CGTGTCTAGCTACCTGGAA 141bp % VNC SYLEGQA FIAWLV GGCCAGGCCGCTAAAGAG AKEFIAW KGRG TTCATCGCCTGGCTGGTGA LVKGRG AGGGCAGAGGC ATGCCTGGCCGGGCCCCTC MPGRAPL TGAGAACAGTGCCCGGCG RTVPGAL CCCTGGGAGCTTGGCTGCT YPYDV GAWLLG GGGCGGACTCTGGGCCTG MPGRAP PDYA+ GLWAWT
[0378] GACCCTGTGCGGCCTGTGT LCGLCSL LRTVPG RGRR+
[0379] EPDR1+11 AGCCTGGGCGCTGTGGGC ALGAWL HGEGT GAVGYPY
[0380] A+FCS+GL TACCCATACGACGTGCCTG 67.08
[0381] 24 LGGLWA DVPDYAR FTSDVS 81aa 243bp P1(7-37) WTLCGL ATTACGCCAGAGGCAGAC %
[0382] SYLEG GRRHGE AXG GGCACGGCGAGGGCACCT QAAKE CSLGAV GTFTSDV TCACCAGCGACGTGTCCTC G FIAWLV SSYLEGQ TTATCTGGAAGGCCAGGCC AAKEFIA KGRG GCCAAGGAGTTCATCGCCT WLVKGR GGCTGGTCAAGGGCAGAG G GA
[0383]
[0384] Example 3
[0385] In the present example, the process outlined in FIG. 3 was followed by extracting proteins from the Human Protein Atlas with the protein class “Plasma protein” and tissue category (RNA) “Adipose tissue” with specificity of “tissue enriched,” “group enriched,” and “tissue enhanced,” or with tissue category (RNA) “Adipose tissue” with specificity of “tissue enriched,” “group enriched,” and “tissue enhanced.” Duplicate entries were removed based on duplicate Uniprot accession numbers and entries were extracted and ranked based on blood concentration as measured by immunoassay.
[0386] The following proteins were obtained: ADIPOQ, ADM, ANGPTL4, C1QTNF1, CETP, CFD, CLEC3B, CXCL3, EDN1, FGFBP2, GHR, IL10, IL20, IL6, ITLN1, LEP, MMP3, PLA2G2A, PTX3, SPX.
[0387] The screen produces the following rank-ordered list of signal sequences:
[0388]
[0389] ATGCTGCTG
[0390] CTCGGCGCC MLLLGAVLL GTGCTGCTG ADIPOQ 18aa 72.22% 54bp LLALPGHDQ CTGCTGGCC CTGCCTGGC CACGACCAG ATGGAGCTG TGGGGCGCC MELWGAYLL TACCTGCTG CLEC3B LCLFSLLTQV 21aa CTGTGCCTG 63.49% 63bp TT TTCAGCCTG CTGACACAG GTGACCACC ATGCTGGCC GCTACAGTG MLAATVLTL CTGACCCTG CETP 17aa 70.59% 51bp ALLGNAHA GCCCTGCTG GGCAACGCC CACGCC ATGCACAGC TGGGAGAGA MHSWERLAV CTGGCCGTG CFD LVLLGAAAC 20aa CTGGTGCTG 73.33% 60bp AA CTGGGCGCC GCTGCCTGC GCCGCC ATGAAGTTC GTGCCTTGT MKFVPCLLL CTGCTGCTG FGFBP2 19aa GTGACCCTG 61.40% 57bp VTLSCLGTLG AGCTGCCTG GGCACACTG GGC ATGGGATCTA GAGGCCAGG GCCTGCTGC MGSRGQGLL TCGCCTACT
[0391] C1QTNF1 LAYCLLLAFA 25aa GCCTGCTGC 66.67% 75bp SGLVLS TGGCTTTCG CCAGCGGCC TGGTGCTGA GC ATGAACCAG CTGAGCTTC MNQLSFLLFL CTGCTGTTC ITLN1 18aa 55.56% 54bp IATTRGWS CTGATCGCC ACAACCAGA
[0392]
[0393] GGCTGGTCTATGCACTGG
[0394] GGCACCCTG MHWGTLCGF TGCGGCTTC X LEP LWLWPYLFY 21aa CTGTGGCTG 63.49% 63bp VQA TGGCCTTAC CTGTTCTACG TGCAGGCC ATGAGCGGC GCCCCTACA GCCGGAGCT MSGAPTAGA GCTCTGATG ANGPTL4 ALMLCAATA 25aa CTGTGCGCC 72.00% 75bp VLLSAQG GCCACCGCC GTGCTGCTG TCTGCCCAG GGC ATGAAGGGC CTGAGATCT CTGGCCGCT MKGLRSLAA ACAACCCTG
[0395] 10 SPX TTLALFLVFV 26aa GCCCTGTTC 58.97% 78bp FLGNSSC CTGGTGTTC GTGTTTCTG GGCAACAGC AGCTGC ATGAAGTCT CTGCCTATCC MKSLPILLLL H TGCTGCTGC MMP3 17aa 60.78% 51bp CVAVCSA TGTGTGTGG CCGTGTGCA GCGCC ATGGATCTGT GGCAGCTGC MDLWQLLLT TGCTGACCC
[0396] 12 GHR 18aa 66.67% 54bp LALAGSSDA TGGCTCTGG CCGGCTCTA GCGACGCC ATGCACCTG CTGGCTATCC MHLLAILFCA TGTTCTGCG
[0397] 13 PTX3 17aa 66.67% 51bp LWSAVLA CCCTGTGGA GCGCCGTGC TGGCC ATGAAGACC CTGCTGCTG MKTLLLLAVI CTGGCCGTG
[0398] 14 PLA2G2A MIFGLLQAH 20aa ATCATGATCT 63.33% 60bp G TCGGCCTGC TGCAGGCCC
[0399]
[0400] ACGGCATGAAGGCC
[0401] TCTAGCCTG GCCTTCAGC MKASSLAFSL CTGCTGTCC
[0402] 15 IL20 LSAAFYLLW 24aa 62.50% 72bp GCCGCTTTC TPSTG TACCTGCTGT GGACACCTA GCACCGGC ATGAAGCTG GTGTCTGTG MKLVSVALM GCCCTGATG
[0403] 16 ADM YLGSLAFLG 21aa TACCTGGGC 63.49% 63bp ADT AGCCTGGCT TTCCTGGGC GCCGACACC ATGGCCCAC GCCACCCTG AGCGCCGCC CCTAGCAAC MAHATLSAA CCCAGACTG PSNPRLLRVA CTGAGAGTG
[0404] 17 CXCL3 34aa 70.59% 102bp LLLLLLVAAS GCCCTGCTG RRAAG CTGCTGCTG CTCGTGGCC GCTTCTAGA CGGGCCGCT GGC ATGGACTAC CTGCTGATG MDYLLMIFSL ATCTTCAGC
[0405] 18 EDN1 17aa 58.82% 51bp LFVACQG CTGCTGTTC GTGGCCTGC CAGGGC ATGAACAGC TTCTCCACC AGCGCCTTC GGCCCAGTG MNSFSTSAFG GCCTTTTCTC
[0406] 19 IL6 PVAFSLGLLL 29aa 65.52% 87bp VLPAAFPAP TGGGCCTGC TGCTGGTGC TGCCTGCCG CTTTCCCCG CCCCT ATGCACTCTA GCGCCCTGC MHSSALLCC TGTGCTGTC
[0407] 20 IL10 LVLLTGVRA 18aa 64.81% 54bp TGGTGCTGC TGACCGGCG
[0408]
[0409] TGAGAGCC
[0410] The full set is screened in vitro using a fluorescent reporter in the supernatant (quantification performed in supernatant collected from 24- well plates transfected with plasmid library hits). Following selection of top in silico hits, 6 combinations of secretion signal / insulin are screened in vitro using primary human adipocytes. Confirmation of secretion into supernatant is optionally confirmed by targeted proteomics. The 3 highest expressing combinations are advanced into in vivo efficacy model verification in C57BL / 6 high-fat diet induced obesity and type 2 diabetes mice.
[0411] Example 4
[0412] In the present example, the process outlined in FIG. 4 was followed using a programming language such as R or python to automatically extract proteins from a database. The raw data is subset to extract proteins that are expressed by and secreted from adipocytes. The signal sequence for each of secreted adipocyte proteins is pulled from Uniprot and attached to a GLP / GIP dual receptor agonist. Optimal cleavage of the signal sequence-dual agonist fusion is verified using PrediSi (using the “score” metric), DeepSig (“Reliability” metric), or a combination thereof. Cleavage sites are optionally optimized after in silico prediction via PrediSi and / or DeepSig.
[0413] The table below displays an unranked list of the scores and cleavage sites from PrediSi and DeepSig for 10 signal sequence-GLPl / GIP dual receptor agonist fusion peptides:
[0414] MLLLG AVL LLLALI3GH
[0415] MLLLGAVL LLLALPGH DQYGEGTF
[0416] ADIPOQ LLLALPGH 18 TSDYSIAM 0.8015 17 0.99 19
[0417] DE DKIAQ] CAF
[0418] VQWLL XGG PSSGAF > PPS
[0419] MELWGAYL LLCLFSLLT QVTTYGEG MELWGAYL TFTSDYSIA CLEC3B LLCLFSLLT 21 0.8225 21 0.99 21
[0420] MDKIAQKA QVTT FVQWLIAG GPSSGAPPP S MLAATVLT LALLGNAH MLAATVLT AYGEGTFT CETP LALLGNAH 17 SDYSIAMD 0.7773 17 0.86 17
[0421] A KIAQKAFV QWLIAGGP
[0422]
[0423] SSGAPPPSMHSWERL
[0424] AVLVLLGA AACAAYGE MHSWERL GTFTSDYSI CFD AVLVLLGA 20 0.9091 22 0.99 19
[0425] AMDKIAQK AACAA AFVQWLIA GGPSSGAPP PS MKFVPCLL LVTLSCLGT MKFVPCLL LGYGEGTF FGFBP2 LVTLSCLGT 19 TSDYSIAM 0.8365 21 0.99 21
[0426] LG DKIAQKAF VQWLIAGG PSSGAPPPS MGSRGQGL LLAYCLLL AFASGLVLS MGSRGQGL YGEGTFTS
[0427] C1QTNF1 LLAYCLLL 25 0.7964 25 0.91 25
[0428] DYSIAMDK AFASGLVLS IAQKAFVQ WLIAGGPS SGAPPPS MNQLSFLL FLIATTRGW MNQLSFLL SYGEGTFTS ITLN1 FLIATTRGW 18 DYSIAMDK 0.6781 18 0.91 16
[0429] S IAQKAFVQ WLIAGGPS SGAPPPS MHWGTLC GFLWLWPY LFYVQAYG MHWGTLC EGTFTSDYS LEP GFLWLWPY 21 0.5637 21 0.79 21
[0430] IAMDKIAQ LFYVQA KAFVQWLI AGGPSSGA PPPS MSGAPTAG AALMLCAA MSGAPTAG TAVLLSAQ AALMLCAA GYGEGTFT ANGPTL4 25 0.6990 27 0.99 23
[0431] TAVLLSAQ SDYSIAMD G KIAQKAFV QWLIAGGP SSGAPPPS MKGLRSLA ATTLALFLV FVFLGNSSC MKGLRSLA YGEGTFTS SPX ATTLALFLV 26 0.7214 24 0.89 28
[0432] DYSIAMDK FVFLGNSSC IAQKAFVQ WLIAGGPS
[0433]
[0434] SGAPPPSThe full set is screened in vitro in primary adipocytes by quantifying the amount of GLP1 / GIP dual agonist present in supernatant collected from 24- well plates transfected with the plasmid library. Orthogonal verification of cytocompatibility and transfection efficiency is performed using an identical library of plasmids with a nanoluciferase reporter in place of the GLP1 / GIP dual agonist. Following selection of top in silico hits, the 3 highest expressing combinations are advanced into in vivo efficacy model verification in C57BL / 6 high-fat diet induced obesity and type 2 diabetes mice.
[0435] Example 5
[0436] In the present example, the process outlined in FIG. 5 was followed by selecting the Enrichment category of Adipose Tissues, and subcategory “Adipocytes, at levels ranging from Moderate to High”. Additional search restriction included secreted and exclusion of the intracellular and immunoglobulin proteins.
[0437] ce enriched: Adipose subcutaneous; Adipocytes (Subcutaneous); high, High, Moderate AND sa_location: Secreted - unknown location, Secreted in brain, Secreted in female reproductive system, Secreted in male reproductive system, Secreted in other tissues, Secreted to
[0438] blood, Secreted to digestive system, Secreted to extracellular matrix
[0439] The following proteins were obtained and intersected with the same set of proteins with known secretion signals from Uniprot:
[0440] ADIPOQ, AOC3, APOB, AZGP1, BTD, CSF2RA, DEFB132, ENPP1, FCN2, FGFBP2, GHR, GPX3, LAMB3, LEP, LPL, MAPT, NMB, NRCAM, PCOLCE2, PRADC1, PRXL2A, PTPRS, PXDN, RBP4, SEMA3G, SPON1, SSC4D, TF, TIMP4, TSKU, VEGFB, ZBED3.
[0441] The library of plasmids is screened in vitro in human multipotent adipose-derived stem cells (hMADs) by quantifying bioluminescent reporter protein present in supernatant collected from 24-well plates transfected with the plasmid library. Orthogonal verification is performed using mass spectrometry. The 3 highest expressing combinations are advanced into in vivo efficacy model verification in C57BL / 6 high-fat diet induced obesity and type 2 diabetes mice, replacing the bioluminescent reporter with insulin.In vitro screening confirmed high expression levels of ADIPOQ, BTD, TIMP4, TSKU, PXDN, VEGFB with the exendin 4 GLP-1 receptor agonist for advancement to in vivo confirmation.
[0442] Example 6
[0443] In the present example the optimal secretion signal / therapeutic protein combination for expression in subcutaneous adipocytes is derived from database mediated identification of secretion signals of naturally expressed and secreted proteins in human adipose tissues, alongside identification of endopeptidase cleavage sites or sequences processable by adipocytes or additionally by extracellular peptidases, and optionally spacer sequences, which prevent endonuclease mediated cleavage of parts of active sites of the therapeutic peptides or proteins. The sequences are stitched together and verified using Signal P 5.0 for the presence of a secretion signal. Once presence is confirmed sequences that impact the N-terminal active site are excluded. Following exclusion the sequences are advanced to DNA plasmid library production and the manufactured plasmids are screened one plasmid per well using primary human adipocytes. Therapeutic protein levels are measured by ELISA and their fusion protein counterparts with nLuc reporters are quantified by the supernatant using a plate reader measuring bioluminescence in the presence of fluorofurimazine substrate (5 mins post substrate addition to the supernatant). Highest expressing sequences are ranked by reporter and ELISA values and the top 6 are advanced to in vivo efficacy for confirmatory screening,
[0444] In vivo, the secretion efficiency is confirmed by measuring the total bioluminescence from the therapeutic protein coupled with nLuc with and without the secretion signal. The total body bioluminescence of the secretion signal coupled sequence constructs delivered subcutaneously is normalized by local bioluminescence from secretion signal-free sequences of the same protein-reporter fusions. Higher signal ratios indicate enhanced processivity and successful optimization of the secretion signal.
[0445] Example 7
[0446] In the present example, a number of optional steps are used in the selection of a secretion signal / linker / therapeutic protein. Proteins are extracted from a database and subset to extract proteins that are expressed by and secreted from adipocytes, and to extract growth factors and hormones. The signal sequence for each of the secreted adipocyte proteins is pulledfrom signalpeptide.de. A list of flexible linkers is extracted from the MEROPS database of proteolytic enzymes and their substrates (EMBL-EBI).
[0447] Combinations of signal sequences, linkers, and therapeutic proteins are generated and optimal cleavage of the therapeutic protein from the signal sequence and / or linker is verified in silico using SignalP, DeepSig, a similar tool, or a combination thereof. Cleavage sites and linker sequences are optionally altered after verification to further optimize the fusion. Sequences that impact the N-terminal active site are excluded. A DNA plasmid library with each plasmid containing a signal sequence, linker, cleavage site, and fluorescent or bioluminescent proteins of interest (e.g. green fluorescent protein, blue fluorescent protein, red fluorescent protein, firefly luciferase, nanoluciferase, red-shifted luciferase) is generated for in vitro screening.
[0448] Plasmids are pooled such that each plasmid in a given pool contains a different fluorescent protein reporter. Pools of plasmids are screened in vitro in SGBS cells and primary adipocytes in 24- well plates. Secreted protein in the supernatant is quantified by fluorescent or bioluminescent spectroscopy and confirmed by Western blot and qPCR of cell lysate. The highest expressing sequences are ranked by reporter and ELISA values and the top 6 are advanced to in vivo efficacy for confirmatory screening. A smaller DNA plasmid library is generated with the top 6 signal sequence / linker / cleavage site combinations paired with exendin-4 (a GLP1 receptor agonist), nanoluciferase, or a fusion protein of exendin-4 and nanoluciferase.
[0449] In vivo secretion efficiency is confirmed by measuring the total bioluminescence with and without the secretion signal in the presence of fluorofurimazine, and confirmed by ELISA and targeted proteomics of plasma, serum, or whole blood.
Claims
CLAIMSWe claim:
1. Compositions for optimized expression of therapeutic proteins in the subcutaneous tissues from DNA or RNA sequences delivered by non-viral vectors, which include:a. A set of natural secretion signals from secreted proteins expressed by human adipocytes or adipocytes of non-human primates, or secretion signals optimized from secretion signals of proteins expressed by human adipocytes or adipocytes of non-human primates in silico and encoded in DNA or RNA cassettes, b. And optionally prioritized in silico, screened and prioritized in vitro, and screened and prioritized in vivo.
2. Secretion signals in Claim 1, which are derived at least in part that contains at least 4 consecutive and up to all consecutive amino acids of the following MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG, MALWMRLLPLLALLALWGPDPAAA, MDLWQLLLTLALAGSSDA, MDYLLMIFSLLFVACQG, MELWGAYLLLCLFSLLTQVTT, METPAQLLFLLLLWLPDTTG, MGPTSGPSLLLLLLTHLPLALG, MGSRGQGLLLAYCLLLAFASGLVLS, MHLLAILFCALWSAVLA, MHSSALLCCLVLLTGVRA, MHSWERLAVLVLLGAAACAA, MHWGTLCGFLWLWPYLFYVQA, MKALCLLLLPVLGLLVSS, MKASSLAFSLLSAAFYLLWTPSTG, MKCLLYLAFLFIGVNC, MKFVPCLLLVTLSCLGTLG, MKGLRSLAATTLALFLVFVFLGNSSC, MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA, MKLVSVALMYLGSLAFLGADT, MKSIYFVAGLFVMLVQGSWQ, MKSLPILLLLCVAVCSA, MKTLLLLAVIMIFGLLQAHG, MKWVTFISLLFLFSSAYSRGVFRRD, MKWVTFISLLFSSAYS, MLAATVLTLALLGNAHA, MLLLGAVLLLLALPGHDQ, MLLLLLL LLGLRL QLSLG, MNQLSFLLFLIATTRGWS, MNSFSTSAFGPVAFSLGLLLVLPAAFPAP, MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG, MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG, MSGAPTAGAALMLCAATAVLLSAQG, MTSKLAVALLAAFLISAAL C,MVFTPQILGLMLFWISASRG, MYSAPSACTCLCLHFLLLCFQVQVLVA, MHWGTLCGFLWLWPYLFYVQA.
3. Natural secretion signals in Claim 1, which are derived from secretion signals obtained from proteins produced and secreted by adipocytes.
4. Natural secretion signals in Claim 1, which are derived from secretion signals obtained from proteins produced and secreted by adipocytes under a given disease state.
5. Natural secretion signals in Claim 1, which are appended to therapeutic proteins such that the first amino acid of the therapeutic protein is identical to the first post-secretion signal amino acid of the original secreted protein.
6. Natural secretion signals in Claim 1, which are appended to therapeutic proteins such that the first two amino acids of the therapeutic protein are sequentially identical to the first two post-secretion signal amino acid of the original secreted protein.
7. Natural secretion signals in Claim 1, which are appended to therapeutic proteins such that the first three amino acids of the therapeutic protein are sequentially identical to the first three post-secretion signal amino acid of the original secreted protein.
8. In silico optimized secretion signals in Claim 1, which include natural secretion signals in combination with synthetic secretion signals, or fully synthetic signals that are optimized completely in silico with or without initial derivation from natural sequences or motifs.
9. Natural secretion signal in Claim 1, which is a secretion signal from proteins produced and secreted by adipocytes, pre-adipocytes, or adipocyte progenitors.
10. Natural secretion signal in Claim 1, which is secretion signal from a pool of known eukaryotic secretion signals.
11. Natural secretion signal in Claim 1, which is a secretion signal from a pool of known mammalian secretion signals.
12. Natural secretion signals in Claim 1, which are derived from secretion signals obtained at least in part from a pool of known human secretion signals.
13. Natural secretion signals in Claim 1, which include matched amino acids such that at least one of the three amino acids in the pre-cleavage site sequence and at least one of the three amino acids in the post-cleavage site sequence, preferentially the n=-l and n=0 position relative to the secretion signal cleavage site of 0, which indicates the future N-terminal of the protein.
14. Natural secretion signals in Claim 1, which includes signals from the following: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), and spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC).
15. Natural secretion signals in Claim 1, which includes signals from the following:adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), and interleukin 10 (MHSSALLCCLVLLTGVRA).
16. Natural secretion signals in Claim 1, which have an identically matching first post cleavage site amino acid of the therapeutic protein of interest or the minimal active sequence of the therapeutic protein of interest to the first post cleavage site of the protein that the secretion signal was obtained from.
17. Secretion signals in Claim 1, which are synthetic, semi- synthetic, or natural, or a combination of synthetic and natural secretion signals.
18. Natural secretion signals in Claim 13, which additionally include an identically matching n=l amino acid of the secretion signal to the secretion signal of the therapeutic protein of interest.
19. Secretion signals in Claim 13, that include synthetic or composite syntheticnatural amino acid sequences.
20. Matching of amino acids, or identity of amino acids in any claims that includes identical matches, or matches based on charge, hydrophilicity, hydrophobicity, pKa, or other physico-chemical parameter.
21. Secretion signals in Claim 1, which include one or more parts of a natural secretion signal.
22. Secretion signals in Claim 1, which includes a natural secretion signal followed by a stretch of modified amino acids derived from natural or in silico designed sequences and then the therapeutic protein to be secreted.
23. Secretion signals in Claim 1, which is derived from secretion signals of predicted proteins or proteins or nucleic acids in bioinformatics databases.
24. Secretion signals in Claim 1, which are codon optimized.
25. Secretion signals in Claim 1, which include at least a part from at least one of the following: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3(MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1 (MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(1-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
26. Signals in Claim 1, which contain motifs designed by modeling, simulation, artificial intelligence, regression, bioinformatics analysis, or combination thereof.
27. Signals in Claim 1, which include sequence homology before and / or after the cleavage site, within up to the first 4 amino acids in each direction inclusive of the therapeutic protein to be secreted and the original protein that the parent signal is derived from.
28. Signals in Claim 1, which include sequences or patterns of sequences common to adipocyte or preadipocyte secretion signals.
29. Signals in Claim 1, which include patterns of nucleic acid or amino acid sequences used by adipocytes to enable secretion, partitioning, cleavage, or export of proteins.
30. Signals in Claim 1, which include pairwise matched natural secretion signals with the secreted sequence of the protein of interest, or a functional fragment of the protein of interest, or a functional fragment or full protein of interest coupled with one or more spacer sequences and identifying sequence sets, which have a cleavage site and predictable n-, h-, and c-regions quantified as likely secreted and cleavable by one or more bioinformatics tools.
31. Secretion signal in Claim 1, which includes a secretion signal sequence and one or more cleavage sites ahead of the therapeutic protein of interest.
32. Secretion signal in Claim 1, which includes a part or a consensus sequence cleaved by an enzyme from a set that includes: serine protease 1 (PRSS1), transmembrane protease serine 2 (TMPRSS2), cathepsin D (CTSD), cathepsin E (CTSE), renin (REN), caspase 3 (CASP3), caspase 9 (CASP9), sentrin-specific protease 1 (SENP1), sentrin-specific protease 6 (SENP6), interstitial collagenase (MMP1), 72 kDa type IV collagenase (MMP2), stromelysin 1 (MMP3), or furin.
33. Secretion signal in Claim 1, which includes a cleavage site from the following set: P-P-T-I-F-F-R-L, K-P-I-E-F-F-R-L, R-[S / K]-[R / S / K]-[R / K]-[]-[]-[]-[G / E], []-[]-[]- [L / F]-[V] -[]-[]-[], P-F-H-L-[L / V / K]-[V / VY]-[Y / H / G]-[S / N], D-E-V-D-[G / S] -[]-[]-[], []-[E / D]-[]-D-[]-[]-[]-[], Q-T-G-G-K-[]-E-[], P-Q-G-I-A-G-Q, D-[]-[]-D, []-R-[]-[K]- R-R-[].
34. Secretion signal in Claim 1, which include a flexible amino acid linker between the secretion signal and the therapeutic protein.
35. Secretion signal in Claim 1, which include a flexible amino acid linker which contains one or more endopeptidase cleavage sites.
36. Secretion signal in Claim 1, which include a flexible linker region or spacer region from a set that includes: QPELQKPFKYTTVTKRSRRIRPTHPA, GGGGS, SGSG, APSVAPEPDGC, AAAAA, PAAAA, GEAAEGPAAA, AAGVGGERSS, GGPSGAGAGDE, VRTHGTLESVNGPKA, DQKVRPNEENNKDADL, GVKDTD, LPVQNGCPESAMEMN.
37. Signal in Claim 1, which includes two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position.
38. Signal in Claim 1, which includes two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site score.
39. Signal in Claim 1, which includes two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein ofinterest, or a fragment or full protein of interest with one or more amino acid spacers based on the cleavage site position and score.
40. Signal in Claim 1, which includes two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal.
41. Signal in Claim 1, which includes two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal and a homology match to a secretion signal from a naturally adipocyte-secreted protein.
42. Signal in Claim 1, which includes two or more pairwise matches between a set of secretion peptides and therapeutic protein of interest, or a fragment of the protein of interest, or a fragment or full protein of interest with one or more amino acid spacers based on the probability of the sequence having a secretion signal, and a homology match to a secretion signal from a naturally adipocyte-secreted protein, and at least one amino acid homology or similarity match to the naturally secreted protein.
43. In silico designed or optimized signal in Claim 1, which includes at least a portion of the following secretion signals: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin (MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCLLLVTLSCLGTLG), Clq and TNF related 1 (MGSRGQGLLLAYCLLLAFASGLVLS), intelectin 1 (MNQLSFLLFLIATTRGWS), leptin (MHWGTLCGFLWLWPYLFYVQA), angiopoietin like 4 (MSGAPTAGAALMLCAATAVLLSAQG), spexin hormone (MKGLRSLAATTLALFLVFVFLGNSSC), matrix metallopeptidase 3 (MKSLPILLLLCVAVCSA), growth hormone receptor (MDLWQLLLTLALAGSSDA), pentraxin 3 (MHLLAILFCALWSAVLA), phospholipase A2 group IIA (MKTLLLLAVIMIFGLLQAHG), interleukin 20 (MKASSLAFSLLSAAFYLLWTPSTG), adrenomedullin (MKLVSVALMYLGSLAFLGADT), C-X-C motif chemokine ligand 3 (MAHATLSAAPSNPRLLRVALLLLLLVAASRRAAG), endothelin 1(MDYLLMIFSLLFVACQG), interleukin 6 (MNSFSTSAFGPVAFSLGLLLVLPAAFPAP), interleukin 10 (MHSSALLCCLVLLTGVRA), angiotensinogen (MRKRAPQSEMAPAGVSLRATILCLLAWAGLAAG), interleukin 8 (MTSKLAVALLAAFLISAAL C), secreted alkaline phosphatase (MLLLLLL LLGLRL QLSLG), exendin-4 (MKIILWLCVFGLFLATLFPISWQMPVESGLSSEDSASSESFA), immunoglobulin kappa light chain (METPAQLLFLLLLWLPDTTG), proglucagon (MKSIYFVAGLFVMLVQGSWQ), insulin (MALWMRLLPLLALLALWGPDPAAA), albumin(1-25) (MKWVTFISLLFLFSSAYSRGVFRRD), resistin (MKALCLLLLPVLGLLVSS), complement C3 (MGPTSGPSLLLLLLTHLPLALG), vesicular stomatitis virus G protein (MKCLLYLAFLFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPLRTVPGALGAWLLGGLWAWTLCGLCSLGAVG).
44. Signal in Claim 1, which contains homology of the secretion signal sequence, and the first three amino acid sequence post cleavage site to the naturally secreted protein section signal sequence and the first three amino acid sequence of the naturally secreted protein.
45. DNA cassettes in Claim 1, are linear double stranded DNA, linear single stranded DNA, circular double stranded DNA, circular single stranded DNA with or without chemical modifications or substitutions of one or more bases.
46. DNA cassettes in Claim 1, which include at least one synthetic DNA containing at least one synthetic DNA encoding a therapeutic protein of interest and a secretion signal.
47. RNA cassettes in Claim 1, which include at least one mRNA containing at least one mRNA encoding a therapeutic protein of interest and a secretion signal.
48. DNA or RNA cassettes in Claim 1, which include synthetic, semi- synthetic, enzymatic, or recombinant DNA or RNA.
49. DNA or RNA cassettes in Claim 1, which include at least one DNA or RNA strand encoding at least a part of the amino acid sequence from the following proteins: adiponectin (MLLLGAVLLLLALPGHDQ), C-type lectin domain family 3 member B (MELWGAYLLLCLFSLLTQVTT), cholesteryl ester transfer protein (MLAATVLTLALLGNAHA), complement factor D / adipsin(MHSWERLAVLVLLGAAACAA), fibroblast growth factor binding protein 2 (MKFVPCEEEVTESCEGTEG), Clq and TNF related 1 (MGSRGQGEEEAYCEEEAFASGEVES), intelectin 1 (MNQESFEEFEIATTRGWS), leptin (MHWGTECGFEWEWPYEFYVQA), angiopoietin like 4 (MSGAPTAGAAEMECAATAVEESAQG), spexin hormone (MKGERSEAATTEAEFEVFVFEGNSSC), matrix metallopeptidase 3 (MKSFPIFkkkCVAVCSA), growth hormone receptor (MDEWQEEETEAEAGSSDA), pentraxin 3 (MHEEAIEFCAEWSAVEA), phospholipase A2 group IIA (MKTFFkkAVIMIFGkkQAHG), interleukin 20 (MKASSEAFSEESAAFYEEWTPSTG), adrenomedullin (MKEVSVAEMYEGSEAFEGADT), C-X-C motif chemokine ligand 3 (MAHATESAAPSNPREERVAEEEEEEVAASRRAAG), endothelin 1 (MDYEEMIFSEEFVACQG), interleukin 6 (MNSFSTSAFGPVAFSEGEEEVEPAAFPAP), interleukin 10 (MHSSAEECCEVEETGVRA), angiotensinogen (MRKRAPQSEMAPAGVSERATIECEEAWAGEAAG), interleukin 8 (MTSKEAVAEEAAFEISAAEC), secreted alkaline phosphatase (MEEEEEEEGEREQESEG), exendin-4 (MKIIEWECVFGEFEATEFPISWQMPVESGESSEDSASSESFA), immunoglobulin kappa light chain (METPAQEEFEEEEWEPDTTG), proglucagon (MKSIYFVAGEFVMEVQGSWQ), insulin (MAEWMREEPEEAEEAEWGPDPAAA), albumin( 1 -25) (MKWVTFISEEFEFSSAYSRGVFRRD), resistin (MKAECEEEEPVEGEEVSS), complement C3 (MGPTSGPSEEEEEETHEPEAEG), vesicular stomatitis virus G protein (MKCEEYEAFEFIGVNC), and mammalian ependymin-related protein 1 (MPGRAPERTVPGAEGAWEEGGEWAWTECGECSEGAVG), or an amino acid sequence of a natural secretion signal.
50. In vitro, in vivo, or ex vivo screened secretion signal from any claim that includes the use of fluorescent or bioluminescent reporter proteins expressed in tandem or as a fusion with the therapeutic protein of interest to quantify expression.
51. Composition in Claim 1 or a substantially similar composition designed with the purpose of increasing expression of therapeutic proteins from subcutaneous adipocytes treated with a non-viral vector for the delivery of DNA.
52. Composition in Claim 1 or a substantially similar composition for the attainment of sufficient efficacy or the reduction of a therapeutic dose of a non-viral vector gene therapy targeting or predominantly targeting subcutaneous adipocytes or preadipocytes.
53. DNA or RNA cassettes in Claim 1, which are comprised of natural chemical structures of DNA or RNA, hybrids, or chemically modified structures of DNA or RNA.
54. DNA cassettes in Claim 1, which are plasmids, double stranded DNA, single stranded DNA, linear or circular DNA, and are used in vitro or in vivo to select optimal secretion signals or cleavage sites.
55. RNA cassettes in Claim 1, which are mRNA, circular RNA, linear RNA, double stranded RNA, single stranded RNA or a combination of more than one RNA type.
56. Compositions derived from methods described in the examples, or combinations of one or more parts thereof.
57. In vitro screening derived compositions and secretion signals in Claim 1, which are derived from or utilized in the process of introducing plasmids, mRNA, circRNA, self-amplifying RNA, nanoplasmids, DNA minicircles, double stranded linear DNA, single stranded linear or circular DNA, covalently closed single stranded DNA or other genetic cargo into primary human adipocytes via one or more means such as electroporation, metal particle bombardment, chemical transfection, viral transduction, or combination of one or more approaches.
58. Compositions or secretion signals in Claim 1, which are obtained via quantification of the secreted therapeutic protein of interest, a reporter protein such as a fluorescent or bioluminescent reporter, or a fusion protein between the therapeutic protein and a reporter protein.
59. Therapeutic proteins in Claim 1 which are natural proteins, human proteins, hybrid proteins, fusion proteins, or parts of natural proteins with or without additional amino acid sequences, motifs, or domains.