Novel nuclear targeting signal peptide and uses thereof

By designing novel signal peptides that bind to mRNA, targeted localization and nuclear expression of payload proteins are achieved, overcoming the shortcomings of post-translational control and targeted localization in existing technologies. This improves the delivery efficiency of therapeutic proteins in the cell nucleus and enhances the efficacy of gene editing therapy.

CN122459320APending Publication Date: 2026-07-24BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 15 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BOARD OF RGT THE UNIV OF TEXAS SYST
Filing Date
2024-09-27
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies lack effective means for post-translational control and targeted localization of encoded proteins, resulting in low delivery and localization efficiency of therapeutic mRNAs.

Method used

The design and use of novel signal peptides to achieve targeted localization and nuclear expression of payload proteins by binding to protein-encoding mRNA, including the fusion or linkage of signal peptides with payload proteins, and the use of nucleic acid molecules to encode and deliver them to the cell nucleus.

Benefits of technology

It improves the targeting and expression efficiency of therapeutic proteins in the cell nucleus, enhancing therapeutic effects, especially the therapeutic potential for gene editing-related diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122459320A_ABST
    Figure CN122459320A_ABST
Patent Text Reader

Abstract

Provided herein are novel signal peptides that direct encoded proteins to the nucleus, recombinant polypeptides comprising the same, nucleic acid molecules encoding the same, pharmaceutical compositions comprising the same, and methods of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Related applications

[0001] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 586,847, filed September 29, 2023, and U.S. Provisional Application Serial No. 63 / 603,603, filed November 28, 2023, each of which is incorporated herein by reference in its entirety.

[0002] sequence list References to amino acid sequences and / or nucleic acid sequences contained in this application have been submitted with this application as a sequence listing XML file entitled "824623_sequencelisting.xml", which is 16,384 bytes in size and was created on September 27, 2024. The entire sequence listing is incorporated herein by reference in accordance with 37 CFR §1.52(e)(5). Technical Field

[0003] This disclosure generally relates to novel signal peptides, and more specifically to novel signal peptides for targeting encoded proteins to the cell nucleus. Background Technology

[0004] Recent scientific discoveries have highlighted the numerous therapeutic applications of mRNA due to its modular nature and ability to provide customizable "instructions" to produce functional proteins. Clinical studies continue to confirm its beneficial efficacy, further emphasizing its potential as a next-generation genetic drug, where scientists are only beginning to scratch the surface of its potential applications. To date, most research has focused on enhancing the intracellular expression of encoded peptides. However, little attention has been paid to the post-translational control of the encoded protein and its targeted localization and / or secretion in vivo.

[0005] Within the cell, various unique pathways and processes exist that allow proteins to shuttle to various organelles and to be exported to the extracellular space via excretion pathways. However, in order to utilize these transport systems, the mRNA encoding each protein must also contain a metaphorical transport marker called a signal peptide (SP) upstream of the protein sequence. The design of novel SPs for therapeutic mRNAs can greatly benefit the delivery, localization, and clearance of the encoded therapeutic proteins. This disclosure addresses these and other needs. Summary of the Invention

[0006] In some embodiments, a signal peptide is provided. In some embodiments, the signal peptide comprises an amino acid sequence selected from the group consisting of formulas I, II, III, IV, V, VI, VII, and VIII; Where equation I is expressed as: A1-A2-A3-A4-A5-A6-A7-A8-A9-A10 -A 11 -A 12 -A 13 -A 14 -A 15 -A 16 (Formula I) And A1-A 16 The identifiers are provided in Table 2; Equation II is expressed as: B1-B2-B3-B4-B5-B6-B7-B8-B9-B 10 -B 11 -B 12 -B 13 -B 14 -B 15 -B 16 -B 17 -B 18 (Formula II) And B1-B 18 The identifiers are provided in Table 3; Equation III is expressed as: C1-C2-C3-C4-C5-C6-C7-C8-C9-C 10 -C 11 -C 12 -C 13 -C 14 -C 15 -C 16 -C 17 -C 18 -C 19 (Formula III) And C1-C 19 The identifiers are provided in Table 4; Where equation IV is represented as: D1-D2-D3-D4-D5-D6-D7-D8-D9-D 10 -D 11 -D 12 -D 13 -D 14 -D 15 -D 16 -D 17 -D 18 -D 19 -D 20 (Formula IV) And D1-D 20 The identifiers are provided in Table 5; Where V is expressed as: E1-E2-E3-E4-E5-E6-E7-E8-E9-E 10 -E11 -E 12 -E 13 -E 14 -E 15 -E 16 -E 17 -E 18 -E 19 -E 20 -E 21 -E 22 (Formula V) And E1-E 22 The identifiers are provided in Table 6; Wherein, VI is represented as: F1-F2-F3-F4-F5-F6-F7-F8-F9-F 10 -F 11 -F 12 -F 13 -F 14 -F 15 -F 16 -F 17 -F 18 -F 19 -F 20 -F 21 -F 22 -F 23 (Form VI) And F1-F 23 The identifiers are provided in Table 7; Equation VII is expressed as: G1-G2-G3-G4-G5-G6-G7-G8-G9-G 10 -G 11 -G 12 -G 13 -G 14 -G 15 -G 16 -G 17 -G 18 -G 19 -G 20 -G 21 -G 22 -G 23 -G 24 -G 25 -G 26 (Equation VII) And G1-G 26 The identifiers are provided in Table 8; and Where equation VIII is expressed as: H1-H2-H3-H4-H5-H6-H7-H8-H9-H 10 -H 11-H 12 -H 13 -H 14 -H 15 -H 16 -H 17 -H 18 -H 19 -H 20 -H 21 -H 22 -H 23 -H 24 -H 25 -H 26 -H 27 -H 28 -H 29 -H 30 (Formula VIII) And H1-H 30 The identifier is provided in Table 9.

[0007] In some embodiments, the signal peptide comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

[0008] In some embodiments, a signal peptide is provided comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

[0009] In some implementations, the signal peptides provided herein are novel signal peptides.

[0010] In some embodiments, a recombinant polypeptide is provided. In some embodiments, the recombinant polypeptide comprises X1-(Y1). a -Z1, where X1 is the signal peptide, Y1 is the peptide linker, and Z1 is the payload protein, where a is an integer selected from 0 and 1.

[0011] In some embodiments, signal peptide X1 comprises an amino acid sequence selected from the group consisting of formulas I, II, III, IV, V, VI, VII, and VIII; Where equation I is expressed as: A1-A2-A3-A4-A5-A6-A7-A8-A9-A 10 -A 11 -A 12 -A 13 -A 14 -A 15 -A 16 (Formula I) And A1-A 16 The identifiers are provided in Table 2; Equation II is expressed as: B1-B2-B3-B4-B5-B6-B7-B8-B9-B 10 -B 11 -B 12 -B 13 -B 14 -B 15 -B 16 -B 17 -B 18 (Formula II) And B1-B 18 The identifiers are provided in Table 3; Equation III is expressed as: C1-C2-C3-C4-C5-C6-C7-C8-C9-C 10 -C 11 -C 12 -C 13 -C 14 -C 15 -C 16 -C 17 -C 18 -C 19 (Formula III) And C1-C 19 The identifiers are provided in Table 4; Where equation IV is represented as: D1-D2-D3-D4-D5-D6-D7-D8-D9-D 10 -D 11 -D 12 -D 13 -D 14 -D 15 -D 16 -D 17 -D 18 -D 19 -D 20 (Formula IV) And D1-D 20 The identifiers are provided in Table 5; Where V is expressed as: E1-E2-E3-E4-E5-E6-E7-E8-E9-E 10 -E 11 -E 12 -E 13 -E 14 -E 15 -E 16 -E 17 -E 18 -E 19 -E 20 -E 21 -E 22 (Formula V) And E1-E 22 The identifiers are provided in Table 6; Wherein, VI is represented as: F1-F2-F3-F4-F5-F6-F7-F8-F9-F 10 -F 11 -F 12 -F 13 -F 14 -F 15 -F 16 -F 17 -F 18 -F 19 -F 20 -F 21 -F 22 -F 23 (Form VI) And F1-F 23 The identifiers are provided in Table 7; Equation VII is expressed as: G1-G2-G3-G4-G5-G6-G7-G8-G9-G 10 -G 11 -G 12 -G 13 -G 14 -G 15 -G 16 -G 17 -G 18 -G 19 -G 20 -G 21 -G 22 -G 23 -G 24 -G 25 -G 26 (Equation VII) And G1-G 26 The identifiers are provided in Table 8; and Where equation VIII is expressed as: H1-H2-H3-H4-H5-H6-H7-H8-H9-H 10 -H 11 -H 12 -H 13 -H 14 -H 15 -H 16 -H 17 -H 18 -H 19 -H 20 -H 21 -H 22 -H 23 -H 24 -H 25 -H 26 -H 27 -H 28 -H 29 -H 30 (Formula VIII) And H1-H 30 The identifier is provided in Table 9.

[0012] In some embodiments, the signal peptide X1 comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

[0013] In some implementations, the payload protein is a therapeutic peptide or protein.

[0014] In some embodiments, a nucleic acid molecule is provided. In some embodiments, the nucleic acid molecule encodes a signal peptide as provided herein.

[0015] In some embodiments, a nucleic acid molecule is provided. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide as provided herein.

[0016] In some embodiments, a vector is provided. In some embodiments, the vector comprises a nucleic acid molecule as provided herein. In some embodiments, the nucleic acid molecule encodes a signal peptide as provided herein. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide as provided herein.

[0017] In some embodiments, a cell is provided. In some embodiments, the cell comprises nucleic acid molecules as provided herein. In some embodiments, the cell comprises a vector as provided herein. In some embodiments, the nucleic acid molecule encodes a signal peptide as provided herein. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide as provided herein.

[0018] In some embodiments, a composition is provided. In some embodiments, the composition comprises a nucleic acid molecule as provided herein. In some embodiments, the composition comprises a vector as provided herein. In some embodiments, the nucleic acid molecule encodes a signal peptide as provided herein. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide as provided herein.

[0019] In some embodiments, a method is provided for treating a disease or condition in a subject in need. In some embodiments, the method includes administering an effective amount of a nucleic acid molecule, as provided herein, to the subject to treat the disease or condition. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide, as provided herein. In some embodiments, the disease or condition is a nuclear-related disease or condition of a cell. In some embodiments, the payload protein is a protein that can be used for gene editing, and the disease or condition is any disease or condition to which gene editing is beneficial.

[0020] In some embodiments, a method is provided for treating a nucleus-related disease or condition in a subject of need. In some embodiments, the method includes administering a vector to the subject, the vector comprising a nucleic acid molecule encoding a signal peptide fused to or linked to a payload protein, wherein the signal peptide is a signal peptide as provided herein, and the payload protein is a therapeutic peptide or protein that can be used for the disease or condition.

[0021] In some embodiments, a method is provided for treating a disease or condition in a subject of need. In some embodiments, the method includes administering to the subject a vector comprising a nucleic acid molecule encoding: i) a signal peptide fused to or linked to a payload protein, wherein the signal peptide is a signal peptide as provided herein, and wherein the payload protein is a protein that can be used for gene editing; and ii) at least one nucleic acid molecule that targets the payload protein to a gene of interest; wherein the signal peptide promotes a translocation of the gene-editable protein to the nucleus, wherein the payload protein edits the gene of interest, thereby treating the disease or condition.

[0022] In some embodiments, a method for editing target nucleic acids in cells is provided. In some embodiments, the method includes administering a nucleic acid molecule encoding: i) a signal peptide fused to or linked to a payload protein, wherein the signal peptide is a signal peptide as provided herein, and wherein the payload protein is a protein that can be used for gene editing; and ii) at least one nucleic acid molecule that targets the payload protein to a gene of interest; wherein the signal peptide promotes the translocation of the gene-editable protein to the nucleus, wherein the payload protein edits the gene of interest, thereby editing the target nucleic acid. Attached Figure Description

[0023] Figure 1A A schematic diagram of an mCherry containing an N-terminus signal peptide is provided, and its use for signal peptide screening via in vitro pDNA transfection is shown.

[0024] Figure 1B Representative images of mCherry tagged with a signal peptide co-localized with Hoechst 33342 dye are shown. Top: SP is Nov-Nuc-16-1 (SEQ ID NO: 1). Bottom: SP is Nov-Nuc-22-1 (SEQ ID NO: 5).

[0025] Figure 1C Representative images of mCherry tagged with a signal peptide co-localized with Hoechst 33342 dye are shown. Top: SP is Nov-Nuc-18-1 (SEQ ID NO: 9). Bottom: SP is Nov-Nuc-26-1 (SEQ ID NO: 7).

[0026] Figure 1D Representative images of mCherry tagged with a signal peptide co-localized with Hoechst 33342 dye are shown. Top: SP is Nov-Nuc-19-1 (SEQ ID NO: 3). Bottom: No SP control.

[0027] Figure 1E Examples are given corresponding to Figure 1B , Figure 1C and Figure 1D Quantitative analysis of representative images from experiments.

[0028] Figure 1F Examples are given corresponding to Figure 1B , Figure 1C and Figure 1D Quantitative analysis of representative images from experiments.

[0029] Figure 2AQuantification of mCherry fluorescence in HeLa cell lysates and culture medium at different time points was provided. mCherry without the signal peptide was used as a control (“WT-mCherry”). gLuc, Gaussian luciferase.

[0030] Figure 2B The results provided quantification of mCherry fluorescence in HeLa cell lysates and culture medium after 72 hours.

[0031] Figure 2C Quantification of mCherry fluorescence in Huh7 cell lysates and culture medium was provided 72 h post-transfection.

[0032] Figure 3A We provided quantification of time-dependent (100 ng mRNA / well) and dose-dependent (at 72 h) mCherry secretion in Huh7 cells, culture medium, and cell lysates.

[0033] Figure 3B The quantification of mCherry signaling in different cell lines is shown. Cells in 96-well plates were treated with mDLNP-mRNA at given time points. Detailed Implementation

[0034] Before describing the compositions and methods of the present invention, it should be understood that the scope of this disclosure is not limited to the specific processes, compositions, or methodologies described herein, as these are subject to variation. It should also be understood that the terminology used herein is for the purpose of describing a particular version or embodiment only and is not intended to limit the scope of this disclosure. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. It should be further understood that any methods and materials similar to or equivalent to those described and materials herein can be used in the practice or testing of embodiments of the methods and systems disclosed herein, and such equivalents are within the scope of this disclosure.

[0035] definition The following explanations of terms and methods are provided to better describe this disclosure and to guide those skilled in the art in practicing this disclosure.

[0036] As used herein, unless the context explicitly indicates otherwise, "comprising" means "including" and the singular forms "a," "an," or "the" include a plural number of indicators. For example, references to "comprising a therapeutic agent" include one or more such therapeutic agents. Unless the context explicitly indicates otherwise, the term "or" refers to a single element among the stated alternative elements. For example, the phrase "A or B" means A alone or B alone. The phrase "A, B, or a combination thereof" means A alone, B alone, or a combination of A and B. Similarly, "one or more of A and B" means A, B, or a combination of both A and B. The phrase "A and B" means a combination of A and B. Furthermore, the various elements, features, and steps discussed herein, as well as other known equivalents of each such element, feature, or step, can be mixed and matched by those skilled in the art to perform methods according to the principles described herein. Among the various elements, features, and steps, some will be specifically included in a particular instance, while others will be specifically excluded from a particular instance.

[0037] In some instances, figures used to describe and claim certain embodiments, representing the amount of an ingredient, properties such as molecular weight, reaction conditions, etc., should be understood to be modified in some cases by the terms “about” or “approximately.” For example, “about” or “approximately” can indicate a variation of + / - 10%, + / - 5%, or + / - 1% in the value it describes. Thus, in some embodiments, the numerical parameters set forth herein are approximate values ​​that can vary depending on the desired characteristics of a particular embodiment. Furthermore, unless the context indicates otherwise, when the phrase enumerates “about x to y,” the term “about” modifies both x and y and can be used interchangeably with the phrase “about x to about y.”

[0038] As used herein, the terms “transformation” and “transfection” are intended to refer to a variety of art-recognized techniques for introducing exogenous nucleic acids into host cells, including calcium phosphate or calcium chloride coprecipitation, DEAE-dextran-mediated transfection, lipid transfection (e.g., using commercially available reagents such as, for example, LIPOFECTIN® (Invitrogen Corp., San Diego, CA), LIPOFECTAMINE® (Invitrogen), FUGENE® (Roche Applied Science, Basel, Switzerland), JETPEI™ (Polyplus-transfection Inc., New York, NY), EFFECTENE® (Qiagen, Valencia, CA), DREAMFECT™ (OZ Biosciences, France)), or electroporation (e.g., in vivo electroporation). Suitable methods for transforming or transfecting host cells can be found in Sambrook et al. ( Molecular Cloning: A Laboratory Manual. (2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989) and other laboratory manuals.

[0039] Methods and materials for non-viral delivery of nucleic acids to cells also include biolistics, virions, liposomes, lipid nanoparticles, immunoliposomes, polycationic or lipid-nucleic acid conjugates, naked DNA, artificial viral particles, and drug-enhanced DNA uptake. Lipid transfection is described in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipid transfection reagents are commercially available (e.g., TRANSFECTAM™ and LIPOFECTIN™). Cationic and neutral lipids suitable for effective receptor recognition of polynucleotides in lipid transfection include those disclosed in WO 91 / 17424 and WO 91 / 16024.

[0040] The chemical composition of peptides will be described herein by a series of single-letter abbreviations of amino acids, or “amino acid sequence” or “sequence”, which are conventional and known to those skilled in the art. Although a reference sequence will be explicitly disclosed, it may be modified in any aspect and embodiment to include conserved amino acid substitutions, as well as variants and fragments, while maintaining the characteristics and functionality of the reference sequence.

[0041] As used herein, a "payload protein" or "protein of interest" refers to a protein that will be generated by the host cell and chaperoned to a desired intracellular location (e.g., chaperoned to the nucleus). After translocation to the appropriate intracellular space, all, some, or no novel signal peptides may be fused to the payload protein. Optionally, the payload protein, still partially or completely linked to the novel signal peptide, may be further processed, for example, to remove any remaining novel signal peptides. The payload protein can be any protein known or yet to be known, such as enzymes, enzyme inhibitors, growth factors, hormones, antibodies, antigens, vaccines, therapeutics, or any combination thereof. More specific examples are given below.

[0042] As used herein, the terms “fused” or “linked” when referring to proteins with different domains or heterologous sequences mean that the protein domains are parts of the same peptide chain linked together by peptide bonds or other covalent bonds. Domains or segments may be directly linked or fused to each other, or another domain or peptide sequence may be between two domains or sequences, and such sequences will still be considered fused or linked together. In some embodiments, the various domains or proteins provided herein are directly linked or fused to each other or with adapter sequences, such as the glycine / serine sequence described herein that links two domains together.

[0043] As used herein, “identity” refers to the subunit sequence identity between two polymeric molecules, such as two nucleic acid or amino acid molecules, or between two polynucleotide or polypeptide molecules. Two amino acid sequences are identical when they have the same residue at the same position; for example, if each of two polypeptide molecules has an arginine residue at that position, then they are identical at that position. The degree of identity, or similarity, between two amino acid or two nucleic acid sequences having the same residue at the same position in an alignment is usually expressed as a percentage. Identity between two amino acid or two nucleic acid sequences is a direct function of the number of matching or identical positions; for example, if half of the positions in two sequences are identical, then the two sequences have 50% identity; if 90% of the positions (e.g., 9 out of 10) match or are identical, then the two amino acid sequences have 90% identity.

[0044] "Substantially identical" means that the polypeptide or nucleic acid molecule has at least 50% identity with a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). In some embodiments, this sequence has at least 60%, 80%, or 85%, or 90%, 95%, or even 99% identity with the sequence used for comparison at the amino acid level or nucleic acid level. Other percentages of identity with respect to a particular sequence are described herein.

[0045] Sequence identity can be measured / determined using sequence analysis software (e.g., sequence analysis packages from the Genetics Computing Group of the University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, including BLAST, BESTFIT, GAP, BLAST-2, ALIGN, MEGALIGN (DNASTAR), CLUSTALW, CLUSTALOMEGA, MUSCLE, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by specifying degrees of homology for various substitutions, deletions, and / or other modifications. Conserved substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In exemplary methods for determining the degree of identity, a BLAST program can be used, where a probability score between e3 and e100 indicates closely related sequences. In some embodiments, sequence identity is determined using BLAST with default settings. In some implementations, Clustal Omega is used to determine sequence identity.

[0046] For the purposes of the embodiments provided herein, compositions comprising various proteins may be included, in some cases, containing amino acid sequences that have sequence identity with the amino acid sequences disclosed herein. Therefore, in some embodiments, depending on the specific sequence, the degree of sequence identity with the SEQ ID NO disclosed herein is preferably greater than 50% (e.g., 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher). In addition to these percentages, other percentages of identity are also provided herein. Identity between peptides can be determined using an affine gap search with parameters vacancy opening penalty – 12 and vacancy extension penalty = 1, as implemented in the MPSRCH (Oxford Molecular) program.

[0047] Compared to publicly disclosed proteins, these proteins may include one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.) conserved amino acid substitutions, i.e., one amino acid is replaced by another amino acid with the relevant side chain. Genetically encoded amino acids are generally classified into four families: (1) acidic, i.e., aspartic acid and glutamic acid; (2) basic, i.e., lysine, arginine, and histidine; (3) nonpolar, i.e., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, and tryptophan; and (4) uncharged polar, i.e., glycine, asparagine, glutamine, cysteine, serine, threonine, and tyrosine. Phenylalanine, tryptophan, and tyrosine are sometimes collectively classified as aromatic amino acids. Generally, the substitution of a single amino acid within these families does not have a significant impact on biological activity. Proteins may have one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.) single amino acid deletions relative to publicly disclosed protein sequences. In addition to the disclosed protein sequence, the protein may also include one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.) insertions (e.g., each of 1, 2, 3, 4, or 5 amino acids).

[0048] "Encoding" refers to the inherent property of a specific nucleotide sequence in a polynucleotide (such as a gene, cDNA, or mRNA) to serve as a template for the synthesis of other polymers and macromolecules in biological processes. These polymers and macromolecules have defined nucleotide sequences (i.e., rRNA, tRNA, and mRNA) or defined amino acid sequences, and the resulting biological properties. Therefore, if the transcription and translation of the mRNA corresponding to a gene produces a protein in a cell or other biological system, then that gene encodes a protein. Both the coding strand, whose nucleotide sequence is identical to the mRNA sequence and is typically provided in sequence listings, and the non-coding strand, which serves as a template for transcription of a gene or cDNA, can be referred to as encoding the protein or other product of that gene or cDNA.

[0049] As used in this article, the following abbreviations for common nucleic acid bases are used: "A" refers to adenosine, "C" refers to cytosine, "G" refers to guanosine, "T" refers to thymidine, and "U" refers to uridine.

[0050] Unless otherwise specified, "nucleotide sequence encoding an amino acid sequence" includes all nucleotide sequences that are degenerate versions of each other and encode the same amino acid sequence. A phrase encoding a protein or RNA nucleotide sequence may also include introns, provided that the nucleotide sequence encoding that protein may contain introns in some versions.

[0051] The term "oligonucleotide" generally refers to short polynucleotides. It should be understood that when a nucleotide sequence is represented by a DNA sequence (i.e., A, T, C, G), this also provides a corresponding RNA sequence (i.e., A, U, C, G), where "U" replaces "T".

[0052] As used herein, the term "polynucleotide" is defined as a nucleotide chain. Furthermore, nucleic acids are polymers of nucleotides. Therefore, as used herein, the terms "nucleic acid" or "nucleic acid molecule" and "polynucleotide" are interchangeable. As used herein, polynucleotides include, but are not limited to, all nucleic acid sequences obtained by any method available in the art and by synthetic means, including but not limited to recombinant methods, i.e., cloning nucleic acid sequences from recombinant libraries or cell genomes using cloning techniques such as PCR.

[0053] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably and refer to compounds comprising a plurality of amino acid residues covalently linked by peptide bonds. As used herein, the term refers to short chains, which are also commonly referred to in the art as, for example, peptides, oligopeptides, and oligomers; and long chains, which are commonly referred to in the art as proteins, which exist in many types. “Polypeptide” includes, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, variants of polypeptides, modified polypeptides, derivatives, analogs, fusion proteins, etc. Polypeptides include natural peptides, recombinant peptides, synthetic peptides, or combinations thereof.

[0054] As used herein, unless otherwise indicated, "antibody fragment" or "antigen-binding fragment" means an antigen-binding fragment of an antibody, i.e., an antibody fragment that retains the ability to specifically bind to the antigen to which the full-length antibody is bound, such as a fragment retaining one or more CDR regions. Examples of antibody-binding fragments include, but are not limited to, Fab, Fab', F(ab')2, and Fv fragments; bivalent antibodies; linear antibodies; single-chain antibody molecules, such as sc-Fv; nanobodies (single-domain antibodies); and multispecific antibodies formed from antibody fragments.

[0055] The "Fab fragment" contains a light chain and a heavy chain of C. H 1. Variable region. The heavy chain of a Fab molecule cannot form a disulfide bond with another heavy chain molecule.

[0056] The “Fc” region contains two C-cells that contain antibodies. H 2 and C H 3. Heavy chain segments of structural domains. Two heavy chain segments are connected by two or more disulfide bonds and by C... H The hydrophobic interactions of the three structural domains are combined.

[0057] A “Fab” fragment contains a portion or segment of a light chain and a heavy chain, the portion or segment containing VH Domain and C H 1. Structural domains and C H 1 and C H The region between the two structural domains allows interchain disulfide bonds to form between the two heavy chains of the two Fab' segments to form the F(ab')2 molecule.

[0058] The “F(ab')2 fragment” contains two light chains and two chains containing C. H 1 and C H The heavy chains, which constitute a portion of the constant region between the two heavy chains, allow interchain disulfide bonds to form between them. Therefore, the F(ab')2 segment consists of two Fab' segments bonded together by disulfide bonds between the two heavy chains.

[0059] The “Fv region” contains variable regions from both the heavy and light chains, but lacks constant regions.

[0060] The term "single-chain Fv" or "scFv" antibody refers to a V antibody containing an antibody. H and V L Antibody fragments containing domains, wherein these domains are present within a single polypeptide chain. Generally, Fv polypeptides also contain V... H With V L A peptide linker between the domains allows scFv to form the structure required for antigen binding. For a review of scFv, see Pluckthun (1994). Volume 113, edited by Rosenburg and Moore, Springer-Verlag, New York, pp. 269-315. See also International Patent Application Publication No. WO 88 / 01649 and U.S. Patent Nos. 4,946,778 and 5,260,203.

[0061] Antibody molecules can be monospecific (e.g., monovalent or bivalent), bispecific (e.g., bivalent, trivalent, tetravalent, pentavalent, or hexavalent), trispecific (e.g., trivalent, tetravalent, pentavalent, or hexavalent), or have a higher order of specificity (e.g., tetraspecific) and / or a higher order of valence than hexavalent. Antibody molecules can contain functional fragments of both the light chain variable region and the heavy chain variable region, or the heavy and light chains can be fused together to form a single polypeptide.

[0062] As used herein, the terms “comprising” (and any form of inclusion, such as “comprise”, “comprises”, and “comprised”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unlisted elements or method steps. Any step or composition using the transitional phrase “comprise” or “comprising” can also be described using the transitional phrase “consisting of” or “consists”.

[0063] As used herein, the term “contact” means bringing together two elements of an in vitro or in vivo system. For example, “contacting” a virus or vector described herein with an individual or patient or cell includes administering the virus to an individual or patient (such as a human), and, for example, introducing a compound into a sample containing a cell preparation or a purified preparation containing such cells.

[0064] As used herein, the interchangeable terms “individual” or “subject” or “patient” mean any animal, including mammals such as mice, rats, other rodents, rabbits, dogs, cats, pigs, cattle, sheep, horses, or primates such as humans. In some embodiments, the subject is a human. A “subject in need” means a subject who has been identified as requiring treatment for a condition to be treated and is being treated with the specific intent to treat such condition. A condition can be any condition described herein, for example.

[0065] The compositions disclosed herein can be administered to a subject in a variety of ways by administering the composition to the subject. As used herein, administering or administering means providing or providing the composition to a subject. As used herein, oral administration means delivery of the active agent through the mouth. As used herein, topical administration means delivery of the active agent to the body surface, such as the skin, mucous membranes (e.g., nasal membranes, vaginal membranes, buccal membranes, etc.).

[0066] "Disease" is a health condition in which an animal is unable to maintain homeostasis, and if the disease is not treated, the animal's health continues to deteriorate. In contrast, an animal's "symptom" is a health condition in which the animal is able to maintain homeostasis, but the animal's health is not as good as it would be in the absence of the symptom. A symptom, if left untreated, does not necessarily lead to further deterioration of the animal's health.

[0067] The terms "effective amount" or "therapeutic effective amount" are used interchangeably herein and refer to the amount of a compound, formulation, material, or composition as described herein that effectively achieves a particular biological outcome or provides a therapeutic or preventative benefit. Such outcomes may include, but are not limited to, the amount of immune cell activation induced at a detectable level when administered to a mammal, compared to immune cell activation detected in the absence of the composition. Immune responses can be readily assessed using a wide range of methods recognized in the art. Those skilled in the art will understand that variations in the amount of the composition administered herein can be readily determined based on a number of factors, such as the disease or symptom being treated, the age and health status and physical condition of the mammal being treated, the severity of the disease, and the specific compound administered.

[0068] Scope: Throughout this disclosure, various aspects of the embodiments may be presented in the form of scope. It should be understood that the description in scope form is merely for convenience and brevity and should not be construed as an immutable limitation. Therefore, the description of scope should be considered as having specifically disclosed all possible sub-scopes and individual numerical values ​​within that scope. For example, a description of a scope such as 1 to 6 should be considered as having specifically disclosed sub-scopes such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and individual numbers within that scope, such as 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the width of the scope. Unless otherwise explicitly stated to the contrary, the disclosed scope also includes the endpoints of the scope.

[0069] Novel signal peptides Unbound by any particular theory, it has been found that the embodiments described herein demonstrate that novel signal peptides can guide the expression of payload proteins, as presented herein, in the cell nucleus. The novel signal peptides provided herein can guide payload proteins to any suitable location on or within the cell nucleus. The payload proteins may or may not naturally contain a signal peptide that guides the protein to a specific cellular compartment (e.g., the nucleus). Therefore, the novel signal peptides provided can guide the enhanced expression of native nuclear proteins in the nucleus, or can guide the expression of payload proteins (e.g., therapeutic proteins) to the nucleus for expression.

[0070] In some embodiments, a signal peptide is provided. In some embodiments, the signal peptide comprises the amino acid sequences listed in Table 1 below: Table 1: Exemplary Novel Signal Peptides

[0071] In some embodiments, the signal peptide comprises an amino acid sequence substantially similar to the amino acid sequences listed in Table 1. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequences listed in Table 1. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequences of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 1. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 2. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 3. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 4. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 5. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 6.In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 7. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with the amino acid sequence of SEQ ID NO: 9.

[0072] In some embodiments, the signal peptide is a variant of the amino acid sequence provided in Table 1. In some embodiments, the signal peptide comprises the amino acid sequence defined by Formulas I-VIII below.

[0073] Variants of SEQ ID NO: 1 (Formula I) In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: A1-A2-A3-A4-A5-A6-A7-A8-A9-A 10 -A 11 -A 12 -A 13 -A 14 -A 15 -A 16 (Formula I); Where A1–A 16 The identifiers for each of them are provided in Table 2 below.

[0074] Table 2

[0075] In some embodiments, A1 is methionine (M). In some embodiments, A2 is an amino acid selected from the group consisting of K, R, and V. In some embodiments, A3 is an amino acid selected from the group consisting of I, W, and T. In some embodiments, A4 is an amino acid selected from the group consisting of S, L, and I. In some embodiments, A5 is an amino acid selected from the group consisting of I, W, and A. In some embodiments, A6 is an amino acid selected from the group consisting of L and F. In some embodiments, A7 is leucine (L). In some embodiments, A8 is an amino acid selected from the group consisting of M, L, and R. In some embodiments, A9 is an amino acid selected from the group consisting of F and L. In some embodiments, A... 10 These are amino acids selected from the group consisting of L, P, and I. In some implementations, A 11 These are amino acids selected from the group consisting of W, D, and T. In some implementations, A 12 These are amino acids selected from the group consisting of G, S, and T. In some implementations, A 13 These are amino acids selected from the group consisting of L and P. In some implementations, A 14 These are amino acids selected from the group consisting of V and S. In some implementations, A 15 These are amino acids selected from the group consisting of C, T, and K. In some implementations, A... 16 These are amino acids selected from the group consisting of A, L, and G.

[0076] It should be understood that variants of SEQ ID NO: 1 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula I. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 2 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 3 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 4 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 5 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 6 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 8 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 9 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 10 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 11 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 12 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 13 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 14 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 15 of SEQ ID NO: 1. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at position 16 of SEQ ID NO: 1. Variants may also comprise any number of substitutions as provided in Formula I. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and a substitution at at least one position corresponding to position 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15 or 16 of SEQ ID NO: 1.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 1 and substitutions at more than one position corresponding to position 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15 or 16 of SEQ ID NO: 1.

[0077] Variant (Formula II) of SEQ ID NO: 2 In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: B1-B2-B3-B4-B5-B6-B7-B8-B9-B 10 -B 11 -B 12 -B 13 -B 14 -B 15 -B 16 -B 17 -B 18 (Formula II); Among them, B1-B 18 The identifiers for each of them are provided in Table 3 below.

[0078] Table 3

[0079] In some embodiments, B1 is methionine (M). In some embodiments, B2 is an amino acid selected from the group consisting of K and S. In some embodiments, B3 is an amino acid selected from the group consisting of F and I. In some embodiments, B4 is an amino acid selected from the group consisting of L and S. In some embodiments, B5 is an amino acid selected from the group consisting of H and L. In some embodiments, B6 is an amino acid selected from the group consisting of W and S. In some embodiments, B7 is an amino acid selected from the group consisting of L and S. In some embodiments, B8 is an amino acid selected from the group consisting of M and L. In some embodiments, B9 is an amino acid selected from the group consisting of S and I. In some embodiments, B... 10 These are amino acids selected from the group consisting of V and L. In some implementations, B... 11 These are amino acids selected from the group consisting of Y and L. In some implementations, B... 12 These are amino acids selected from the group consisting of V and P. In some implementations, B... 13 These are amino acids selected from the group consisting of V and I. In some implementations, B... 14 These are amino acids selected from the group consisting of E and W. In some implementations, B... 15 These are amino acids selected from the group consisting of L and I. In some implementations, B... 16These are amino acids selected from the group consisting of L and N. In some implementations, B... 17 These are amino acids selected from the group consisting of R and M. In some embodiments, B... 18 It is an amino acid group selected from those composed of S and A.

[0080] It should be understood that variants of SEQ ID NO: 2 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula II. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 2 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 3 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 4 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 5 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 6 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 7 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 8 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 9 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 10 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 11 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 12 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 13 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 14 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 15 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 16 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 17 of SEQ ID NO: 2.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at position 18 of SEQ ID NO: 2. Variations may also comprise any number of substitutions as provided in Formula II. Thus, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and a substitution at at least one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of SEQ ID NO: 2. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 2 and substitutions at more than one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of SEQ ID NO: 2.

[0081] In some non-limiting embodiments, variants of SEQ ID NO: 2 comprise the amino acid sequence of SEQ ID NO: 9. In some embodiments, the amino acid sequence of SEQ ID NO: 9 is derived from Formula II as follows: B1 is methionine (M), B2 is lysine (K), B3 is phenylalanine (F), B4 is leucine (L), B5 is histidine (H), B6 ​​is tryptophan (W), B7 is leucine (L), B8 is methionine (M), B9 is serine (S), B... 10 It is valine (V), B 11 It is tyrosine (Y), B 12 It is valine (V), B 13 It is valine (V), B 14 It is glutamic acid (E), B 15 It is leucine (L), B 16 It is leucine (L), B 17 It is arginine (R), and B 18 It is serine (S).

[0082] Variant (Formula III) of SEQ ID NO: 3 In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: C1-C2-C3-C4-C5-C6-C7-C8-C9-C 10 -C 11 -C 12 -C 13 -C 14 -C 15 -C 16 -C 17 -C 18 -C 19 (Formula III); Where C1-C19 The identifiers for each of them are provided in Table 4 below.

[0083] Table 4

[0084] In some embodiments, C1 is methionine (M). In some embodiments, C2 is an amino acid selected from the group consisting of A, R, T, and W. In some embodiments, C3 is an amino acid selected from the group consisting of A, M, L, and T. In some embodiments, C4 is an amino acid selected from the group consisting of Q, M, G, and L. In some embodiments, C5 is an amino acid selected from the group consisting of A, V, and K. In some embodiments, C6 is an amino acid selected from the group consisting of A, V, and S. In some embodiments, C7 is an amino acid selected from the group consisting of S, A, and G. In some embodiments, C8 is an amino acid selected from the group consisting of A and L. In some embodiments, C9 is an amino acid selected from the group consisting of V, A, and F. In some embodiments, C... 10 These are amino acids selected from the group consisting of Q, H, S, and L. In some implementations, C... 11 These are amino acids selected from the group consisting of G, A, and L. In some embodiments, C... 12 These are amino acids selected from the group consisting of L, A, and S. In some embodiments, C... 13 These are amino acids selected from the group consisting of A, F, W, and C. In some implementations, C... 14 These are amino acids selected from the group consisting of A, T, G, and L. In some implementations, C... 15 These are amino acids selected from the group consisting of Q, A, G, and T. In some implementations, C... 16 These are amino acids selected from the group consisting of C, A, and S. In some embodiments, C... 17 These are amino acids selected from the group consisting of A and S. In some embodiments, C... 18 These are amino acids selected from the group consisting of Q, A, L, and Y. In some implementations, C... 19 It is an amino acid group selected from A and P.

[0085] It should be understood that variants of SEQ ID NO: 3 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula III. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 2 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 3 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 4 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 5 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 6 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 7 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 8 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 9 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 10 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 11 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 12 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 13 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 14 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 15 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 16 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 17 of SEQ ID NO: 3.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 18 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at position 19 of SEQ ID NO: 3. Variants may also comprise any number of substitutions as provided in Formula III. Thus, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and a substitution at at least one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 of SEQ ID NO: 3. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 3 and substitutions at more than one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 of SEQ ID NO: 3.

[0086] Variant of SEQ ID NO: 4 (Formula IV) In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: D1-D2-D3-D4-D5-D6-D7-D8-D9-D 10 -D 11 -D 12 -D 13 -D 14 -D 15 -D 16 -D 17 -D 18 -D 19 -D 20 (Form IV); Where D1-D 20 The identifiers for each of them are provided in Table 5 below.

[0087] Table 5

[0088] In some embodiments, D1 is methionine (M). In some embodiments, D2 is serine (S). In some embodiments, D3 is an amino acid selected from the group consisting of K and R. In some embodiments, D4 is an amino acid selected from the group consisting of C and E. In some embodiments, D5 is leucine (L). In some embodiments, D6 is an amino acid selected from the group consisting of P and A. In some embodiments, D7 is an amino acid selected from the group consisting of A and P. In some embodiments, D8 is an amino acid selected from the group consisting of V and L. In some embodiments, D9 is an amino acid selected from the group consisting of F and L. In some embodiments, D... 10 It is leucine (L). In some implementations, D 11 These are amino acids selected from the group consisting of A and L. In some implementations, D... 12 These are amino acids selected from the group consisting of H and L. In some embodiments, D... 13 These are amino acids selected from the group consisting of W and L. In some implementations, D... 14 These are amino acids selected from the group consisting of V and S. In some implementations, D... 15 These are amino acids selected from the group consisting of L and I. In some embodiments, D... 16 These are amino acids selected from the group consisting of L and H. In some embodiments, D... 17 These are amino acids selected from the group consisting of L and S. In some embodiments, D... 18 These are amino acids selected from the group consisting of V and A. In some implementations, D... 19 It is leucine (L). In some implementations, D 20 It is an amino acid selected from the group composed of R and A.

[0089] It should be understood that variants of SEQ ID NO: 4 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula IV. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 3 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 4 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 6 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 7 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 8 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 9 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 11 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 12 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 13 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 14 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 15 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 16 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 17 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 18 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and a substitution at position 20 of SEQ ID NO: 4. Variations may also include any number of substitutions as provided in Formula IV.Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and substitutions at at least one position corresponding to position 3, 4, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, or 20 of SEQ ID NO: 4. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 4 and substitutions at more than one position corresponding to position 3, 4, 6, 7, 8, 9, 11, 12, 13, 14, 15, 16, 17, 18, or 20 of SEQ ID NO: 4.

[0090] Variant (Formula V) of SEQ ID NO: 5 In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: E1-E2-E3-E4-E5-E6-E7-E8-E9-E 10 -E 11 -E 12 -E 13 -E 14 -E 15 -E 16 -E 17 -E 18 -E 19 -E 20 -E 21 -E 22 (Formula V); Where E1-E 22 The identifiers for each of them are provided in Table 6 below.

[0091] Table 6

[0092] In some embodiments, E1 is methionine (M). In some embodiments, E2 is an amino acid selected from the group consisting of D, A, F, and P. In some embodiments, E3 is an amino acid selected from the group consisting of L, P, and G. In some embodiments, E4 is an amino acid selected from the group consisting of P, L, E, and F. In some embodiments, E5 is an amino acid selected from the group consisting of C, L, and Q. In some embodiments, E6 is an amino acid selected from the group consisting of C, F, A, Y, and L. In some embodiments, E7 is an amino acid selected from the group consisting of L, V, and Q. In some embodiments, E8 is an amino acid selected from the group consisting of W, A, L, N, and T. In some embodiments, E9 is an amino acid selected from the group consisting of L, G, F, and C. In some embodiments, E... 10 These are amino acids selected from the group consisting of V, L, and P. In some implementations, E... 11These are amino acids selected from the group consisting of F, V, A, and L. In some implementations, E... 12 These are amino acids selected from the group consisting of F, V, L, Y, and I. In some implementations, E... 13 These are amino acids selected from the group consisting of F, L, C, and T. In some implementations, E... 14 These are amino acids selected from the group consisting of S, L, A, V, and G. In some implementations, E... 15 These are amino acids selected from the group consisting of C, L, V, F, and T. In some implementations, E... 16 These are amino acids selected from the group consisting of L, S, and W. In some implementations, E... 17 These are amino acids selected from the group consisting of V, S, G, and L. In some implementations, E... 18 These are amino acids selected from the group consisting of P, M, D, H, and S. In some embodiments, E... 19 These are amino acids selected from the group consisting of P, A, G, and W. In some implementations, E... 20 These are amino acids selected from the group consisting of G, A, R, T, and V. In some embodiments, E... 21 These are amino acids selected from the group consisting of H, T, Q, N, and A. In some implementations, E... 22 It is an amino acid group selected from G, P, C and L.

[0093] It should be understood that variants of SEQ ID NO: 5 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the range of Formula V. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 2 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 3 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 4 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 5 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 6 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 7 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 8 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 9 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 10 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 11 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 12 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 13 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 14 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 15 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 16 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 17 of SEQ ID NO: 5.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 18 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 19 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 20 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 21 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and a substitution at position 22 of SEQ ID NO: 5. Variations may also comprise any number of substitutions as provided in Formula V. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and substitutions at at least one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 of SEQ ID NO: 5. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 5 and substitutions at more than one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 of SEQ ID NO: 5.

[0094] Variant (Formula VI) of SEQ ID NO: 6 In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: F1-F2-F3-F4-F5-F6-F7-F8-F9-F 10 -F 11 -F 12 -F 13 -F 14 -F 15 -F 16 -F 17 -F 18 -F 19 -F 20 -F 21 -F 22 -F 23 (Formula VI); Among them F1-F 23 The identifiers for each of them are provided in Table 7 below.

[0095] Table 7

[0096] In some embodiments, F1 is methionine (M). In some embodiments, F2 is an amino acid selected from the group consisting of A and P. In some embodiments, F3 is an amino acid selected from the group consisting of G and S. In some embodiments, F4 is glycine (G). In some embodiments, F5 is an amino acid selected from the group consisting of R and T. In some embodiments, F6 is an amino acid selected from the group consisting of C and G. In some embodiments, F7 is an amino acid selected from the group consisting of G and K. In some embodiments, F8 is an amino acid selected from the group consisting of P and T. In some embodiments, F9 is an amino acid selected from the group consisting of Q and V. In some embodiments, F... 10 These are amino acids selected from the group consisting of L and S. In some implementations, F... 11 These are amino acids selected from the group consisting of T and L. In some implementations, F... 12 These are amino acids selected from the group consisting of A and L. In some implementations, F... 13 These are amino acids selected from the group consisting of L and A. In some implementations, F... 14 It is leucine (L). In some implementations, F 15 These are amino acids selected from the group consisting of A and I. In some implementations, F... 16 These are amino acids selected from the group consisting of A and M. In some implementations, F... 17 These are amino acids selected from the group consisting of W and A. In some implementations, F... 18 These are amino acids selected from the group consisting of I and Y. In some implementations, F 19 These are amino acids selected from the group consisting of A and Q. In some implementations, F... 20 These are amino acids selected from the group consisting of A and R. In some embodiments, F... 21 These are amino acids selected from the group consisting of V and A. In some implementations, F... 22 These are amino acids selected from the group consisting of A and Y. In some implementations, F 23 It is an amino acid group selected from A and P.

[0097] It should be understood that variants of SEQ ID NO: 6 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula VI. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 2 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 3 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 5 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 6 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 7 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 8 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 9 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 10 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 11 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 12 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 13 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 15 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 16 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 17 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 18 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 19 of SEQ ID NO: 6.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 20 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 21 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 22 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at position 23 of SEQ ID NO: 6. Variants may also comprise any number of substitutions as provided in Formula VI. Thus, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and a substitution at at least one position corresponding to position 2, 3, 5, 6, 7, 8, 9, 10, 11, 12, 13, 15, 16, 17, 18, 19, 20, 21, 22, or 23 of SEQ ID NO: 6. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 6 and substitutions at more than one position corresponding to SEQ ID NO: 6 at positions 2, 3, 5, 6, 7, 8, 9, 10, 11, 12, 13, 15, 16, 17, 18, 19, 20, 21, 22 or 23.

[0098] Variant of SEQ ID NO: 7 (Formula VII) In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: G1-G2-G3-G4-G5-G6-G7-G8-G9-G 10 -G 11 -G 12 -G 13 -G 14 -G 15 -G 16 -G 17 -G 18 -G 19 -G 20 -G 21 -G 22 -G 23 -G 24 -G 25 -G 26 (Equation VII); Among them G1-G 26 The identifiers for each of them are provided in Table 8 below.

[0099] Table 8

[0100] In some embodiments, G1 is methionine (M). In some embodiments, G2 is an amino acid selected from the group consisting of A and L. In some embodiments, G3 is an amino acid selected from the group consisting of A, V, and M. In some embodiments, G4 is an amino acid selected from the group consisting of R, L, and K. In some embodiments, G5 is an amino acid selected from the group consisting of G, F, and R. In some embodiments, G6 is an amino acid selected from the group consisting of R, P, and L. In some embodiments, G7 is an amino acid selected from the group consisting of G, L, and A. In some embodiments, G8 is an amino acid selected from the group consisting of L, S, and A. In some embodiments, G9 is an amino acid selected from the group consisting of L, S, and R. In some embodiments, G... 10 These are amino acids selected from the group consisting of C and L. In some implementations, G... 11 These are amino acids selected from the group consisting of L, G, and F. In some implementations, G... 12 These are amino acids selected from the group consisting of T, W, and A. In some implementations, G... 13 These are amino acids selected from the group consisting of L, A, and G. In some embodiments, G... 14 These are amino acids selected from the group consisting of S, A, and L. In some implementations, G... 15 These are amino acids selected from the group consisting of V, I, and L. In some implementations, G... 16 These are amino acids selected from the group consisting of L, S, and I. In some implementations, G... 17 It is leucine (L). In some implementations, G 18 These are amino acids selected from the group consisting of A, F, and S. In some implementations, G... 19 These are amino acids selected from the group consisting of A, L, and P. In some implementations, G... 20 These are amino acids selected from the group consisting of G, S, and L. In some implementations, G... 21 These are amino acids selected from the group consisting of P, A, and T. In some implementations, G... 22 These are amino acids selected from the group consisting of S, Q, and V. In some implementations, G... 23 These are amino acids selected from the group consisting of A, S, and I. In some implementations, G... 24 These are amino acids selected from the group consisting of A, C, and S. In some implementations, G... 25 These are amino acids selected from the group consisting of A, Y, and D. In some implementations, G... 26 It is an amino acid group selected from those composed of S and A.

[0101] It should be understood that variants of SEQ ID NO: 7 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula VII. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 2 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 3 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 4 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 5 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 6 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 7 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 8 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 9 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 10 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 11 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 12 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 13 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 14 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 15 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 16 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 18 of SEQ ID NO: 7.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 19 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 20 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 21 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 22 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 23 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 24 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 25 of SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and a substitution at position 26 of SEQ ID NO: 7. Variants may also contain any number of substitutions as provided in Formula VII. Thus, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and substitutions at at least one of the positions 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 22, 23, 24, 25, or 26 corresponding to SEQ ID NO: 7. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 7 and substitutions at more than one of the positions 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 18, 19, 20, 21, 22, 23, 24, 25, or 26 corresponding to SEQ ID NO: 7.

[0102] Variant of SEQ ID NO: 8 (Formula VIII) In some implementations, the signal peptide comprises an amino acid sequence represented by the following formula: H1-H2-H3-H4-H5-H6-H7-H8-H9-H 10 -H 11 -H 12 -H 13 -H 14 -H 15 -H 16 -H 17 -H 18 -H 19-H 20 -H 21 -H 22 -H 23 -H 24 -H 25 -H 26 -H 27 -H 28 -H 29 -H 30 (Formula VIII); Where H1-H 30 The identifiers for each of them are provided in Table 9 below.

[0103] Table 9

[0104] In some embodiments, H1 is methionine (M). In some embodiments, H2 is an amino acid selected from the group consisting of A and G. In some embodiments, H3 is an amino acid selected from the group consisting of P and A. In some embodiments, H4 is an amino acid selected from the group consisting of H and A. In some embodiments, H5 is an amino acid selected from the group consisting of D and G. In some embodiments, H6 is an amino acid selected from the group consisting of P and R. In some embodiments, H7 is an amino acid selected from the group consisting of G and Q. In some embodiments, H8 is an amino acid selected from the group consisting of S and D. In some embodiments, H9 is an amino acid selected from the group consisting of L and F. In some embodiments, H... 10 These are amino acids selected from the group consisting of T and L. In some embodiments, H... 11 These are amino acids selected from the group consisting of T and F. In some implementations, H... 12 These are amino acids selected from the group consisting of L and K. In some embodiments, H... 13 These are amino acids selected from the group consisting of V and A. In some embodiments, H... 14 These are amino acids selected from the group consisting of P and M. In some embodiments, H... 15 These are amino acids selected from the group consisting of W and L. In some embodiments, H... 16 These are amino acids selected from the group consisting of A and T. In some implementations, H... 17 These are amino acids selected from the group consisting of A and I. In some embodiments, H... 18 These are amino acids selected from the group consisting of A and S. In some embodiments, H... 19 These are amino acids selected from the group consisting of L and W. In some embodiments, H... 20 It is leucine (L). In some implementations, H... 21These are amino acids selected from the group consisting of L and T. In some embodiments, H... 22 These are amino acids selected from the group consisting of A and L. In some embodiments, H... 23 These are amino acids selected from the group consisting of L and T. In some embodiments, H... 24 These are amino acids selected from the group consisting of G and C. In some implementations, H... 25 These are amino acids selected from the group consisting of V and F. In some embodiments, H... 26 These are amino acids selected from the group consisting of E and P. In some embodiments, H... 27 These are amino acids selected from the group consisting of R and G. In some embodiments, H... 28 It is alanine (A). In some implementations, H... 29 These are amino acids selected from the group consisting of L and T. In some embodiments, H... 30 It is an amino acid group selected from A and S.

[0105] It should be understood that variants of SEQ ID NO: 8 do not need to contain a substitution at every position providing the substitution, but only need to contain one substitution to be within the scope of Formula VIII. Therefore, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 2 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 3 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 4 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 5 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 6 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 7 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 8 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 9 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 10 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 11 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 12 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 13 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 14 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 15 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 16 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 17 of SEQ ID NO: 8.In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 18 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 19 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 21 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 22 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 23 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 24 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 25 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 26 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 27 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 29 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at position 30 of SEQ ID NO: 8. Variants may also comprise any number of substitutions as provided in Formula VIII. Thus, in some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and a substitution at at least one position corresponding to position 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 29, or 30 of SEQ ID NO: 8. In some embodiments, the signal peptide comprises the amino acid sequence of SEQ ID NO: 8 and substitutions at more than one position corresponding to SEQ ID NO: 8, namely positions 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 29, or 30.

[0106] In some embodiments, the signal peptide is a novel signal peptide. In the context of this disclosure, a "novel" signal peptide should be understood as containing an amino acid sequence that has not been found in nature or is not yet known to exist in nature. Therefore, if an embodiment refers to a novel signal peptide, it should be understood to exclude any corresponding natural sequence. For example, in some embodiments, a novel signal peptide is provided, wherein the novel signal peptide comprises an amino acid sequence corresponding to Formula V. In such embodiments, if a natural signal peptide can be interpreted by Formula V, the embodiment is understood to exclude said natural signal peptide.

[0107] In some embodiments, a novel signal peptide is provided. In some embodiments, the novel signal peptide comprises an amino acid sequence selected from the group consisting of Formula I, Formula II, Formula III, Formula IV, Formula V, Formula VI, Formula VII, or Formula VIII. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula I. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula II. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula III. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula IV. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula I. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula VI. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula VII. In some embodiments, the novel signal peptide comprises an amino acid sequence of Formula VIII.

[0108] In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity with the amino acid sequence selected from the group consisting of SEQ ID NO: 1. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 1. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 2. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 3. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 4. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 5. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 6. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 7. In some embodiments, the novel signal peptide comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 8.

[0109] Recombinant peptides In some embodiments, a recombinant polypeptide is provided comprising the formula X1-Z1, wherein X1 is a signal peptide as provided herein, and Z1 is a payload protein. In some embodiments, the recombinant polypeptide comprises the formula X1-Z1, wherein X1 is a novel signal peptide as provided herein, and Z1 is a payload protein.

[0110] In some embodiments, X1 comprises an amino acid sequence selected from the group consisting of Formula I, Formula II, Formula III, Formula IV, Formula V, Formula VI, Formula VII, or Formula VIII. In some embodiments, X1 comprises an amino acid sequence of Formula I. In some embodiments, X1 comprises an amino acid sequence of Formula II. In some embodiments, X1 comprises an amino acid sequence of Formula III. In some embodiments, X1 comprises an amino acid sequence of Formula IV. In some embodiments, X1 comprises an amino acid sequence of Formula V. In some embodiments, X1 comprises an amino acid sequence of Formula VI. In some embodiments, X1 comprises an amino acid sequence of Formula VII. In some embodiments, X1 comprises an amino acid sequence of Formula VIII.

[0111] In some embodiments, X1 comprises an amino acid sequence having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 4, SEQ ID NO: 4, SEQ ID NO: 1 ... In some embodiments, X1 comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 5. In some embodiments, X1 comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 6. In some embodiments, X1 comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 7. In some embodiments, X1 comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, X1 comprises an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 9.

[0112] In some embodiments, payload protein Z1 can be any peptide or protein. In some embodiments, payload protein Z1 is any peptide or protein that can be used to treat a disease or condition. In some embodiments, payload protein Z1 is any peptide or protein that can be used to treat a disease or condition related to the nucleus of a cell. In some embodiments, payload protein Z1 is any peptide or protein that can be used for gene editing, such as, but not limited to, Cas9, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), or a wide range of nucleases.

[0113] In some embodiments, X1 is directly linked to Z1. In other embodiments, X1 is indirectly linked to Z1 via, for example, a peptide linker. Peptide linkers are known in the art, and any such linker can be incorporated into the recombinant peptides of this disclosure. Therefore, in some embodiments, the recombinant peptide can be of formula X1-(Y1). a -Z1 indicates that X1 is a signal peptide as provided herein, Y1 is a linker, such as but not limited to a peptide linker, Z1 is a payload protein as provided herein, and a is an integer selected from 0 or 1. In some embodiments, X1 is a novel signal peptide as provided herein.

[0114] Nucleic acid molecules In some embodiments, a nucleic acid molecule is provided. In some embodiments, the nucleic acid molecule encodes a signal peptide as provided herein. In some embodiments, the nucleic acid molecule encodes a novel signal peptide as provided herein. In some embodiments, the nucleic acid molecule is a deoxyribonucleotide sequence (DNA). In some embodiments, the nucleic acid molecule is a ribonucleic acid sequence (RNA).

[0115] In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence selected from the group consisting of Formula I, II, III, IV, V, VI, VII, or VIII. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula I. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula II. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula III. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula IV. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula V. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula VI. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula VII. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence of Formula VIII.

[0116] In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 1. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 2. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 3. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 4. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 5. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 6. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 7. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 8. In some embodiments, the nucleic acid molecule encodes a signal peptide (e.g., a novel signal peptide) comprising an amino acid sequence having at least 70% identity with the amino acid sequence of SEQ ID NO: 9.

[0117] In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide as provided herein. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide comprising formulas X1-Z1, wherein X1 is a signal peptide as provided herein, and Z1 is a payload protein as provided herein. In some embodiments, the recombinant polypeptide comprises formulas X1-Z1, wherein X1 is a novel signal peptide as provided herein, and Z1 is a payload protein as provided herein. In some embodiments, the nucleic acid molecule encodes a formula comprising X1-(Y1). a -Z1 is a recombinant polypeptide, wherein X1 is a signal peptide (e.g., a novel signal peptide) as provided herein, Y1 is a linker, such as, but not limited to, a polypeptide linker, Z1 is a payload protein as provided herein, and a is an integer selected from 0 or 1.

[0118] In embodiments where Z1 is a peptide or protein that can be used for gene editing, the nucleic acid molecule may also contain any additional necessary components for gene editing, such as guide RNA. Alternatively, additional necessary components for gene editing may be provided in one or more additional nucleic acid molecules.

[0119] Those skilled in the art will readily be able to deduce appropriate DNA or RNA sequences based on the amino acid sequences and variants provided herein. Furthermore, those skilled in the art will readily recognize that, due to codon degeneracy, multiple nucleic acid molecules can be used to encode the same amino acid sequence. Therefore, any nucleic acid sequence capable of encoding the amino acid sequences provided herein is within the scope of this disclosure.

[0120] In some embodiments, the nucleic acid molecules disclosed herein may further comprise one or more elements selected from, but not limited to, promoters, enhancers, leader sequences, transcription start sites (TSS), adapters, 5' and 3' untranslated regions (UTRs), Kozak sequences, introns, polyadenylation signals, cap sequences, enhancers, viral sequences, IRES sequences, or termination regions, or any element suitable for regulating or allowing the expression of the recombinant polypeptides disclosed herein in cells, or necessary for regulating or allowing such expression. By definition, the wild-type untranslated region (UTR) of a gene is transcribed but not translated. In mRNA, the 5' UTR begins at the transcription start site and continues to the start codon, but does not include the start codon; however, the 3' UTR begins immediately after the stop codon and continues until the transcription termination signal. Regulatory features of the UTR may be incorporated into the polynucleotides of this disclosure to, for example, enhance the stability of the molecule. Specific features may also be incorporated to ensure controlled downregulation of the transcript to prevent its misdirection to undesirable organ sites. In some embodiments, any suitable naturally occurring or synthetic UTR sequence may be incorporated into the nucleic acid molecules disclosed herein. Other non-UTR sequences may also be incorporated into the nucleic acid molecules. For example, introns or portions of intron sequences may be incorporated into regions of the nucleic acid molecules disclosed herein. Incorporation of intron sequences may increase protein yield and polynucleotide levels. Combinations of features may be included in flanking regions and may be contained within other features. For example, the flanking region of an ORF may be a 5' UTR that may contain a strong Kozak translation initiation signal and / or a 3' UTR that may include an oligo(dT) sequence for template addition of a poly-A tail. The 5' UTR may contain a first polynucleotide fragment and a second polynucleotide fragment from the same and / or different genes.

[0121] In some embodiments, the nucleic acid molecules disclosed herein can be assembled intracellularly. In some embodiments, the nucleic acid molecules can be synthesized in vivo. In some embodiments, the nucleic acid molecules can be synthesized in vitro using methods known in the art, such as in vitro transcription, DNA, RNA, and cDNA synthesis methods. In some embodiments, the nucleic acid molecules disclosed herein can be incorporated into suitable viral vectors, expression cassettes, expression vectors, transposons, extrachromosomal elements, and integrated into chromosomes, host cells, or delivery systems.

[0122] In some embodiments, the nucleic acid molecule is a chemically modified nucleic acid molecule. In some embodiments, the nucleic acid molecule may comprise one or more modified nucleosides containing modified sugar moieties. Such compounds containing one or more sugar-modified nucleosides may possess desired properties, such as enhanced nuclease stability, compared to oligonucleotides containing only nucleosides with naturally occurring sugar moieties. In some embodiments, the modified sugar moieties are substituted sugar moieties. In some embodiments, the modified sugar moieties are sugar substitutes. Such sugar substitutes may contain one or more substitutions corresponding to those substitutions in the substituted sugar moieties. In some embodiments, the modified nucleic acid molecule may comprise a modified backbone, such as a thiophosphate, a triphosphate, a morpholino, a methylphosphonate, short-chain alkyl or cycloalkyl sugar inter-linked or short-chain heteroatom or heterocyclic sugar inter-linked.

[0123] In some embodiments, the modified sugar moiety is a substituted sugar moiety comprising one or more non-bridged sugar substituents, including but not limited to substituents at the 2' and / or 5' positions. Examples of suitable sugar substituents at the 2' position include, but are not limited to, 2'-F, 2-OCH3 (“OMe” or “O-methyl”), and 2'-O(CH2)2OCH3 (“MOE”). In some aspects, the sugar substituent at the 2' position is selected from allyl, amino, azide, thio, O-allyl, O-C1-C10 alkyl, O-C1-C10 substituted alkyl; OCF3, O(CH2)2SCH3, O(CH2)2-O-N(Rm)(Rn), and O-CH2-C(=O)-N(Rm)(Rn), wherein each Rm and Rn is independently H or a substituted or unsubstituted C1-C10 alkyl group. Examples of sugar substituents at the 5'-position include, but are not limited to: 5'-methyl (R or S); 5'-vinyl and 5'-methoxy. In some embodiments, the substituted sugar comprises more than one non-bridged sugar substituent, such as the TF-5'-methyl sugar moiety.

[0124] Nucleosides containing a 2'-substituted sugar moiety are called 2'-substituted nucleosides. In some embodiments, the 2'-substituted nucleosides contain a 2'-substituent group selected from the following: halogroup, allyl, amino, azide, SH, CN, OCN, CF3, OCF3, O, S, or N(Rm)-alkyl; O, S, or N(Rm)-alkenyl; O, S, or N(Rm)-ynyl; O-alkylene-O-alkyl, ynyl, alkylaryl, aralkyl, O-alkylaryl, O-aralkyl, O(CH2)2SCH3, O(CH2)2—O—N(Rm)(Rn) or O—CH2—C(=O)—N(Rm)(Rn), wherein each Rm and Rn is independently H, an amino protecting group, or a substituted or unsubstituted C10 alkyl group. These 2'-substituent groups may be further substituted independently by one or more substituent groups selected from the following: hydroxyl, amino, alkoxy, carboxyl, benzyl, phenyl, nitro (NO2), thiol, thioalkoxy (S-alkyl), halogen, alkyl, aryl, alkenyl, and alkynyl.

[0125] In some embodiments, the 2'-substituted nucleoside comprises a 2'-substituent group selected from the following: F, NH2, N3OCF3, O—CH3, O(CH2)3NH2, CH2—CH=CH2, O—CH2—CH=CH2, OCH2CH2OCH3, O(CH2)2SCH3, O—(CH2)2—O—N(Rm)(Rn), O(CH2)2O(CH2)2N(CH3)2, and N-substituted acetamide (O—CH2—C(=O)—N(Rm)(Rn), wherein each Rm and R is independently H, an amino protecting group, or substituted or unsubstituted. The 2'-substituted nucleoside comprises a sugar moiety containing a 2'-substituent group selected from the following: F, OCF3, O-CH3, O2CH2OCH3, O(CH2)2SCH3, O(CH2)2-O-N(CH3)2, -O(CH2)2O(CH2)2N(CH3)2, and O-CH2-C(=O)-N(H)CH3. In some embodiments, the 2'-substituted nucleoside comprises a sugar moiety containing a 2'-substituent group selected from the following: F, O-CH3, and OCH2CH2OCH3.

[0126] Some modified sugar moieties contain bridging sugar substituents that form a second ring, thereby producing a bicyclic sugar moiety. In some such aspects, the bicyclic sugar moieties contain a bridge between the 4' furanose ring atom and the 2' furanose ring atom. Examples of such 4' to 2' sugar substituents include, but are not limited to: —[C(Ra)(Rb)]—, —[C(Ra)(Rb)]n—O—, —C(RaRb)—N(R)—O— or —C(RaRb)—O—N(R)—; 4'-CH2-2', 4'-(CH2)2-2', 4'-(CH2)—O-2' (LNA); 4'-(CH2)—S-2'; 4'-(CH2)2—O-2' (ENA); 4'-CH(CH3)—O-2' (cEt) and 4'-CH(CH2OCH3)—O-2' and their analogues (see, for example, U.S. Patent No. 7,399,845); 4'-C(CH3)(CH3)—O-2' and their analogues (see, for example, WO 4'-CH2—N(OCH3)-2' and its analogues (see, for example, WO2008 / 150729); 4'-CH2—O—N(CH3)-2' (see, for example, US2004 / 0171570, published September 2, 2004); 4'-CH2—O—N(R)-2' and 4-CH2—N(R)—0-2'-, wherein each R is independently H, a protecting group, or a C1-C12 alkyl group; 4-CH2—N(R)—0-2', wherein R is H, a C1-C12 alkyl group, or a protecting group (see US Patent No. 7,427,672); 4'-CH2—C(H)(CH3)-2' (see, for example, Chattopadhyaya et al., J. Org. Chern., 2009, 74, 118-134); and 4-CH2—C(=CH2)-2' and its analogues (see PCT International Application WO 2008 / 154401).

[0127] In some embodiments, such 4' to 2' bridges independently comprise 1 to 4 linked groups, which are independently selected from —[C(Ra)(Rb)]n—, —C(Ra)=C(Rb)—, —C(Ra)=N—, —C(=NRa)—, —C(=O)—, —C(=S)—, —O—, —Si(Ra)2—S(=O)x— and —N(Ra)—; wherein: x is 0, 1, or 2; n is 1, 2, 3, or 4; each Ra and Rb is independently H, a protecting group, a hydroxyl group, a C1-C12 alkyl group, a substituted C1-C12 alkyl group, a C2-C12 alkenyl group, a substituted C2-C12 alkenyl group, a C2-C12 ynyl group, a substituted C2-C12 ynyl group, a C5-C20 aryl group, a substituted C5-C20 aryl group, a heterocyclic group, a substituted heterocyclic group, or a substituted heterocyclic group. Cycloyl, heteroaryl, substituted heteroaryl, C5-C7 alicyclic, substituted C5-C7 alicyclic, halogen, OJ1, NJ1J2, SJ1, N3, C00J1, acyl (C(=O)—H), substituted acyl, CN, sulfonyl (S(=O)2-J1) or sulfoxide (S(=O)-J1); and each J1 and J2 is independently H, C1-C12 alkyl. Substituted C1-C12 alkyl, C2-C12 alkenyl, substituted C2-C12 alkenyl, C2-C12 alkynyl, substituted C2-C12 alkynyl, C5-C20 aryl, substituted C5-C20 aryl, acyl (C(=O)—H), substituted acyl, heterocyclic, substituted heterocyclic, 01-012 aminoalkyl, substituted C1-C12 aminoalkyl or protecting group.

[0128] Nucleosides containing a bicyclic sugar moiety are called bicyclic nucleosides or BNAs. Bicyclic nucleosides include, but are not limited to, (A) α-L-methyleneoxy(4-CH2—O-2') BNA, (B) β-D-methyleneoxy(4-CH2—O-2') BNA (also known as locked nucleic acid or LNA), (C) ethyleneoxy(4'-(CH2)2—O-2') BNA, (D) aminooxy(4'-CH2—O-N(R)-2') BNA, (E) oxyamino(4'-CH2—N(R)—O-2') BNA, (F) methyl(methyleneoxy)(4'-CH(CH3)—O-2') BNA (also known as restricted ethyl or cEt), (G) methylene-thio(4-CH2—S-2') BNA, (H) methyleneamino(4'-CH2—N(R)-2') BNA, and (I) methylcarbocyclic(4'-CH2—CH(CH3)-2') BNA, (J)propene carbocyclic (4'-(CH2)3-2') BNA, and (K)methoxy(ethyleneoxy) (4'-CH(CH2OMe)-O-2') BNA (also known as restricted MOE or cMOE). Additional bicyclic sugar moieties are known in the art, and any such bicyclic sugar moieties are within the scope of this disclosure.

[0129] In some embodiments, the bicyclic sugar moiety and the nucleoside incorporated into such a bicyclic sugar moiety are further defined by isomer configuration. For example, a nucleoside containing a 4'-2' methylene-oxygen bridge can be in the α-L or β-D configuration. α-L-methyleneoxy(4-CH2-O-2') bicyclic nucleosides have previously been incorporated into antisense polynucleotides exhibiting antisense activity.

[0130] In some embodiments, the substituted sugar moiety comprises one or more non-bridging sugar substituents and one or more bridging sugar substituents.

[0131] In some embodiments, the modified sugar moiety is a sugar substitute. In some embodiments, the oxygen atom of a naturally occurring sugar is replaced by, for example, a sulfur, carbon, or nitrogen atom. In some embodiments, such a modified sugar moiety also contains bridging substituents and / or non-bridging substituents as described above. For example, some sugar substitutes contain a 4'-sulfur atom and substitutions at the 2'- and / or 5' positions. As an additional example, carbocyclic bicyclic nucleosides with 4-2' bridges have been described.

[0132] In some embodiments, the sugar substitute comprises a ring having more than five atoms. For example, in some embodiments, the sugar substitute comprises a six-membered tetrahydropyran (THP). Such tetrahydropyrans may be further modified or substituted. Nucleosides containing such modified tetrahydropyrans include, but are not limited to, hexitol nucleic acid (HNA), anitol nucleic acid (ANA), mannitol nucleic acid (MNA), and fluoroHNA (F-HNA).

[0133] Many other bicyclic and tricyclic sugar substitute ring systems are also known in the art, which can be used to modify nucleosides for incorporation into antisense compounds, and any such sugar substitute ring system is within the scope of this disclosure.

[0134] Unrestricted combinations of modifications are also provided, such as, but not limited to, 2-F-5'-methyl-substituted nucleosides and ribosyl epoxy atoms replaced with S and further substituted at the 2'-position, or alternatively, 5'-substituted bicyclic nucleic acids. The synthesis and preparation of carbocyclic bicyclic nucleosides, along with their oligomerization and biochemical studies, have also been described.

[0135] In some embodiments, this disclosure provides nucleic acid molecules comprising modified nucleosides. Those modified nucleotides may include modified sugars, modified nucleobases, and / or modified linkages. Specific modifications are chosen such that the resulting polynucleotide possesses desired properties. In some embodiments, the nucleic acid molecule comprises one or more RNA-like nucleosides. In some embodiments, the nucleic acid molecule comprises one or more DNA-like nucleotides.

[0136] In some embodiments, the nucleosides of this disclosure comprise one or more unmodified nucleobases. In some embodiments, the nucleosides of this disclosure comprise one or more modified nucleobases.

[0137] In some embodiments, the modified nucleobases are selected from: universal bases as defined herein, hydrophobic bases, hybrid bases, size-expanded bases, and fluorinated bases. 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil; 5-propynylcytosine; 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, adenine, and guanine's 6-methyl and other alkyl derivatives, adenine and guanine's 2-propyl and other alkyl derivatives, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynylCH3, uracil and cytosine, and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine, and thymine. 5-Uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenine and guanine, 5-halogenated specifically 5-bromine, 5-trifluoromethyl and other 5-substituted uracil and cytosine, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deadenine and 7-deadenine, 3-deadenine and 3-deadenine, general bases, hydrophobic bases, mixed bases, size-expanded bases and fluorinated bases, as defined herein. Other modified nucleobases include tricyclic pyrimidines such as phenoxazincytidine ([5,4-b][1,4]benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido[5,4-b][1,4]benzothiazin-2(3H)-one), G-clasps such as substituted phenoxazincytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido[5,4-13][1,4]benzoxazin-2(3H)-one), carbazocytidine (2H-pyrimido[4,5-b]indole-2-one), and pyridoindolecytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimido-2-one). The modified nucleobases may also include those in which the purine or pyrimidine bases are replaced by other heterocycles such as 7-deadenine, 7-deadenine, 2-aminopyridine, and 2-pyridone.

[0138] In some embodiments, this disclosure provides nucleic acid molecules comprising linked nucleosides. In some embodiments, the nucleosides can be linked together using any nucleoside linker. Two main classes of linker groups are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing linker groups include, but are not limited to, phosphodiester (P=O), phosphotriester, methylphosphonate, aminophosphate, and thiophosphate (P=S). Representative non-phosphoside linker groups include, but are not limited to, methylenemethylimino (—CH2—N(CH3)—O—CH2—), thiodiester (—O—C(O)—S—), thionocarbamate (—O—C(O)(NH)—S—); siloxane (—O—Si(H)2—O—); and N,N'-dimethylhydrazine (—CH2—N(CH3)—N(CH3)—). Compared to natural phosphodiester linkers, modified linkers can be used to alter (typically increase) the nuclease resistance of polynucleotides. In some embodiments, internucleotide linkages with chiral atoms can be prepared as racemic mixtures or as individual enantiomers. Representative chiral linkages include, but are not limited to, alkyl phosphonates and thiophosphates. Methods for preparing phosphorus-containing and phosphorus-free internucleotide linkages are well known to those skilled in the art.

[0139] The nucleic acid molecules described herein may contain one or more asymmetric centers and thus produce enantiomers, diastereomers, and other stereoisomers that can be defined by absolute stereochemistry as (R) or (S), such as for glycoterminal isomers, or defined as (D) or (L), such as for amino acids. All of the possible isomers described herein, as well as their racemic and optically pure forms, are included in the antisense compounds provided herein.

[0140] Neutral nucleoside linkages include, but are not limited to, triphosphates, methylphosphonates, MMI (3-CH2—N(CH3)—O-5'), amide-3 (3-CH2—C(=O)—N(H)-5'), amide-4 (3'-CH2—N(H)—C(=O)-5'), methyl acetal (3'-O—CH2—O-5'), and thiomethyl ethyl acetal (3'-5-CH2—O-5'). Other neutral nucleoside linkages include nonionic linkages comprising siloxanes (dialkylsiloxanes), carboxylic esters, carboxamides, sulfides, sulfonates, and amides. Further neutral nucleoside linkages include nonionic linkages comprising a mixture of N, O, S, and CH2 components.

[0141] Additional modifications can also be made at other positions on the nucleic acid molecule, particularly at the 3' position of the sugar at the 3' end of the nucleotide and the 5' position of the 5' end of the nucleotide. For example, one additional modification of the nucleic acid molecule disclosed herein involves chemically linking one or more additional moieties or conjugates to a polynucleotide, which enhances the activity, cellular distribution, or cellular uptake of the polynucleotide. Such moieties include, but are not limited to, lipid moieties such as cholesterol moieties, bile acids, thioethers such as hexyl-5-triphenylmethylthiol, thiocholesterol, aliphatic chains such as dodecyl glycol or undecyl residues, phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycerol-3-H-phosphonate, polyamines or polyethylene glycol chains, or adamantaneacetic acid, palmityl moieties, or octadecylamine or hexano-carbonyl-oxycholesterol moieties.

[0142] In some embodiments, a vector is provided. In some embodiments, the vector comprises a nucleic acid molecule as provided herein. In some embodiments, the nucleic acid molecule encodes a signal peptide as provided herein. In some embodiments, the nucleic acid molecule encodes a novel signal peptide as provided herein. In some embodiments, the nucleic acid molecule encodes a recombinant polypeptide as provided herein.

[0143] cell In some embodiments, a cell is provided. In some embodiments, the cell comprises a nucleic acid molecule encoding a signal peptide as provided herein. In some embodiments, the cell comprises a nucleic acid molecule encoding a recombinant polypeptide as provided herein. In some embodiments, the cell comprises a vector as provided herein. In some embodiments, the cell is any suitable cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a mammalian cell used in a method of producing a protein product. In some embodiments, the cell is any in vitro cell line. In some embodiments, the cell is any ex vivo cell. In some embodiments, the cell is present in vivo.

[0144] Composition In some embodiments, a composition is provided. In some embodiments, the composition comprises a vector encoding a signal peptide as provided herein and a delivery system. In some embodiments, the composition comprises a vector encoding a recombinant polypeptide as provided herein and a delivery system.

[0145] In some embodiments, the delivery system may be a viral vector. In some embodiments, the viral vector is an RNA viral vector. In some embodiments, the viral vector is a DNA viral vector. Non-limiting examples of suitable viral vectors include adenovirus, adeno-associated virus (AAV), retrovirus, herpesvirus, lentivirus, poxvirus, or papillomavirus vectors.

[0146] In some implementations, the delivery system is a non-viral delivery system. Non-limiting examples of non-viral delivery systems include polymers, polyplexes, lipids, lipid-like substances, lipoplexes, liposomes, lipid fusion constructs, polymer nanoparticles, nanoparticles, lipid nanoparticles (LNPs), core-shell nanoparticles, solid lipid nanoparticles, metal nanoparticles, self-assembled nucleic acid nanoparticles, hyaluronidase, nanoparticle mimics, ribonucleoproteins, positively charged peptides, small RNA conjugates, aptamer-RNA chimeras, RNA-fusion protein complexes, and any combination thereof.

[0147] The exemplary delivery system described above should not be construed as limiting in any way. Suitable delivery systems are known in the art, and any such delivery system is within the scope of this disclosure.

[0148] method This paper also considers methods using the novel signal peptides and recombinant peptides presented herein.

[0149] In some embodiments, a method for treating a disease or condition is provided. In some embodiments, the method includes administering to a subject in need a vector comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein, thereby treating the disease or condition. In some embodiments, the payload protein of the recombinant polypeptide is a peptide or protein that can be used to treat the disease or condition. In some embodiments, a signal peptide guides the expression of the payload protein at a specific cellular location useful for treating the disease or condition. In some embodiments, an effective amount of a vector comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein is administered to a subject in need. In some embodiments, a therapeutically effective amount of a vector comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein is administered to a subject in need. In some embodiments, the disease or condition is a nuclear-related disease or condition. In some embodiments, the disease or condition is a genetic disease or condition. In some embodiments, the disease or condition is any disease or condition that would benefit from gene editing or genetic modification. In some embodiments, the disease or condition is a DNA-related disease or condition. In some embodiments, the disease or condition is any disease or condition related to transcription factors, initiation factors, DNA-related complexes, DNA-related proteins, etc.

[0150] In some embodiments, a method for treating a nuclear-related disease or condition of a cell is provided, the method comprising administering to a subject in need a carrier comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein, thereby treating the disease or condition. In some embodiments, the payload protein of the recombinant polypeptide is a peptide or protein that can be used to treat a nuclear-related disease or condition of a cell. In some embodiments, a signal peptide guides the expression of the payload protein in the nucleus of a cell, thereby treating the disease or condition. In some embodiments, an effective amount of a carrier comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein is administered to a subject in need. In some embodiments, a therapeutically effective amount of a carrier comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein is administered to a subject in need.

[0151] In some embodiments, a method for treating a disease or condition is provided. In some embodiments, the method includes administering to a subject in need a composition comprising a vector, said vector containing a nucleic acid molecule encoding a recombinant polypeptide as provided herein, thereby treating the disease or condition. In some embodiments, the payload protein of the recombinant polypeptide is a peptide or protein that can be used to treat the disease or condition. In some embodiments, a signal peptide guides the expression of the payload protein at a specific cellular location useful for treating the disease or condition. In some embodiments, an effective amount of the composition is administered to the subject in need. In some embodiments, a therapeutically effective amount of the composition is administered to the subject in need. In some embodiments, the disease or condition is a disease or condition related to the nucleus of a cell. In some embodiments, the disease or condition is a genetic disease or condition. In some embodiments, the disease or condition is any disease or condition to which gene editing may be beneficial.

[0152] In some embodiments, a method for treating a nuclear-related disease or condition of a cell is provided, the method comprising administering to a subject in need a composition comprising a carrier containing a nucleic acid molecule encoding a recombinant polypeptide as provided herein, thereby treating the disease or condition. In some embodiments, the payload protein of the recombinant polypeptide is a peptide or protein that can be used to treat a nuclear-related disease or condition of a cell. In some embodiments, a signal peptide directs the expression of the payload protein in the nucleus, thereby treating the disease or condition. In some embodiments, an effective amount of the composition is administered to a subject in need. In some embodiments, a therapeutically effective amount of the composition is administered to a subject in need.

[0153] In some embodiments, a method for generating a payload protein is provided, the method comprising administering a carrier comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein to cells, and culturing the cells under conditions sufficient to generate the payload protein. In some embodiments, the payload protein is secreted from the cells, and the method further comprises collecting a cell supernatant containing the payload protein and purifying the payload protein from the cell supernatant. In some embodiments, the payload protein is not secreted, and the method further comprises collecting cells containing the payload protein, lysing the cells, and purifying the payload protein from the cell lysate.

[0154] In some embodiments, a method for producing a payload protein is provided, the method comprising administering a carrier comprising a nucleic acid molecule encoding a recombinant polypeptide as provided herein to cells, and culturing the cells under conditions sufficient to produce the payload protein. In some embodiments, the payload protein is secreted from the cells, and the method further comprises collecting a cell supernatant containing the payload protein and purifying the payload protein from the cell supernatant. In some embodiments, the payload protein is not secreted, and the method further comprises collecting cells containing the payload protein, lysing the cells, and purifying the payload protein from the cell lysate.

[0155] Unbound by any particular theory, the novel signal peptides of this disclosure can be used to treat diseases or conditions related to the cell nucleus by incorporating a recombinant polypeptide containing a payload protein for treating the disease or condition. Correctly, or in some cases, enhanced targeting of the payload protein to the nucleus is beneficial for the treatment of the disease or condition. The payload protein is not limited to a specific payload but includes any therapeutic agent (e.g., a therapeutic protein) that functions in the cell nucleus. Furthermore, by utilizing cells affected by diseases or conditions related to the cell nucleus to generate the payload protein, the method of this invention avoids the need for timely and inefficient recombinant protein infusions, instead allowing the patient's own cells to generate the therapeutic molecule.

[0156] This disclosure also relates in part to methods for generating or producing payload proteins using novel signal peptides as provided herein. Without being bound by any particular theory, the novel signal peptides provided herein guide payload proteins to specific cellular or extracellular locations, such as, but not limited to, the cell nucleus. Using novel signal peptides to guide payload proteins to specific cellular or extracellular locations can enhance the expression and / or production of payload proteins.

[0157] The signal peptide disclosed herein can also be used to deliver gene editing technology to the nucleus of a cell. Therefore, in some embodiments, a method for editing a target nucleic acid molecule in a cell is provided, the method comprising administering a signal peptide encoding a molecule of formula X1-(Y1). aA carrier of a nucleic acid molecule containing a recombinant polypeptide of type Z1, wherein X1 is a signal peptide as provided herein, Y1 is a peptide linker as provided herein, a is an integer selected from 0 and 1, and Z1 is a payload protein as provided herein for gene editing; and at least one nucleic acid molecule that targets the payload protein to a target nucleic acid molecule, thereby editing the target nucleic acid molecule. In some embodiments, the payload protein Z1 is Cas9, a zinc finger nuclease, TALEN, or a broad-spectrum nuclease. In some embodiments, at least one nucleic acid molecule that targets the payload protein to the target nucleic acid molecule is a guide RNA. In some embodiments, the vector encodes a protein having the formula X1-(Y1). a The recombinant polypeptide of type -Z1 and at least one nucleic acid molecule that targets the payload protein to the target nucleic acid molecule are housed in the same vector. In some embodiments, the vector encodes a protein having the formula X1-(Y1). a The nucleic acid molecule of the recombinant polypeptide of -Z1 and at least one nucleic acid molecule that targets the payload protein to the target nucleic acid molecule are located in different vectors.

[0158] In some implementations, a method is provided for editing a target nucleic acid molecule in the cells of a subject in need, the method comprising administering to the subject a molecule containing an encoding of the formula X1-(Y1). a A carrier of a nucleic acid molecule containing a recombinant polypeptide of type Z1, wherein X1 is a signal peptide as provided herein, Y1 is a peptide linker as provided herein, a is an integer selected from 0 and 1, and Z1 is a payload protein as provided herein for gene editing; and at least one nucleic acid molecule that targets the payload protein to a target nucleic acid molecule, thereby editing the target nucleic acid molecule. In some embodiments, the payload protein Z1 is Cas9, a zinc finger nuclease, TALEN, or a broad-spectrum nuclease. In some embodiments, at least one nucleic acid molecule that targets the payload protein to the target nucleic acid molecule is a guide RNA. In some embodiments, the vector encodes a protein having the formula X1-(Y1). a The recombinant polypeptide of type -Z1 and at least one nucleic acid molecule that targets the payload protein to the target nucleic acid molecule are housed in the same vector. In some embodiments, the vector encodes a protein having the formula X1-(Y1). a The nucleic acid molecule of the recombinant polypeptide of -Z1 and at least one nucleic acid molecule that targets the payload protein to the target nucleic acid molecule are located in different vectors.

[0159] Unbound by any particular theory, the novel signal peptide disclosed herein can be used for the editing of target nucleic acid molecules by incorporating the signal peptide into a recombinant polypeptide containing a payload protein (such as Cas9) for gene editing. The novel signal peptide guides the payload protein to the cell nucleus, where it can facilitate the editing of the target nucleic acid molecule. Therefore, the novel signal peptide disclosed herein can also be used to treat any disease or condition for which gene editing may be useful, as provided herein.

[0160] Example The following examples illustrate the compounds, compositions, particles, peptides, and methods described herein and should not be construed as limiting in any way.

[0161] method SP-mCherry plasmid (pDNA) construction With Cheng, Qiang and others Proceedings of the National Academy of Sciences of the United States of America Plasmids were constructed in a manner similar to that described in Volume 120, 52 (2023) and International Patent Publication Serial No. WO2024064874A2, each of which is incorporated herein by reference in its entirety. In short, the SP-mCherry coding region was obtained directly by PCR using carefully designed primers. The SPs for the following examples are provided in Table 10 below: Table 10

[0162] Following a standard protocol, the enzymatically digested SP-mCherry product was cloned into the pCS2-MT vector. After sequencing verification, the SP-mCherry plasmid was prepared for in vitro screening.

[0163] In vitro SP screening via pDNA transfection For SP selection, pDNA transfection was performed in the cells. pDNA was encapsulated in Lipfectamine 2000 and administered at 100 ng / well to HeLa cells in 6-well plates, allowing incubation for 24 hours. After incubation, the DNA in the nuclei of live cells was stained with Hoechst 33342 dye according to the manufacturer's protocol. Cells were then imaged using confocal microscopy targeting Hoechst 33342 (nuclear) and mCherry signals. Colocalization and quantification of nuclear and mCherry signals were assessed using Fiji software.

[0164] Example 1: Delivering mCherry to the cell nucleus using a novel signal peptide.

[0165] The ability of the novel signal peptide presented in this paper to deliver the payload protein to the cell nucleus was tested by constructing an mCherry pDNA plasmid modified with a signal peptide containing a nuclear-targeting signal peptide located directly upstream of the mCherry sequence. As a control, an unmodified mCherry plasmid without the signal peptide was also used. General plasmid construction and general methods are as described above. Figure 1A Provided in [the document / document].

[0166] like Figures 1B to 1F As shown, four of the five novel signal peptides tested were able to drive mCherry proteins to the nucleus, while the mCherry proteins without a nuclear-targeting signal peptide did not localize to the nucleus. Compared with the no-SP control, the results of each of Nov-Nuc 18-1 (SEQ ID NO: 9), Nov-Nuc 19-1 (SEQ ID NO: 3), Nov-Nuc 22-1 (SEQ ID NO: 5), and Nov-Nuc 26-1 (SEQ ID NO: 7) showed a statistically significant increase in nuclear-localized mCherry, while Nov-Nuc 16-1 (SEQ ID NO: 1) showed a trend toward increasing nuclear-localized mCherry, although not a significant increase.

[0167] The data from this embodiment show that the nuclear-targeting signal peptide sequence encoded directly upstream of the mCherry sequence drives the mCherry protein to localize to the nucleus.

[0168] Example 2: Delivery of Cas9 to the nucleus using a novel signal peptide.

[0169] Example 2 was performed using a method similar to that of Example 1, except that the payload protein was Cas9. As a control, a plasmid encoding Cas9 without a nuclear targeting signal peptide was also used. Instead of mCherry assessment, the correct Cas9 targeting was determined by staining Cas9 with an appropriate antibody and evaluating it via confocal microscopy.

[0170] Example 3: Using a novel signal peptide to deliver antibodies to the nucleus.

[0171] Example 3 was performed using a method similar to that of Example 1, except that the payload protein was an antibody. As a control, a plasmid encoding a payload antibody without a nuclear targeting signal peptide was also used. Instead of mCherry assessment, the correct antibody target was determined by staining with an appropriate secondary antibody and evaluation via confocal microscopy.

[0172] Example 4: mRNA synthesis and mRNA-nanoparticle formulation mRNA is generated via in vitro transcription as described below: Cheng, Qiang et al. Proceedings of the National Academy of Sciences of the United States of America Volume 120, 52 (2023); Cheng, Qiang et al. Nature Nanotechnology References 15, 313-320 (2020); and International Patent Publication Serial No. WO2024064874A2, each of which is incorporated herein by reference in its entirety. In summary, linear pDNA with an optimized 5'(3')-untranslated region (UTR) and polyA sequence was first obtained by enzymatic digestion, followed by IVT preparation using N1-methylpseudouridine-5'-triphosphate modification according to a standard protocol. Finally, the mRNA was capped (Cap-1) using a vaccinia capping enzyme and 2'-O-methyltransferase (NEB).

[0173] LNP formulations loaded with mRNA were prepared using the ethanol dilution method described below: Cheng, Qiang et al. Proceedings of the National Academy of Sciences of the United States of America Volume 120, 52 (2023); Cheng, Qiang et al. Nature Nanotechnology References 15, 313-320 (2020); and International Patent Publication Serial No. WO2024064874A2, each of which is incorporated herein by reference in its entirety. In summary, all lipids were dissolved in ethanol at a specified molar ratio, and RNA was first dissolved in 10 mM citrate buffer (pH 4.0). The two solutions were then rapidly mixed at a 3:1 volume ratio of aqueous solution to ethanol (3:1, aqueous solution:ethanol, vol:vol) to achieve a final weight ratio of 40:1 (total lipids:mRNA). After incubation at room temperature for 10 min, the mRNA LNP formulation was immediately added to cells, or dialyzed against PBS for 2 h for in vivo experiments.

[0174] Example 5: Appropriate secretion of signal peptide-targeted mCherry from cells transfected with pDNA or treated with LNP loaded with mRNA.

[0175] The SP of this embodiment is provided in Table 11 below: Table 11

[0176] Such as Cheng, Qiang, etc. Proceedings of the National Academy of Sciences of the United States of AmericaVolume 120, 52 (2023); and International Patent Publication Serial No. WO2024064874A2 detail the generation and testing of pDNA constructs encoding mCherry immediately downstream of the secretory signal peptide, each of which is incorporated herein by reference in its entirety. In summary, HeLa cells were transfected with SP-free wild-type (WT) mCherry pDNA and gLuc-mCherry pDNA via Lipofectamine 2000, and intracellular and extracellular fluorescence was quantified by fluorescence microscopy at 24, 48, and 72 hours post-transfection, where gLuc SP induced high levels of mCherry expression into the culture medium. Furthermore, the mCherry protein content present in the cell culture medium and cell lysates was quantified by fluorescence plate reader at 24, 48, and 72 hours, revealing an increase in mCherry fluorescence in the gLuc SP group. Figure 2A The SP group was then expanded to include a negative control (scrambled sequence) other than gLuc, hAlb, hApoB, and hFVII. HeLa cells were then re-transfected with pDNA constructs using Lipofectamine 2000. Images acquired by fluorescence microscopy and IVIS at 72 h post-transfection showed that SP hApoB, gLuc, and hFVII all generated high levels of mCherry protein secretion, while the NC and hAlb constructs effectively mediated intracellular mCherry expression but did not promote significant extracellular secretion. Figure 2B The same group of SP cells was evaluated in the Huh7 hepatocellular carcinoma cell line, where the mCherry secretion trend observed in HeLa cells persisted. Figure 2C ).

[0177] Based on the above, we investigated whether mRNA containing the integrated SP sequence would produce similar observations to those observed with pDNA. hFVII-mCherry mRNA from the FVII-mCherry-pCS2-MT plasmid was generated via in vitro transcription (IVT). mDLNP lipid nanoparticles were tested as the initial carrier for RNA. Transfection of multiple different cell lines with mDLNPs containing FVII-mCherry mRNA showed a positive correlation between protein efflux into the culture medium and post-transfection time and dose, with greater fluorescence intensity observed throughout the cell lines at longer time intervals and higher doses. Figure 3A and Figure 3B ).

[0178] The data from this embodiment demonstrate that pDNA can be transcribed into mRNA via IVT, and that transfection of cells with pDNA via lipofectamine or with mRNA generated from the same pDNA via LNP produces comparable experimental results. Therefore, the data from this embodiment support the conclusion that, in embodiments where pDNA is first transcribed into mRNA via IVT, the construct tested in pDNA form in Example 1 will be expected to produce similar results.

Claims

1. A signal peptide comprising an amino acid sequence selected from the group consisting of formulas I, II, III, IV, V, VI, VII, and VIII; Where equation I is expressed as: A1 - A2 - A3 - A4 - A5 - A6 - A7 - A8 - A9 - A 10 -A 11 -A 12 -A 13 -A 14 -A 15 -A 16 (Formula I) And A1-A 16 The identifiers are provided in Table 2; Equation II is expressed as: B1 - B2 - B3 - B4 - B5 - B6 - B7 - B8 - B9 - B 10 -B 11 -B 12 -B 13 -B 14 -B 15 -B 16 -B 17 -B 18 (Formula II) And B1-B 18 The identifiers are provided in Table 3; Equation III is expressed as: C1-C2-C3-C4-C5-C6-C7-C8-C9-C 10 -C 11 -C 12 -C 13 -C 14 -C 15 -C 16 -C 17 -C 18 -C 19 (Formula III) And C1-C 19 The identifiers are provided in Table 4; Where equation IV is represented as: D1 - D2 - D3 - D4 - D5 - D6 - D7 - D8 - D9 - D 10 -D 11 -D 12 -D 13 -D 14 -D 15 -D 16 -D 17 -D 18 -D 19 -D 20 (Formula IV) And D1-D 20 The identifiers are provided in Table 5; Where V is expressed as: E1-E2-E3-E4-E5-E6-E7-E8-E9-E 10 -E 11 -E 12 -E 13 -E 14 -E 15 -E 16 -E 17 -E 18 -E 19 -E 20 -E 21 -E 22 (Formula V) And E1-E 22 The identifiers are provided in Table 6; Wherein, VI is represented as: F1-F2-F3-F4-F5-F6-F7-F8-F9-F 10 -F 11 -F 12 -F 13 -F 14 -F 15 -F 16 -F 17 -F 18 -F 19 -F 20 -F 21 -F 22 -F 23 (Formula VI) And F1-F 23 The identifiers are provided in Table 7; Equation VII is expressed as: G1 - G2 - G3 - G4 - G5 - G6 - G7 - G8 - G9 - G 10 -G 11 -G 12 -G 13 -G 14 -G 15 -G 16 -G 17 -G 18 -G 19 -G 20 -G 21 -G[[ID=2S]] 22 -G 23 -G 24 -G 25 -G 26 (Formula VII) And G1-G 26 The identifiers are provided in Table 8; and Where equation VIII is expressed as: H1-H2-H3-H4-H5-H6-H7-H8-H9-H 10 -H 11 -H 12 -H 13 -H 14 -H 15 -H 16 -H 17 -H 18 -H 19 -H 20 -H 21 -H 22 -H 23 -H 24 -H 25 -H 26 -H 27 -H 28 -H 29 )]]END]]-H 30 (Formula VIII) And H1-H 30 The identifier is provided in Table 9.

2. The signal peptide of claim 1, comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

3. A signal peptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

4. The signal peptide according to any one of claims 1 to 3, wherein the signal peptide is a novel signal peptide.

5. A subset X1-(Y1) a -Z1 recombinant polypeptide, wherein: X1 is a signal peptide. Y1 is a peptide linker, and Z1 is the payload protein. Where a is an integer selected from 0 and 1.

6. The recombinant polypeptide of claim 5, wherein X1 comprises an amino acid sequence selected from the group consisting of formulas I, II, III, IV, V, VI, VII, and VIII; Where equation I is expressed as: A1 - A2 - A3 - A4 - A5 - A6 - A7 - A8 - A9 - A 10 -A 11 -A 12 -A 13 -A 14 -A 15 -A 16 (Formula I) And A1-A 16 The identifiers are provided in Table 2; Equation II is expressed as: B1 - B2 - B3 - B4 - B5 - B6 - B7 - B8 - B9 - B 10 -B 11 -B 12 -B 13 -B 14 -B 15 -B 16 -B 17 -B 18 (Formula II) And B1-B 18 The identifiers are provided in Table 3; Equation III is expressed as: C1-C2-C3-C4-C5-C6-C7-C8-C9-C 10 -C 11 -C 12 -C 13 -C 14 -C 15 -C 16 -C 17 -C 18 -C 19 (Formula III) And C1-C 19 The identifiers are provided in Table 4; Where equation IV is represented as: D1 - D2 - D3 - D4 - D5 - D6 - D7 - D8 - D9 - D 10 -D 11 -D 12 -D 13 -D 14 -D 15 -D 16 -D 17 -D 18 -D 19 -D 20 (Formula IV) And D1-D 20 The identifiers are provided in Table 5; Where V is expressed as: E1-E2-E3-E4-E5-E6-E7-E8-E9-E 10 -E 11 -E 12 -E 13 -E 14 -E 15 -E 16 -E 17 -E 18 -E 19 -E 20 -E 21 -E 22 (Formula V) And E1-E 22 The identifiers are provided in Table 6; Wherein, VI is represented as: F1 - F2 - F3 - F4 - F5 - F6 - F7 - F8 - F9 - F 10 -F 11 -F 12 -F 13 -F 14 -F 15 -F 16 -F 17 -F 18 -F 19 -F 20 -F 21 -F 22 -F 23 (Formula VI) And F1-F 23 The identifiers are provided in Table 7; Equation VII is expressed as: G1 - G2 - G3 - G4 - G5 - G6 - G7 - G8 - G9 - G 10 -G 11 -G 12 -G 13 -G 14 -G 15 -G 16 -G 17 -G 18 -G 19 -G 20 -G 21 -G 22 -G 23 -G 24 -G 25 -G 26 (Formula VII) And G1-G 26 The identifiers are provided in Table 8; and Where equation VIII is expressed as: H1-H2-H3-H4-H5-H6-H7-H8-H9-H 10 -H 11 -H 12 -H 13 -H 14 -H 15 -H 16 -H 17 -H 18 -H 19 -H 20 -H 21 -H 22 -H 23 -H 24 -H 25 -H 26 -H 27 -H 28 -H 29 -H 30 (Formula VIII) And H1-H 30 The identifier is provided in Table 9.

7. The recombinant polypeptide of claim 5 or claim 6, wherein X1 comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with an amino acid sequence selected from the group consisting of SEQ ID NO: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

8. The recombinant polypeptide according to any one of claims 5 to 7, wherein the payload protein is a therapeutic peptide or protein.

9. A nucleic acid molecule encoding a signal peptide as described in any one of claims 1 to 4.

10. A nucleic acid molecule encoding a recombinant polypeptide as described in any one of claims 5 to 8.

11. A vector comprising the nucleic acid molecule as described in claim 9 or claim 10.

12. A cell comprising a nucleic acid molecule as claimed in claim 9 or claim 10, or a carrier as claimed in claim 11.

13. A composition comprising a nucleic acid molecule as claimed in claim 9 or claim 10, or a carrier as claimed in claim 11.

14. A method for treating a disease or condition in a subject in need, the method comprising administering to the subject an effective amount of the nucleic acid molecule as claimed in claim 10, thereby treating the disease or condition.

15. The method of claim 14, wherein the payload protein is a protein that can be used for gene editing, and the disease or condition is any disease or condition for which gene editing is useful.

16. The method of claim 14, wherein the disease or symptom is a disease or symptom related to the nucleus of a cell.

17. A method for treating a nucleus-related disease or condition in a subject in need, the method comprising administering a carrier to the subject, the carrier comprising a nucleic acid molecule encoding a signal peptide fused to or linked to a payload protein, wherein the signal peptide is a signal peptide as claimed in any one of claims 1 to 4, and wherein the payload protein is a therapeutic peptide or protein that can be used for the disease or condition.

18. A method for treating a disease or condition in a subject in need, the method comprising administering to the subject a vector comprising a nucleic acid molecule encoding: i) A signal peptide fused to or linked to a payload protein, wherein the signal peptide is the signal peptide as described in any one of claims 1 to 4, and wherein the payload protein is a protein that can be used for gene editing; and ii) At least one nucleic acid molecule that targets the payload protein to a gene of interest; The signal peptide promotes the translocation of the protein, which can be used for gene editing, to the nucleus, and the payload protein edits the gene of interest, thereby treating the disease or condition.

19. A method for editing a target nucleic acid in a cell, the method comprising administering a nucleic acid molecule encoding: i) A signal peptide fused to or linked to a payload protein, wherein the signal peptide is the signal peptide as described in any one of claims 1 to 4, and wherein the payload protein is a protein that can be used for gene editing; and ii) At least one nucleic acid molecule that targets the payload protein to a gene of interest; The signal peptide facilitates the translocation of the protein, which can be used for gene editing, to the nucleus, wherein the payload protein edits the gene of interest, thereby editing the target nucleic acid.

Citation Information

Patent Citations

  • Polycyclic sugar surrogate-containing oligomeric compounds and compositions for use in gene modulation

    US20040171570A1

  • N[ omega ,( omega -1)-dialkyloxy]- and N-[ omega ,( omega -1)-dialkenyloxy]-alk-1-yl-N,N,N-tetrasubstituted ammonium lipids and uses therefor

    US4897355A

  • Single polypeptide chain binding molecules

    US4946778A

  • N-( omega ,( omega -1)-dialkyloxy)- and N-( omega ,( omega -1)-dialkenyloxy)-alk-1-yl-N,N,N-tetrasubstituted ammonium lipids and uses therefor

    US4946787A

  • N- omega ,( omega -1)-dialkyloxy)- and N-( omega ,( omega -1)-dialkenyloxy)Alk-1-YL-N,N,N-tetrasubstituted ammonium lipids and uses therefor

    US5049386A