Pilot sequence and application thereof
By providing a lead sequence containing an N-terminal amino acid sequence derived from an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria, the problem of limited PVC lead sequences is solved, improving the loading efficiency and delivery effect of PVC complexes and expanding their application range.
Patent Information
- Application Number
- CN202511212256.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-06-06
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-02
AI Technical Summary
In the existing technology, the limited number of leader sequences for PVC restricts its widespread application, and it is unclear whether leader sequences from different bacterial sources are universally applicable to PVC.
A lead sequence comprising at least 20 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria is provided for guiding the packaging of a payload into a protein complex, particularly a PVC complex, the payload comprising a peptide, nucleic acid, or a combination thereof.
This has enriched our understanding of the structure and function of PVC composites, improved loading efficiency and delivery effect, expanded the application range of PVC, and laid the foundation for clinical applications.
Smart Images

Figure HDA0005570222250000011 
Figure HDA0005570222250000012 
Figure HDA0005570222250000021
Abstract
Description
Technical Field
[0001] This application relates to the field of biotechnology, and more specifically, to a lead sequence, a payload delivery system based on the lead sequence, and its applications. Background Technology
[0002] Microorganisms have continuously acquired stronger abilities to alter their environment during evolution. Among these evolutionary directions is the release of toxins, signals, and other substances to specifically inhibit the growth of certain organisms or cells. Extracellular contractile injection systems (eCIS) are bacterial in vitro protein injection systems similar to the retractable tail of T4 bacteriophages. Currently, 631 possible eCIS have been identified in archaea, Gram-negative bacteria, and Gram-positive bacteria. Common eCIS include *Photorhabdus virulence cassette* (PVC) from *Photorhabdus*, antifeedingprophage (AFP) from *Serratia entomophila*, R-typepyocin secreted by *Pseudomonas aeruginosa*, and metamorphosis-associated contractile structure (MAC) secreted by *Pseudoalteromonas luteoviolacea*.
[0003] PVC is an eCIS particle produced by the non-symbiotic luminescent bacterium *Photorhabdus asymbiotica*, first discovered in genome sequencing in 2006. Recent studies have shown that by fusing a portion of the N-terminal amino acid sequence (or "lead sequence") of the PVC effector with a heterologous protein, the protein can be packaged within PVC and accurately injected into cells to exert its function. This discovery demonstrates that PVC is a modifiable protein delivery tool.
[0004] Despite the immense application potential of PVC, the limited number of identified lead sequences currently available restricts its widespread use. Furthermore, while similar injection systems to PVC exist in other bacteria, it remains unclear whether potential lead sequences from these different bacterial sources are equally applicable to PVC. Therefore, further research and validation of the diversity and universality of lead sequences are crucial for expanding the applications of PVC. Summary of the Invention
[0005] In view of the problems existing in the prior art, the purpose of this application is to provide a lead sequence, a payload delivery system constructed based on the lead sequence, and its application.
[0006] 1. A lead sequence for guiding the packaging of a payload into a protein complex, wherein the lead sequence comprises at least 20 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria;
[0007] The payload is a polypeptide, nucleic acid, or a combination thereof.
[0008] 2. The lead sequence according to item 1, wherein the protein complex is derived from archaea, Gram-negative bacteria, or Gram-positive bacteria;
[0009] Preferably, the protein complex is derived from non-symbiotic luminescent bacillus (Photorhabdus asymbiotica), Serratia entomophila, Pseudoalteromonas luteoviolacea, or Xenorhabdus khoisanae.
[0010] More preferably, the protein complex comprises a non-symbiotic luminescent bacillus virulence box (PVC);
[0011] More preferably, the protein complex is a PVC-V complex.
[0012] 3. The lead sequence according to claim 1 or 2, wherein the lead sequence comprises at least 20 amino acids at the N-terminus of an effector derived from *Xenorhabdus bovienii*, *Non-symbiotic luminescent bacteria*, *Yersinia pekkanenii*, *Yersinia similis*, *Yersinia ruckeri*, or *Xenorhabdus bovienii*.
[0013] Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 at the N-terminus.
[0014] 4. The leader sequence according to any one of claims 1-3, wherein the leader sequence comprises an amino acid sequence as shown in any one of SEQ ID NO:1-3, SEQ ID NO:5-165, or comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with an amino acid sequence as shown in any one of SEQ ID NO:1-3, SEQ ID NO:5-165.
[0015] 5. The leader sequence according to any one of claims 1-4, wherein the payload is a polypeptide, the polypeptide comprising one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone regulatory molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene editing proteins.
[0016] 6. An isolated nucleic acid encoding a leader sequence as described in any one of items 1-5.
[0017] 7. A fusion protein comprising a leader sequence as described in any one of items 1-5 and a polypeptide linked thereto, wherein the linkage is a covalent or non-covalent linkage.
[0018] 8. The fusion protein according to claim 7, wherein the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cellular metabolism, and gene-editing proteins.
[0019] 9. An isolated nucleic acid encoding a fusion protein as described in item 7 or 8.
[0020] 10. An expression vector comprising isolated nucleic acids as described in item 6 or 9.
[0021] 11. A recombinant host cell comprising isolated nucleic acids as described in item 6 or 9, or an expression vector as described in item 10.
[0022] 12. A payload delivery system comprising a protein complex and a leader sequence as described in any one of items 1-5;
[0023] The payload is a polypeptide, a nucleic acid, or a combination thereof.
[0024] Preferably, the payload delivery system is an extracellular retractable injection system (eCIS).
[0025] 13. The payload delivery system according to claim 12, wherein the protein complex is derived from archaea, Gram-negative bacteria, or Gram-positive bacteria;
[0026] Preferably, the protein complex is derived from non-symbiotic luminescent bacillus (Photorhabdus asymbiotica), Serratia entomophila, Pseudoalteromonas luteoviolacea, or Xenorhabdus khoisanae.
[0027] More preferably, the protein complex comprises a non-symbiotic luminescent bacillus virulence box (PVC);
[0028] More preferably, the protein complex is a PVC-V complex.
[0029] 14. The payload delivery system according to item 12 or 13, wherein the payload is a polypeptide, and the leader sequence is covalently linked to the polypeptide;
[0030] Preferably, the leader sequence is covalently linked to the N-terminus of the polypeptide.
[0031] 15. The payload delivery system according to claim 14, wherein the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cellular metabolism, and gene-editing proteins.
[0032] 16. An isolated nucleic acid encoding a payload delivery system as described in any one of items 12-15.
[0033] 17. The isolated nucleic acid according to claim 16, wherein the isolated nucleic acid comprises one or more of the following:
[0034] (i) The nucleotide sequence encoding the protein complex in the payload delivery system;
[0035] (ii) A nucleotide sequence encoding the leader sequence of the payload delivery system;
[0036] (iii) The nucleotide sequence encoding the polypeptide.
[0037] 18. An expression vector comprising isolated nucleic acids as described in item 16 or 17.
[0038] 19. A recombinant host cell comprising a payload delivery system as described in any one of items 12-15, isolated nucleic acid as described in item 16 or 17, or an expression vector as described in item 18.
[0039] 20. A method for preparing a payload delivery system as described in any one of items 12-15, comprising culturing the recombinant host cells described in item 19 to produce the payload delivery system.
[0040] 21. The use of a lead sequence as described in any one of items 1-5, a payload delivery system as described in any one of items 12-15, an isolated nucleic acid as described in item 16 or 17, an expression vector as described in item 18, or a recombinant host cell as described in item 19 in the preparation of a drug, reagent, or kit for delivering a payload.
[0041] 22. A pharmaceutical composition comprising a payload delivery system as described in any one of items 12-15, or isolated nucleic acid as described in item 16 or 17.
[0042] 23. The pharmaceutical composition according to claim 22, wherein the pharmaceutical composition further comprises a pharmaceutically acceptable carrier.
[0043] 24. A method for transferring a payload into target cells, comprising:
[0044] The payload delivery system described in any one of items 12-15 is brought into contact with the target cell, such that the payload delivery system delivers the payload into the target cell;
[0045] Optionally, the target cells are eukaryotic cells;
[0046] Preferably, the eukaryotic cell is a yeast cell, insect cell, mammalian cell, plant cell, or fungal cell;
[0047] More preferably, the eukaryotic cell is a human cell.
[0048] 25. The method of claim 24, wherein the payload is a polypeptide, the polypeptide comprising one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone regulatory molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene editing proteins.
[0049] Beneficial effects:
[0050] (1) This application found that the lead sequences of effector factors from various bacterial sources can guide the packaging of peptides in PVC complexes, which indicates that the loading mechanism of PVC complexes has strong universality and expands more options and resources for the development of novel biological agents.
[0051] (2) The research in this application has enriched our understanding of the structure and function of PVC composites, provided more options and possibilities for optimizing and improving the loading efficiency of PVC composites, and helped to conduct in-depth research on the biological mechanisms of PVC composites.
[0052] (3) By comparing the packaging capabilities of different lengths of leader sequences, this application has determined the most suitable leader sequence length range. This finding helps to improve the loading efficiency and delivery effect of PVC compounds, enabling them to play a more efficient role in practical applications.
[0053] (4) This application further confirms that the PVC complex has actual biological effects at the cellular level. This discovery lays the foundation for subsequent clinical applications and provides important experimental evidence for the application of the PVC complex in the biomedical field. Attached Figure Description
[0054] Figure 1 This study identifies the loading of TcsT protein into the PVC complex guided by the N70 leader sequences (amino acids 1-70 from the N-terminus) of the effector factors XK1 (encoded by AB204_RS00510), XK2 (encoded by AB204_RS00495), and XK3 (encoded by AB204_RS00490) from the Xenorhabdus khoisanae MCB strain. In the figures, 1 represents a negative control; 2 represents the loading of TcsT protein into the PVC complex guided by the XK2-N70 leader sequence; 3 represents the loading of TcsT protein into the PVC complex guided by the XK3-N70 leader sequence; and 4 represents the loading of TcsT protein into the PVC complex guided by the XK1-N70 leader sequence. Anti-Pvc16 antibody was used to detect the PVC-V structural protein.
[0055] Figures 2A-2B The identification results of TcsT protein loading into the PVC complex were obtained by using the N20, N25, N30, N35, N45, N50, N55, N60, N65, N70, N75, N80, N85, N90, N95, N100, N105, N110, N115, and N120 leader sequences of the effector factor XK3. Among these, Figure 2A The identification results of TcsT protein loading into the PVC complex guided by the N20, N25, N30, N35, N45, N50, N55, N60, N65, and N70 leader sequences of the effector factor XK3; Figure 2B The identification results of TcsT protein loading into the PVC complex were obtained by using the N60, N65, N70, N75, N80, N85, N90, N95, N100, N105, N110, N115, and N120 leader sequences of the effector factor XK3. Anti-Pvc16 antibody was used to detect PVC-V structural proteins, with Pdp1-N50 as a positive control.
[0056] Figure 3 The results of the CCK8 assay were used to detect the killing effect of PVC / XK3-N70TcsT on mouse monocyte-macrophage J774A.1 cells. Empty PVC complex and PBS were used as controls.
[0057] Figures 4A-4K The identification results of TcsT protein loading into the PVC complex guided by the leader sequences (N20, N30, N40, N50, N60, N70) of P. asymbiotica effector factors F1, F2, F3, F5, F7, F8, F10, F11, F12, F13, and F16. Figure 4A The identification results show that the leader sequence of the effector factor F1 guides the loading of TcsT protein into the PVC complex; Figure 4B The identification results show that the leader sequence of the effector factor F2 guides the loading of TcsT protein into the PVC complex; Figure 4C The identification results show that the leader sequence of the effector factor F3 guides the loading of TcsT protein into the PVC complex; Figure 4D The identification results show that the leader sequence of the effector factor F5 guides the loading of TcsT protein into the PVC complex; Figure 4E The identification results show that the leader sequence of the effector factor F7 guides the loading of TcsT protein into the PVC complex; Figure 4F The identification results show that the leader sequence of the effector factor F8 guides the loading of TcsT protein into the PVC complex; Figure 4G The identification results show that the leader sequence of the effector factor F10 guides the loading of TcsT protein into the PVC complex; Figure 4H The identification results show that the leader sequence of the effector factor F11 guides the loading of TcsT protein into the PVC complex; Figure 4I The identification results show that the leader sequence of the effector factor F12 guides the loading of TcsT protein into the PVC complex; Figure 4J The identification results show that the leader sequence of the effector factor F13 guides the loading of TcsT protein into the PVC complex; Figure 4K The identification results were obtained by using the leader sequence of the effector factor F16 to guide the loading of TcsT protein into the PVC complex. An empty PVC complex was used as a control, and an anti-Pvc16 antibody was used to detect PVC-V structural proteins.
[0058] Figure 5The identification results were based on the leader sequences of other effector factors of *P. asymbiotica* guiding the loading of TcsT protein into the PVC complex. Anti-Pvc16 antibody was used to detect PVC-V structural proteins.
[0059] Figure 6 The identification results were obtained by using the leader sequence of effector factors derived from other strains to guide the loading of TcsT protein into PVC complexes. Empty PVC complexes were used as controls, Pdp1-N50 as a positive control, and anti-Pvc16 antibody was used to detect PVC-V structural proteins. Detailed Implementation
[0060] The present application is further illustrated below with reference to embodiments. It should be understood that the embodiments are only used to further illustrate and explain the present application and are not intended to limit the present application.
[0061] Unless otherwise defined, technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art. While similar or identical methods and materials may be applied in experimental or practical applications, materials and methods are described herein. In case of conflict, the definitions included herein shall prevail. Furthermore, materials, methods, and examples are for illustrative purposes only and are not intended to be limiting. The present application is further described below with reference to specific embodiments, but is not intended to limit the scope of the application.
[0062] definition
[0063] The term “contractile injection system (CIS)” as used in this paper refers to a diverse range of evolutionarily related macromolecular devices that utilize retractable sheaths to deliver nucleic acids and proteins. Contractile injection systems can help microorganisms transport a variety of effector molecules extracellularly to gain a survival advantage (NMITaylor, MJvan Raaij, and PGLeiman, Contractile injection systems of bacteria and related systems. Mol Microbiol, 2018, 108(1): p.6-15.). Typical CISs include the retractable tails of bacteriophages T4, P2, and Mu. Besides retractable bacteriophages, retractable injection systems similar to bacteriophage tails are also prevalent in bacteria and archaea. For example, the type VI secretion system (T6SS) mediates intercellular communication and plays a role in cellular defense (NMITaylor, MJvan Raaij, and PGLeiman, Contractile injection systems of bacteria and related systems. Mol Microbiol, 2018, 108(1):p.6-15.). Contractile injection systems also include extracellular contractile injection systems (eCIS).
[0064] The term “extracellular retractable injection system” or “eCIS” as used in this article refers to a type of retractable injection device that can be secreted into the extracellular space of bacteria and act on target cells through the extracellular environment. Extracellular contractile injection systems include: bacterial tailin / pyocin; Photorhabdus virulence cassette (PVC) (G. Yang, et al., Photorhabdus virulence cassettes confer injectable insecticidal activity against the wax moth. J Bacteriol, 2006. 188(6): p. 2254-61.); Antifeeding prophage (Afp) (A. Desfosses, et al., Atomic structures of an entire contractile injection system in both the extended and contracted states. Nature Microbiology, 2019. 4(11): p. 1885-1894.); and Metamorphosis-associated contractile structure (MAC) (NJ Shikuma, et al., Marine tubeworm metamorphosis induced by arrays of bacterial phage). tail-likestructures. Science, 2014.343(6170):p.529-33.) etc.
[0065] The term “Photorhabdus asymbiotica” as used in this article belongs to the genus Photorhabdus. While bacteria in this genus are generally considered insect pathogens, Photorhabdus asymbiotica possesses unique human pathogenicity (P. Wilkinson, et al., Comparative genomics of the emerging human pathogen Photorhabdus asymbiotica with the insect pathogen Photorhabdusluminescens. BMC Genomics, 2009.10:p.302).
[0066] In this paper, the term "PVC" generally refers to a retractable injection system produced by the genus *Photorhabdus*. The PVC of the non-symbiotic *Photorhabdus asymbiotica* ATCC43949 is a protein complex device with a molecular weight exceeding 10 MDa. Its structure is similar to a simplified T4 phage tail, comprising a hexagonal baseplate complex with six fibers, a 117 nm long sheath trunk with a cap structure, and an inner tube inside the sheath in which effector proteins are loaded. PVC can be released extracellularly by the bacteria to exert its effects. PVC is generally considered to be toxic to eukaryotic cells because it can transfer effector proteins into insect hemocytes and promote actin aggregation (G. Yang, et al., *Photorhabdus virulence cassettes confer injectable insecticidal activity against the wax moth. J Bacteriol, 2006, 188(6): p. 2254-61.). Specifically, in the context of this application, PVC refers to a retractable injection system isolated from the non-symbiotic luminescent bacterium Photorhabdusasymbiotica ATCC43949 or the host cells described herein, capable of delivering polypeptides / proteins across the eukaryotic cell membrane to the cytoplasm. The P. asymbiotica genome encodes five PVC subtypes (PVC-I, PVC-II, PVC-III, PVC-IV, and PVC-V).
[0067] In this paper, the term "PVC gene cluster" refers to a multi-gene cluster encoding PVC structural proteins present in the genome of the non-symbiotic luminescent bacterium *Photorhabdus asymbiotica* (e.g., *Photorhabdus asymbiotica* ATCC43949). This strain's genome encodes five PVC gene clusters: PVC-I, PVC-II, PVC-III, PVC-IV, and PVC-V. Each PVC gene cluster also contains one or more potential effector gene encodings downstream of it.
[0068] In this paper, the term "effect factor" refers to bacterial secretory proteins produced by bacteria, transported to host cells via specific secretion systems, and involved in host recognition or pathogenicity. Structurally, effector factors can be divided into signaling regions and functional regions.
[0069] As used herein, the term “amino acid” or “amino acid sequence” refers to an oligopeptide, peptide, polypeptide, or protein sequence, or any fragment thereof, and refers to a naturally occurring or synthetic molecule. When “amino acid sequence” is described herein as referring to the amino acid sequence of a naturally occurring protein molecule, “amino acid sequence” and similar terms are not intended to limit the amino acid sequence to the complete naturally occurring amino acid sequence associated with the described protein molecule.
[0070] The term “amino acid” as used in this article may be referred to by its name, its commonly known three-letter symbol, or a single-letter symbol recommended by the IUPAC-IUB Biochemical Nomenclature Commission.
[0071] As used herein, the terms “nucleic acid,” “nucleic acid sequence,” “nucleotide sequence,” “polynucleotide,” “polynucleotide sequence,” “RNA sequence,” or “DNA sequence” refer to oligonucleotides, nucleotides, or polynucleotides, and fragments or portions thereof, and refer to DNA or RNA of genotype or synthetic origin, which may be single-stranded or double-stranded and represent sense or antisense strands. Sequences may be non-coding sequences, coding sequences, or mixtures of both. The nucleic acid sequences of this application may be prepared using standard techniques well known to those skilled in the art.
[0072] The percentage of "identity" used herein, such as 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, and 99.5%, refers to the degree of similarity between amino acid sequences or nucleotide sequences determined by sequence alignment, specifically 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, and 99.5%. For example, it is the percentage of positions with identical bases or amino acid residues determined after two sequences have as many identical residues as possible by introducing vacancies, etc. The percentage of "identity" can be determined using software programs known in the art. It is preferred to use default parameters for alignment. A preferred alignment program is BLAST. Preferred programs are BLASTN and BLASTP. Details of these procedures can be found at the following internet address: ncbi.nlm.nih.gov / cgi-bin / BLAST.
[0073] As used herein, the term "expression vector" refers to a linear or circular DNA molecule containing a polynucleotide encoding a polypeptide, which is effectively linked to a control sequence for its expression.
[0074] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.
[0075] As used herein, the term "recombinant expression vector" refers to a DNA structure containing a polynucleotide encoding, for example, a desired polypeptide. A recombinant expression vector may include, for example, a collection of genetic elements that regulate gene expression, such as promoters and enhancers; ii) a structural or coding sequence transcribed into mRNA and translated into a protein; and iii) a transcriptional subunit containing appropriate transcription and translation initiation and termination sequences. Recombinant expression vectors are constructed in any suitable manner. The nature of the vector is not important, and any vector, including plasmids, viruses, bacteriophages, and transposons, may be used. Possible vectors used in this disclosure include, but are not limited to, chromosomal, non-chromosomal, and synthetic DNA sequences, such as bacterial plasmids, bacteriophage DNA, yeast plasmids, and vectors derived from combinations of plasmids and bacteriophage DNA, from viruses such as vaccinia, adenovirus, fowlpox, baculovirus, SV40, and pseudorabies DNA.
[0076] As used herein, the term “host cell” refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0077] As used herein, the term "recombinant host cell" encompasses a host cell that differs from the parent cell after the introduction of a polynucleotide or recombinant expression vector encoding a mutant polypeptide, specifically achieved through transformation. The host cell disclosed herein can be a prokaryotic or eukaryotic cell. In one embodiment, the host cell refers to a prokaryotic cell, specifically derived from microorganisms suitable for protein fermentation production, such as those derived from the genera *Escherichia*, *Erwinia*, *Serratia*, *Providencia*, *Enterobacteria*, *Salmonella*, *Streptomyces*, *Pseudomonas*, *Brevibacterium*, *Bacillus*, or *Corynebacterium*.
[0078] As used in this article, the term "cancer" or "tumor" generally refers to a physiological condition in mammals characterized by uncontrolled cell growth / proliferation. Examples of cancer include, but are not limited to, lymphomas (such as Hodgkin's and non-Hodgkin's lymphomas), blastomas, sarcomas, and leukemias. More specific examples of cancer include squamous cell carcinoma, small cell lung cancer, non-small cell lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, peritoneal carcinoma, hepatocellular carcinoma, gastrointestinal cancer, pancreatic cancer, glioma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatocellular carcinoma, breast cancer, colon cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, kidney cancer, liver cancer, prostate cancer, vulvar cancer, thyroid cancer, leukemia, and other lymphoproliferative disorders, as well as various types of head and neck cancers.
[0079] As used in this paper, the term "fusion protein" refers to a hybrid protein expressed by a nucleic acid molecule containing nucleotide sequences of at least two genes.
[0080] The term “codon optimization” as used in this paper refers to the configuration of the nucleotide sequence encoding a polypeptide to contain codons preferred by the host cell or organism in order to improve gene expression and translation efficiency in the host cell or organism.
[0081] As used herein, the term "lead sequence" (which is interchangeable with "leading sequence," "leading peptide," "signal peptide," "signal sequence," "targeting signal," "localization signal," "localization sequence," and "transporting peptide") refers to a peptide chain that guides the transfer of a synthesized polypeptide or protein toward a target. In the context of this application, a "lead sequence" is capable of guiding a polypeptide or protein to be delivered into the lumen of a PVC sheath.
[0082] As used herein, the term “isolated” means a substance in a form or environment not naturally occurring. Non-limiting examples of isolated substances include (1) any substance not naturally occurring, (2) any substance, including but not limited to any enzyme, mutant, nucleic acid, protein, peptide, or cofactor, which is at least partially removed from one or more naturally occurring components associated with it; (3) any substance artificially modified relative to a naturally found substance; or (4) any substance modified by increasing the amount of the substance relative to other components naturally associated with it (e.g., recombinant generation in a host cell; multiple copies of the gene encoding the substance; and the use of a promoter stronger than that naturally associated with the gene encoding the substance). Isolated substances may be present in fermentation broth samples. For example, host cells may be genetically modified to express the polypeptides disclosed herein. Fermentation broth from host cells will contain isolated polypeptides. “Recombinant polynucleotide” is a type of “polynucleotide.”
[0083] As used herein, the term "linker" can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that connects two molecules or parts (e.g., two components of a protein complex), such as two domains of a fusion protein. A linker can be located between or on either side of two groups, molecules, or other parts and is connected to each of them via a covalent bond or non-covalent interaction, thereby connecting the two. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can be one or more amino acids (e.g., a peptide or protein). In some embodiments, the length of the linker can be from about 5 to 100 amino acids, for example, from about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, or 90 to 100 amino acids. In some embodiments, the length of the connector can be about 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, 400 to 450, or 450 to 500 amino acids. Longer or shorter connectors may also be considered.
[0084] As used herein, the term "payload" refers to a molecule packaged in a protein complex and capable of being delivered to (target) cells. For example, in wild-type *Bacillus*, the payload is a PVC effector, typically encoded by a gene downstream (3') of the structural gene of the PVC operon.
[0085] Preceding sequence
[0086] In a first aspect, this application provides a lead sequence for guiding the packaging of a payload into a protein complex, wherein the lead sequence comprises at least 20 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria.
[0087] The phrase "at least 20 amino acids at the N-terminus" refers to the first 20 or more consecutive amino acids starting from the first amino acid at the N-terminus. For example, the leader sequence may contain the first 20 to 150 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria.
[0088] In this application, the term "derived from" means originating from a parent and capable of being modified (substitution, deletion, and / or insertion) while retaining the function of the parental sequence. In the context, "the leader sequence comprising at least 20 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria" means that the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria, or a leader sequence modified (substitution, deletion, and / or insertion) thereon while retaining the function of the parental leader sequence.
[0089] In some embodiments, the lead sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, from the N-terminus of an effector derived from Xenorhabduskhoisanae, Photorhabdus asymbiotica, Yersinia (Yersinia pekkanenii, Yersinia similis, or Yersinia ruckeri), or Xenorhabdus bovienii.
[0090] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from *Klebsiella pneumoniae*. *Klebsiella pneumoniae* is, for example, *Klebsiella pneumoniae* strain MCB.
[0091] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of the effector XK1 (encoded by AB204_RS00510) derived from the Klebsiella pneumoniae MCB strain, for example, 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids at the N-terminus. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor XK1 (encoded by AB204_RS00510). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector factor XK1 (encoded by AB204_RS00510). The amino acid sequence of the effector factor XK1 (encoded by AB204_RS00510) is shown in SEQ ID NO:166.
[0092] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, of the effector XK2 (encoded by AB204_RS00495) derived from the pathogenic bacterium Klebsiella pneumoniae MCB strain. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor XK2 (encoded by AB204_RS00495). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector factor XK2 (encoded by AB204_RS00495). The amino acid sequence of the effector factor XK2 (encoded by AB204_RS00495) is shown in SEQ ID NO:167.
[0093] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, of the effector Xk3 (encoded by AB204_RS00490) derived from the pathogenic bacterium Klebsiella pneumoniae MCB strain. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor XK3 (encoded by AB204_RS00490). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector factor XK3 (encoded by AB204_RS00490). The amino acid sequence of the effector factor XK3 (encoded by AB204_RS00490) is shown in SEQ ID NO:168.
[0094] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from a non-symbiotic luminescent bacterium. The non-symbiotic luminescent bacterium is, for example, *Photorhabdus asymbiotica* ATCC43949.
[0095] The genome of non-symbiotic luminescent bacteria contains five luminescent virulence cassette gene clusters: PVC-I, PVC-II, PVC-III, PVC-IV, and PVC-V. Each PVC gene cluster contains one or more genes encoding potential effectors.
[0096] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F1 (encoded by PAU_RS09715) in the non-symbiotic luminescent bacterium PVC-I gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F1 (encoded by PAU_RS09715). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F1 (encoded by PAU_RS09715). The amino acid sequence of the effector F1 (encoded by PAU_RS09715) is shown in SEQ ID NO:169.
[0097] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F2 (encoded by PAU_RS09720) in the non-symbiotic luminescent bacterium PVC-I gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor F2 (encoded by PAU_RS09720). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F2 (encoded by PAU_RS09720). The amino acid sequence of the effector F2 (encoded by PAU_RS09720) is shown in SEQ ID NO:170.
[0098] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F3 (encoded by PAU_RS09725) in the non-symbiotic luminescent bacterium PVC-I gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor F3 (encoded by PAU_RS09725). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F3 (encoded by PAU_RS09725). The amino acid sequence of the effector F3 (encoded by PAU_RS09725) is shown in SEQ ID NO:171.
[0099] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F5 (encoded by PAU_RS10135) in the non-symbiotic luminescent bacterium PVC-II gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F5 (encoded by PAU_RS10135). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F5 (encoded by PAU_RS10135). The amino acid sequence of the effector F5 (encoded by PAU_RS10135) is shown in SEQ ID NO:172.
[0100] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F7 (encoded by PAU_RS10125) in the non-symbiotic luminescent bacterium PVC-II gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F7 (encoded by PAU_RS10125). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F7 (encoded by PAU_RS10125). The amino acid sequence of the effector F7 (encoded by PAU_RS10125) is shown in SEQ ID NO:173.
[0101] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F8 (encoded by PAU_RS10120) in the non-symbiotic luminescent bacterium PVC-II gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F8 (encoded by PAU_RS10120). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F8 (encoded by PAU_RS10120). The amino acid sequence of the effector F8 (encoded by PAU_RS10120) is shown in SEQ ID NO:174.
[0102] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F10 (encoded by PAU_RS13645) in the non-symbiotic luminescent bacterium PVC-IV gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F10 (encoded by PAU_RS13645). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F10 (encoded by PAU_RS13645). The amino acid sequence of the effector F10 (encoded by PAU_RS13645) is shown in SEQ ID NO:175.
[0103] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F11 (encoded by PAU_RS22355) in the non-symbiotic luminescent bacterium PVC-IV gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F11 (encoded by PAU_RS22355). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F11 (encoded by PAU_RS22355). The amino acid sequence of the effector F11 (encoded by PAU_RS22355) is shown in SEQ ID NO:176.
[0104] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F12 (encoded by PAU_RS13655) in the non-symbiotic luminescent bacterium PVC-IV gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F12 (encoded by PAU_RS13655). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F12 (encoded by PAU_RS13655). The amino acid sequence of the effector F12 (encoded by PAU_RS13655) is shown in SEQ ID NO:177.
[0105] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F13 (encoded by PAU_RS13660) in the non-symbiotic luminescent bacterium PVC-IV gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector F13 (encoded by PAU_RS13660). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F13 (encoded by PAU_RS13660). The amino acid sequence of the effector F13 (encoded by PAU_RS13660) is shown in SEQ ID NO:178.
[0106] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, derived from the effector F16 (encoded by PAU_RS16545) in the non-symbiotic luminescent bacterium PVC-V gene cluster. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor F16 (encoded by PAU_RS16545). More preferably, the leader sequence comprises amino acids 1-20, 1-30, 1-40, 1-50, 1-60, or 1-70 from the N-terminus of the effector F16 (encoded by PAU_RS16545). The amino acid sequence of the effector F16 (encoded by PAU_RS16545) is shown in SEQ ID NO:179.
[0107] In some embodiments, the lead sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, from the N-terminus of an effector derived from a non-symbiotic luminescent bacterium. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of an effector derived from a non-symbiotic luminescent bacterium.The effector factors are encoded by the following genes: PAU_RS09615, PAU_RS09620, PAU_RS09625, PAU_RS09630, yciH (PAU_RS10065), PAU_RS10070, PAU_RS10075, PAU_RS10080, PAU_RS10085, PAU_RS10100, PAU_RS25580, PAU_RS10205, nth (PAU_RS10210), rsxC (PAU_RS10230), rsxB (PAU_RS10235), PAU_RS1 0715, PAU_RS10720, PAU_RS10725, PAU_RS10730, PAU_RS10855, PAU_RS10870, PAU_RS13665, PAU_RS13670, PAU_RS23900, cmoM (PAU_RS13 690), elyC(PAU_RS13695), kdsB(PAU_RS13700), PAU_RS13560, PAU_RS13520, PAU_RS22390, PAU_RS24740, PAU_RS16505, PAU_RS16515, P AU_RS25720, PAU_RS16660, PAU_RS16665, PAU_RS16690, PAU_RS16695, PAU_RS09595, PAU_RS09735, PAU_RS09740, PAU_RS09750, PAU_RS 09755, PAU_RS10105, rsxG(PAU_RS10220), rsxD(PAU_RS10225), PAU_RS10245, PAU_RS10865, mukB(PAU_RS13675), PAU_RS13540, pilV(P AU_RS13535), PAU_RS13530, PAU_RS13525, PAU_RS16495, PAU_RS09715, PAU_RS09725, PAU_RS09730, PAU_RS10135, PAU_RS10125, PAU_RS 10120, PAU_RS10765, PAU_RS22355, PAU_RS13655, PAU_RS13660, PAU_RS16545, PAU_RS24760, PAU_RS10735, PAU_RS10115 or PAU_RS09745.
[0108] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from Yersinia similis. In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of the effector YP2 (encoded by BF17_RS02855) derived from Yersinia similis strain 228, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector YP2 (encoded by BF17_RS02855). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector YP2 (encoded by BF17_RS02855). The amino acid sequence of YP2 (encoded by BF17_RS02855) is shown in SEQ ID NO:180.
[0109] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from Yersinia pekkanenii.
[0110] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, of the effector YPe1 (encoded by AEP37_RS09605) derived from the Yersinia pekkanenii A125KOH2 strain. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector YPe1 (encoded by AEP37_RS09605). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector YPe1 (encoded by AEP37_RS09605). The amino acid sequence of the effector YPe1 (encoded by AEP37_RS09605) is shown in SEQ ID NO:181.
[0111] In some embodiments, the leader sequence comprises at least 20 amino acids, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids, of the effector YPe2 (encoded by AEP37_RS09610) derived from the Yersinia pekkanenii A125KOH2 strain. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector YPe2 (encoded by AEP37_RS09610). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector YPe2 (encoded by AEP37_RS09610). The amino acid sequence of the effector YPe2 (encoded by AEP37_RS09610) is shown in SEQ ID NO:182.
[0112] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from *Yersinia ruckeri*. In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of the effector Yr2 (encoded by LGL87_RS15575) derived from the *Yersinia ruckeri* strain, for example, 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector Yr2 (encoded by LGL87_RS15575). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector Yr2 (encoded by LGL87_RS15575). The amino acid sequence of the effector Yr2 (encoded by LGL87_RS15575) is shown in SEQ ID NO:183.
[0113] In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector derived from *Burkholderia gravidii*. In some embodiments, the leader sequence comprises at least 20 amino acids at the N-terminus of an effector Xb3 (encoded by AACW61_RS01295) derived from *Xenorhabdus bovienii* strain, such as 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 amino acids. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 of the N-terminus of the effector factor Xb3 (encoded by AACW61_RS01295). More preferably, the leader sequence comprises amino acids 1-70 from the N-terminus of the effector factor Xb3 (encoded by AACW61_RS01295). The amino acid sequence of the effector factor Xb3 (encoded by AACW61_RS01295) is shown in SEQ ID NO:184.
[0114] In some embodiments, the leader sequence comprises an amino acid sequence as shown in any one of SEQ ID NO:1-3 and SEQ ID NO:5-165, or comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NO:1-3 and SEQ ID NO:5-165.
[0115] In some embodiments, the leader sequence comprises an amino acid sequence having 1 to 5, for example 1, 2 or 3 amino acid substitutions compared to the amino acid sequences shown in any of SEQ ID NO:1-3 or SEQ ID NO:5-165.
[0116] In some embodiments, the amino acid sequence of the leader sequence is as shown in any one of SEQ ID NO:1-3, SEQ ID NO:5-165, or has at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with the amino acid sequence shown in any one of SEQ ID NO:1-3, SEQ ID NO:5-165.
[0117] The lead sequence provided in this application enables the packaging of payloads into protein complexes. Based on this, this application also provides the use of the lead sequence described in the first aspect of this application for guiding the packaging of payloads into protein complexes.
[0118] The protein complex is derived, for example, from archaea, Gram-negative bacteria, or Gram-positive bacteria. Preferably, the protein complex is derived from non-symbiotic luminescent bacilli, Serratia entomophila, Pseudoalteromonas luteoviolacea, or Klebsiella pneumoniae. More preferably, the protein complex is derived from an extracellular retractable injection system (eCIS) of non-symbiotic luminescent bacilli, Serratia entomophila, Pseudoalteromonas luteoviolacea, or Klebsiella pneumoniae.
[0119] In this application, the terms "derived from" or "originating from" mean that a particular whole is available from a particular source, but not necessarily directly from that source. Protein complexes derived from archaea, Gram-negative bacteria, or Gram-positive bacteria can be genetically manipulated.
[0120] In some embodiments, the protein complex comprises a bioluminescent virulence box (PVC).
[0121] In some embodiments, the protein complex comprises a non-symbiotic luminescent bacterium virulence box (PVC). In some embodiments, the protein complex comprises PVC-I, PVC-II, PVC-III, PVC-IV, or PVC-V. In some embodiments, the PVC is derived from the non-symbiotic luminescent bacterium Photorhabdus asymbiotica ATCC43949.
[0122] In the genome of non-symbiotic luminescent bacteria, the PVC-V structural proteins in the PVC-V gene cluster include Pvc1, Pvc2, Pvc3, Pvc4, Pvc5, Pvc6, Pvc7, Pvc8, Pvc9, Pvc10, Pvc11, Pvc12, Pvc13, Pvc14, Pvc15, and Pvc16.
[0123] In some embodiments, the protein complex is a PVC-V complex, which includes PVC-V structural proteins Pvc1, Pvc2, Pvc3, Pvc4, Pvc5, Pvc6, Pvc7, Pvc8, Pvc9, Pvc10, Pvc11, Pvc12, Pvc13, Pvc14, Pvc15, and Pvc16, wherein the PVC-V structural proteins are derived from the non-symbiotic luminescent bacterium Photorhabdusasymbiotica ATCC43949.
[0124] In some embodiments, the protein complex is a modified PVC-V complex. In some embodiments, the structural protein Pvc13 in the PVC-V complex is modified to recognize cell surface molecules. For example, a protein that recognizes surface molecules is inserted into the receptor-binding domain of the structural protein Pvc13 (the receptor-binding domain has been described in the literature, Kreitz J, Friedrich MJ, Guru A, Lash B, Saito M, Macrae RK, Zhang F. Programmable protein delivery with a bacterial contractile injection system[J]. Nature. 2023; 616:357)). In some embodiments, the receptor-binding domain of the structural protein Pvc13 in the PVC-V complex is modified to recognize cell surface tumor-specific antigens or tumor-associated antigens.
[0125] The payload is a polypeptide, a nucleic acid, or a combination thereof.
[0126] In some embodiments, the payload is a polypeptide.
[0127] In some embodiments, the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene-editing proteins.
[0128] In this application, the term "cytotoxin" refers to a polypeptide or protein that has a toxic effect on specific cells and can cause cell damage or death.
[0129] In this application, the term "antimicrobial peptide" refers to a peptide with broad-spectrum antipathogenic activity that can rapidly kill pathogens. Antimicrobial peptides typically consist of 20-60 amino acid residues. Targets of antimicrobial peptides include Gram-negative bacteria, Gram-positive bacteria, fungi, parasites, tumor cells, etc. Depending on their origin, antimicrobial peptides can include insect antimicrobial peptides, mammalian antimicrobial peptides, amphibian antimicrobial peptides, antimicrobial peptides derived from fish, mollusks, crustaceans, plants, and bacteria. The antimicrobial peptides can be defensins, Cecropin A and its analogues, Magainins, Melitiin, Cecropins, cathelicidin, apidaecins, drosocin, coleoptericin, hemipteracin, bactenecin, Cecropin, etc.
[0130] In some embodiments, the polypeptide is a hormone or a hormone-regulating molecule that can regulate the expression and / or secretion of a hormone.
[0131] In some embodiments, the polypeptide is a cytotoxin, such as a protein toxin. In some embodiments, the polypeptide is a TcsT protein (Trichosanthes kirilowii pollen protein (TcsT), whose nucleotide sequence is shown in SEQ ID NO:4).
[0132] In some embodiments, the polypeptide is an anti-angiogenic inhibitor.
[0133] In some embodiments, the polypeptide is a pathogen-specific antigen. The pathogen-specific antigen may be a pathogen-specific antigen of bacteria, viruses, fungi, mycoplasma, chlamydia, parasites, or other pathogens.
[0134] In some embodiments, the polypeptide is a malaria-specific antigen, such as the Plasmodium falciparum-specific antigens PfSir2a and PfRH5.
[0135] In some embodiments, the polypeptide is a specific antigen or related antigen of a certain type of cell in the organism itself. In some embodiments, the polypeptide is a tumor-specific antigen or tumor-associated antigen.
[0136] In some embodiments, the polypeptide is an antigen derived from a protein that causes or progresses a disease. In some embodiments, the polypeptide is GSDMD (gasdermin D) or GSDMD-NT (obtained by cleaving GSDMD after Asp276 and Asp275), a protein involved in pyroptosis (or inflammatory apoptosis) that causes renal tubular damage. In some embodiments, the polypeptide is an antigen derived from a protein that causes or progresses a tumor. For example, RhoA can regulate actin polymerization, cell adhesion, cell transformation, and cell movement, proliferation, and migration closely related to tumor cell invasion and metastasis.
[0137] In some embodiments, the polypeptide is a tag protein or a reporter protein. For example, the tag protein or reporter protein may be BlaM, luciferase, CyaA, β-galactosidase, chloramphenicol acetyltransferase (CAT), secretory phosphatase (SEAP), fluorescent protein, etc. In some embodiments, the tag protein or reporter protein is, for example, BlaM, Renilla Luciferase, or mRFP.
[0138] In some embodiments, the polypeptide is an antimicrobial peptide. For example, the antimicrobial peptide may be defensins, cecropin A and its analogues, magainins, melittin, cecropins, cathelicidin, apidaecins, drosocin, coleoptericin, hemipteracin, bactenecin, cecropin, human neutrophil peptide-1 (HNP-1), hBD-1, hepcidin25, E50-52, PK34, etc.
[0139] In some embodiments, the polypeptide is an enzyme. In some embodiments, the enzyme is an enzyme involved in cellular metabolism. In some embodiments, the enzyme may be selected from Aritilysin, recombinant phage lysin LysSAP26, ribonuclease RNase 3, and RNase 7.
[0140] In some embodiments, the polypeptide is a gene-editing protein. The gene-editing protein may be, for example, a zinc finger nuclease, a TALEN nuclease, Cas9, Cas12 (formerly known as Cpf1), Cas12a, Cas13, Cas13a, Cas13b, etc.
[0141] In some embodiments, the polypeptide is an effector of the bacterial secretion system. In some embodiments, the polypeptide is a T3SS effector, such as EcEspN, EcCif, PaExoU, EcNleC, SfOspB, SfOspF, YpYopT. In some embodiments, the polypeptide is a T4SS effector, such as LpAnkB, BaBspB, HpCagA. In some embodiments, the polypeptide is a T6SS effector, such as EtEvpP, BcTecA, PpTge2, YpYezP, PaTse1, PaTse3, PaPldA, PaPldB.
[0142] Specifically, the lead sequence provided in this application enables the packaging of a payload into a luminescent bacteria virulence box (PVC). Based on this, this application also provides the application of the lead sequence described in the first aspect of this application in guiding the packaging of a payload into a luminescent bacteria virulence box (PVC).
[0143] Secondly, this application provides an isolated nucleic acid that encodes the lead sequence described in the first aspect of this application.
[0144] The nucleic acids include sequences that have been isolated from their natural environment, isolates of recombinant or cloned (e.g., DNA), chemically synthesized analogs, or analogs biosynthesized in heterologous systems.
[0145] The nucleic acid can be prepared by any method known in the art. For example, it can be produced by replication and / or expression in a suitable host cell. Typically, a natural or synthetic DNA fragment encoding the desired segment is incorporated into a recombinant nucleic acid construct (usually a DNA construct) capable of being introduced into and replicated in prokaryotic or eukaryotic cells. Typically, the DNA construct will be adapted for autonomous replication in a single-celled host (such as yeast or bacteria), but can also be used to introduce and integrate into the genome of cultured bacteria, insects, mammals, plants, or other eukaryotic cell lines. The nucleic acid can also be produced by chemical synthesis.
[0146] Thirdly, this application provides an expression vector comprising the isolated nucleic acid described in the second aspect of this application.
[0147] In some embodiments, the expression vector is a plasmid, granule, bacteriophage, or viral vector, preferably a plasmid.
[0148] Fourthly, this application provides a recombinant host cell comprising the isolated nucleic acid described in the second aspect of this application or the expression vector described in the third aspect of this application.
[0149] In some embodiments, the host cell may include bacteria, microorganisms, plant or animal cells. Easily transformable bacteria include members of the Enterobacteriaceae family, such as strains of *Escherichia coli* or *Salmonella*; and members of the Bacillaceae family, such as *Bacillus subtilis*. Suitable microorganisms include *Saccharomyces cerevisiae* and *Pichia pastoris*. In some embodiments, the host cell is derived from the genus *Escherichia*, more preferably *Escherichia coli*, including *Escherichia coli* DH10B, *Escherichia coli* Top10, *Escherichia coli* Trans T1, *Escherichia coli* EPI300, *Escherichia coli* BL21(DE3), etc.
[0150] Fusion protein
[0151] Fifthly, this application provides a fusion protein comprising the lead sequence described in the first aspect of this application and a polypeptide linked thereto, wherein the linkage is a covalent link or a non-covalent link.
[0152] In some embodiments, the connection is a covalent connection. In some embodiments, the leader sequence is attached to the N-terminus of the polypeptide. In some embodiments, by attaching the leader sequence to the N-terminus of the polypeptide, the resulting fusion protein can be recognized by the PVC-V complex and the leader sequence can be partially or completely cleaved, thereby loading the polypeptide into the PVC-V complex.
[0153] In some embodiments, the leader sequence further includes a linker sequence between the polypeptide and the peptide.
[0154] In some embodiments, the fusion protein is not naturally occurring, meaning that the fusion protein formed by the fusion of the leader sequence and polypeptide as described in this application does not exist in nature under natural conditions; that is, the fusion protein is artificially synthesized through recombinant technology.
[0155] In some embodiments, the polypeptide in the fusion protein comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene-editing proteins.
[0156] Sixthly, this application provides an isolated nucleic acid that encodes the fusion protein described in the fifth aspect of this application.
[0157] The nucleic acids include sequences that have been isolated from their natural environment, isolates of recombinant or cloned (e.g., DNA), chemically synthesized analogs, or analogs biosynthesized in heterologous systems.
[0158] The nucleic acid can be prepared by any method known in the art. For example, it can be produced by replication and / or expression in a suitable host cell. Typically, a natural or synthetic DNA fragment encoding the desired segment is incorporated into a recombinant nucleic acid construct (usually a DNA construct) capable of being introduced into and replicated in prokaryotic or eukaryotic cells. Typically, the DNA construct will be adapted for autonomous replication in a single-celled host (such as yeast or bacteria), but can also be used to introduce and integrate into the genome of cultured bacteria, insects, mammals, plants, or other eukaryotic cell lines. The nucleic acid can also be produced by chemical synthesis.
[0159] In a seventh aspect, this application provides an expression vector comprising the isolated nucleic acid described in the sixth aspect of this application.
[0160] In some embodiments, the expression vector is a plasmid, granule, bacteriophage, or viral vector, preferably a plasmid.
[0161] Eighthly, this application provides a recombinant host cell comprising the isolated nucleic acid described in the sixth aspect of this application or the expression vector described in the seventh aspect of this application.
[0162] In some embodiments, the host cell may include bacteria, microorganisms, plant or animal cells. Easily transformable bacteria include members of the Enterobacteriaceae family, such as strains of *Escherichia coli* or *Salmonella*; and members of the Bacillaceae family, such as *Bacillus subtilis*. Suitable microorganisms include *Saccharomyces cerevisiae* and *Pichia pastoris*. In some embodiments, the host cell is derived from the genus *Escherichia*, more preferably *Escherichia coli*, including *Escherichia coli* DH10B, *Escherichia coli* Top10, *Escherichia coli* Trans T1, *Escherichia coli* EPI300, *Escherichia coli* BL21(DE3), etc.
[0163] Payload delivery system
[0164] Ninthly, this application provides a payload delivery system comprising a protein complex and the lead sequence described in the first aspect of this application.
[0165] The payload delivery system is, for example, an extracellular retractable injection system (eCIS).
[0166] The protein complex is derived, for example, from archaea, Gram-negative bacteria, or Gram-positive bacteria. Preferably, the protein complex is derived from non-symbiotic luminescent bacilli, Serratia entomophila, Pseudoalteromonas luteoviolacea, or Klebsiella pneumoniae. More preferably, the protein complex is derived from an extracellular retractable injection system (eCIS) of non-symbiotic luminescent bacilli, Serratia entomophila, Pseudoalteromonas luteoviolacea, or Klebsiella pneumoniae.
[0167] In some embodiments, the protein complex comprises a bioluminescent virulence box (PVC).
[0168] In some embodiments, the protein complex comprises a non-symbiotic luminescent bacterium virulence box (PVC). In some embodiments, the protein complex comprises PVC-I, PVC-II, PVC-III, PVC-IV, or PVC-V. In some embodiments, the PVC is derived from the non-symbiotic luminescent bacterium Photorhabdus asymbiotica ATCC43949.
[0169] In some embodiments, the protein complex is a PVC-V complex, which includes PVC-V structural proteins Pvc1, Pvc2, Pvc3, Pvc4, Pvc5, Pvc6, Pvc7, Pvc8, Pvc9, Pvc10, Pvc11, Pvc12, Pvc13, Pvc14, Pvc15, and Pvc16, wherein the PVC-V structural proteins are derived from the non-symbiotic luminescent bacterium Photorhabdusasymbiotica ATCC43949.
[0170] In some embodiments, the protein complex is a modified PVC-V complex. In some embodiments, the structural protein Pvc13 in the PVC-V complex is modified to recognize cell surface molecules. For example, a protein that recognizes surface molecules is inserted into the receptor-binding domain of the structural protein Pvc13 (the receptor domain has been described in the literature, Kreitz J, Friedrich MJ, Guru A, Lash B, Saito M, Macrae RK, Zhang F. Programmable protein delivery with a bacterial contractile injection system[J]. Nature. 2023; 616:357)). In some embodiments, the receptor-binding domain of the structural protein Pvc13 in the PVC-V complex is modified to recognize cell surface tumor-specific antigens or tumor-associated antigens.
[0171] The payload is a polypeptide, a nucleic acid, or a combination thereof. The payload is loaded into the protein complex under the guidance of the leader sequence.
[0172] In some embodiments, the payload is a polypeptide.
[0173] In some embodiments, the leader sequence is covalently linked to the polypeptide. In some embodiments, the signal peptide is covalently linked to the N-terminus of the polypeptide. In some embodiments, a linker sequence is further included between the leader sequence and the polypeptide.
[0174] In some embodiments, the molecular weight of the polypeptide is 5kDa-200kDa, for example 10kDa-200kDa, 15kDa-200kDa, 20kDa-200kDa, 25kDa-200kDa, 30kDa-200kDa, 40kDa-200kDa, 5kDa-180kDa, 10kDa-180kDa, 15kDa-180kDa, 20kDa-180kDa, 25kDa-180kDa, 30kDa-180kDa, 40kDa-180kDa, 5kDa-160kDa, 10kDa-160kDa, 15kDa-160kDa. Da, 20kDa-160kDa, 25kDa-160kDa, 30kDa-160kDa, 40kDa-160kDa, 5kDa-150kDa, 10kDa-150kDa, 15kDa-150kDa, 20kDa-150kDa, 25kDa-150kDa , 30kDa-150kDa, 40kDa-150kDa, 5kDa-140kDa, 10kDa-140kDa, 15kDa-140kDa, 20kDa-140kDa, 25kDa-140kDa, 30kDa-140kDa, or 40kDa-140kDa.
[0175] In some embodiments, the isoelectric point of the polypeptide is 2-12, for example 2-11.5, 2-11, 2-10.5, 2-10, 2-9.5, 2-9.10, 2.5-11.5, 2.5-11, 2.5-10.5, 2.5-10, 2.5-9.5, 2.5-9.10, 3-11.5, 3-11, 3-10.5, 3-10, 3-9.5, 3-9.10, 3.5-11.5, 3.5-11, 3. 5-10.5, 3.5-10, 3.5-9.5, 3.5-9.10, 4-11.5, 4-11, 4-10.5, 4-10, 4-9.5, 4-9.10, 4.5-11.5, 4.5-11, 4.5-10.5, 4.5-10, 4.5-9.5, 4.5-9.10, 4.6-11.5, 4.6-11, 4.6-10.5, 4.6-10, 4.6-9.5, or 4.6-9.10.
[0176] In some embodiments, the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene-editing proteins.
[0177] In some embodiments, the polypeptide is a hormone or a hormone-regulating molecule that can regulate the expression and / or secretion of a hormone.
[0178] In some embodiments, the polypeptide is a cytotoxin, such as a protein toxin. In some embodiments, the polypeptide is a TcsT protein.
[0179] In some embodiments, the polypeptide in the fusion protein is an anti-angiogenic inhibitor.
[0180] In some embodiments, the polypeptide is a pathogen-specific antigen. The pathogen-specific antigen can be a specific antigen of bacteria, viruses, fungi, mycoplasma, chlamydia, parasites, or other pathogens. The delivery system of this application can be used to incubate with cells in vitro to deliver the pathogen-specific antigen in the delivery system into the cells, thereby producing a vaccine corresponding to the pathogen-specific antigen. After administration to a subject, the vaccine can be used to induce an immune response against the pathogen-specific antigen polypeptide in the subject, thereby preventing and / or treating infection with the pathogen and related diseases.
[0181] In some embodiments, the polypeptide is a malaria-specific antigen, such as the Plasmodium falciparum-specific antigens PfSir2a and PfRH5.
[0182] In some embodiments, the polypeptide is a specific antigen or related antigen of a certain type of cell in the organism itself. In some embodiments, the polypeptide is a tumor-specific antigen or tumor-associated antigen. The delivery system of this application can be used to incubate with cells in vitro to deliver the tumor-specific antigen or tumor-associated antigen in the delivery system into the cells, thereby producing a vaccine against a tumor corresponding to the tumor-specific antigen or tumor-associated antigen. After administration to a subject, the vaccine can induce an immune response against the polypeptide in the subject, thereby preventing and / or treating the associated tumor. The cells may be human embryonic kidney cells HEK293, antigen-presenting cells (e.g., dendritic cells), etc.
[0183] In some embodiments, the polypeptide is an antigen derived from a protein that causes or progresses a disease. This antigen can be delivered to antigen-presenting cells using the delivery system of this application, thereby creating a vaccine. Upon administration to a subject, the vaccine can elicit an immune response against the protein that causes or progresses the disease. In some embodiments, the polypeptide is GSDMD (gasdermin D) or GSDMD-NT (obtained by cleaving GSDMD after Asp276 and Asp275), a protein involved in pyroptosis (or inflammatory apoptosis) that causes renal tubular damage. In some embodiments, the polypeptide is an antigen derived from a protein that causes or progresses tumors. For example, RhoA, which can regulate actin polymerization, cell adhesion, cell transformation, and cell movement, proliferation, and migration closely related to tumor cell invasion and metastasis.
[0184] In some embodiments, the polypeptide is a tag protein or a reporter protein. The tag protein or reporter protein can be delivered to cells using the delivery system of this application, and the cells can be characterized by detecting the signal of the tag protein or reporter protein. For example, the tag protein or reporter protein may be BlaM, luciferase, CyaA, β-galactosidase, chloramphenicol acetyltransferase (CAT), secretory phosphatase (SEAP), fluorescent protein, etc. In some embodiments, the tag protein or reporter protein is, for example, BlaM, Renilla Luciferase, or mRFP.
[0185] In some embodiments, the polypeptide is an antimicrobial peptide. The antimicrobial peptide can be delivered into cells by the delivery system of this application, thereby exerting an antipathogenic effect within the cells. For example, the antimicrobial peptide may be defensins, cecropin A and its analogues, magazineins, melittin, cecropins, cathelicidin, apidaecins, drosocin, coleoptericin, hemipteracin, bactenecin, cecropin, human neutrophil peptide-1 (HNP-1), hBD-1, hepcidin25, E50-52, PK34, etc.
[0186] In some embodiments, the polypeptide is an enzyme. The enzyme can be delivered into cells using the delivery system of this application, thereby exerting a corresponding catalytic effect within the cells. In some embodiments, the enzyme is an enzyme involved in cellular metabolism. In some embodiments, the enzyme may be selected from Aritilysin, recombinant phage lyase LysSAP26, ribonucleases RNase 3 and RNase 7.
[0187] In some embodiments, the polypeptide is a gene-editing protein. The gene-editing protein can be delivered into cells using the delivery system of this application, thereby exerting gene-editing functions within the cells. The gene-editing protein may be, for example, a zinc finger nuclease, a TALEN nuclease, Cas9, Cas12 (formerly known as Cpf1), Cas12a, Cas13, Cas13a, Cas13b, etc.
[0188] In some embodiments, the polypeptide is an effector of the bacterial secretion system. In some embodiments, the polypeptide is a T3SS effector, such as EcEspN, EcCif, PaExoU, EcNleC, SfOspB, SfOspF, YpYopT. In some embodiments, the polypeptide is a T4SS effector, such as LpAnkB, BaBspB, HpCagA. In some embodiments, the polypeptide is a T6SS effector, such as EtEvpP, BcTecA, PpTge2, YpYezP, PaTse1, PaTse3, PaPldA, PaPldB.
[0189] Specifically, for example, the payload delivery system includes a luminescent bacillus virulence box (PVC) and the lead sequence described in the first aspect of this application.
[0190] In a tenth aspect, this application provides an isolated nucleic acid that encodes the payload delivery system described in the ninth aspect of this application.
[0191] In some embodiments, the isolated nucleic acid comprises one or more of the following:
[0192] (i) The nucleotide sequence encoding the protein complex in the payload delivery system;
[0193] (ii) A nucleotide sequence encoding the leader sequence of the payload delivery system;
[0194] (iii) The nucleotide sequence encoding the polypeptide.
[0195] The nucleic acids include sequences that have been isolated from their natural environment, isolates of recombinant or cloned (e.g., DNA), chemically synthesized analogs, or analogs biosynthesized in heterologous systems.
[0196] The nucleic acid can be prepared by any method known in the art. For example, it can be produced by replication and / or expression in a suitable host cell. Typically, a natural or synthetic DNA fragment encoding the desired segment is incorporated into a recombinant nucleic acid construct (usually a DNA construct) capable of being introduced into and replicated in prokaryotic or eukaryotic cells. Typically, the DNA construct will be adapted for autonomous replication in a single-celled host (such as yeast or bacteria), but can also be used to introduce and integrate into the genome of cultured bacteria, insects, mammals, plants, or other eukaryotic cell lines. The nucleic acid can also be produced by chemical synthesis.
[0197] In some implementations, the nucleotide sequence encoding the payload delivery system is codon-optimized. This type of optimization may require mutations in the nucleotide sequence encoding the payload delivery system to mimic the codon preferences of the intended host organism or cell while encoding the same protein.
[0198] In the eleventh aspect, this application provides an expression vector comprising the isolated nucleic acid described in the tenth aspect of this application.
[0199] In some embodiments, the expression vector comprises one or more of the following:
[0200] (i) The nucleotide sequence encoding the protein complex in the payload delivery system;
[0201] (ii) A nucleotide sequence encoding the leader sequence of the payload delivery system;
[0202] (iii) The nucleotide sequence encoding the payload.
[0203] In some embodiments, any two or all three of (i) to (iii) above can be in the same expression vector. In some embodiments, any two or all three of (i) to (iii) above can be in different expression vectors.
[0204] In some embodiments, the expression vector is a plasmid, granule, bacteriophage, or viral vector, preferably a plasmid.
[0205] In a twelfth aspect, this application provides a recombinant host cell comprising the payload delivery system described in the ninth aspect of this application, the isolated nucleic acid described in the tenth aspect of this application, or the expression vector described in the eleventh aspect of this application.
[0206] In some embodiments, the host cell may include bacteria, microorganisms, plant or animal cells. Easily transformable bacteria include members of the Enterobacteriaceae family, such as strains of *Escherichia coli* or *Salmonella*; and members of the Bacillaceae family, such as *Bacillus subtilis*. Suitable microorganisms include *Saccharomyces cerevisiae* and *Pichia pastoris*. In some embodiments, the host cell is derived from the genus *Escherichia*, more preferably *Escherichia coli*, including *Escherichia coli* DH10B, *Escherichia coli* Top10, *Escherichia coli* EPI300, *Escherichia coli* Trans T1, *Escherichia coli* BL21(DE3), etc.
[0207] In some embodiments, the recombinant host cell is capable of expressing PVC-V structural proteins and assembling into a complete PVC-V complex.
[0208] In some embodiments, the recombinant host cell comprises:
[0209] (i) A first vector containing a gene encoding a LysR regulator, preferably under the control of its natural promoter, the first vector being preferably a plasmid, more preferably pBR60;
[0210] (ii) A second vector comprising a nucleotide sequence encoding a PVC-V structural protein, wherein the vector is preferably a plasmid, more preferably a pCNM3 plasmid, and the PVC-V structural protein comprises Pvc1, Pvc2, Pvc3, Pvc4, Pvc5, Pvc6, Pvc7, Pvc8, Pvc9, Pvc10, Pvc11, Pvc12, Pvc13, Pvc14, Pvc15 and Pvc16, wherein the PVC-V structural protein is derived from the non-symbiotic luminescent bacterium Photorhabdus asymbiotica ATCC43949;
[0211] (iii) A third carrier comprising a nucleotide sequence encoding a polypeptide and a lead sequence in the payload delivery system of this application.
[0212] In some embodiments, the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene-editing proteins.
[0213] In a thirteenth aspect, this application provides a method for preparing the payload delivery system described in the ninth aspect of this application, comprising culturing the recombinant host cells described in the twelfth aspect of this application to generate the payload delivery system.
[0214] In some embodiments, the method further includes the step of isolating proteins from cultured recombinant host cells. In some embodiments, the method further includes the steps of purifying the isolated proteins and removing endotoxins.
[0215] Pharmaceutical compositions, reagents or kits
[0216] In a fourteenth aspect, this application provides a pharmaceutical composition comprising the payload delivery system described in the ninth aspect of this application.
[0217] In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier.
[0218] In this application, the term "pharmaceutical acceptable" means that when the molecular basis and the composition are properly administered to animals or humans, they do not produce adverse, allergic, or other adverse reactions.
[0219] In this application, the term "pharmaceuticalally acceptable carrier" means a carrier that is pharmacologically and / or physiologically compatible with the subject and the active ingredient, which is well known in the art (see, for example, Remington's Pharmaceutical Sciences. Edited by Gennaro AR, 19th ed. Pennsylvania: Mack Publishing Company, 1995), including but not limited to binders, diluents, excipients, adjuvants, fillers, disintegrants, wetting agents, lubricants, colorants, flavoring agents, solubilizers, osmolarizers, or other conventional additives.
[0220] Examples of substances that can serve as pharmaceutically acceptable carriers or components thereof include sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives such as sodium carboxymethyl cellulose, ethyl cellulose, and methyl cellulose; tragacanth gum powder; malt; gelatin; talc; solid lubricants such as stearic acid and magnesium stearate; calcium sulfate; vegetable oils such as peanut oil, cottonseed oil, sesame oil, olive oil, corn oil, and cocoa butter; polyols such as propylene glycol, glycerin, sorbitol, mannitol, and polyethylene glycol; alginic acid; emulsifiers such as wetting agents such as sodium lauryl sulfate; colorants; flavoring agents; tableting agents; stabilizers; antioxidants; preservatives; pyrogen-free water; isotonic salt solutions and phosphate buffers, etc.
[0221] In some embodiments, the pharmaceutical composition may be formulated into various dosage forms as needed.
[0222] In the fifteenth aspect, this application provides the use of the lead sequence described in the first aspect of this application, the fusion protein described in the fifth aspect of this application, the payload delivery system described in the ninth aspect of this application, the isolated nucleic acid described in the second / sixth / tenth aspects of this application, the expression vector described in the third / seventh / eleventh aspects of this application, and the recombinant host cell described in the fourth / eighth / twelfth aspects of this application in the preparation of drugs, reagents, or kits for delivering payloads.
[0223] method
[0224] In a sixteenth aspect, this application provides a method for transferring a payload into a target cell, comprising contacting the payload delivery system described in the ninth aspect of this application with the target cell, such that the payload delivery system delivers the payload into the target cell.
[0225] The payload is a polypeptide, a nucleic acid, or a combination thereof.
[0226] In some embodiments, the payload is a polypeptide.
[0227] In some embodiments, the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene-editing proteins.
[0228] In a seventeenth aspect, this application provides a method for regulating cell signal transduction, the method comprising contacting a payload delivery system as described in the ninth aspect of this application with a target cell, thereby causing the payload delivery system to deliver a payload into the target cell.
[0229] The payload is a signaling pathway regulatory protein.
[0230] In an eighteenth aspect, this application provides a method for regulating cell metabolism, the method comprising contacting the payload delivery system described in the ninth aspect of this application with target cells, thereby causing the payload delivery system to deliver a payload into the target cells.
[0231] The payload is an enzyme involved in cell metabolism.
[0232] In a nineteenth aspect, this application provides a method for regulating molecular transport and / or secretion in cells, the method comprising contacting a payload delivery system as described in the ninth aspect of this application with a target cell, thereby causing the payload delivery system to deliver a payload into the target cell.
[0233] The payload is a transport protein.
[0234] In a twentieth aspect, this application provides a method for gene editing of cells, the method comprising contacting the payload delivery system described in the ninth aspect of this application with a target cell, thereby causing the payload delivery system to deliver the payload into the target cell.
[0235] The payload is a gene-editing protein.
[0236] In a twentieth aspect, this application provides a method for labeling cells, the method comprising contacting a payload delivery system as described in the ninth aspect of this application with a target cell, thereby causing the payload delivery system to deliver the payload into the target cell.
[0237] The payload is a tag protein or a reporter protein.
[0238] In a twentieth aspect, this application provides a method for enhancing the ability of cells to resist pathogens, the method comprising contacting the payload delivery system described in the ninth aspect of this application with target cells, thereby causing the payload delivery system to deliver the payload into the target cells.
[0239] The payload is an antimicrobial peptide.
[0240] In a twentieth aspect, this application provides a method for preparing a cell vaccine, the method comprising contacting a payload delivery system as described in the ninth aspect of this application with target cells, thereby causing the payload delivery system to deliver the payload into the target cells.
[0241] The payload is an antigen or an immunogen.
[0242] The target cells mentioned above are, for example, eukaryotic cells, including but not limited to yeast cells, insect cells, mammalian cells, plant cells, or fungal cells. In some embodiments, the target cells are human cells.
[0243] In a twentieth aspect, this application provides a method for treating and / or preventing disease, comprising administering the payload delivery system described in the ninth aspect of this application to a subject.
[0244] The payload delivery system carries a therapeutically effective amount of biomolecules.
[0245] The disease can be determined based on the function of the biological macromolecules used.
[0246] In a twentieth aspect, this application provides the use of the payload delivery system described in the ninth aspect of this application in the preparation of medicaments for treating and / or preventing diseases.
[0247] The payload delivery system carries a therapeutically effective amount of biomolecules.
[0248] The disease can be determined based on the function of the biological macromolecules used.
[0249] Example
[0250] The following description, in conjunction with specific embodiments, illustrates the content of this application, but the scope of this application is not limited thereto. Unless otherwise specified, the reagents and instruments used in the following embodiments are all conventional reagents and instruments in the art and can be obtained commercially. The methods used are all conventional experimental methods, and those skilled in the art can undoubtedly implement the described schemes and obtain corresponding results based on the embodiments.
[0251] Example 1: The XK1, XK2, and XK3 leader sequences can be used for packaging peptides in PVC complexes.
[0252] (1) Strains construction
[0253] Plasmid construction: In the Xenorhabdus khoisanae MCB strain, the XkCIS structural genes gene cluster contains three effectors: XK1 (encoded by AB204_RS00510), XK2 (encoded by AB204_RS00495), and XK3 (encoded by AB204_RS00490).
[0254] The leader sequences (XK1-N70, XK2-N70, and XK3-N70, amino acid sequences shown in SEQ ID NO:1-3) of the N-terminus of the effector factors (XK1, XK2, and XK3) were amplified by PCR and fused with the TcsT gene (nucleotide sequence shown in SEQ ID NO:4) with a Flag tag. The fusion results were then inserted into the pBBR-MCS5 vector (described in the literature, ME, Kovach, PH, Elzer, DS, Hill et al. Four new derivatives of the broad-host-range cloning vector pBBR1MCS, carrying different antibiotic-resistance cassettes. [J]. Gene, 1995, 166:0.) to construct recombinant plasmids XK1-N70-TcsT-Flag, XK2-N70-TcsT-Flag, and XK3-N70-TcsT-Flag.
[0255] The constructed recombinant plasmids were transformed into EPI300 strains expressing PVC particles (derived from non-symbiotic luminescent bacterium Photorhabdus asymbiotica ATCC43949), with TcsT strains without any leader sequence serving as negative controls.
[0256] (2) Strain culture
[0257] The transformed strain was cultured overnight at 37°C in TB medium, then transferred to 200 ml of TB liquid medium and cultured at 30°C for 24 hours. The bacterial pellet was collected by centrifugation and stored at -80°C overnight.
[0258] (3) Extraction and purification of PVC complex (i.e. PVC-V complex)
[0259] Referring to the extraction and purification methods for PVC complexes described in patents CN116554283A and CN119735656A, PVC complexes loaded with TcsT protein were extracted.
[0260] (4) Western Blot Analysis
[0261] The extracted PVC complex was separated by SDS-PAGE electrophoresis. The loading of TcsT protein was detected by Flag antibody (purchased from Beijing Huaxing Bochuang Gene Technology Co., Ltd., HX1801). Western blotting analysis was used to verify whether the XK1, XK2 and XK3 leader sequences could guide the loading of TcsT protein into the PVC complex. Anti-PVC16 antibody (custom-made by Nanjing Genscript Biotech Co., Ltd.) was used to detect PVC-V structural protein. Secondary antibodies were purchased from Beijing Huaxing Bochuang Gene Technology Co., Ltd. (HX2032, HX2031).
[0262] The results are as follows Figure 1 As shown, the N70 leader sequences of effector factors XK1, XK2, and XK3 can all effectively guide the loading of TcsT protein into the PVC complex.
[0263] Example 2: Comparison and testing of the packaging capabilities of XK3 leader sequences of different lengths
[0264] (1) Strains construction
[0265] Different lengths of leader sequences at the N-terminus of the effector factor XK3 were amplified by PCR, with amino acid sequences shown in SEQ ID NO:3 (XK3-N70) and SEQ ID NO:5-24 (XK3-N20, XK3-N25, XK3-N30, XK3-N35, XK3-N45, XK3-N50, XK3-N55, XK3-N60, XK3-N65, XK3-N70, XK3-N75, XK3-N80, XK3-N85, XK3-N90, XK3-N95, XK3-N100, XK3-N105, XK3-N110, XK3-N115, XK3-N120). These leader sequences were then compared with the TcsT gene (with a Flag tag, the nucleotide sequence of which is shown in SEQ ID NO:3) to XK3-N70. (As shown in NO:4) The plasmid was fused and inserted into the pBBR-MCS5 vector to construct the recombinant plasmid.
[0266] The constructed recombinant plasmids were transformed into EPI300 strains expressing PVC particles. The N-terminal 50 amino acids (pdp1-N50, the amino acid sequence of which is shown in SEQ ID NO:185) of the effector pdp1 (PAU_RS16575, derived from non-symbiotic luminescent bacterium Photorhabdusasymbiotica ATCC43949), which is downstream of the PVC-V gene cluster of the luminescent bacterium virulence box, were used as a positive control.
[0267] (2) Strain culture
[0268] The transformed strain was cultured overnight at 37°C in TB medium, then transferred to 200 ml of TB liquid medium and cultured at 30°C for 24 hours. The bacterial pellet was collected by centrifugation and stored at -80°C overnight.
[0269] (3) Extraction and purification of PVC complex (i.e. PVC-V complex)
[0270] Referring to the extraction and purification methods for PVC complexes described in patents CN116554283A and CN119735656A, PVC complexes loaded with TcsT protein were extracted.
[0271] (4) Western Blot Analysis
[0272] The extracted PVC complex was separated by SDS-PAGE electrophoresis, and the loading of TcsT protein was detected by Flag antibody. Western blotting analysis was used to verify whether the N-terminal leader sequences of different lengths of XK3 could guide the loading of TcsT protein into the PVC complex.
[0273] As shown in Figure 2, the N20-N120 leader sequences of the effector factor XK3 were able to guide the loading of TcsT protein into the PVC complex, indicating that the leader sequence has good packaging ability within a certain range.
[0274] Example 3: Cell-level testing
[0275] The PVC complex loaded with XK3-N70-TcsT (with a flag) (PVC / XK3-N70TcsT) and the empty PVC complex without any protein were extracted and purified, respectively (construction, extraction and purification methods were the same as in Example 1); mouse monocyte-macrophage J774A.1 (purchased from Peking Union Medical College Cell Resource Center) were seeded in DMEM medium (2 x 10⁻⁶ cells / year). 4 Cells were cultured in 96-well plates at 37°C and 5% CO2. Purified empty PVC complex and PVC / XK3-N70TcsT were added to J774A.1 cells at different concentrations. After incubation for 24 hours, cell viability was detected by CCK8 assay (CCK8 kit purchased from APExBIO, K1018), and the cell-killing effect of PVC was evaluated.
[0276] The results are as follows Figure 3As shown, PVC / XK3-N70TcsT significantly reduced the survival rate of J774A.1 cells, indicating that TcsT protein has a cytotoxic effect at the cellular level. Furthermore, this result also confirms that the N70 leader sequence of the effector factor XK3 can guide the loading of TcsT protein into the PVC complex, and that PVC can inject TcsT protein into cells.
[0277] Example 4: Testing the packaging capacity of N20-N70 leader sequences of P. asymbiotica effector factors F1, F2, F3, F5, F7, F8, F10, F11, F12, F13, and F16.
[0278] (1) Strains construction
[0279] Plasmid construction: Using the genome of the non-symbiotic luminescent bacterium Photorhabdus asymbiotica ATCC43949 as a template, the N-terminal 20-70 amino acid leader sequences of the following effector factors F1 (encoded by PAU_RS09715), F2 (encoded by PAU_RS09720), F3 (encoded by PAU_RS09725), F5 (encoded by PAU_RS10135), F7 (encoded by PAU_RS10125), F8 (encoded by PAU_RS10120), F10 (encoded by PAU_RS13645), F11 (encoded by PAU_RS22355), F12 (encoded by PAU_RS13655), F13 (encoded by PAU_RS13660), and F16 (encoded by PAU_RS16545) were amplified by PCR (the amino acid sequences are shown in SEQ ID). (as shown in NO:25-90), and fused with the TcsT gene with the Flag tag (the nucleotide sequence of the TcsT gene is shown in SEQ ID NO:4), and inserted into the pBBR-MCS5 vector to construct the recombinant plasmid.
[0280] The constructed recombinant plasmids were transformed into EPI300 strains expressing PVC particles.
[0281] (2) Strain culture
[0282] The transformed strain was cultured overnight at 37°C in TB medium, then transferred to 200 ml of TB liquid medium and cultured at 30°C for 24 hours. The bacterial pellet was collected by centrifugation and stored at -80°C overnight.
[0283] (3) Extraction and purification of PVC complex (i.e. PVC-V complex)
[0284] Referring to the extraction and purification methods for PVC complexes described in patents CN116554283A and CN119735656A, PVC complexes loaded with TcsT protein were extracted.
[0285] (4) Western Blot Analysis
[0286] The extracted PVC complex was separated by SDS-PAGE electrophoresis, and the loading of TcsT protein was detected by Flag antibody. Western blotting analysis was used to verify whether the N-terminal leader sequences of different lengths of F1, F2, F3, F5, F7, F8, F10, F11, F12, F13 and F16 could guide the loading of TcsT protein into the PVC complex.
[0287] The results are as follows Figures 4A-4K As shown, among the P. asymbiotica effector leader sequences, except for F2-N20, F3-N20, F8-N20, F11-N20, F13-N30 and F16-N20, the other leader sequences can guide the loading of TcsT protein into the PVC complex and have good packaging ability.
[0288] Example 5: Testing the packaging ability of P. asymbiotica leader sequences for other effector factors.
[0289] (1) Strains construction
[0290] Plasmid construction: Using the genome of the non-symbiotic luminescent bacterium Photorhabdus asymbiotica ATCC43949 as a template, the N-terminal leader sequences of the following effectors were amplified by PCR: PAU_RS09615-N70 (SEQ ID NO:91), PAU_RS09620-N70 (SEQ ID NO:92), PAU_RS09625-N70 (SEQ ID NO:93), PAU_RS09630-N70 (SEQ ID NO:94), yciH(PAU_RS10065)-N70 (SEQ ID NO:95), PAU_RS10070-N70 (SEQ ID NO:96), PAU_RS10075-N70 (SEQ ID NO:97), PAU_RS10080-N70 (SEQ ID NO:98), PAU_RS10085-N70 (SEQ ID NO:98), and PAU_RS10085-N70 (SEQ ID NO:99). NO:99), PAU_RS10100-N70 (SEQ ID NO:100), PAU_RS25580-N70 (SEQ ID NO:101), PAU_RS10205-N70 (SEQ ID NO:102), nth (PAU_RS10210)-N70 (SEQ ID NO:103), rsxC(PAU_RS10230)-N70(SEQ ID NO:104), rsxB(PAU_RS10235)-N70(SEQ ID NO:105), PAU_RS10715-N70(SEQ ID NO:106), PAU_RS10720-N70(SEQ ID NO:107), PAU_RS10725-N70 (SEQ ID NO:108), PAU_RS10730-N70 (SEQ ID NO:109), PAU_RS10855-N70 (SEQ ID NO:110), PAU_RS10870-N77 (SEQ ID NO:111), PAU_RS13665-N70 (SEQ ID NO:112), PAU_RS13670-N79 (SEQID NO:113), PAU_RS23900-N62 (SEQ ID NO:114), cmoM (PAU_RS13690)-N70 (SEQ ID NO:115), elyC (PAU_RS13695)-N70 (SEQ ID NO:116), kdsB (PAU_RS13700)-N70 (SEQ ID NO:117), PAU_RS13560-N70 (SEQ ID NO:118), PAU_RS13520-N70 (SEQ IDNO:119)、PAU_RS22390-N70(SEQ ID NO:120)、PAU_RS16505-N70(SEQ ID NO:121)、PAU_RS16515-N100(SEQ ID NO:122)、PAU_RS25720-N59(SEQ ID NO:123)、PAU_RS16660-N102(SEQ ID NO:124)、PAU_RS16665-N70(SEQ ID NO:125)、PAU_RS16690-N70(SEQ ID NO:126)、PAU_RS16695-N70(SEQ ID NO:127)、PAU_RS09595-N70(SEQ ID NO:128)、PAU_RS09735-N70(SEQ ID NO:129)、PAU_RS09740-N70(SEQ ID NO:130)、PAU_RS09750-N70(SEQ ID NO:131)、PAU_RS09755-N70(SEQ ID NO:132)、PAU_RS10105-N70(SEQ ID NO:133)、rsxG(PAU_RS10220)-N70(SEQ ID NO:134)、rsxD(PAU_RS10225)-N70(SEQ ID NO:135)、PAU_RS10245-N70(SEQID NO:136)、PAU_RS10865-N70(SEQ ID NO:137)、mukB(PAU_RS13675)-N70(SEQ ID NO:138)、PAU_RS13540-N70(SEQ ID NO:139), pilV(PAU_RS13535)-N70(SEQ ID NO:140), PAU_RS13530-N70(SEQ ID NO:141), PAU_RS13525-N70(SEQ ID NO:142), PAU_RS16495-N70(SEQ ID NO:143), PAU_RS09715-N70(SEQ ID NO:30), PAU_RS09725-N70(SEQ ID NO:42), PAU_RS09730-N70(SEQ ID NO:144), PAU_RS10135-N70(SEQ ID NO:48), PAU_RS10125-N70(SEQ ID NO:54), PAU_RS10120-N70(SEQ ID NO:55) NO:60)、PAU_RS10765-N70(SEQ IDThe recombinant plasmids were constructed by fusing the above-mentioned leader sequences with the TcsT gene (the nucleotide sequence of the TcsT gene is shown in SEQ ID NO:4) with a Flag tag, and inserting them into the pBBR-MCS5 vector.
[0291] The constructed recombinant plasmids were transformed into EPI300 strains expressing PVC particles.
[0292] (2) Strain culture
[0293] The transformed strain was cultured overnight at 37°C in TB medium, then transferred to 200 ml of TB liquid medium and cultured at 30°C for 24 hours. The bacterial pellet was collected by centrifugation and stored at -80°C overnight.
[0294] (3) Extraction and purification of PVC complex (i.e. PVC-V complex)
[0295] Referring to the extraction and purification methods for PVC complexes described in patents CN116554283A and CN119735656A, PVC complexes loaded with TcsT protein were extracted.
[0296] (4) Western Blot Analysis
[0297] The extracted PVC complex was separated by SDS-PAGE electrophoresis, and the loading of TcsT protein was detected by Flag antibody. Western blotting analysis was used to verify whether the lead sequences of other effector factors of P. asymbiotica could guide the loading of TcsT protein into the PVC complex.
[0298] The results are as follows Figure 5 As shown, the N-terminal leader sequences of various effector factors can guide the loading of TcsT protein into the PVC complex, which further verifies the diversity of leader sequences.
[0299] Example 6: Testing the packaging ability of other strains of effector leader sequences
[0300] (1) Strains construction
[0301] This embodiment detected the lead sequences of effector factors from other strains (Yersinia similis (YP), Yersinia pekkanenii (YPe), Yersinia ruckeri (Yr), and burgdorferi Xenorhabdus bovienii (Xb)):
[0302] The following sequences were synthesized by Beijing Qingke Technology Co., Ltd.: YP1 (encoded by BF17_RS02745)-N70 (SEQ ID NO:150), YP2 (encoded by BF17_RS02855)-N70 (SEQ ID NO:151), YP3 (encoded between BF17_RS02845 and BF17_RS24790)-N66 (SEQ ID NO:152), and YP3 (encoded by BF17_RS24790)-N50 (SEQ ID NO:153) from Yersinia similis 228 strain; YPe1 (encoded by AEP37_RS09605)-N70 (SEQ ID NO:154) and YPe2 (encoded by AEP37_RS09610)-N70 (SEQ ID NO:154) from Yersinia pekkanenii A125KOH2 strain. NO:155); Yr2 (encoded by LGL87_RS15575)-N70 (SEQ ID NO:156), Yr4 (encoded by LGL87_RS15545)-N70 (SEQ ID NO:157), Yr5 (encoded by LGL87_RS15540)-N70 (SEQ ID NO:158) from the Yersinia ruckeri strain; Xb3 (encoded by AACW61_RS01295)-N70 (SEQ ID NO:159) leader sequences from the Xenorhabdus bovienii strain.
[0303] Plasmid construction: The above-mentioned leader sequence was fused with the TcsT gene with the Flag tag (the nucleotide sequence of the TcsT gene is shown in SEQ ID NO:4) and inserted into the pBBR-MCS5 vector to construct the recombinant plasmid.
[0304] The constructed recombinant plasmids were transformed into EPI300 strain expressing PVC particles, with pdp1-N50 (whose amino acid sequence is shown in SEQ ID NO:185) as a positive control.
[0305] (2) Strain culture
[0306] The transformed strain was cultured overnight at 37°C in TB medium, then transferred to 200 ml of TB liquid medium and cultured at 30°C for 24 hours. The bacterial pellet was collected by centrifugation and stored at -80°C overnight.
[0307] (3) Extraction and purification of PVC complex (i.e. PVC-V complex)
[0308] Referring to the extraction and purification methods for PVC complexes described in patents CN116554283A and CN119735656A, PVC complexes loaded with TcsT protein were extracted.
[0309] (4) Western Blot Analysis
[0310] The extracted PVC complex was separated by SDS-PAGE electrophoresis, and the loading of TcsT protein was detected by Flag antibody. Western blotting analysis was used to verify whether the leader sequence from other strains could guide the loading of TcsT protein into the PVC complex.
[0311] The results are as follows Figure 6 As shown, YP2-N70, YPe1-N70, YPe2-N70, Yr2-N70 and Xb3-N70 can all guide the loading of TcsT protein into the PVC complex, and the loading efficiency of XK2-N70 and XK3-N70 is better than that of Pdp1-N50.
[0312] The above description is merely a preferred embodiment of this application and is not intended to limit the application in any other way. Any person skilled in the art may make changes or modifications to the disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the protection scope of this application.
[0313] The sequences involved in this application are as follows:
[0314] The amino acid sequence of XK1-N70:
[0315] MYSLKQKEEKKNHSPRPTQNTVNGDHAITPRIAELRKNFQQLNSQTVRPRVAPRPLHSLSKTGPFQQATK(SEQ ID NO:1);
[0316] The amino acid sequence of XK2-N70:
[0317] MIYGYDNKMVYRKRQESNTALEALHLDIPDTAEVSAAYQRDLSSVMTRFQNKTDSILEESKQPSPLKQKP(SEQ ID NO:2);
[0318] The amino acid sequence of XK3-N70:
[0319] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQGHYVNHSKLPAGTKLSEIAM(SEQ ID NO:3);
[0320] The nucleotide sequence of the TcsT gene:
[0321] GATGTTAGCTTTCGTCTGAGCGGGGCAACCAGCAGCAGCTATGGTGTTTTTATTTCTAATCTGCGTAAAGCACTGCCGAATGAACGTAAACTGTATGATATTCCGCTGCTGCGTAGTAGCCTGCCGGGTAGTCAACGTTATGCACTGATTCATCTGACCAATTATGCCGACGAAACCATCAGCGTGGCAATTGACGTTACCAATGTATATATTATGGGCTATCGTGCCGGTGATACAAGCTATTTTTTTAATGAAGCAAGCGCAACCGAAGCAGCAAAATATGTTTTTAAAGACGCCATGCGTAAAGTGACCCTGCCGTATAGTGGTAACTACGAACGCCTGCAAACAGCAGCAGGAAAAATTCGTGAAAATATCCCGCTGGGCCTGCCGGCACTGGATAGCGCAATCACCACCCTGTTTTATTATAACGCCAATTCTGCAGCAAGCGCCCTGATGGTTCTGATTCAGAGTACAAGCGAAGCAGCACGTTACAAATTTATTGAACAGCAGATTGGGAAACGTGTCGACAAAACCTTTCTGCCGTCACTGGCAATTATTAGTCTGGAAAATAGCTGGTCTGCACTGAGCAAACAAATCCAAATTGCAAGCACCAATAACGGGCAGTTTGAAAGCCCGGTTGTTCTGATTAACGCCCAAAATCAGCGTGTGACCATCACCAATGTGGACGCAGGTGTCGTTACCAGCAACATCGCCCTGCTGCTGAATCGCAATAATATGGCA(SEQ ID NO:4);
[0322] Amino acid sequence of XK3-N20:
[0323] MPVYENKEKRKYSKNPLHNT(SEQ ID NO:5);
[0324] Amino acid sequence of XK3-N25:
[0325] MPVYENKEKRKYSKNPLHNTSTINQ(SEQ ID NO:6);
[0326] Amino acid sequence of XK3-N30:
[0327] MPVYENKEKRKYSKNPLHNTSTINQFTVQD(SEQ ID NO:7);
[0328] The amino acid sequence of XK3-N35:
[0329] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSAS(SEQ ID NO:8);
[0330] The amino acid sequence of XK3-N45:
[0331] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLT(SEQ ID NO:9);
[0332] The amino acid sequence of XK3-N50:
[0333] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ(SEQ ID NO:10);
[0334] The amino acid sequence of XK3-N55:
[0335] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVN(SEQ ID NO:11);
[0336] The amino acid sequence of XK3-N60:
[0337] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLP(SEQ IDNO:12);
[0338] The amino acid sequence of XK3-N65:
[0339] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKL(SEQ ID NO:13);
[0340] The amino acid sequence of XK3-N70:
[0341] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQGHYVNHSKLPAGTKLSEIAM(SEQ ID NO:14);
[0342] The amino acid sequence of XK3-N75:
[0343] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQGHYVNHSKLPAGTKLSEIAMLGSHD(SEQ ID NO:15);
[0344] The amino acid sequence of XK3-N80:
[0345] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQGHYVNHSKLPAGTKLSEIAMLGSHDAGTYA(SEQ ID NO:16);
[0346] The amino acid sequence of XK3-N85:
[0347] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRK(SEQ ID NO:17);
[0348] The amino acid sequence of XK3-N90:
[0349] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFAS(SEQ ID NO:18);
[0350] The amino acid sequence of XK3-N95:
[0351] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSL(SEQ ID NO:19);
[0352] The amino acid sequence of XK3-N100:
[0353] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSLLPAAF(SEQ ID NO:20);
[0354] The amino acid sequence of XK3-N105:
[0355] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSLLPAAFKTQ NR(SEQ ID NO:21);
[0356] The amino acid sequence of XK3-N110:
[0357] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSLLPAAFKTQ NRTLRQQ (SEQ ID NO: 22);
[0358] The amino acid sequence of XK3-N115:
[0359] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSLLPAAFKTQ NRTLRQQAEAGA (SEQ ID NO: 23);
[0360] The amino acid sequence of XK3-N120:
[0361] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQ GHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSLLPAAFKTQ NRTLRQQAEAGARYFDI(SEQ ID NO:24);
[0362] The amino acid sequence of F1-N20:
[0363] MMREYSKEDDCVKEKTNLAE(SEQ ID NO:25);
[0364] The amino acid sequence of F1-N30:
[0365] MMREYSKEDDCVKEKTNLAESENVEADNYL(SEQ ID NO:26);
[0366] The amino acid sequence of F1-N40:
[0367] MMREYSKEDDCVKEKTNLAESENVEADNYLEMDCLNYLAK(SEQ ID NO:27);
[0368] The amino acid sequence of F1-N50:
[0369] MMREYSKEDDCVKEKTNLAESENVEADNYLEMDCLNYLAKLNGMPER KDH(SEQ ID NO:28);
[0370] The amino acid sequence of F1-N60:
[0371] MMREYSKEDDCVKEKTNLAESENVEADNYLEMDCLNYLAKLNGMPER KDHSLNSTKLIDD(SEQ IDNO:29);
[0372] The amino acid sequence of F1-N70:
[0373] MMREYSKEDDCVKEKTNLAESENVEADNYLEMDCLNYLAKLNGMPERKDHSLNSTKLIDDIIKLHNDRKG(SEQ ID NO:30);
[0374] The amino acid sequence of F2-N20:
[0375] MPIIGHKEDLIRTERSSVDL(SEQ ID NO:31);
[0376] The amino acid sequence of F2-N30:
[0377] MPIIGHKEDLIRTERSSVDLTRSSNNRQTD(SEQ ID NO:32);
[0378] The amino acid sequence of F2-N40:
[0379] MPIIGHKEDLIRTERSSVDLTRSSNNRQTDNLELNIPQHK (SEQ ID NO:33); Amino acid sequence of F2-N50:
[0380] MPIIGHKEDLIRTERSSVDLTRSSNNRQTDNLELNIPQHKRDNKDIEHAV(SEQ ID NO:34);
[0381] The amino acid sequence of F2-N60:
[0382] MPIIGHKEDLIRTERSSVDLTRSSNNRQTDNLELNIPQHKRDNKDIEHAVIY GFSQHRGP(SEQ IDNO:35);
[0383] The amino acid sequence of F2-N70:
[0384] MPIIGHKEDLIRTERSSVDLTRSSNNRQTDNLELNIPQHKRDNKDIEHAVIYGFSQHRGPEMQKAFADNK(SEQ ID NO:36);
[0385] The amino acid sequence of F3-N20:
[0386] VKYDPRLRTWVEDDFDYEKN(SEQ ID NO:37);
[0387] The amino acid sequence of F3-N30:
[0388] VKYDPRLRTWVEDDFDYEKNFKKQMDYINY(SEQ ID NO:38);
[0389] The amino acid sequence of F3-N40:
[0390] VKYDPRLRTWVEDDFDYEKNFKKQMDYINYKDLEKQLKEN(SEQ ID NO:39);
[0391] The amino acid sequence of F3-N50:
[0392] VKYDPRLRTWVEDDFDYEKNFKKQMDYINYKDLEKQLKENVDYYALLD EN(SEQ ID NO:40);
[0393] The amino acid sequence of F3-N60:
[0394] VKYDPRLRTWVEDDFDYEKNFKKQMDYINYKDLEKQLKENVDYYALLD ENEAIIFLKELG(SEQ IDNO:41);
[0395] The amino acid sequence of F3-N70:
[0396] VKYDPRLRTWVEDDFDYEKNFKKQMDYINYKDLEKQLKENVDYYALLDENEAIIFLKELGCDIKSLLNGT(SEQ ID NO:42);
[0397] The amino acid sequence of F5-N20:
[0398] MVFEHDKTVERKRKPSIQLG(SEQ ID NO:43);
[0399] The amino acid sequence of F5-N30:
[0400] MVFEHDKTVERKRKPSIQLGNDKEKSSEQA(SEQ ID NO:44);
[0401] The amino acid sequence of F5-N40:
[0402] MVFEHDKTVERKRKPSIQLGNDKEKSSEQALELPQSKQNN(SEQ ID NO:45);
[0403] The amino acid sequence of F5-N50:
[0404] MVFEHDKTVERKRKPSIQLGNDKEKSSEQALELPQSKQNNPLLHDLITSN(SEQ ID NO:46);
[0405] The amino acid sequence of F5-N60:
[0406] MVFEHDKTVERKRKPSIQLGNDKEKSSEQALELPQSKQNNPLLHDLITSN NLRKEAAVFA(SEQ IDNO:47);
[0407] The amino acid sequence of F5-N70:
[0408] MVFEHDKTVERKRKPSIQLGNDKEKSSEQALELPQSKQNNPLLHDLITSNNLRKEAAVFAKQIGPSYQGI(SEQ ID NO:48);
[0409] The amino acid sequence of F7-N20:
[0410] MEREYNKKEKQKKSAIKLDD(SEQ ID NO:49);
[0411] The amino acid sequence of F7-N30:
[0412] MEREYNKKEKQKKSAIKLDDAVGNNEENMD(SEQ ID NO:50);
[0413] The amino acid sequence of F7-N40:
[0414] MEREYNKKEKQKKSAIKLDDAVGNNEENMDMTSPLELNSQ(SEQ ID NO:51);
[0415] The amino acid sequence of F7-N50:
[0416] MEREYNKKEKQKKSAIKLDDAVGNNEENMDMTSPLELNSQYTNRKRPGLR(SEQ ID NO:52);
[0417] The amino acid sequence of F7-N60:
[0418] MEREYNKKEKQKKSAIKLDDAVGNNEENMDMTSPLELNSQYTNRKRPG LRERFSATLQRN(SEQ IDNO:53);
[0419] The amino acid sequence of F7-N70:
[0420] MEREYNKKEKQKKSAIKLDDAVGNNEENMDMTSPLELNSQYTNRKRPGLRERFSATLQRNLPGHSMLDRE(SEQ ID NO:54);
[0421] The amino acid sequence of F8-N20:
[0422] MISTFDPAICAGTPTVTVLD(SEQ ID NO:55);
[0423] The amino acid sequence of F8-N30:
[0424] MISTFDPAICAGTPTVTVLDNRNLTVREIV (SEQ ID NO:56);
[0425] The amino acid sequence of F8-N40:
[0426] MISTFDPAICAGTPTVTVLDNRNLTVREIVFHRAKAGGDT (SEQ ID NO:57); Amino acid sequence of F8-N50:
[0427] MISTFDPAICAGTPTVTVLDNRNLTVREIVFHRAKAGGDTDTLITRHQYD(SEQ ID NO:58);
[0428] The amino acid sequence of F8-N60:
[0429] MISTFDPAICAGTPTVTVLDNRNLTVREIVFHRAKAGGDTDTLITRHQYDL RGNLTQSLD(SEQ IDNO:59);
[0430] The amino acid sequence of F8-N70:
[0431] MISTFDPAICAGTPTVTVLDNRNLTVREIVFHRAKAGGDTDTLITRHQYDLRGNLTQSLDPRLYDLMQKD(SEQ ID NO:60);
[0432] The amino acid sequence of F10-N20:
[0433] MPNKKYSENTHQGKKPLIKS(SEQ ID NO:61);
[0434] The amino acid sequence of F10-N30:
[0435] MPNKKYSENTHQGKKPLIKSEANNEHAIDN(SEQ ID NO:62);
[0436] The amino acid sequence of F10-N40:
[0437] MPNKKYSENTHQGKKPLIKSEANNEHAIDNSPLGIGLDLN(SEQ ID NO:63);
[0438] The amino acid sequence of F10-N50:
[0439] MPNKKYSENTHQGKKPLIKSEANNEHAIDNSPLGIGLDLNSILGNNSASL(SEQ ID NO:64);
[0440] The amino acid sequence of F10-N60:
[0441] MPNKKYSENTHQGKKPLIKSEANNEHAIDNSPLGIGLDLNSILGNNSASLS QIHDYSFWK(SEQ IDNO:65);
[0442] The amino acid sequence of F10-N70:
[0443] MPNKKYSENTHQGKKPLIKSEANNEHAIDNSPLGIGLDLNSILGNNSASLSQIHDYSFWKENISEYYKWM(SEQ ID NO:66);
[0444] The amino acid sequence of F11-N20:
[0445] MEREYSEKEKHKKHPIQLRD(SEQ ID NO:67);
[0446] The amino acid sequence of F11-N30:
[0447] MEREYSEKEKHKKHPIQLRDAIEQHAEETA(SEQ ID NO:68);
[0448] The amino acid sequence of F11-N40:
[0449] MEREYSEKEKHKKHPIQLRDAIEQHAEETANNSLGLGLDL(SEQ ID NO:69);
[0450] The amino acid sequence of F11-N50:
[0451] MEREYSEKEKHKKHPIQLRDAIEQHAEETANNSLGLGLDLHQAINTPKVP(SEQ ID NO:70);
[0452] The amino acid sequence of F11-N60:
[0453] MEREYSEKEKHKKHPIQLRDAIEQHAEETANNSLGLGLDLHQAINTPKVP KDNYNEENGD(SEQ IDNO:71);
[0454] The amino acid sequence of F11-N70:
[0455] MEREYSEKEKHKKHPIQLRDAIEQHAEETANNSLGLGLDLHQAINTPKVPKDNYNEENGDLFYGLAAQRG(SEQ ID NO:72);
[0456] The amino acid sequence of F12-N20:
[0457] MVHEYSINDRQKRHSFSSAN(SEQ ID NO:73);
[0458] The amino acid sequence of F12-N30:
[0459] MVHEYSINDRQKRHSFSSANPIDPEVTNRE(SEQ ID NO:74);
[0460] The amino acid sequence of F12-N40:
[0461] MVHEYSINDRQKRHSFSSANPIDPEVTNRENSRHRFPKDN(SEQ ID NO:75);
[0462] The amino acid sequence of F12-N50:
[0463] MVHEYSINDRQKRHSFSSANPIDPEVTNRENSRHRFPKDNYNKGHGDLFY(SEQ ID NO:76);
[0464] The amino acid sequence of F12-N60:
[0465] MVHEYSINDRQKRHSFSSANPIDPEVTNRENSRHRFPKDNYNKGHGDLF YGLAPERGKYI(SEQ IDNO:77);
[0466] The amino acid sequence of F12-N70:
[0467] MVHEYSINDRQKRHSFSSANPIDPEVTNRENSRHRFPKDNYNKGHGDLFYGLAPERGKYIKEANPKFDPN(SEQ ID NO:78);
[0468] The amino acid sequence of F13-N20:
[0469] MNISSYFFLNEENIKFNNQY(SEQ ID NO:79);
[0470] The amino acid sequence of F13-N30:
[0471] MNISSYFFLNEENIKFNNQYLNTYIGYPQP(SEQ ID NO:80);
[0472] The amino acid sequence of F13-N40:
[0473] MNISSYFFLNEENIKFNNQYLNTYIGYPQPISKDWPNLPA (SEQ ID NO:81); Amino acid sequence of F13-N50:
[0474] MNISSYFFLNEENIKFNNQYLNTYIGYPQPISKDWPNLPAGFQRYIDDVI(SEQ ID NO:82);
[0475] The amino acid sequence of F13-N60:
[0476] MNISSYFFLNEENIKFNNQYLNTYIGYPQPISKDWPNLPAGFQRYIDDVINL NGFLYFFK(SEQ IDNO:83);
[0477] The amino acid sequence of F13-N70:
[0478] MNISSYFFLNEENIKFNNQYLNTYIGYPQPISKDWPNLPAGFQRYIDDVINLNGFLYFFKGSLYLKFDIA(SEQ ID NO:84);
[0479] The amino acid sequence of F16-N20:
[0480] MKLSEKGFELIKHFEGLRLH (SEQ ID NO:85);
[0481] The amino acid sequence of F16-N30:
[0482] MKLSEKGFELIKHFEGLRLHAYQCSANVWT(SEQ ID NO:86);
[0483] The amino acid sequence of F16-N40:
[0484] MKLSEKGFELIKHFEGLRLHAYQCSANVWTIGYGHTAGVR(SEQ ID NO:87);
[0485] The amino acid sequence of F16-N50:
[0486] MKLSEKGFELIKHFEGLRLHAYQCSANVWTIGYGHTAGVRLGDVISAEK A(SEQ ID NO:88);
[0487] The amino acid sequence of F16-N60:
[0488] MKLSEKGFELIKHFEGLRLHAYQCSANVWTIGYGHTAGVRLGDVISAEKADAFLRRRDVAD(SEQ IDNO:89);
[0489] The amino acid sequence of F16-N70:
[0490] MKLSEKGFELIKHFEGLRLHAYQCSANVWTIGYGHTAGVRLGDVISAEKADAFLRRDVADAERTVNNAVS (SEQ ID NO:90);
[0491] The amino acid sequence of PAU_RS09615-N70:
[0492] LNKRRILFILLSVIAIYALWRWYFPDPYHLNLTEKEKKVTTEMLANMQTRCVGRYLIDIPEAFGNIIHDG (SEQ ID NO:91);
[0493] The amino acid sequence of PAU_RS09620-N70:
[0494] MEKDKQTRICRCRLNKRNILFVLLSILAIYTLWRWYFPDPYHPNLTEKEKKVTTEMLANMQTRCVGRYLI(SEQ ID NO:92);
[0495] The amino acid sequence of PAU_RS09625-N70:
[0496] MSWPIPEIPKLKAVAPLRYRRWFIVLILMLVVGTLVGAFLNGTTDLFHIAIYGSLPAFFVWLMLFGAVLN (SEQ ID NO:93);
[0497] The amino acid sequence of PAU_RS09630-N70:
[0498] MLKITRDVYLDCDSNKIIHGKNMFIITEKEKRLLLALWEHASQGNILNREQLIPLVWPERKNGVAETNLL(SEQ ID NO:94);
[0499] The amino acid sequence of yciH(PAU_RS10065)-N70:
[0500] MQDNNSRLVYSTDRGRIQEEVVKPQRPKGDGIVRIQRQTSGRKGKGVCVITGIDADQTLIKLAAELKKK(SEQ ID NO:95);
[0501] The amino acid sequence of PAU_RS10070-N70:
[0502] MENDRSTMKNMEGEQVPSIRENMLDAVEELIYTRGIAATGIDLITRTTGSSRKTIYRYFGTKDGLVEEVL(SEQ ID NO:96);
[0503] The amino acid sequence of PAU_RS10075-N70:
[0504] MSDIKEIRQPVPPFTRETAIQKIRLAEDGWNSRDPEKVSLAYSLDTHWRNRSEFVYSRKEAQEFLTRKWE(SEQ ID NO:97);
[0505] The amino acid sequence of PAU_RS10080-N70:
[0506] MTTRNFQYRSLDRLQVIYNGRGEQWVLGHLAESTVHGRYLFEYTHKALDSGIEFSPLLLPLSTDTYANFE(SEQ ID NO:98);
[0507] The amino acid sequence of PAU_RS10085-N70:
[0508] MALILYTESELLAELGHRLREHRLRRNMLQTELASRSGISVSALKKIEGSRGTLENLMKVVFALRLENE (SEQ ID NO:99);
[0509] The amino acid sequence of PAU_RS10100-N70:
[0510] MSNSEYDIIDKNDTYQVKNKEYTVVNGKYWQYEREGNENNNKTFISLMKGNQNDPIWITSDIKEMSLYII(SEQ ID NO:100);
[0511] The amino acid sequence of PAU_RS25580-N70:
[0512] LGIYLKDSYDFVNPDEFLGVWSRDGVLSKVKTVVYLGFYKDNMWRELATGEYSKYVPVYNVNFREWQRKY(SEQ ID NO:101);
[0513] The amino acid sequence of PAU_RS10205-N70:
[0514] MFQNNPLLAQLKQQLHSQTLRVEGLVKGTEKGFGFLEVDGQKSYFIPPPQMKKVMHGDRITAAIHTDKER(SEQ ID NO:102);
[0515] The amino acid sequence of nth(PAU_RS10210)-N70:
[0516] MNTQKRIEILTRLRDNNPKPTTELVFTTPFELLISVLLSAQATDVSVNKATAKLYPVANTPQAILNLGVD(SEQ ID NO:103);
[0517] The amino acid sequence of rsxC(PAU_RS10230)-N70:
[0518] MFNLFNLLNKNKIWDFDGGIHPPEMKLQSSQTPLRIAPLPEELIIPLQQHLGTEGELLVKTGDKVLKGQS(SEQ ID NO:104);
[0519] The amino acid sequence of rsxB(PAU_RS10235)-N70:
[0520] MISLWIAIGALSALGLIFGLILGFAARRFQVEEDPIVEKVDNILPQSQCGQCGYPGCRPYAEAVANNGEM(SEQ ID NO:105);
[0521] The amino acid sequence of PAU_RS10715-N70:
[0522] LNELEHVTDGSIEAFNHSPLFKTLGFTIVEWGVDFVDLRIQHGNPCHAGSFGKNTAVGINGAIVTAALES(SEQ ID NO:106);
[0523] The amino acid sequence of PAU_RS10720-N70:
[0524] MEVNKSAREGSDQDILSSECNVDDLRFWAALLSGRDNFSIPVDRLVRDFGENRVLTSSFDIDEKLSLRLR(SEQ ID NO:107);
[0525] The amino acid sequence of PAU_RS10725-N70:
[0526] MSSVPFKQVDVFTSKRFKGNPVAVVMDASELSTEQMQNIASWTNLSETTFVLPTVDAAADYRVRIFTPGS (SEQ ID NO: 108);
[0527] The amino acid sequence of PAU_RS10730-N70:
[0528] MREAIEEAKKSAKYPFGAVIVNRSSGEILSRGVNSWVKNPILHGEIQAINHYVTLYGNQGWNNVALYTTA (SEQ ID NO: 109);
[0529] The amino acid sequence of PAU_RS10855-N70:
[0530] MNTAQEIINRLSGRAVTLGWDVVIAYDRKKINTLLEQQYVEKVKNGENFPLINWENQRKTLQFKDLQLGV(SEQ ID NO:110);
[0531] The amino acid sequence of PAU_RS10870-N77:
[0532] MKSYNHFLAESLITYSHLTDPDACYEMITEACQELGDELGNLFYAQLVLLLVNHIGDSNILREAITVTRKGLMTAAN(SEQ ID NO:111);
[0533] The amino acid sequence of PAU_RS13665-N70:
[0534] MKRKNTPTPHDAVFKQFLSHIDIARDFLEIHLPATLRAVCDLDTLRLESGSFIEDNLRAHYSDILYSLKT(SEQ ID NO:112);
[0535] The amino acid sequence of PAU_RS13670-N79:
[0536] MRKTTYSVAKANLVSIMDQAIQDGVPILITRQNGGDCVIISSAEYVSLEETAYLLRSSANAKHLLKSLERVSSGNVQER(SEQ ID NO:113);
[0537] The amino acid sequence of PAU_RS23900-N62:
[0538] IQLLSLNGACHHVKNTPFQTFPSETFSFDELVSKNHLVRQVDAAIDFEFIYP RISCSIVLPQ(SEQID NO:114);
[0539] The amino acid sequence of cmoM(PAU_RS13690)-N70:
[0540] MRDRNFDDIADKFSRNIYGTTKGKIRQAVVWQDISELLAQLPQRPLRILDAGGGGEGNMACQLAELGHQVI(SEQ ID NO:115);
[0541] The amino acid sequence of elyC(PAU_RS13695)-N70:
[0542] MLFLFKKYLGAMLMPLSLILLISLVALLLLWFTRWQKTGKTILTFSWLALLLLSLQPIADKLLLPLESIY(SEQ ID NO:116);
[0543] The amino acid sequence of kdsB(PAU_RS13700)-N70:
[0544] MFTVIIPARFASTRLPGKPLADIHGKPMIVRVMERAMRSGAKRVIVATDNHDVVAAVIAAGGEACLTNEN (SEQ ID NO: 117);
[0545] The amino acid sequence of PAU_RS13560-N70:
[0546] MTIIDVKNLTKVYKYHEKEPGFVGSIKSLINRKYEYNNAVNNISLKINKGEFVGLIGLNGAGKTTTLKML(SEQ ID NO:118);
[0547] The amino acid sequence of PAU_RS13520-N70:
[0548] MCIQFLPFCRLILLPCIVAGLSGCVAISKVDHRADTEMARATDIIKELRKTPAAAVVHVHEHGYWVSTKE(SEQ ID NO:119);
[0549] The amino acid sequence of PAU_RS22390-N70:
[0550] MGRLFGSEGAKHFEGAGGSLAQKMVQSGATPEEINAALTHYLKGQVPDGQDPATGWSQYGKYRNEKGDWNW(SEQ ID NO:120);
[0551] The amino acid sequence of PAU_RS16505-N70:
[0552] MSMRRIALSNKKLRGIPRRLRSLKIWSESYKAYFPVITENDYSYGYWNVKIPVHSALVQGKQTNKNIQSI(SEQ ID NO:121);
[0553] The amino acid sequence of PAU_RS16515-N100:
[0554] MSRSRYTPEQKQQHVTQWRHSDLTRKQYCEQHQLNFSSFRDWIADSNK MRQPLSQTLPALLPVSLQPDDAHTVTLHTPDGYAIACPLTLLPDVMRVLARC(SEQ ID NO:122);
[0555] The amino acid sequence of PAU_RS25720-N59:
[0556] MLKPQQLFLVREPVDMRRGIDALTQHLAGLNLRWDRHGVWLCTRRLHR AQFDWLIRGIH(SEQ IDNO:123);
[0557] The amino acid sequence of PAU_RS16660-N102:
[0558] MSFKIPNDEDVIIWIIIGTFSTWGGVVRYIIDMKKDKIRWDCKEATSQMIV SSFTGFLGGFLSFESGVSLYMTFVISGLFGTMGSTGISYLWGKFFGGEDKK (SEQ ID NO: 124);
[0559] The amino acid sequence of PAU_RS16665-N70:
[0560] MSINNSFYFSKRSEQNLIGVHPDLVKVTRLALQLSNTDFCVIEGLRTAERQRQMFADGHSQTLNSRHLTG (SEQ ID NO: 125);
[0561] The amino acid sequence of PAU_RS16690-N70:
[0562] MTDDLDIYQKIGQLLVDAGPSDAQKMIVRAKLFPENDGCKYEFDYIDENGDLGWFSPDSKASGDLTELLV(SEQ ID NO:126);
[0563] The amino acid sequence of PAU_RS16695-N70:
[0564] LDSSVSRQSFSANGNKTYHIDGQGRSSNIEASLSPSRNDRNTYQQCKAGKCGNTGDEGGHLIASIFNGPG (SEQ ID NO: 127);
[0565] The amino acid sequence of PAU_RS09595-N70:
[0566] MTLKQFTFATLLFSLLSTPVLAEKPQNFLYTSSDDLNQLRSLLERQDIDGVQIIYNWKQLESAPGKYDFS(SEQ ID NO:128);
[0567] The amino acid sequence of PAU_RS09735-N70:
[0568] MELYIPVVGLGIAVFALIFLVLRTRVHALLAMLIAAAIAGISGGLTAANTIDVITKGFGSTLGSIGIVIG (SEQ ID NO: 129);
[0569] The amino acid sequence of PAU_RS09740-N70:
[0570] MKVIITGAAGFLGQQLASALLTNNQELNIEQLILTDIHPPISPVNDPRVQCLALDLTQPDAAEKLIDEDS(SEQ ID NO:130);
[0571] The amino acid sequence of PAU_RS09750-N70:
[0572] MAEFAANLSTMFNDVPFKERFARAAKAGFKGVEYLFPYEETAEDLAALLQQHQLTQVLFNMPAGNWAANE(SEQ ID NO:131);
[0573] The amino acid sequence of PAU_RS09755-N70:
[0574] MSEQALRTELVEWARSMFYRGYSSGGAGNISAKLDDGTIIITPTNSSFGDLQADRLSKLDIEGNWLSGDK (SEQ ID NO: 132);
[0575] The amino acid sequence of PAU_RS10105-N70:
[0576] MSTNTTYRVAAVQAAPVFLDLEEATVAKTITLIESAANNGAKLIAFSETWIPGYPWFIWLDSPLWGMQFLK(SEQ ID NO:133);
[0577] The amino acid sequence of rsxG(PAU_RS10220)-N70:
[0578] MLETMRRHGITLAIFAAFTTGLTAIVNSLTQNTIAEQAALQHKSLLDQVIPPELYDNDIQNECYLVSADA (SEQ ID NO: 134);
[0579] The amino acid sequence of rsxD(PAU_RS10225)-N70:
[0580] MKFRPVHTDNKKLKVASSPFTHNQQSTSRIMLWVALAAIPGIAVQTYFFSYGTLFQLLLAMITALLAESL(SEQ ID NO:135);
[0581] The amino acid sequence of PAU_RS10245-N70:
[0582] MYQTHDYHRINGWILAPAAYLIMTFLSASVLGLYIMAFFSQNNMLGSTTNHFTLMWFLSVAITAAMWCF(SEQ ID NO:136);
[0583] The amino acid sequence of PAU_RS10865-N70:
[0584] MILDTSYRLESRYRRGIDEPQGKIFISGQQALVRMLLAQSALDRSVGLNTAGFVSGYRGSPLGGVDKELW(SEQ ID NO:137);
[0585] The amino acid sequence of mukB(PAU_RS13675)-N70:
[0586] MIERGKFRSLTLVNWNGFFARTFDLDELVTTLSGGNGAGKSTTMAAFVTALIPDLTLLHFRNTTEAGATS (SEQ ID NO: 138);
[0587] The amino acid sequence of PAU_RS13540-N70:
[0588] MMGWLQSSIGCLVLTSGLLGLCVGSFLNVVICRLPIMIISESNNETTGFNLCFPSSHCPRCGRVLAVRDN (SEQ ID NO: 139);
[0589] The amino acid sequence of pilV(PAU_RS13535)-N70:
[0590] MHDESISTLTIRRHDAGFTLLEVTVALIVLASMMVVGVVYLNRQSDMLVNQVVAGQIQHLGDAVVDYVND (SEQ ID NO: 140);
[0591] The amino acid sequence of PAU_RS13530-N70:
[0592] MWPLAIVGIAVIFFMQLIFGELLYEVRVQSVRARTDIAVLFRCYASAGVHLLNDHVQGAVRRRDMQQALT(SEQ ID NO:141);
[0593] The amino acid sequence of PAU_RS13525-N70:
[0594] MNKIIAYVLAAFFCMSNTAVYAAFTPISVPLAAVWIATPGQTLRAVTQEWANKSGYQVIWDASYDFPIRA (SEQ ID NO: 142);
[0595] The amino acid sequence of PAU_RS16495-N70:
[0596] MRLERDLAENILAKVVNMKLPLDEIFHEINRLLSDHGVMDDVYALNQPDIDDKYCLHLEEGFWVSYYSER (SEQ ID NO: 143);
[0597] The amino acid sequence of PAU_RS09730-N70:
[0598] MGNKNTPSRVKIFISALIFMSASVIGVLASTIDYRGFLTRSDIITSSTISAWFIWCSPLCMYISILLFKS (SEQ ID NO: 144);
[0599] The amino acid sequence of PAU_RS10765-N70:
[0600] MKGIEGVIMLSHDILPEKLLVSEKKHENVGSYFSDDIGEQSEQTEVSHFNLSLDDAFDIYADISIENQQE(SEQ ID NO:145);
[0601] The amino acid sequence of PAU_RS24760-N70:
[0602] MRPVPVCLFIQRFVGQKDENARCQTTGWPIDTDHVFNNIPAAVVVGLAITSALSDDISAHNLWTNGMETG (SEQ ID NO: 146);
[0603] The amino acid sequence of PAU_RS10735-N70:
[0604] MYTSLLLTLADDTFCLRLQRAARMWRKVSDEELSKLNLSEATTTPLWLINKLGEGLRQRTLADHMGLEGQS(SEQ ID NO:147);
[0605] The amino acid sequence of PAU_RS10115-N70:
[0606] MKIFITDEKAELEHLHHTCHDKRECDRIKAVLLASEGWSVMIAQALRLHETTVNRHISDYLNHRKLKPE(SEQ ID NO:148);
[0607] The amino acid sequence of PAU_RS09745-N70:
[0608] MLPSERRDFIYRYVHEYRTISISALVELMNVSHMTVRRDIRTLEEEGKVISISGGVQLSDALRQELPWNE (SEQ ID NO: 149);
[0609] The amino acid sequence of YP1(BF17_RS02745)-N70:
[0610] MTTLKLDTLSARINAHKNILIHIVKPPVCTERAQHYTQIYQENMDKPMPVRRALALAYHLANRTIWIKHD(SEQ ID NO:150);
[0611] The amino acid sequence of YP2(BF17_RS02855)-N70:
[0612] MLKTCEKPKKNKGSDTTESNSSYEYNTPKKCLKTDDGPSSSTTEDSQKTKVLVVGYSPTGGGHTGRTLDI (SEQ ID NO: 151);
[0613] The amino acid sequence of YP3 (between BF17_RS02845 and BF17_RS24790)-N66:
[0614] MGRILSYPDSITHNKLFILNTQTNVIRTMSKKNRLIFDSFIYANHSSHIGHR YEHHPLLGSTGGQA (SEQ ID NO: 152);
[0615] The amino acid sequence of YP3(BF17_RS24790)-N50:
[0616] MPNRAILFTGLLLLLALIGSVLSTRHYHRLATDWKNSARQSQQALTTANT(SEQ ID NO:153);
[0617] The amino acid sequence of YPe1(AEP37_RS09605)-N70:
[0618] MLYSSESKEKKTHSKETERDNAGHLFQQVSQGSVGVSPPDEGGKLSSGGGYGRLFAFIREAHLEEAQEFR (SEQ ID NO: 154);
[0619] The amino acid sequence of YPe2(AEP37_RS09610)-N70:
[0620] MLVNYENELTKLSLIDMESLINYLKAIGDEDTIAKMKSYKPKGNYEDVIFAANLAQELISGMKNFSEKFS(SEQ ID NO:155);
[0621] The amino acid sequence of Yr2(LGL87_RS15575)-N70:
[0622] MPYFNKSKKNEIRPEKSKEEVGGVLFDDSAIHENIDHNMEPQTGDSVATFPDNSDEVVGGDLAALRARLQ (SEQ ID NO: 156);
[0623] The amino acid sequence of Yr4(LGL87_RS15545)-N70:
[0624] MKNKIFKLSPAGKLAASLAIILASQGSAYTADIVGAGDSAHQPGISNAANGAAVVNIVTPSASGLSHNQY(SEQ ID NO:157);
[0625] The amino acid sequence of Yr5(LGL87_RS15540)-N70:
[0626] MIKKYAAFLLLLSAYSQGETLPNMGSFSPMSESRRALQDSSRTVQALMEERRYQQLKKQQLLNTQVSAQH(SEQ ID NO:158);
[0627] The amino acid sequence of Xb3(AACW61_RS01295)-N70:
[0628] MTIAVEKTAGSYMPYATYLNNKDVNSKNASSYHDNSEVLGDTDDWLLVDKQPVPPPPMLKSSIPIPPPPL(SEQ ID NO:159);
[0629] The amino acid sequence of F4 (PAU_RS09730, derived from the non-symbiotic luminescent bacterium Photorhabdus asymbioticaATCC43949)-N70 is as follows:
[0630] VGNKNTPSRVKIFISALIFMSASVIGVLASTIDYRGFLTRSDIITSSTISAWFIWCSPLCMYISILLFKS(SEQ ID NO:160);
[0631] The amino acid sequence of Pnf(PAU_RS16555, derived from the non-symbiotic luminescent bacterium Photorhabdus asymbioticaATCC43949)-N70 is as follows:
[0632] MLKYANPQTVATQRTKNTAKKPPSSTSFDGHLELSNGENQPYEGHKIRKIKGLRQHLADRSLNKGHISPL(SEQ ID NO:161);
[0633] The amino acid sequence of Xb1(AACW61_RS01395, derived from Xenorhabdus bovienii)-N70 is as follows:
[0634] MDNILASPFTRLFWNEWILEPDSVKYNMVMDQTIQGDLDERKLRSAIQGLMNQYPLFQYQLDEENGELYW(SEQ ID NO:162);
[0635] The amino acid sequence of Xb2(XAACW61_RS01300, derived from Xenorhabdus bovieni)-N70 is as follows:
[0636] MPIPSNKTIDEKIDMFRSNHSTDEDIDSATSEIGWDTNAMFSNGSVRLDQIKDPGFWKENIADYYKWMVV(SEQ ID NO:163);
[0637] The amino acid sequence of Yr3 (LGL87_15575, derived from Yersinia ruckeri)-N70 is as follows:
[0638] MSTSFTRHPASRRVKIIAWGNILFQLLFPLSLSFTPVMAAAPTSLTSNASIPTTEPYVLGVGENVDTVAR(SEQ ID NO:164);
[0639] The amino acid sequence of PAU_RS24740-N70:
[0640] VDNNRLSVDQNHSRSEELVNYHGNTSCRSAVRDKYRPEYDKVQTRITSCTGAEQCVAVAKELRELQGDYS(SEQ ID NO:165);
[0641] The amino acid sequence of XK1(AB204_RS00510):
[0642] MYSLKQKEEK KNHSPRPTQN TVNGDHAITP RIAELRKNFQ QLNSQ TVRPRVAPRPLHSLSKTGPFQQATK APIASEIITP VSAATKTFTS NTQNPKVTPTSTS APTNGVKKSFTLYRADNRSFEELQKSS PEGFKAWVHLDSEKARKFASVFLG ND NVES LPKHIIDEINKWKKGGTPKLSDLSTFIKYT KDRSTVWVSTAVNTEA GGQSSGAPLYEISM ELYEFGVDKGKLVPLPNGRTGNMKPSIHL DTQNLADAT IIALNHGPVNDA EMSFLTTIPM SNLKPYRR(SEQ ID NO:166);
[0643] Amino acid sequence of XK2 (AB204_RS00495):
[0644] MIYGYDNKMVYRKRQESNTALEALHLDIPDTAEVSAAYQRDLSSVMTRFQNKTDSILEESKQPSPLKQKPSEKKKYAKKALKNFAAHAGYSHINSYQDEFVNFKDNNYNLAPGKLFWGVELIPQKTIEINRPIGPWRNKLSYDIEDIFRNNMANEYRVTNTFEFIKGLQDMYKKNGQTLHPMTQQLVEEHIKHNGNILPTMAGIAGLHAEVQALNQLFIKADEDAGQSPEPFSMRYIKAMLQSSIFTKRLTTANAGQDFPACHNCSGIIQSPANVITGTVSSAGSNFSKQVAKRDRSQSLSQ(SEQ ID NO:167); Amino acid sequence of XK3 (AB204_RS00490):
[0645] MPVYENKEKRKYSKNPLHNTSTINQFTVQDLDSASQQETATPLLTNEFKQGHYVNHSKLPAGTKLSEIAMLGSHDAGTYAYSRRKSGFASSLGSLLPAAFKTQNRTLRQQAEAGARYFDIRVAQNKDGSFSFFHGPSVAGSDAVSDVKSLLEYAAGDTNHFYLLKLVFKGEKGQSSAAASSNTFLQSILEGHQQRLITLDDTSSLGEAMVDLLDKGKNIGIMVDRKKYDGTEPHWGYKESVNTKWANRANAEETANFLLKFHENPVPDKLNIMQTNIPVASVGRGQFTSGVKRYLFGNRDPLIKAVAQLPAGIISADYVGDSGSATSKFMETINQYNRSLMERDEETRL(SEQ ID NO:168);
[0646] Amino acid sequence of F1 (PAU_RS09715):
[0647] MMREYSKEDDCVKEKTNLAESENVEADNYLEMDCLNYLAKLNGMPERKDHSLNSTKLIDDIIKLHNDRKGNKLLWNDNWQDKIIDRDLESIFKKIDEMVSEFGGIEIYKDIVGENPYDPTEPVCGYSAQNIFKLMTEGEHAVDPVKMAQTGKINGNEFAEKLEQLNSSNNYVALINDHRLGHMFLVDIPSTNREKVGYIYQSDLGDGALPALKIADWLKSRGKESINVNKLKKFLSNEFTMLSESEQKELIAEIFDINKDIANVKLGKIKKDKAVDVYLREYDLNDFISNIEKLKTKLV(SEQ ID NO:169);
[0648] Amino acid sequence of F2 (PAU_RS09720):
[0649] MPIIGHKEDLIRTERSSVDLTRSSNNRQTDNLELNIPQHKRDNKDIEHAVIYGFSQHRGPEMQKAFADNKNPVTIDEYNAGLGIMGELSLSDYFRISQDLKENRLPELNEKNIQNHSLKYFDAMGVNMKSADPNVKEEAKEQQRAYTRSWGFYMMENKEKLDIQSKINNLIPKKKSFFSKSPGEDEYKKLDEFILKNSNGSNLTIPKQRKILMKFASAKNAVDVTKNLSGEEQTWLKDIIATAFFRQTSKLGMSWFIEQLASPDFRFVIVGFNGEELTTDQIRSNKPWKHGNRRKEGASEYAEPITFSEIRHAHRKGYDSKINFIKK(SEQ ID NO:170);
[0650] Amino acid sequence of F3 (PAU_RS09725):
[0651] VKYDPRLRTWVEDDFDYEKNFKKQMDYINYKDLEKQLKENVDYYALLDENEAIIFLKELGCDIKSLLNGTVFPVTDVLSNFSGNIKDTPGVFKVAKNFKPINIGIFTHIINELKGNGIKAIEYLGKNGERYIKLTDCPGIRKYLNATRYLINNKKIMEVGIGSVAMEGSIVKGARFGVIYSAAYRSVELMFKSEYDLTNFFVNLSMDMAKMIVATIIAKTTVAAATSFVVTAALSTTAIAIGVFIIGALVVWGLMWLDDEFKISETIIKRLKEHKAKTPISTYHPDQIFNAWGRDYRG(SEQ ID NO:171);
[0652] Amino acid sequence of F5 (PAU_RS10135):
[0653] It should be noted that the "
[0650] " in the translation of line may need to be further confirmed according to the actual situation, as it seems to be an incorrect encoding in the original text. It may be a misrepresentation or an encoding issue that needs to be corrected in the source if possible.MVFEHDKTVERKRKPSIQLGNDKEKSSEQALELPQSKQNNPLLHDLITSNNLRKEAAVFAKQIGPSYQGILDGLEHLHNLSGNEQLTAGFELHRRITRYLEEHPDSKRNAALRRTQTQLGDLMFTGTLQEVRHPLLEMAETRPAMASQIYQIARDEAKGNTPGLTDLMVRWVKEDPYLAAKSGYQGKIPNDLPFEPKFHVELGDQFGEFKTWLDTAQNQGLLTHTRLDEQNKQVHLGYSYNELLDMTGGVESVKMAVYFLKEAAKQAEPGSAKSQEAILLNRFANPAYLTQLEQGRLAQMEAIYHSSHNTDVAAWDQQFSPDALTQFNHQLDNSVDLNSQLSFLLKDRQGLLIGESHGSDLNGLRFVEEQMDALKAHGVTVIGLEHLRSDLAQPLIDKFLTSENEPMPAELAAMLKTKHLSVNLFEQARSKQMKIIALDNNSTTRPAEGEHSLMYRAGAANNVAVERLQQLPAEEKFVAIYGNAHLQSHEGIDHFLPGITHRLGLPALKVDENNRFTAQADNINQRKCYDDVVEVSRIQLTS(SEQ ID NO:172);
[0654] Amino acid sequence of F7(PAU_RS10125):
[0655] MEREYNKKEKQKKSAIKLDDAVGNNEENMDMTSPLELNSQYTNRKRPGLRERFSATLQRNLPGHSMLDRELTTDGQKNQESRFSPGMIMDRIMHLGVRTRLGKVRNSASKYGGQVTFKFAQTKGTFLDQIMKHKDTSGGVCESISAHWISAHAKGESIFDQLYVGGQKGKFHIDTLFSIKQLQMDGYLDDEQSTMTEYWLGTQGIQPNRQKNDNMNEHSSKIVGETGTRGTKDLLRAILDTGDKGSGYKKISFLGKMAGHTVAAYVDDQKGVTFFDPNFGEFNFPDKVSFSHWFTDDFWPKSWYSLEIGLGQEFEVFNYEPKEP(SEQ ID NO:173);
[0656] Amino acid sequence of F8(PAU_RS10120):
[0657] MISTFDPAICAGTPTVTVLDNRNLTVREIVFHRAKAGGDTDTLITRHQYDLRGNLTQSLDPRLYDLMQKDNTVQPNFYWQHDLLGRVLHTVSIDAGGTVTLSDIEDRPALNVNAMGVVKTWQYEANSLPGRLLSVSEQSANEAVPRVIEHFIWAGNSQAEKDLNLAGQYMRHYDTAGLDQLNSLSLTGAHLSQSLQLLKDDQMPDWAGDNESVWQNKLKNEVHTTQSTTDATGAPLTQTDAKENMQRLAYNVTGQLKSSWLTLNGQLEQIIVKSLAYSESGQKIREEHGNGVVTKYSYEPDTQRLINITTQRSKGHVFSEKLLQDLLYEYDPVGNIVSILNRAEATHFWRNQKVSPRNTYTYDSLYQLIQSTGREMADIGQQNNKMPTPLVPLSSDDKVYTTYTRTYSYDRGNNLTKIQHRAPASHNIYTTEITVSNRSNRAVLSHNGLTPREVDAQFDASGHQISLPTGQNLSWNQRGELQQATTINRDNSATDREWYRYNAGSARILKVSEQQTGNSTQQQQVTYLPGLELRTTKSGTNTTEDLQVITMVETERTQVRILHWSAGKPNDIANNQVRYSYDNLIESNVMELDTKGKIISQEEYYPYGGTAIWTARNQIEASYKTVRYSGKERDKTGLYYYRHRYYQPWLGRWLSADPAGTVDGLNLYRMVKNNPIRYQDESGTNANDKAQAIFKEGKKIAINQLKIASNFLKDSKNSENALEIYRIFFGGHQDIEQLPQWKKRIDSVIYGLDKLKTTKHVHYQQDKSGSSSTVADLNVDEYKKWSEGNKSIYVNVYADALKRVYEDPLLGREHVAHIAIHELSHGVLRTQDHKYIGVLSSPGSHDLTDLLSILMPPANEQDRTEKQRRATGARKALENADSFTLSARYLYYTAQDPNFLSSLRKAHRDFNNKKTDRLIIRPPERR(SEQ ID NO:174);
[0658] Amino acid sequence of F10(PAU_RS13645):
[0659] MPNKKYSENTHQGKKPLIKSEANNEHAIDNSPLGIGLDLNSILGNNSASLSQIHDYSFWKENISEYYKWMVVVKAHLKQLDWTLKSMDSPESAGANIAKNIGTTTLQTLLNTGGSIAGGAIGGAIGSAIAPGVGTIAGMGIGALAGTGLNYLNDTAIEKLNEKLEIAYPYPKTRNMIFDINNYDKNPLIKAIKKKTKKDNLKVMAGSSLTSQLLGRITPIKIPAYKLADLAVSHHRALAGLSSDKARHILDFTNSIREVLNESHSDAVAFMRKNYGDNAMGLSGLSSKIKGDKLTLDTLARTRNKIENRINSINKQTLKLSSKNSNE(SEQ ID NO:175);
[0660] Amino acid sequence of F11(PAU_RS22355):
[0661] MEREYSEKEKHKKHPIQLRDAIEQHAEETANNSLGLGLDLHQAINTPKVPKDNYNEENGDLFYGLAAQRGRYIKSVNPNFDPDKTNSSPMVIDVYNNHVSNTILNKYPLDKLGKLYGNPQKYAKDIKVTNSLQQDVAASKRGWYPLWNDYFKAGNENKKFNIADIYKETRNQYGSDYYHTWHEPTGAAPKLLWKRGSKLGIAMAASNEKTKIHFVLDGLNIQEVVNKQKGSTPLEQGRGESITASELRYAYRNRERLAGKIHFYENDQETIAPWEKSPELWQNYIPKNKSQNESSTPQRNNGALYRLGGPFRKLRASLRKRS(SEQ ID NO:176);MVHEYSINDRQKRHSFSSANPIDPEVTNRENSRHRFPKDNYNKGHGDLFYGLAPERGKYIKEANPKFDPNNPENAAMIIDVYNDEISRVILNNNANKISTNRLLNFIYNFRKNRLENLMKNPEKYAKDIKVKDNLRENISPKKIEKYPLWNDYFEAGIRNKKFNIAEIFKETASQYNSDYYHAWHIGGNSAPRLLWKRGSKLGIEIAASNQRTKIHFILDGLKIEDVVNKTKGPAPLKAGPGESITASELRYAYRNRARLAGRIHFYENGKETIAPWDKDPELWQKYTPKNRSGMEL (SEQ ID NO:177);
[0664] Amino acid sequence of F13 (PAU_RS13660):
[0665] MNISSYFFLNEENIKFNNQYLNTYIGYPQPISKDWPNLPAGFQRYIDDVINLNGFLYFFKGSLYLKFDIAKAQVVDGPNFIIDGWPGLKGTELENGIDAAIELTTNTVCFFKGEHCVDYTIDLHTIKTSTISDRWGMTGKYAAFSRNLDAIILWPDIDGNFIYFFKGDSFIRFDPNLNALDAGPIIISNDNNGWRGLVFKNVQAAVSVDTDLLGSHRDNNGGNSKVCNGTCGTNDTGKYCFQLPQSIRFGLIAYANTDKPQTVKVYIDDLLVDTLTSTSKGQNNLMATKAYTSGTGKICIEIAGDGKPCKLRYFDNIFDGNPGTAIISAENGTNNHYNDSVVFLNWPLT (SEQ ID NO:178);
[0666] Amino acid sequence of F16 (PAU_RS16545):
[0667] MKLSEKGFELIKHFEGLRLHAYQCSANVWTIGYGHTAGVRLGDVISAEK ADAFLRRDVADAERTVNNAVSVSINQHQFDALVSFVFNLGAGNFRSSVLLKKL NAGDYAGAAGELLRWVNAGGQKLAGLVRRREAEKMLFETPV (SEQ ID NO:179);
[0668] Amino acid sequence of YP2 (BF17_RS02855, derived from Yersinia similis 228):
[0669] MLKTCEKPKKNKGSDTTESNSSYEYNTPKKCLKTDDGPSSSTTEDSQKTKVLVVGYSPTGGGHTGRTLDIVKYALAQGTLDEYKLVVFYVPPKWEDKDRPEGLNNLAKNILDREIRVKIIESEKPVYGYLNEDGSSNDAKIITRIALQPLYRKVNIMLSKDNITEEEKNKIKEENNFLVKYYGELTPTIEKVNDYKEGNNINQLQRMTANTLIKSLKDEYKITVLSDMDPALQKAAKNHGISDNQRLDQQNHAILLDLENTNHNLNMKNAVLAKVLGGRGEKISHISLGGKNTLAGALSTLNQFGDENSKMEDVRNAIYEKIFNAAINANDINHQKKSPYVGVVKSENLDKSSDIKNVIYIYAHNKTNTILNYIINKINSEDYKNKVFIFCGPGAIPDYNAMHLAYLIDADGITTSGAGTSGEFVYLHKNAGAKSNLLSLPIEGHNEQEKITDILYKDIDTNKHMVPKHRKIKYLEKNINKSGKNINELEEDIDELVKKSGHDKNNGTESCQKLRRALENEETYVKQAHDIIFTGEELDSYASKYKEIEEKMYKDPTLKATRKYLKLVFQSLSYLTKKTKEEPMNIYFSGKEEPIKFKDIDELRKQFNEPESLQKILGLKSKDDVKSLPLFKEVRTLINSKDYKDESKFFKLAKLFGKDMTTGF (SEQ ID NO:180);
[0670] Amino acid sequence of YPe1 (AEP37_RS09605, derived from Yersinia pekkanenii A125KOH2):
[0671]
[0672] The amino acid sequence of YPe2 (AEP37_RS09610, derived from Yersinia pekkanenii A125KOH2):
[0673]
[0674] The amino acid sequence of Yr2 (LGL87_RS15575, derived from Yersinia ruckeri):
[0675]
[0676] Amino acid sequence of Xb3 (AACW61_RS01295, derived from Xenorhabdus bovienii):
[0677] MTIAVEKTAGSYMPYATYLNNKDVNSKNASSYHDNSEVLGDTDDWLLVDKQPVPPPPMLKSSIPIPPPPLGAGMPPPPMLSASTSKSATGATKKWRIKEDPALYKQQGGEYPYYSSFTITLQNMGIKHYSNILESDIGNFVSEWKKVGTEMKPPARNISEDKFNDVKQTLTRMDSEWDTYTKANVKETFRGDTEAVIKSYPWLSEFIQKTNGMNQTYSEPVNQEIASPTVMSTAKDPMMSYVNQKTIMWHFSLEKGHAGVSEGLYAAEGEVTFPLYNRIHIDSLHFIPKGSAFKISDQS (SEQ ID NO:184);
[0678] Amino acid sequence of pdp1-N50:
[0679] MPRYANYQINPKQNIKNSHGKSSSSDFSSGYLSFSNNSLDDPFIRQQVKR (SEQ ID NO:185).
Claims
1. A lead sequence for guiding the packaging of a payload into a protein complex, wherein the lead sequence comprises at least 20 amino acids at the N-terminus of an effector derived from archaea, Gram-negative bacteria, or Gram-positive bacteria; The payload is a polypeptide, nucleic acid, or a combination thereof.
2. The lead sequence according to claim 1, wherein the protein complex is derived from archaea, Gram-negative bacteria, or Gram-positive bacteria; Preferably, the protein complex is derived from non-symbiotic luminescent bacillus (Photorhabdus asymbiotica), Serratia entomophila, Pseudoalteromonas luteoviolacea, or Xenorhabdus khoisanae. More preferably, the protein complex comprises a non-symbiotic luminescent bacillus virulence box (PVC); More preferably, the protein complex is a PVC-V complex.
3. The lead sequence according to claim 1 or 2, wherein the lead sequence comprises at least 20 amino acids at the N-terminus of an effector derived from *Xenorhabdus bovienii*, *Non-symbiotic luminescent bacteria*, *Yersinia pekkanenii*, *Yersinia similis*, *Yersinia ruckeri*, or *Xenorhabdus bovienii*. Preferably, the leader sequence comprises amino acids at positions 1-20, 1-25, 1-30, 1-35, 1-40, 1-45, 1-50, 1-55, 1-60, 1-65, 1-70, 1-75, 1-80, 1-85, 1-90, 1-95, 1-100, 1-105, 1-110, 1-115, 1-120, 1-125, 1-130, 1-135, 1-140, 1-145, or 1-150 at the N-terminus.
4. The leader sequence according to any one of claims 1-3, wherein the leader sequence comprises an amino acid sequence as shown in any one of SEQ ID NO:1-3, SEQ ID NO:5-165, or comprises an amino acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with an amino acid sequence as shown in any one of SEQ ID NO:1-3, SEQ ID NO:5-165.
5. The leader sequence according to any one of claims 1-4, wherein the payload is a polypeptide, the polypeptide comprising one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone regulatory molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene editing proteins.
6. An isolated nucleic acid encoding a leader sequence as described in any one of claims 1-5.
7. A fusion protein comprising a leader sequence as described in any one of claims 1-5 and a polypeptide linked thereto, wherein the linkage is a covalent or non-covalent linkage.
8. The fusion protein according to claim 7, wherein the polypeptide comprises any one or more of the following: signaling pathway regulatory proteins, structural proteins, transport proteins, hormones or hormone-regulating molecules, cytotoxins, antigens or immunogens, antibody proteins or fragments thereof, tag proteins or reporter proteins, antimicrobial peptides, enzymes involved in cell metabolism, and gene-editing proteins.
9. An isolated nucleic acid encoding a fusion protein as described in claim 7 or 8.
10. An expression vector comprising the isolated nucleic acid as described in claim 6 or 9.
Citation Information
Patent Citations
Polypeptide injection system and application thereof
CN116554283A
Preparation, purification and application of protein compound
CN119735656A