Virus-like protein particles for delivering nucleic acids and proteins, and preparation method therefor
By encapsulating nucleic acids or proteins by self-assembled viroid-like protein particles, the off-target risk and toxicity problems of delivery technology in gene therapy are solved, and efficient and safe delivery of biological macromolecules is achieved.
Patent Information
- Application Number
- PCT/CN2025/072920
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-24
AI Technical Summary
The existing gene therapy delivery technology has problems such as high off-target risk, liver toxicity, low capacity and neutralizing antibody production, making it difficult to achieve efficient, specific and safe delivery of biological macromolecules.
Virus-like protein particles containing envelope protein ENV, matrix protein MA, capsid protein CA, nucleocapsid protein NC and proteolytic enzyme PR are used to form delivery particles that encapsulate the target nucleic acid or protein by autonomous assembly, avoid integrase IN genes, and use protein elements that can form dimers to improve stability and targeting.
It realizes efficient, specific and safe delivery of biological macromolecules, reduces off-target events, improves delivery efficiency and reduces toxicity risks.
Smart Images

Figure PCTCN2025072920-FTAPPB-I100001 
Figure PCTCN2025072920-FTAPPB-I100002 
Figure PCTCN2025072920-FTAPPB-I100003
Abstract
Description
Virus-like protein particles for delivering nucleic acids and proteins and preparation methods thereof Technical Field
[0001] The present invention relates to the field of drug delivery, in particular to the technical field of virus-like particles (VLPs). Specifically, the present invention relates to virus-like protein particles for delivering nucleic acids and proteins and methods for preparing the same, including such protein particles, as well as nucleic acids, vectors, cells, and pharmaceutical compositions encoding the same. The present invention also relates to methods for preparing the virus-like protein particles. Background Art
[0002] Biomacromolecules, such as proteins and nucleic acids, often serve as important therapeutic targets in the prevention and treatment of various major diseases. However, traditional chemical drug molecules have difficulty specifically targeting these targets, especially for some difficult-to-drug target genes. The CRISPR-Cas system and its derivatives, a series of next-generation gene-editing technologies, have driven rapid developments in gene therapy. Among them, epigenetic editors based on CRISPRoff and CRISPRon are designed to silence or activate target gene expression by epigenetically modifying their target sites without cutting or altering the original DNA sequence. This technology is believed to be able to treat most genetic diseases and some difficult-to-treat diseases caused by the combined effects of multiple genes. In addition, CRISPR-Cas-based single-base editing tools (such as ABE and CBE) can achieve precise conversion of bases from A>G and C>T, respectively, theoretically treating nearly half of genetic diseases caused by DNA point mutations, thus holding broad application prospects in gene therapy.
[0003] In addition to gene editing tools, another key technical challenge in the field of gene therapy is effective delivery and transportation technology. Currently, the delivery vectors used in clinical practice can be mainly divided into viral vectors (lentivirus, AAV, etc.) and non-viral vectors (LNP, etc.). In recent years, virus-like particles (VLPs) technology, which can achieve effective delivery of large molecules, has made many breakthroughs and has been used in the delivery and gene therapy research of CRISPR-Cas related gene editing systems. Although the CRISPR system can achieve precise editing of target genes, its high off-target risk still restricts the application of this technology in clinical human research. In particular, the CRISPR-Cas system delivered by AAV can survive for a long time in the body, which undoubtedly greatly increases the risk of irreversible safety hazards such as gene mutations. However, the ribonucleoprotein (RNP) complex of Cas protein and gRNA directly delivered by VLPs has a short half-life in the cell, which not only effectively achieves target editing, but also greatly reduces the probability of off-target events.
[0004] Currently, the most mature and widely used delivery technologies are AAV technology and LNP technology, each of which has its own advantages, such as the maturity of AAV technology and the ability to achieve long-term expression; LNP has low cost and lower integration risk. However, these in vivo delivery technologies all have some shortcomings. Lentiviral vectors still have a certain potential risk of genomic integration, and there are currently not many cases of their use for in vivo delivery of gene therapy tools; AAV vectors have a low capacity (can only accommodate 4.7kb in length) and have certain acute toxicity to the liver and spleen. Neutralizing antibodies will gradually be produced in the body, and subsequent injections are less effective; LNPs can currently only target a few organs such as the liver, and have obvious dose-dependent liver toxicity. In general, the research on these new technologies for the effective in vivo delivery of therapeutic proteins / nucleic acids has been slow, and how to achieve efficient, specific and safe delivery of biological macromolecules (proteins / nucleic acids) remains one of the three key technical challenges in the field. Summary of the Invention
[0005] One aspect of the present invention provides a virus-like protein particle comprising: an envelope protein ENV, a matrix protein MA, a capsid protein CA, a nucleocapsid protein NC, a proteolytic enzyme PR, and a protein element capable of forming a dimer;
[0006] The matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, the proteolytic enzyme PR, the dimer-forming protein element and the target nucleic acid or target protein autonomously assemble to form protein particles containing the target nucleic acid or target protein, and the envelope protein ENV packages the protein particles.
[0007] Specifically, the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, the proteolytic enzyme PR and the dimer-forming protein element are sequentially connected to form a fusion protein (which can autonomously assemble to form protein particles), the target nucleic acid or target protein is covalently linked to one end of the fusion protein to be packaged inside the protein particle, and the envelope protein ENV is used to package the protein delivery particles containing the target nucleic acid or target protein.
[0008] Specifically, the virus-like protein particle does not contain the integrase IN gene, such as does not contain the MLV integrase IN (SEQ ID NO. 6).
[0009] In a preferred embodiment, the dimer-forming protein element comprises a homologous dimerization domain or / and a heterologous dimerization domain; it can be any dimerization motif that can form a dimer, such as a dimerization motif on a homologous or heterologous reverse transcriptase, a dimerization region on a dimeric protein; preferably, the dimer-forming protein element is a dimerization motif on a homologous or heterologous reverse transcriptase or a dimerization region on a dimeric protein; preferably, the dimer-forming protein element comprises a dimerization motif in reverse transcriptase RT or an enzyme-inactive mutant of reverse transcriptase RT; preferably, the dimer-forming protein element comprises a hydrophobic amino acid dimerization motif; preferably, the hydrophobic amino acid dimerization motif comprises a leucine zipper (LZ) dimerization motif, a leucine repeat motif (LRM), a leucine-rich repeat motif (LRR) or / and a tryptophan repeat motif (TRM).
[0010] In a preferred embodiment, the dimer-forming protein element comprises at least one complete or incomplete leucine zipper (LZ) dimerization motif, leucine repeat motif (LRM), leucine-rich repeat motif (LRR) or / and tryptophan repeat motif (TRM). In a more preferred embodiment, the dimer-forming protein element comprises one, two, three, four, five or six complete leucine zipper (LZ) dimerization motifs, leucine repeat motifs (LRM), leucine-rich repeat motifs (LRR) or / and tryptophan repeat motifs (TRM).
[0011] In some embodiments, the dimer-forming protein element can be derived from a dimerization motif in reverse transcriptase RT or an inactive mutant of reverse transcriptase RT from human immunodeficiency virus-1 (HIV-1), AMV (avian myeloblastosis virus), M-MuLV (mouse M-MuLV gene), murine leukemia virus (MLV), Moloney murine leukemia virus (MoLV), simian immunodeficiency virus (SIV), or simian-human immunodeficiency virus (SHIV). Specifically, the inactive mutant of reverse transcriptase RT includes the following mutation sites: Y63A, D113A, R115A, and Y221A / S in the inactive mutant of reverse transcriptase RT full-length (SEQ ID NO. 5).
[0012] In some embodiments, the dimer-forming protein element can also be derived from a protein containing a reverse transcriptase domain, such as a dimerization region on a dimer protein in the pol protein of the Terrabacteria group (NCBI Reference Sequence Accession No. WP_229790312.1), a dimerization region on a dimer protein in a protein containing a reverse transcriptase domain of Enterococcus alcedinis (NCBI GenBank Accession No. GGI66836.1), a dimerization region on a dimer protein in a protein containing a reverse transcriptase domain of the Terrabacteria group (NCBI Reference Sequence Accession No. WP_229671171.1), and a dimerization region on a dimer protein in a protein containing a reverse transcriptase domain of Actinoplanes campanulatus (NCBI GenBank: GGN52109.1).
[0013] In some embodiments, the dimer-forming protein element comprises (1) an amino acid sequence having at least 95% sequence identity to the amino acid sequence shown in any one of SEQ ID NOs. 7 to 13; or (2) an amino acid sequence having at least 95% sequence identity to the amino acid sequence shown in any one of SEQ ID NOs. 7 to 10; and, according to the sequence numbering of any one of SEQ ID NOs. 7 to 10, an amino acid substitution is present at at least one of the three positions Y63, D113, and R115, preferably substituted with alanine; or, according to the sequence numbering of any one of SEQ ID NOs. 7 to 10, an amino acid substitution is present at one position Y221, preferably substituted with alanine or serine. In a preferred embodiment, the dimer-forming protein element comprises an amino acid sequence having at least 95% sequence identity to the amino acid sequence shown in SEQ ID NO. 7, SEQ ID NO. 9, SEQ ID NO. 10, or SEQ ID NO. 12. In a preferred embodiment, the dimer-forming protein element comprises the amino acid sequence shown in SEQ ID NO.7, SEQ ID NO.9, SEQ ID NO.10 or SEQ ID NO.12.
[0014] Specifically, the dimer-forming protein element (1) comprises the amino acid sequence shown in any one of SEQ ID NO.7, 8, 9, 10, 11, 12 or 13; or (2) comprises the amino acid sequence shown in SEQ ID NO.7, and, according to the sequence numbering shown in SEQ ID NO.7, at least one of the following positions undergoes amino acid substitution: Y63A, D113A, R115A, Y221A and Y221S; or (3) comprises the amino acid sequence shown in SEQ ID NO.8, and, according to the sequence numbering shown in SEQ ID NO.8, at least one of the following positions undergoes amino acid substitution: Y63A, D113A, R115A, Y221A and Y221S; or (4) comprises the amino acid sequence shown in SEQ ID NO.9, and, according to the sequence numbering shown in SEQ ID NO.9, at least one of the following positions undergoes amino acid substitution: Y63A, D113A, R115A, Y221A and Y221S; or (5) comprises the amino acid sequence shown in SEQ ID NO. NO.10, and, according to the sequence numbering of SEQ ID NO.10, specifically, an amino acid substitution occurs at at least one of the following positions: Y63A, D113A, R115A, Y221A, and Y221S; or (6) comprises the amino acid sequence of SEQ ID NO.11, and, according to the sequence numbering of SEQ ID NO.11, specifically, an amino acid substitution occurs at at least one of the following positions: Y63A, D113A, R115A, Y221A, and Y221S; or (7) comprises the amino acid sequence of SEQ ID NO.12, and, according to the sequence numbering of SEQ ID NO.12, specifically, an amino acid substitution occurs at at least one of the following positions: Y63A, D113A, R115A, Y221A, and Y221S; or (8) comprises the amino acid sequence of SEQ ID NO.13, and, according to the sequence numbering of SEQ ID In the sequence numbering shown in NO.13, amino acid substitution occurs in at least one of the following positions: Y63A, D113A, R115A, Y221A and Y221S.
[0015] In some embodiments, the dimer-forming protein element (1) comprises the amino acid sequence shown in any one of SEQ ID NOs. 7 to 13; or (2) comprises the amino acid sequence shown in any one of SEQ ID NOs. 7 to 10; and, according to the sequence numbering shown in any one of SEQ ID NOs. 7 to 10, has an amino acid substitution at at least one of the three positions Y63, D113 and R115, preferably substituted with alanine; or, according to the sequence numbering shown in any one of SEQ ID NOs. 7 to 10, has an amino acid substitution at one position Y221, preferably substituted with alanine or serine.
[0016] Specifically, according to the sequence numbering shown in any one of SEQ ID NOs. 7 to 10, at least one of the three positions Y63, D113 and R115 has an amino acid substitution, preferably substituted by alanine, and the dimer protein element has no reverse transcriptase activity; or, according to the sequence numbering shown in any one of SEQ ID NOs. 7 to 10, the dimer protein element has an amino acid substitution at one position Y221, preferably substituted by alanine or serine, and has no reverse transcriptase activity.
[0017] In some embodiments, the dimer-forming protein element is a dimerization motif of reverse transcriptase RT, and the dimerization motif of reverse transcriptase RT can be: RT trucapep_01 (SEQ ID NO.7), RT trucapep_02 (SEQ ID NO.8), RT trucapep_03 (SEQ ID NO.9), RT trucapep_04 (SEQ ID NO.10), RT trucapep_05 (SEQ ID NO.11), RT trucapep_06 (SEQ ID NO.12), and RT trucapep_07 (SEQ ID NO.13).
[0018] In a preferred embodiment, the dimer-forming protein element may be one or more of the above-mentioned RT trucapep_01, RT trucapep_03, RT trucapep_04 and RT trucapep_06.
[0019] In a preferred embodiment, the dimer-forming protein element is RT trucapep_04; both ends of the RT trucapep_04 contain complete LRMs, and the two dimerization domains interact with each other to trigger the activation of the proteolytic enzyme PR by promoting Gag-pol dimerization. Therefore, PR activation is normal, the viral packaging titer is high, and the virus-like protein particles have a good delivery effect.
[0020] In some embodiments, the envelope protein ENV, the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, and the proteolytic enzyme PR can be cloned from genes encoding viral polypeptides capable of self-assembly into defective, non-reproductive virus particles (which can be obtained from the genomic DNA of DNA viruses or the genomic cDNA of RNA viruses or from available subgenomes containing these genes).
[0021] In some embodiments, the matrix protein MA, the capsid protein CA and the nucleocapsid protein NC can be structural domains of the population-specific antigen gene (Gag) derived from any viral genome, and the precursors of these structural proteins (matrix protein MA, capsid protein CA and nucleocapsid protein NC) can be obtained by translating the Gag fusion protein, and the precursors are processed by the proteolytic enzyme PR and a protein element that can form a dimer to obtain mature matrix protein MA, capsid protein CA and nucleocapsid protein NC with correct structure.
[0022] In some embodiments, the envelope protein ENV, the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, and the proteolytic enzyme PR may be derived from retrovirus Gag or lentivirus Gag.
[0023] In some embodiments, the envelope protein ENV, the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, and the proteolytic enzyme PR can be derived from alpha retrovirus, beta retrovirus, gamma retrovirus, delta retrovirus, epsilon retrovirus, foamy virus, human immunodeficiency virus-1 (HIV-1), AMV (avian myeloblastosis virus), M-MuLV (mouse M-MuLV gene), murine leukemia virus (MLV), Moloney murine leukemia virus (MoLV), simian immunodeficiency virus (SIV), vi SNA / MAEDI virus (VMV), caprine arthritis encephalitis virus (CAEV), equine infectious anemia virus (EIAV), feline immunodeficiency virus (FIV), bovine immunodeficiency virus (BIV), human foamy virus (HFV), Friedreich's virus (FV), Abelson murine leukemia virus (A-MLV), mouse stem cell virus (MSCV), mouse mammary tumor virus (MMTV), Moloney murine sarcoma virus (MoMSV), Rous sarcoma virus (RSV), Fujinami sarcoma virus (FuSV), FBR murine osteosarcoma virus (FBR MSV), avian myelocytoma virus 29 (MC29), avian erythroblastosis virus (AEV), human T-cell leukemia virus (HTLV), FriendMLV (FrMLV), avian sarcoma virus (ASV), avian leukosis virus, avian myeloblastosis virus, UR2 sarcoma virus, Y73 sarcoma virus, Jaagsiekte sheep retrovirus, leaf monkey virus, Mason-pfizer monkey virus, squirrel monkey retrovirus, avian oncomillary virus 2, bovine leukemia virus, primate T-lymphotropic virus 1, primate T-lymphotropic virus 2, primate T-lymphotropic virus 3, bigeye shad dermal sarcoma virus, bigeye shad epidermoproliferative virus 1, bigeye shad epidermoproliferative virus 2, chicken syncytial virus, feline leukemia virus, Finkel-Biskis-Jinkins murine sarcoma virus, Gardner-Arnstein feline sarcoma virus, gibbon ape leukemia virus, guinea pig type C tumor virus, Hardy-Zuckerman feline sarcoma virus, Harvey murine sarcoma virus, Kir Sten rat sarcoma virus, koala retrovirus, Moloney rat sarcoma virus, porcine tumor virus type C, reticuloendotheliosis virus, Snyder-Theilen feline sarcoma virus, Trager duck spleen necrosis virus, viper retrovirus, hairy monkey sarcoma virus, Jembrana disease virus, puma lentivirus, bovine foamy virus, equine foamy virus, feline foamy virus, brown bushbaby simian foamy virus, Bornean orangutan simian foamy virus, central chimpanzee simian foamy virus, cynomolgus macaque simian foamy virus, eastern chimpanzee simian foamy virus, green monkey simian foamy virus, long-tailed monkey simian foamy virus, Japanese macaque simian foamy virus, rhesus macaque simian foamy virus, spider monkey simian foamy virus, squirrel monkey simian foamy virus, Formosan macaque simian foamy virus, western chimpanzee simian foamy virus, western lowland gorilla simian foamy virus, white-eared marmoset simian foamy virus, and yellow-breasted capuchin simian foamy virus.
[0024] In a preferred embodiment, the fusion protein of the matrix protein MA, the capsid protein CA and the nucleocapsid protein NC can be derived from human immunodeficiency virus-1 (HIV-1) Gag. Specifically, HIV-1 Gag encodes the main structural proteins: matrix protein (p17, MA), capsid protein (p24, CA), and nucleocapsid protein (p7, NC).
[0025] In a preferred embodiment, the fusion protein of the matrix protein MA, the capsid protein CA and the nucleocapsid protein NC can be derived from Moloney murine leukemia virus (MLV) Gag. Specifically, MLVGag encodes the main structural proteins: matrix protein (p15, MA, SEQ ID NO.1), capsid protein (p30, CA, SEQ ID NO.3), and nucleocapsid protein (p10, NC, SEQ ID NO.2).
[0026] In a preferred embodiment, the proteolytic enzyme PR can be derived from the proteolytic enzyme PR (SEQ ID NO. 4) of Moloney murine leukemia virus (MLV), and the proteolytic enzyme PR is used to cleave the Gag fusion protein of the mature virus to produce infectious virus particles.
[0027] In some embodiments, described envelope protein ENV can be derived from any retrovirus.It is understandable that ENV can be the double tropism envelope protein that allows the cell transduction of the mankind and other species, or can be the ecotropic envelope protein (Env not only mediates virus entry cell, and also is the main target of cell and antibody reaction) that can only transduce mouse and rat cells.
[0028] In some embodiments, by being connected to the specific ligand of the receptor of envelope protein ENV and antibody or being used to target specific cell type to carry out targeted recombinant virus.Described virus-like protein particle can have target specificity by inserting for example glycolipid or protein.Targeted operation is usually realized by using antibody that retroviral vector is targeted to the antigen (cell type found in some tissue, or cancer cell type) on specific cell type.
[0029] In some embodiments, the envelope protein ENV is derived from a non-retrovirus (e.g., CMV or VSV). Examples of ENV genes that may also be derived from retroviral sources include, but are not limited to, Moloney murine leukemia virus (MoMuLV), Harvey murine sarcoma virus (HaMuSV), murine mammary oncovirus (MuMTV), gibbon ape leukemia virus (GALV), human immunodeficiency virus (HIV), and Rous sarcoma virus (RSV). Other ENV genes may also be used, such as vesicular stomatitis virus (VSV-G, Vesicular stomatitis virus, G is Glycoprotein glycoprotein, and coat protein VSV-G mediates significant strong infectivity and pantropy), cytomegalovirus envelope (CMV), or influenza virus hemagglutinin (HA).
[0030] In some embodiments, the target nucleic acid or target protein is a therapeutic RNA or therapeutic protein; preferably, the therapeutic RNA or therapeutic protein comprises CRISPR / Cas nuclease, base editor, epigenetic editor, recombinase, transcription factor, reverse transcriptase, lead editor (PE editor), antibody and other functional proteins.
[0031] In some embodiments, the CRISPR / Cas nuclease can be a type I CRISPR / Cas protein, a type II CRISPR / Cas protein, a type V CRISPR / Cas protein, or a type VI CRISPR / Cas protein; in some embodiments, the CRISPR / Cas nuclease can also be a CRISPR / Cas nuclease variant with reduced nucleic acid cleavage activity or a CRISPR / Cas nuclease variant without nucleic acid cleavage activity.
[0032] In some embodiments, the base editor can be a CBE editor (for C-to-T transitions), an ABE editor (for A-to-G transitions), a CGBE editor (for C-to-G transitions), an ACBE editor (for C-to-T and A-to-G transitions), and an AGBE editor (for C-to-G, C-to-T, C-to-A, and A-to-G transitions).
[0033] In some embodiments, the CRISPR / Cas nuclease can be fused with a heterologous polypeptide, and the fused heterologous polypeptide can be a reverse transcriptase, a protein modification enzyme, a nucleic acid modification enzyme, and a transcriptional regulator.
[0034] In some embodiments, the CRISPR / Cas nuclease can be fused with a transcriptional regulator to form an epigenetic editor, and the transcriptional regulator can be a transcriptional activation domain or a transcriptional repression domain. In some embodiments, the transcriptional activation domain can be a transcriptional activator, a histone lysine methyltransferase, a histone lysine demethylase, a histone acetyltransferase, and a DNA demethylase. In a preferred embodiment, the transcriptional activation domain includes VP64; P65; RTA; truncated P65; truncated RTA; or one or more fusion forms thereof. In some embodiments, the transcriptional repression domain can be a transcriptional repressor, a ZIM3 domain, a KOX1 repression domain, a Mad mSIN3 interaction domain (SID), an ERF repressor domain (ERD), an SRDX repression domain, a histone lysine methyltransferase, a histone lysine demethylase, a histone lysine deacetylase, a DNA methylase, and a peripheral recruitment element. In a preferred embodiment, the transcriptional repression domain is selected from a KRAB catalytic domain, a DNA methyltransferase, or a combination thereof. In a preferred embodiment, the DNA methyltransferase may be Dnmt1, Dnmt3A, Dnmt3B, Dnmt3L or a combination thereof.
[0035] In a preferred embodiment, the epigenetic editor can be CRISPR / Cas nuclease-KRAB catalytic domain, KRAB-CRISPR / Cas nuclease-Dnmt3A-Dnmt3L (N-terminus → C-terminus), CRISPR / Cas nuclease-VPR (N-terminus → C-terminus), CRISPR / Cas nuclease-VP64 (N-terminus → C-terminus).
[0036] In some embodiments, the transcription factor is a CRISPR / Cas polypeptide fusion polypeptide comprising a transcriptional regulator, a zinc finger protein transcription factor (ZFP-TF) and a transcription activator-like effector transcription factor (TALE-TF).
[0037] In a preferred embodiment, the recombinase may be Cre recombinase, Hin recombinase, Tre recombinase, or / and FLP recombinase.
[0038] In some embodiments, the antibody can be a single-chain antibody such as a nanobody, a single-chain Fv antibody; a diabody; a minibody, etc.; and the antibody can bind to an intracellular antigen, an antigen present on the cell surface, or an extracellular antigen.
[0039] In some embodiments, the transcription factor may be a cell reprogramming factor such as c-Myc, SOX2, OCT4, KLF4, FOXA, HNF1α, HNF4α, GATA4, etc.
[0040] In some embodiments, the other functional proteins include telomerase protein, telomerase holoenzyme protein, telomere-associated protein and telomerase fusion protein.
[0041] Another aspect of the present invention provides a nucleic acid comprising: (i) a first polynucleotide comprising a nucleic acid sequence encoding a backbone fusion protein; and (ii) a second polynucleotide comprising a nucleic acid sequence encoding a Gag fusion protein fused to a target nucleic acid or target protein; and (iii) a third polynucleotide comprising a nucleic acid sequence encoding an envelope protein ENV; wherein the backbone fusion protein comprises: (a) a Gag fusion protein composed of a fusion of a matrix protein MA, a capsid protein CA, and a nucleocapsid protein NC; and (b) a proteolytic enzyme PR, wherein the proteolytic enzyme PR is covalently linked to the Gag fusion protein by cleaving a short peptide by a heterologous protease; and (c) a protein element capable of forming a dimer, wherein the protein element capable of forming a dimer is covalently linked to the proteolytic enzyme PR by cleaving a short peptide by a heterologous protease; the protein element capable of forming a dimer, (1) comprising an amino acid sequence having at least 95% sequence identity to the amino acid sequence shown in any one of SEQ ID NOs. 7 to 13; or (2) comprising an amino acid sequence having at least 95% sequence identity to the amino acid sequence shown in any one of SEQ ID NOs. NO.7 to 10 compared to an amino acid sequence having at least 95% sequence identity; and, according to the sequence numbering shown in SEQ ID NO.10, at least one of the three positions Y63, D113 and R115 has an amino acid substitution, preferably substituted with alanine; or, according to the sequence numbering shown in SEQ ID NO.10, at one position Y221 has an amino acid substitution, preferably substituted with alanine or serine; wherein the target nucleic acid or target protein is covalently linked to the Gag fusion protein by cleavage of a short peptide by a heterologous protease.
[0042] In some embodiments, the first polynucleotide, the second polynucleotide, and / or the third polynucleotide are constructed into conventional eukaryotic vectors; the expression amount of each protein is regulated by controlling the ratio of these vectors to form a large number of mature virus-like protein particles for delivering nucleic acids and proteins.
[0043] In a preferred embodiment, in the first polynucleotide, the backbone fusion protein structure is matrix protein MA-capsid protein CA-nucleocapsid protein NC-heterologous protease cleavage short peptide-proteolytic enzyme PR-heterologous protease cleavage short peptide-protein element that can form a dimer in the direction from the N-terminus to the C-terminus of the protein; in the second polynucleotide, the target nucleic acid or target protein is covalently linked to the C-terminus of the Gag fusion protein (composed of a fusion of matrix protein MA, capsid protein CA and nucleocapsid protein NC) through a heterologous protease cleavage short peptide.
[0044] In some embodiments, the heterologous protease cleaving peptide can be a tobacco etch virus (TEV) protease cleaving peptide, a matrix metalloproteinase cleaving peptide, a plasminogen activator cleaving peptide, a furin cleaving peptide, an HIV-1 protease cleaving peptide, a PreScission cleaving peptide, a human rhinovirus 3C protease cleaving peptide, an enterokinase cleaving peptide, an Epstein-Barr virus protease cleaving peptide, a cathepsin D cleaving peptide, a thrombin cleavage peptide, or variants thereof.
[0045] Specifically, the amino acid sequence of the heterologous protease cleaving peptide is: TSTLLI (SEQ ID NO.14), TSTLLMENSS (SEQ ID NO.15), PRSSLYPALTP (SEQ ID NO.16), VQALVLTQ (SEQ ID NO.17), PLQVLT (SEQ ID NO.18), PLQVLTLNIERR (SEQ ID NO.19), PLQVLTLNIE (SEQ ID NO.20), MSKLLATVVS (SEQ ID NO.21) or / and IRKIFLDG (SEQ ID NO.22).
[0046] In some embodiments, the second polynucleotide further comprises one or more heterologous polypeptides, and the target nucleic acid or target protein is covalently linked to the Gag fusion protein via the one or more heterologous polypeptides and the heterologous protease cleavage short peptide; the one or more heterologous polypeptides are independently a nuclear export signal, an epitope tag, a reporter gene sequence, an enzyme with a detectable signal, or / and a subcellular localization sequence.
[0047] In some embodiments, the C-terminus of the Gag fusion protein is connected to the one or more heterologous polypeptides, the heterologous protease cleavage peptide and the target nucleic acid or target protein through a connecting peptide; in some embodiments, the C-terminus of the Gag fusion protein is connected to the one or more heterologous polypeptides, the heterologous protease cleavage peptide, the one or more heterologous polypeptides and the target nucleic acid or target protein through a connecting peptide; in a preferred embodiment, the C-terminus of the Gag fusion protein is connected to multiple nuclear export signals, the heterologous protease cleavage peptide and the target nucleic acid or target protein through a connecting peptide; in a preferred embodiment, the C-terminus of the Gag fusion protein is connected to multiple nuclear export signals, the heterologous protease cleavage peptide, the subcellular localization sequence and the target nucleic acid or target protein through a connecting peptide; in a preferred embodiment, the C-terminus of the Gag fusion protein is connected to multiple nuclear export signals, the heterologous protease cleavage peptide, the epitope tag, the subcellular localization sequence and the target nucleic acid or target protein through a connecting peptide.
[0048] In some embodiments, the heterologous polypeptide is selected from a nuclear export sequence (NES). A nuclear export signal (NES) is a short targeting peptide containing four hydrophobic residues in a protein that targets it for export from the nucleus to the cytoplasm via nuclear transport through the nuclear pore complex. The common amino acid sequence of an NES is LXXXLXXLXL, where L is a hydrophobic residue (usually leucine), X is another amino acid, and XXX and XX vary slightly in length. The spatial arrangement of these amino acid residues can be explained by the structure of proteins containing the NES. These key residues are often on the same face of the protein secondary structure, enabling them to interact with nuclear transport proteins. Specifically, the sequence of the nuclear export sequence (NES) is an existing conventional nuclear export sequence, and the number of NES can be one or more, and can be 2 NES, 3 NES, 4 NES, 5 NES, 6 NES, 7 NES, 8 NES, 9 NES or 10 NES; the amino acid sequence of the nuclear export sequence NES is: LQLPPLERLTL (SEQ ID NO.23), and the number of the nuclear export sequence NES can be multiple, for example, 3×NES (SEQ ID NO.24).
[0049] In some embodiments, the heterologous polypeptide is selected from an epitope tag. Such epitope tags are conventional tags, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art know how to select an appropriate epitope tag based on the desired purpose (e.g., purification, detection, or tracing).
[0050] In some embodiments, the heterologous polypeptide is selected from a reporter gene sequence. Such reporter genes are well known to those skilled in the art, and examples thereof include but are not limited to GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0051] In some embodiments, the heterologous polypeptide can also be an enzyme, a radioisotope, a member of a specific binding pair, a fluorophore, a fluorescent protein, a quantum dot, etc. that can detect a signal.
[0052] In some embodiments, the heterologous polypeptide provides subcellular localization, i.e., the heterologous polypeptide contains a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence for retaining the fusion protein outside the nucleus (e.g., a nuclear export sequence (NES)), a sequence for retaining the fusion protein in the cytoplasm, a mitochondrial localization signal for targeting to mitochondria, a chloroplast localization signal for targeting to chloroplasts, an ER retention signal, etc.).
[0053] In some embodiments, the second polynucleotide comprises (is fused with) a nuclear localization signal (NLS) (e.g., in some embodiments, 2 or more, 3 or more, 4 or more, or 5 or more NLS). Thus, in some embodiments, the second polynucleotide includes one or more NLSs (e.g., 2 or more, 3 or more, 4 or more, or 5 or more NLS). In some embodiments, one or more NLSs (2 or more, 3 or more, 4 or more, or 5 or more NLS) are positioned at or near the N-terminus and / or C-terminus (e.g., within 50 amino acids). In some embodiments, one or more NLSs (2 or more, 3 or more, 4 or more, or 5 or more NLS) are positioned at or near the N-terminus (e.g., within 50 amino acids). In some embodiments, one or more NLSs (2 or more, 3 or more, 4 or more, or 5 or more NLS) are positioned at or near the C-terminus (e.g., within 50 amino acids). In some embodiments, one or more NLS (3 or more, 4 or more, or 5 or more NLS) are positioned at or near both the N-terminus and the C-terminus (e.g., within 50 amino acids). In some embodiments, one or more NLS are positioned at the N-terminus and one or more NLS are positioned at the C-terminus.
[0054] In some embodiments, the second polynucleotide comprises (is fused with) 1 to 10 NLSs (e.g., 1-9, 1-8, 1-7, 1-6, 1-5, 2-10, 2-9, 2-8, 2-7, 2-6, or 2-5 NLSs). In some embodiments, the second polynucleotide comprises (is fused with) 2 to 5 NLSs (e.g., 2-4 or 2-3 NLSs). One or more NLSs are linked to the N-terminus of the target nucleic acid or target protein, and one or more NLSs are linked to the C-terminus of the target nucleic acid or target protein. More specifically, the sequence of the NLS is an existing conventional nuclear localization signal sequence, including but not limited to the NLS of the SV40 virus large T antigen (SEQ ID NO.38), nucleoplasmin NLS (SEQ ID NO.39), c-myc NLS, hRNPA1 M9 NLS, the IBB domain of importin-α, myoma T protein, human p53, mouse c-abl IV, influenza virus NS1, hepatitis virus δ antigen, mouse Mx1 protein, human poly (ADP-ribose) polymerase, steroid hormone receptor (human) glucocorticoid and other commonly used NLSs.
[0055] In some embodiments, the connecting peptide (also known as linker polypeptide) of the second polynucleotide generally has flexible properties, but does not exclude other chemical bonds. Suitable joints include polypeptides with a length between 4 to 40 amino acids or a length between 4 to 25 amino acids. These joints can be produced with coupled proteins by using synthetic oligonucleotides encoding joints, or can be encoded by the nucleic acid sequence encoding the fusion protein. Peptide joints with a certain degree of flexibility can be used. The connecting peptide can actually have any amino acid sequence, and it should be remembered that preferred joints will have sequences that produce generally flexible peptides. The purposes of small amino acids (such as glycine and alanine) are used to produce flexible peptides. For those skilled in the art, it is conventional to produce such sequences. A variety of different joints are commercially available and are considered to be suitable for use.
[0056] Specifically, examples of linker polypeptides include glycine polymers (G)n, glycine-serine polymers, glycine-alanine polymers, alanine-serine polymers. Exemplary linkers can comprise amino acid sequences including, but not limited to, GS linker peptides (e.g., GS (SEQ ID NO. 25), GGSGG (SEQ ID NO. 26), GSGSG (SEQ ID NO. 27), GSGGG (SEQ ID NO. 28), GGGSG (SEQ ID NO. 29), GSSSG (SEQ ID NO. 30)), XTEN linker, SGGS (SEQ ID NO. 31), (SGGS)2, GGS (SEQ ID NO. 32), (GGS)3, (GGS)7, or a combination of XTEN and SGGS (SGSETPGTSESATPES (SEQ ID NO. 33), SGGSSGSETPGTSESATPESSGGS (SEQ ID NO. 34), SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO. 35)), and the like. More specifically, the connecting peptide includes but is not limited to the amino acid sequences shown in SEQ ID NOs. 25 to 35. One skilled in the art will recognize that the design of the peptide conjugated to any desired element may include a fully or partially flexible linker, such that the linker may include a flexible linker and one or more portions that confer a less flexible structure.
[0057] In some embodiments, when the nucleic acid of interest or protein of interest comprises a CRISPR / Cas nuclease, the nucleic acid further comprises (iv) a fourth polynucleotide comprising a nucleic acid sequence encoding a guide RNA (gRNA), wherein the gRNA binds to the CRISPR / Cas nuclease.
[0058] In a preferred embodiment, the CRISPR / Cas nuclease can be a type II CRISPR-Cas effector protein, a type V CRISPR-Cas effector protein, or a type VI CRISPR-Cas effector protein, and the fourth polynucleotide is a nucleotide encoding a gRNA that can bind to the corresponding CRISPR-Cas effector protein, and the gRNA can bind to a type II CRISPR-Cas effector protein, a type V CRISPR-Cas effector protein, or a type VI CRISPR-Cas effector protein.
[0059] Another aspect of the present invention provides a viral vector system comprising the nucleic acid; wherein the first polynucleotide, the second polynucleotide and the third polynucleotide are each located on a separate vector; or one or more of the first polynucleotide, the second polynucleotide and the third polynucleotide are located on the same vector; or it comprises the nucleic acid; wherein the first polynucleotide, the second polynucleotide, the third polynucleotide and the fourth polynucleotide are each located on a separate vector; or one or more of the first polynucleotide, the second polynucleotide, the third polynucleotide and the fourth polynucleotide are located on the same vector.
[0060] Specifically, the virus-like protein particles are encoded by one, two, three or four eukaryotic expression vectors, such as the pCMV eukaryotic expression vector.
[0061] In some embodiments, the viral vector system comprises one vector, two vectors, three vectors, or four vectors. In some embodiments, when the viral vector system comprises one vector, the vector comprises a first expression cassette containing an envelope protein ENV, a protein element containing a matrix protein MA, a capsid protein CA, a nucleocapsid protein NC, a proteolytic enzyme PR, and a dimer-forming protein, and a target nucleic acid or target protein are sequentially connected (N-terminal → C-terminal) to form a second expression cassette of a fusion protein; in some embodiments, when the viral vector system comprises two vectors, they are vector A and vector B, vector A encodes the envelope protein ENV, and vector B encodes the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, the proteolytic enzyme PR, and a dimer-forming protein. A dimer protein element, and a target nucleic acid or target protein are sequentially connected (N-terminus → C-terminus) to form a fusion protein; in some embodiments, when the viral vector system comprises three vectors, namely vector A, vector C and vector D, vector A encodes the envelope protein ENV, vector C encodes the matrix protein MA, capsid protein CA, nucleocapsid protein NC, proteolytic enzyme PR, and a dimer-forming protein element are sequentially connected (N-terminus → C-terminus) to form a fusion protein, vector D encodes the matrix protein MA, capsid protein CA, nucleocapsid protein NC, and a target nucleic acid or target protein are sequentially connected (N-terminus → C-terminus) to form a fusion protein.
[0062] In a preferred embodiment, the viral vector system comprises three vectors (including vector A, vector C and vector D), encoding three proteins respectively, vector A encodes the envelope protein ENV; the backbone fusion protein sequence of vector C (N-terminus → C-terminus) is: matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-protein element that can form a dimer; the fusion protein sequence of vector D (N-terminus → C-terminus) is various: ① matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(NES)n-heterologous protease cleavage short peptide-target nucleic acid or target protein ; ② Matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(NES)n-heterologous protease cleavage short peptide-target nucleic acid or target protein-NLS; ③ Matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(NES)n-heterologous protease cleavage short peptide-NLS-target nucleic acid or target protein; ④ Matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(NES)n-heterologous protease cleavage short peptide-NLS-target nucleic acid or target protein-NLS, wherein n can be 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
[0063] In a preferred embodiment, the fusion protein encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(NES)3-heterologous protease cleavage short peptide in the vector D comprises an amino acid sequence having at least 95% sequence identity compared to the amino acid sequence shown in SEQ ID NO.40; in a more preferred embodiment, the fusion protein encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(NES)3-heterologous protease cleavage short peptide in the vector D comprises the amino acid sequence shown in SEQ ID NO.40.
[0064] In a more preferred embodiment, the viral vector system comprises three vectors (including vector A, vector C and vector D), each encoding three proteins, and having seven combinations according to the different sequences of the proteins. Combination ①: vector A encoding the envelope protein ENV, vector C encoding the matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding the matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage peptide-RT trucapep_01 fusion protein (SEQ ID NO.44). Combination ②: vector A encoding envelope protein ENV, vector C encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage short peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-RT trucapep_02 fusion protein (SEQ ID NO.45). Combination ③: vector A encoding envelope protein ENV, vector C encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage short peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-RT trucapep_03 fusion protein (SEQ ID NO.46). Combination ④: vector A encoding envelope protein ENV, vector C encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage short peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-RT trucapep_04 fusion protein (SEQ ID NO.47).Combination ⑤: vector A encoding envelope protein ENV, vector C encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage short peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-RT trucapep_05 fusion protein (SEQ ID NO.48). Combination ⑥: vector A encoding envelope protein ENV, vector C encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage short peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-RT trucapep_06 fusion protein (SEQ ID NO.49). Combination ⑦: vector A encoding envelope protein ENV, vector C encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-connecting peptide-(3×NES)-heterologous protease cleavage short peptide-NLS-SpCas9 / ABE8e-NLS fusion protein (SEQ ID NO.36 and SEQ ID NO.37, respectively), and vector D encoding matrix protein MA-capsid protein CA-nucleocapsid protein NC-proteolytic enzyme PR-heterologous protease cleavage short peptide-RT trucapep_07 fusion protein (SEQ ID NO.50).
[0065] In some embodiments, the backbone fusion protein in the first polynucleotide comprises an amino acid sequence having at least 95% sequence identity with the amino acid sequence shown in any one of SEQ ID NO.44 to 50; the fusion protein in the second polynucleotide comprises an amino acid sequence having at least 95% sequence identity with the amino acid sequence shown in SEQ ID NO.36 or 37; in some embodiments, the backbone fusion protein in the first polynucleotide comprises the amino acid sequence shown in any one of SEQ ID NO.44 to 50; the fusion protein in the second polynucleotide comprises the amino acid sequence shown in SEQ ID NO.36 or 37.
[0066] In some embodiments, when the delivered target nucleic acid or target protein contains a CRISPR / Cas system, the viral vector system comprises four vectors, namely vector A, vector C, vector D and vector E, wherein vector A encodes the envelope protein ENV, vector C encodes the matrix protein MA, capsid protein CA, nucleocapsid protein NC, proteolytic enzyme PR, and a protein element that can form a dimer, which is sequentially connected (N-terminus → C-terminus) to form a backbone fusion protein, vector D encodes the matrix protein MA, capsid protein CA, nucleocapsid protein NC, and the target nucleic acid or target protein, which are sequentially connected (N-terminus → C-terminus) to form a fusion protein, and vector E encodes an expression cassette for expressing gRNA.
[0067] In some embodiments, the gRNA comprises the nucleotide sequence shown in SEQ ID NO.53 or a nucleotide sequence having 1 to 10 nucleotide substitutions, deletions and / or insertions compared to the nucleotide sequence shown in SEQ ID NO.53.
[0068] Specifically, in order to package and form a large number of virus-like protein particles containing the target nucleic acid or target protein, the expression level of each protein needs to be controlled. A large number of virus-like protein particles containing the target nucleic acid or target protein can be obtained by regulating the ratio of the above-mentioned vectors.
[0069] Another aspect of the present invention provides a cell comprising the virus-like protein particle, the nucleic acid, and the viral vector system; preferably, the cell is a eukaryotic cell; more preferably, the cell is a human cell.
[0070] Another aspect of the present invention provides a pharmaceutical composition comprising the virus-like protein particle.
[0071] In some embodiments, the pharmaceutical composition further comprises one or more pharmaceutically acceptable additives; the one or more additives are buffers, surfactants, antioxidants, hydrophilic polymers, dextrins, chelating agents, suspending agents, solubilizers, thickening agents, stabilizers, bacteriostats, wetting agents, preservatives, protease inhibitors, nuclease inhibitors, and reagents required for developing or visualizing detectable labels.
[0072] Another aspect of the present invention provides a method for treating a subject diagnosed with a disease associated with or caused by a point mutation, which point mutation can be corrected by a virus-like protein particle containing a base editor provided herein. For example, in some embodiments, a method is provided, comprising administering an effective amount of a virus-like protein particle containing an adenosine base editor to a subject suffering from the disease (e.g., a cancer associated with a point mutation as described above), the editor correcting the point mutation or introducing an inactivating mutation into a disease-associated gene. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a tumor disease. In some embodiments, the disease is a metabolic disease.
[0073] In some embodiments, the purpose of the methods provided herein is to restore the function of dysfunctional genes by genome editing. The virus-like protein particles containing base editors provided herein can be validated in vitro for human therapy based on gene editing, for example, by correcting disease-associated mutations in human cell culture. Those skilled in the art will understand that the virus-like protein particles containing base editors provided herein can be used to correct any single-point G to A or C to T mutation.
[0074] In some embodiments, the pharmaceutical compositions provided herein are formulated for delivery to a subject, e.g., to a human subject to achieve targeted genome modification within the subject. In some embodiments, cells are obtained from the subject and contacted with any pharmaceutical composition provided herein. In some embodiments, cells removed from the subject and contacted ex vivo with the pharmaceutical composition are reintroduced into the subject, optionally after the desired genome modification is achieved or detected in the cell.
[0075] The formulations of the pharmaceutical compositions described herein can be prepared by any method known in the art of pharmacology. Generally, such preparation methods include combining the active ingredient with an excipient and / or one or more other auxiliary ingredients, and then, if necessary and / or desired, shaping and / or packaging the product into the desired single-dose or multi-dose units.
[0076] Specifically, the fusion protein provided herein can be used to treat various rare diseases, tumors, cancers, inflammation, viral infections, genetic diseases, central nervous system diseases, aging and multiple autoimmune diseases and common and chronic diseases. More specifically, the disease for treatment can be hypertension, hyperlipidemia, idiopathic fibrosis (IPF), liver fibrosis, hepatitis B virus (HBV), hepatocellular carcinoma (HCC), scapulohumeral muscular dystrophy (FSHD), heterozygous familial hypercholesterolemia (HeFH), alpha-1 antitrypsin deficiency (A1AD), non-arteritic anterior ischemic optic neuropathy (NAION), retinitis pigmentosa (RP) or Duchenne muscular dystrophy (DMD).
[0077] Another aspect of the present invention provides a method for preparing virus-like protein particles, which comprises: (a) introducing the nucleic acid or the viral vector system into production cells; and (b) harvesting the supernatant containing the virus-like protein particles produced by the packaging cells, concentrating the supernatant, and obtaining virus-like protein particles.
[0078] In some embodiments, the production cell is selected from HEK293, 293T or 293FT. The concentration can be carried out by high-speed centrifugation or HPLC.
[0079] In some embodiments, the method of the present invention for delivering virus-like protein particles of nucleic acids and proteins comprises the following steps:
[0080] 1) Transfect a single plasmid containing a first expression cassette of the envelope protein ENV and a second expression cassette of the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, the proteolytic enzyme PR, a dimer-forming protein element, and the target nucleic acid or target protein in sequence (N-terminus → C-terminus) to form a fusion protein into the production cell; or,
[0081] Plasmid A encoding envelope protein ENV, plasmid B encoding matrix protein MA, capsid protein CA, nucleocapsid protein NC, proteolytic enzyme PR, dimer-forming protein element, and target nucleic acid or target protein are sequentially linked (N-terminus → C-terminus) to form a fusion protein, and these two plasmids are co-transfected into production cells; or,
[0082] Plasmid A encoding envelope protein ENV, plasmid C encoding matrix protein MA, capsid protein CA, nucleocapsid protein NC, proteolytic enzyme PR and dimer-forming protein elements are sequentially linked (N-terminus → C-terminus) to form a backbone fusion protein, and plasmid D encoding matrix protein MA, capsid protein CA, nucleocapsid protein NC, and target nucleic acid or target protein are sequentially linked (N-terminus → C-terminus) to form a fusion protein, and these three plasmids are co-transfected into production cells; or,
[0083] When the delivered target nucleic acid or target protein contains the CRISPR / Cas system, plasmid A encoding the envelope protein ENV, plasmid C encoding the matrix protein MA, capsid protein CA, nucleocapsid protein NC, proteolytic enzyme PR and a protein element capable of forming a dimer are sequentially connected (N-terminus → C-terminus) to form a backbone fusion protein, plasmid D encoding the matrix protein MA, capsid protein CA, nucleocapsid protein NC, and the target nucleic acid or target protein are sequentially connected (N-terminus → C-terminus) to form a fusion protein, and plasmid E encoding the gRNA nucleotide sequence are co-transfected into the production cells;
[0084] 2) Collecting the cell supernatant containing the virus-like protein particles of the target nucleic acid or target protein, and concentrating the supernatant to obtain the virus-like protein particles.
[0085] In some embodiments, the viral vector system comprises the two plasmids (plasmid A and plasmid B) described above, and the ratio of plasmid A: plasmid B is (1-50%): (50-99%).
[0086] In some embodiments, the viral vector system comprises the three plasmids mentioned above (plasmid A, plasmid C and plasmid D), plasmid A accounts for 1-20% of the total proportion of plasmid A, plasmid C and plasmid D; the ratio of plasmid C: plasmid D is (50-80%): (20-50%).
[0087] In some embodiments, the viral vector system comprises the above-mentioned four plasmids (plasmid A, plasmid C, plasmid D and plasmid E), plasmid A accounts for 1 to 20% of the total proportion of plasmid A, plasmid C and plasmid D; the ratio of plasmid C: plasmid D is (50 to 80%): (20 to 50%); plasmid E accounts for 1 to 50% of plasmid D.
[0088] Another aspect of the present invention provides a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with the virus-like protein particles provided by the present invention, wherein the contact causes the target nucleic acid to be modified. In a preferred embodiment, the modification comprises increasing or decreasing the expression of the target sequence in the target nucleic acid. In a preferred embodiment, the modification comprises deaminating the target adenine or target cytosine in the target nucleic acid to achieve base pair conversion. In a preferred embodiment, the target nucleic acid is selected from the group consisting of double-stranded DNA, single-stranded DNA, RNA, genomic DNA, and extrachromosomal DNA. In a preferred embodiment, the contact occurs outside the cell in vitro, inside the cultured cell, or inside the cell in vivo. In a preferred embodiment, the cell is a eukaryotic cell, more preferably a human cell.
[0089] Another aspect of the present invention provides the use of a polypeptide having an amino acid sequence as set forth in SEQ ID NO. 36 or 37, in combination with a polypeptide having an amino acid sequence as set forth in SEQ ID NO. 44, 45, 46, 47, 48, 49, or 50, in the preparation of a medicament for modifying nucleic acids. In a preferred embodiment, the modification comprises increasing or decreasing the expression of a target sequence in the target nucleic acid. In a preferred embodiment, the modification comprises deaminating a target adenine or cytosine in the target nucleic acid to effect a base pair conversion.
[0090] It should be noted that the present invention discovered that the Gag fusion protein precursor of the matrix protein MA, capsid protein CA and nucleocapsid protein NC is transcribed and translated in the production cells and transported to the surface of the production cells; the protein element that can form a dimer can interact with the proteolytic enzyme PR, activate the proteolytic enzyme PR, and mediate the processing and maturation of the Gag fusion protein precursor through the activated proteolytic enzyme PR, so as to facilitate the processing of the correct mature matrix protein MA, capsid protein CA and nucleocapsid protein NC on the surface of the production cells; at the same time, the envelope protein ENV is transcribed and translated in the production cells and transported to the surface of the production cells, and the components of the virus-like protein particles (including the matrix protein MA, capsid protein CA and nucleocapsid protein NC) and the envelope protein ENV are autonomously assembled, budded and secreted at specific locations outside the production cells, and finally form mature virus-like protein particles that can package the target nucleic acid or target protein.
[0091] Secondly, the virus-like protein particles provided by the present invention have a relatively small molecular weight, which is approximately 45%-75% of the molecular weight of conventional virus-like particles, and can package target nucleic acids or target proteins with larger molecular weights; moreover, the skeleton fusion protein with a smaller molecular weight can effectively improve the packaging titer of the virus-like protein particles; in addition, the virus-like protein particles do not contain the genetic material with replication function in the virus, which greatly improves safety.
[0092] In addition, the structural proteins of the virus-like protein particles provided by the present invention can be translated in the production cells and transported to the cell surface. The proteins can autonomously assemble to form virus-like protein particles that package and carry proteins or long-chain RNA (can carry CRISPR / Cas, base editors, epigenetic editors, antibodies, tumor antigens and viral antigens, cell reprogramming genes and chimeric antigen receptors and other proteins or mRNA), which can be used to achieve gene editing and gene therapy, and can also be used to prepare vaccines, and can also be used to produce pluripotent stem cells and transform cell functions, greatly improving its application value.
[0093] References: ①, Oscorbin IP, Filipenko ML.M-MuLV reverse transcriptase:Selected properties and improved mutants.Comput Struct Biotechnol J.2021;19:6315-6327.Published 2021 Nov 22.doi:10.1016 / j.csbj.2021.11.030;
[0094] ②、CotéML,Roth MJ.Murine leukemia virus reverse transcriptase:structural comparison with HIV-1 reverse transcriptase.Virus Res.2008 Jun;134(1-2):186-202.doi:10.1016 / j.virusres.2008.01.001.Epub 2008 Feb 21. PMID: 18294720; PMCID: PMC2443788. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1. Vectors used to prepare virus-like protein particles packaging the ABE8e editor and SpCas9.
[0096] Figure 2. Editing and sequencing diagram of the EGFP target after delivering the ABE8e editor to the GFP stably transfected strain using different virus-like protein particles.
[0097] Figure 3. Comparison of A5 editing efficiency of the EGFP target after delivery of the ABE8e editor to a GFP stably transfected strain using different virus-like protein particles.
[0098] Figure 4. EGFP fluorescence and TIDE analysis of the EGFP target in the cell line after SpCas9 was delivered to the GFP stably transfected cell line using different virus-like protein particles. DETAILED DESCRIPTION
[0099] Sequence Listing
[0100] Example 1. Establishment of a 293T cell line stably expressing GFP, comprising the following steps:
[0101] 1) Use commercially available conventional lentivirus (LV) plasmid to construct a recombinant plasmid containing the GFP gene.
[0102] 2) Use a transfection reagent to transform the recombinant plasmid and the lentiviral packaging plasmid into HEK-293T cells.
[0103] 3) Collect the virus supernatant from the above step, filter and collect the filtrate containing the virus, and then infect HEK-293T cells.
[0104] 4) The virus-infected HEK-293T cells were digested and inoculated into new culture dishes. The corresponding antibiotics were added for screening. Positive cells were screened until a cell cluster appeared, forming a cell pool that stably expressed GFP.
[0105] 5) After digesting the cell pool, inoculate it into a cell well plate and continue screening with antibiotics. When the single clone in the well grows to a certain degree of confluence, pick the single clone and expand it into a new cell well plate to finally obtain a GFP stable transfectant.
[0106] Example 2. Construction of plasmids encoding different proteins for preparing virus-like protein particles, comprising the following steps:
[0107] As shown in Figure 1 , using genetic engineering technology, the VSV-G coding sequence was constructed downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid A encoding VSV-G;
[0108] As shown in Figure 1, using genetic engineering technology, MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA and nucleocapsid protein NC), connecting peptide (SEQ ID NO.31), 3×NES (SEQ ID NO.24), heterologous protease cleavage peptide (SEQ ID NO.15), NLS (SEQ ID NO.38), and a fusion protein (SEQ ID NO.37) of ABE8e and NLS (SEQ ID NO.38) were constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid C1 encoding MLV gag-3×NES-heterologous protease cleavage peptide-NLS-ABE8e-NLS fusion protein;
[0109] As shown in Figure 1, using genetic engineering technology, the MLV gag expression sequence (comprising matrix protein MA, inner shell protein p12, capsid protein CA and nucleocapsid protein NC), connecting peptide (SEQ ID NO.31), 3×NES (SEQ ID NO.24), heterologous protease cleavage short peptide (SEQ ID NO.15), NLS (SEQ ID NO.38), SpCas9 and nucleoplasmin NLS (SEQ ID NO.39) fusion protein (SEQ ID NO.36) were constructed from N-terminus to C-terminus downstream of the CMV promoter of the eukaryotic expression vector to obtain plasmid C2 encoding MLV gag-3×NES-heterologous protease cleavage short peptide-NLS-SpCas9-NLS fusion protein;
[0110] As shown in Figure 1, using genetic engineering technology, a fusion protein of MLV gag nucleotides (comprising matrix protein MA, capsid protein CA, and nucleocapsid protein NC) and proteolytic enzyme PR (SEQ ID NO. 41) was constructed from the N-terminus to the C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D1 encoding the MLV gag-PR fusion protein;
[0111] As shown in FIG1 , using genetic engineering technology, the MLV gag-pol nucleotide, which is derived from the Moloney murine leukemia virus genome and contains a fusion protein (SEQ ID NO. 42) of matrix protein MA, capsid protein CA, nucleocapsid protein NC, proteolytic enzyme PR, full-length reverse transcriptase RT, and integrase IN, was constructed into the downstream of the CMV promoter of a eukaryotic expression vector from N-terminus to C-terminus to obtain plasmid D2 encoding the MLV gag-pol fusion protein;
[0112] As shown in FIG1 , using genetic engineering technology, a fusion protein (SEQ ID NO. 43) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage short peptide (SEQ ID NO. 18), and reverse transcriptase RT full-length (SEQ ID NO. 5) was constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D3 encoding the MLV gag-PR-reverse transcriptase RT full-length fusion protein;
[0113] As shown in Figure 1, using genetic engineering technology, the fusion protein (SEQ ID NO. 44) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage short peptide (SEQ ID NO. 18), and RT trucapep_01 (SEQ ID NO. 7) was constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D4 encoding the MLV gag-PR-RT trucapep_01 fusion protein, wherein RT trucapep_01 contains an incomplete LRM;
[0114] As shown in Figure 1, using genetic engineering technology, a fusion protein (SEQ ID NO. 45) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage peptide (SEQ ID NO. 18), and RT trucapep_02 (SEQ ID NO. 8) was constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D5 encoding the MLV gag-PR-heterologous protease cleavage peptide-RT trucapep_02 fusion protein, wherein RT trucapep_02 contains an incomplete LRM;
[0115] As shown in Figure 1, using genetic engineering technology, a fusion protein (SEQ ID NO. 46) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage peptide (SEQ ID NO. 18), and RT trucapep_03 (SEQ ID NO. 9) was constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D6 encoding the MLV gag-PR-heterologous protease cleavage peptide-RT trucapep_03 fusion protein, wherein RT trucapep_03 contains a complete LRM;
[0116] As shown in FIG1 , using genetic engineering technology, a fusion protein (SEQ ID NO. 47) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage peptide (SEQ ID NO. 18), and RT trucapep_04 (SEQ ID NO. 10) was constructed from N-terminus to C-terminus and inserted downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D7 encoding the MLV gag-PR-heterologous protease cleavage peptide-RT trucapep_04 fusion protein, wherein RT trucapep_04 contains at least one complete LRM;
[0117] As shown in Figure 1, using genetic engineering technology, a fusion protein (SEQ ID NO. 48) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage peptide (SEQ ID NO. 18), and RT trucapep_05 (SEQ ID NO. 11) was constructed from N-terminus to C-terminus and inserted downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D8 encoding the MLV gag-PR-heterologous protease cleavage peptide-RT trucapep_05 fusion protein, wherein RT trucapep_05 contains an incomplete LRM;
[0118] As shown in Figure 1, using genetic engineering technology, a fusion protein (SEQ ID NO. 49) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage peptide (SEQ ID NO. 18), and RT trucapep_06 (SEQ ID NO. 12) was constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D9 encoding the MLV gag-PR-heterologous protease cleavage peptide-RT trucapep_06 fusion protein, wherein RT trucapep_06 contains an incomplete LRM;
[0119] As shown in Figure 1, using genetic engineering technology, a fusion protein (SEQ ID NO. 50) of MLV gag nucleotides (comprising matrix protein MA, inner shell protein p12, capsid protein CA, and nucleocapsid protein NC), proteolytic enzyme PR, heterologous protease cleavage peptide (SEQ ID NO. 18), and RT trucapep_07 (SEQ ID NO. 13) was constructed from N-terminus to C-terminus downstream of the CMV promoter of a eukaryotic expression vector to obtain plasmid D10 encoding the MLV gag-PR-heterologous protease cleavage peptide-RT trucapep_07 fusion protein, wherein RT trucapep_07 contains an incomplete LRM;
[0120] As shown in Figure 1, genetic engineering technology was used to construct the guide RNA (gRNA was formed by covalently linking EGFP Target gRNA (SEQ ID NO.52) and SpCas9 gRNA scaffold (SEQ ID NO.53) in sequence, SEQ ID NO.51) downstream of the U6 promoter of the eukaryotic expression vector to obtain plasmid E encoding sgRNA.
[0121] Example 3. Using the virus-like protein particles of the present invention to deliver the ABE8e editor in a GFP-stable transgenic strain, comprising the following steps:
[0122] Wild-type adherent HEK293T cells were used as producer cells and co-transfected with different plasmid combinations to obtain corresponding virus-like protein particles (VLPs). The VLPs were packaged with ABE8e and its gRNA targeting GFP. GFP-stable transfected cells were grouped and infected with different VLPs, specifically:
[0123] Control 1: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D1, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then infected with a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2A shows the sequencing results of the EGFP target site, which show that A at positions 4, 5, and 10 of the EGFP gene were edited to G. The editing efficiency at position 5 was the highest, reaching 20%, while the editing efficiencies at positions 4 and 10 were lower, at 4% and 3%, respectively. These results indicate that the presence of the proteolytic enzyme PR alone is insufficient to process the precursors of the matrix protein MA, capsid protein CA, and nucleocapsid protein NC, resulting in only a small amount of mature VLPs.
[0124] Control 2: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D2, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then infected with a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2B shows the sequencing results of the EGFP target site, which showed that A at positions 4, 5, and 10 of the EGFP gene were edited to G. The editing efficiency at position 5 was the highest, reaching 67%, while the editing efficiencies at positions 4 and 10 were lower, at 26% and 24%, respectively. Plasmid C1 contains MLV gag-pol, which results in an excessively large pol protein in the VLPs, reducing the packaging titer. Furthermore, the pol protein contains the integrase IN, which allows the VLPs to integrate nucleic acids, posing a serious safety concern.
[0125] Control 3: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D3, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for ABE8e editing efficiency, as shown in Figure 2. Figure 2C shows the sequencing results of the EGFP target site, which show that A at positions 4, 5, and 10 of the EGFP gene were edited to G. The editing efficiency at position 5 was the highest, reaching 69%, while the editing efficiencies at positions 4 and 10 were lower, at 29% and 22%, respectively.
[0126] Experimental Group 1: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D4, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2D shows the sequencing results of the EGFP target site, which show that A at positions 4, 5, and 10 of the EGFP gene were edited to G. The editing efficiency at position 5 was the highest, reaching 41%, while the editing efficiencies at positions 4 and 10 were lower, at 11% and 10%, respectively.
[0127] Experimental Group 2: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D5, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2E shows the sequencing results of the EGFP target site, which show that the fourth, fifth, and tenth A positions in the EGFP gene were edited to Gs. The fifth position had the highest editing efficiency, reaching 32%, while the fourth and tenth positions had lower editing efficiencies, both at 7%.
[0128] Experimental Group 3: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D6, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2F shows the sequencing results of the EGFP target site, which show that the fourth, fifth, and tenth A positions in the EGFP gene were edited to Gs. The fifth position had the highest editing efficiency, reaching 47%, while the fourth and tenth positions had lower editing efficiencies, at 15% and 12%, respectively.
[0129] Experimental Group 4: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D7, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2G shows the sequencing results of the EGFP target site, which show that the A at positions 4, 5, and 10 of the EGFP gene were edited to G. The editing efficiency at position 5 was the highest, reaching 77%, while the editing efficiencies at positions 4 and 10 were lower, at 32% and 28%, respectively.
[0130] Experimental Group 5: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D8, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. The VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2H shows the sequencing results of the EGFP target site, which show that the fourth, fifth, and tenth A positions in the EGFP gene were edited to Gs. The fifth position had the highest editing efficiency, reaching 21%, while the fourth and tenth positions had lower editing efficiencies, at % and 2%, respectively.
[0131] Experimental Group 6: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D9, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e. The results are shown in Figure 2. Figure 2I shows the sequencing results of the EGFP target site, which show that the A at positions 4, 5, and 10 of the EGFP gene were edited to G. The editing efficiency at position 5 was the highest, reaching 29%, while the editing efficiency at positions 4 and 10 was lower, both at 8%.
[0132] Experimental Group 7: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C1, plasmid D10, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. The VLPs were then used to infect a GFP-stable strain. After 6 days of culture, the EGFP target site was amplified by PCR and analyzed by Sanger sequencing for the editing efficiency of ABE8e (Figure 2). Figure 2J shows the sequencing results of the EGFP target site, which show that the fourth, fifth, and tenth A positions in the EGFP gene were edited to Gs. The fifth position had the highest editing efficiency, reaching 22%, while the fourth and tenth positions had lower editing efficiencies, at 5% and 4%, respectively.
[0133] Based on Figure 2, the editing efficiency of the fifth position A5 in control groups 1-3 and experimental groups 1-7 was statistically analyzed. The results are shown in Figure 3. The results show that RT trucapep_01 to RT trucapep_07 contain 1-2 complete or incomplete protein elements LRM that can form dimers. 1-2 complete or incomplete LRMs can interact to a certain extent to trigger PR activation by promoting Gag-pol dimerization, especially RT trucapep_01, RT trucapep_03, RT trucapep_04 and RT trucapep_06. These four elements can mediate the processing and maturation of Gag fusion protein precursors through the activated proteolytic enzyme PR. These four elements cause differences in the enzymatic activity of the proteolytic enzyme PR, resulting in abnormal VLPs maturation and affecting the VLPs delivery efficiency. This shows that the complete hydrophobic amino acid dimerization motif plays a key role in Gag-pol dimerization.
[0134] Example 4. Using the virus-like protein particles of the present invention to deliver SpCas9 in a GFP-stably transfected strain, comprising the following steps:
[0135] Wild-type adherent HEK293T cells were used as producer cells and co-transfected with different plasmid combinations to obtain corresponding virus-like protein particles (VLPs). These VLPs were packaged with SpCas9 and its gRNA targeting GFP. GFP-stable transfected cells were grouped and infected with different VLPs, specifically:
[0136] Blank control: The GFP stably transfected strain was not infected. After 48 hours of culture, GFP fluorescence was analyzed by flow cytometry, and the EGFP target site was amplified by PCR and the editing efficiency of SpCas9 was analyzed by Sanger sequencing. The results are shown in Figures 4A and 4G. In Figure 4A, 99.3% of the GFP stably transfected strains expressed EGFP (the green box in Figure 4A indicates the percentage of cells that did not express EGFP fluorescence), and Figure 4G shows that the EGFP target site was not knocked out.
[0137] Control 1: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C2, plasmid D1, and plasmid E. The ratio of plasmids A, C2, D1, and E was 1:1:1:1. After 48 hours, the VLPs supernatant was harvested. A 293T cell line expressing stable GFP was infected with VLPs. After 6 days of culture, EGFP fluorescence was analyzed by flow cytometry. The EGFP target site was amplified by PCR and sequenced. The results are shown in Figures 4B (the green box in Figure 4B indicates the percentage of cells that do not express EGFP fluorescence) and 4H. In Figure 4B, plasmids A, C2, and D1 were translated into corresponding proteins in approximately 41.9% of the cells. These proteins autonomously assembled into VLPs and packaged the SpCas9 protein inside the VLPs. Plasmid E was transcribed into gRNA, which recognized the target sequence of EGFP and guided the SpCas9 nuclease to effectively cut the EGFP gene. The EGFP gene was knocked out by SpCas9, resulting in these cells not expressing EGFP fluorescence. Figure 4H is TIDE analysis of the target sequencing results. TIDE showed that the EGFP gene knockout efficiency was only 28.7%. These results indicate that a single proteolytic enzyme PR cannot process the precursors of the matrix protein MA, capsid protein CA, and nucleocapsid protein NC alone, and therefore can only form a small number of mature VLPs, requiring the assistance of other auxiliary proteins.
[0138] Control 2: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C2, plasmid D2, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLPs supernatant was harvested. A 293T cell line expressing stable GFP was infected with the VLPs. After 6 days of culture, EGFP fluorescence was analyzed by flow cytometry, and the EGFP target site was amplified by PCR and sequenced. The results are shown in Figures 4C (the green box in Figure 4C indicates the percentage of cells that do not express EGFP fluorescence) and 4I. In Figure 4C, plasmids A, C2, and D2 were translated into corresponding proteins in approximately 99.56% of the cells. These proteins autonomously assembled into VLPs and packaged the SpCas9 protein inside the VLPs. Plasmid E was transcribed into sgRNA, which recognized the target sequence of EGFP and guided the SpCas9 nuclease to effectively cut the EGFP gene. The EGFP gene was knocked out by SpCas9, resulting in these cells not expressing EGFP fluorescence; Figure 4I is TIDE analysis of the target sequencing results. TIDE showed that the EGFP gene knockout efficiency was 90.8%. These results indicate that the MLV gag-pol full protein (including the proteolytic enzyme PR, the full-length reverse transcriptase RT, and the integrase IN) interacts with the p14 PR, effectively processes the precursors of MA, CA, and NC, matures them to form VLPs, and knocks out EGFP in the GFP stably transfected strain. However, the molecular weight of the pol protein is too large, which will limit the molecular size of the packaged target nucleic acid and / or target protein and reduce the packaging titer of VLPs. Moreover, the pol protein contains integrase IN, which makes the VLPs capable of nucleic acid integration, posing serious safety issues.
[0139] Control 3: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C2, plasmid D3, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLPs supernatant was harvested. A 293T cell line expressing stable GFP was infected with the VLPs. After 6 days of culture, EGFP fluorescence was analyzed by flow cytometry, and the EGFP target site was amplified by PCR and sequenced. The results are shown in Figures 4D (the green box in Figure 4D indicates the percentage of cells that do not express EGFP fluorescence) and 4J. In Figure 4D, plasmids A, C2, and D3 were translated into corresponding proteins in approximately 98.46% of the cells. These proteins autonomously assembled into VLPs and packaged the SpCas9 protein inside the VLPs. Plasmid E was transcribed into gRNA, which recognized the target sequence of EGFP and guided the SpCas9 nuclease to effectively cut the EGFP gene. The EGFP gene was knocked out by SpCas9, resulting in these cells not expressing GFP fluorescence. Figure 4J is TIDE analysis of the target sequencing results. TIDE showed that the EGFP gene knockout efficiency was 84.8%. These results indicate that the proteolytic enzyme PR can interact with the full-length reverse transcriptase RT, process the precursors of MA, CA, and NC, mature them to form VLPs, and edit EGFP in the GFP stably transfected strain.
[0140] Experimental Group 1: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C2, plasmid D4, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLPs supernatant was harvested. A 293T cell line expressing stable GFP was infected with the VLPs. After 6 days of culture, EGFP fluorescence was analyzed by flow cytometry, and the EGFP target site was amplified by PCR and sequenced. The results are shown in Figures 4E (the green box in Figure 4E indicates the percentage of cells that do not express EGFP fluorescence) and 4K. In Figure 4E, plasmids A, C2, and D4 were translated into corresponding proteins in approximately 89.91% of the cells. These proteins autonomously assembled into VLPs and packaged the SpCas9 protein inside the VLPs. Plasmid E was transcribed into gRNA, which recognized the target sequence of EGFP and guided the SpCas9 nuclease to effectively cut the EGFP gene. The EGFP gene was knocked out by SpCas9, resulting in these cells not expressing GFP fluorescence. Figure 4K is TIDE analysis of the target sequencing results. TIDE showed that the EGFP gene knockout efficiency was 66.2%. These results indicate that the proteolytic enzyme PR can interact with RT trucapep_01, process the precursors of MA, CA, and NC, mature them to form VLPs, and knock out EGFP in the GFP stably transfected strain.
[0141] Experimental Group 2: Producer cells were co-transfected with four plasmids: plasmid A, plasmid C2, plasmid D7, and plasmid E, at a ratio of 1:1:1:1. After 48 hours, the VLP supernatant was harvested. A 293T cell line expressing stable GFP was infected with the VLPs. After 6 days of culture, EGFP fluorescence was analyzed by flow cytometry, and the EGFP target site was amplified by PCR and sequenced. The results are shown in Figures 4F (the green box in Figure 4F indicates the percentage of cells that do not express EGFP fluorescence) and 4L. In Figure 4F, plasmids A, C2, and D7 were translated into corresponding proteins in approximately 99.74% of the cells. These proteins autonomously assembled into VLPs and packaged the SpCas9 protein inside the VLPs. Plasmid E was transcribed into gRNA, which recognized the target sequence of EGFP and guided the SpCas9 nuclease to effectively cut the EGFP gene. The EGFP gene was knocked out by SpCas9, resulting in these cells not expressing GFP fluorescence. Figure 4L is TIDE analysis of the target sequencing results. TIDE showed that the EGFP gene knockout efficiency was 94%. These results indicate that the proteolytic enzyme PR can interact with RT trucapep_04, process the precursors of MA, CA, and NC, mature them to form a large number of VLPs, and knock out EGFP in the GFP stably transfected strain.
[0142] The results of Examples 3 and 4 show that the proteolytic enzyme PR can interact with the dimerization motif (trucapep_01, RT trucapep_03, RT trucapep_04 or RT trucapep_06) in the reverse transcriptase RT, activate the proteolytic enzyme PR, and mediate the precursor processing and maturation of MA, CA and NC through the activated proteolytic enzyme PR. These fusion proteins of the present invention form virus-like particles packaged with target proteins or target nucleic acids in host cells, deliver Cas nucleases and base editors in the host cells, and correctly knock out and base edit the target sequences in the cells.
[0143] In summary, the present invention can deliver CRISPR / Cas nucleases or mRNA and gRNA for gene editing and gene therapy; or deliver tumor or viral antigen mRNA or protein for immunotherapy; or deliver cell reprogramming factor mRNA or protein to produce pluripotent stem cells and transform cell function; or deliver chimeric antigen receptor mRNA or protein for cellular immunotherapy.
Claims
1. A viroid - like protein particle, comprising: an envelope protein ENV, a matrix protein MA, a capsid protein CA, a nucleocapsid protein NC, a protease PR, and a protein element capable of forming a dimer; the matrix protein MA, the capsid protein CA, the nucleocapsid protein NC, the protease PR, the protein element capable of forming a dimer self - assemble with a target nucleic acid or a target protein to form a protein particle containing the target nucleic acid or the target protein, and the envelope protein ENV packages the protein particle.
2. The viroid - like protein particle according to claim 1, wherein the protein element capable of forming a dimer comprises a homodimeric domain or / and a heterodimeric domain; preferably, the protein element capable of forming a dimer is a dimerization motif on a homologous or heterologous reverse transcriptase or a dimerization region on a dimer protein; preferably, the protein element capable of forming a dimer comprises a dimerization motif in reverse transcriptase RT or its enzymatically inactivated mutant; preferably, the protein element capable of forming a dimer comprises a hydrophobic amino acid dimerization motif; preferably, the hydrophobic amino acid dimerization motif comprises a leucine zipper (LZ) dimerization motif, a leucine repeat motif (LRM), a leucine - rich repeat motif (LRR), or / and a tryptophan repeat motif (TRM).
3. The viroid - like protein particle according to claim 1, wherein the protein element capable of forming a dimer, (1) comprises an amino acid sequence having at least 95% sequence identity compared to any one of the amino acid sequences shown in SEQ ID NOs. 7 to 13; or (2) comprises an amino acid sequence having at least 95% sequence identity compared to any one of the amino acid sequences shown in SEQ ID NOs. 7 to 10; and, according to the sequence numbering shown in SEQ ID NO. 10, has an amino acid substitution at at least one of the three positions of Y63, D113, and R115, preferably substituted with alanine; or, according to the sequence numbering shown in SEQ ID NO. 10, has an amino acid substitution at the position of Y221, preferably substituted with alanine or serine.
4. The viroid - like protein particle according to claim 1, wherein the target nucleic acid or the target protein is a therapeutic RNA or a therapeutic protein; preferably, the therapeutic RNA or the therapeutic protein comprises a CRISPR / Cas nuclease, a base editor, an epigenetic editor, a recombinase, a transcription factor, a reverse transcriptase, a prime editor (PE editor), an antibody, and other functional proteins.
5. A nucleic acid, comprising encoding: (i) a first polynucleotide, which comprises a nucleic acid sequence encoding a scaffold fusion protein; and (ii) a second polynucleotide, which comprises a nucleic acid sequence encoding a Gag fusion protein and a nucleic acid sequence fused with a target nucleic acid or a target protein; and (iii) a third polynucleotide, which comprises a nucleic acid sequence encoding an envelope protein ENV; Among them, the scaffold fusion protein comprises: (a) a Gag fusion protein formed by fusing a matrix protein MA, a capsid protein CA, and a nucleocapsid protein NC; and (b) A proteolytic enzyme PR, which is covalently linked to the Gag fusion protein through a heterologous protease-cleavable short peptide; and (c) A protein element capable of forming a dimer, which is covalently linked to the proteolytic enzyme PR through a heterologous protease-cleavable short peptide; the protein element capable of forming a dimer, (1) comprises an amino acid sequence having at least 95% sequence identity compared to the amino acid sequence shown in any one of SEQ ID NOs. 7 to 13; or (2) comprises an amino acid sequence having at least 95% sequence identity compared to the amino acid sequence shown in any one of SEQ ID NOs. 7 to 10; and, according to the sequence numbering shown in SEQ ID NO. 10, has an amino acid substitution at at least one of the three positions of Y63, D113, and R115, preferably substituted with alanine; or, according to the sequence numbering shown in SEQ ID NO. 10, has an amino acid substitution at the position of Y221, preferably substituted with alanine or serine; Wherein, the target nucleic acid or target protein is covalently linked to the Gag fusion protein through a heterologous protease-cleavable short peptide.
6. The nucleic acid according to claim 5, wherein the second polynucleotide further comprises one or more heterologous polypeptides, and the target nucleic acid or target protein is covalently linked to the Gag fusion protein through the one or more heterologous polypeptides and the heterologous protease-cleavable short peptide; The one or more heterologous polypeptides are independently a nuclear export signal, an epitope tag, a reporter gene sequence, an enzyme for a detectable signal, and / or a subcellular localization sequence.
7. The nucleic acid according to claim 5, wherein the first polynucleotide comprises an amino acid sequence having at least 95% sequence identity compared to the amino acid sequence shown in any one of SEQ ID NOs. 44 to 50; the second polynucleotide comprises an amino acid sequence having at least 95% sequence identity compared to the amino acid sequence shown in SEQ ID NO. 36 or 37.
8. The nucleic acid according to claim 5, wherein the target nucleic acid or target protein comprises a CRISPR / Cas nuclease, and the nucleic acid further comprises (iv) a fourth polynucleotide, which comprises a nucleic acid sequence encoding a guide RNA (gRNA), wherein the gRNA binds to the CRISPR / Ca nuclease.
9. A viral vector system comprising the nucleic acid according to claim 5 or 6; wherein, The first polynucleotide, the second polynucleotide, and the third polynucleotide are each located on a separate vector; or one or more of the first polynucleotide, the second polynucleotide, and the third polynucleotide are located on the same vector; or It comprises the nucleic acid according to claim 7; wherein, the first polynucleotide, the second polynucleotide, the third polynucleotide, and the fourth polynucleotide are each located on a separate vector; or one or more of the first polynucleotide, the second polynucleotide, the third polynucleotide, and the fourth polynucleotide are located on the same vector.
10. A cell, which comprises the viroid-like protein particle according to any one of claims 1 to 4, the nucleic acid according to any one of claims 5 to 7, and the viral vector system according to claim 8; preferably, the cell is a eukaryotic cell; more preferably, the cell is a human cell.
11. A pharmaceutical composition, which comprises the viroid-like protein particle according to any one of claims 1 to 4.
12. A method for preparing a viroid-like protein particle, the preparation method comprising: (a) introducing the nucleic acid according to any one of claims 5 to 7, or the viral vector system according to claim 8 into a production cell; and (b) harvesting the supernatant containing the viroid-like protein particle produced by the packaging cell, and concentrating to obtain the viroid-like protein particle.
Citation Information
Patent Citations
Particle for the encapsidation of a genome engineering system
CN109415415A
Virus-like protein particles for delivery of nucleic acids and proteins and methods of making same
CN118126137A
Virus-like particles and use thereof
US20210269790A1
Compositions and methods for delivering crispr / cas effector polypeptides
US20230193255A1
Self-assembling virus-like particles for delivery of nucleic acid programmable fusion proteins and methods of making and using same
WO2023102537A2