Expression vectors, bacterial sequence free vectors and methods for making and using them
Patent Information
- Application Number
- JP2023577551
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-15
- Filing Date
- 2022-06-16
- Publication Date
- 2025-06-24
AI Technical Summary
Current gene therapy delivery systems, such as viral vectors, face challenges including inefficiencies in gene transfer and transgene expression, immune responses, and high production and storage costs, while non-viral vectors have lower effectiveness and persistence issues.
Development of expression vectors with specific backbone sequences, recombinase target sites, endonuclease targets, and additional elements like CMV enhancers, 5'UTRs, chromatin insulators, and WPREs to enhance transgene expression and reduce bacterial sequence contamination.
The improved vectors achieve higher transgene expression levels and persistence, reducing the risk of immune responses and production costs, and eliminate bacterial sequences, ensuring safer and more efficient gene delivery.
Smart Images

Figure 00000116_0000 
Figure 00000116_0001 
Figure 00000117_0000
Abstract
Description
[Technical Field]
[0001] Reference to sequence listings submitted electronically via EFS-WEB The contents of the electronically submitted sequence listing (Name: 4471_007PC03_Seqlisting_ST25.txt; Size: 216,898 bytes; and Creation Date: June 15, 2022) are incorporated by reference in their entirety into the present specification.
[0002] FIELD OF THE INVENTION The present invention provides expression vectors, bacterial sequence-free vectors, vector production systems for producing bacterial sequence-free vectors and uses thereof. [Background technology]
[0003] background Gene therapy holds significant therapeutic promise, but challenges remain in realizing its potential.
[0004] Most clinical trials utilize viral delivery systems, such as adenoviral vectors, lentiviral vectors, and adeno-associated viral vectors. Despite progress, viral systems have variable gene transfer and transgene expression efficiencies, and concerns remain regarding undesirable effects such as inflammatory and immune responses or insertional mutagenesis. Furthermore, the production, purification, and storage of viral vectors are often costly, variable, and inefficient. See, for example, Lingelbach, D., Drug Development & Delivery 20(5): 50-54 (2020); Wright, JF, Gene Therapy 15:840-848 (2008).
[0005] Non-viral vectors have also been studied as gene therapy delivery systems.Non-viral vectors are safer than viral vectors, but their effectiveness may be limited due to reasons such as low transgene expression level and persistence of expression.See, for example, Kay, M., Nature Reviews Genetics 12: 316-328 (2011).
[0006] There is a need for improved vectors such as those described herein. Summary of the Invention
[0007] overview The present invention provides a nucleic acid sequence comprising: (a) a backbone sequence; (b) a sequence comprising: (i) an expression cassette comprising a nucleic acid sequence of interest; (ii) a first target sequence for a first recombinase adjacent to the 5' end of the expression cassette; (iii) a second target sequence for the first recombinase adjacent to the 3' end of the expression cassette; and (iv) one or more additional target sequences for one or more additional recombinases integrated within the non-binding regions of the first and second target sequences for the first recombinase; and (c) (i) an endonuclease target sequence integrated within the first and / or second target sequence for the first recombinase in the non-binding regions for the first recombinase and the one or more additional recombinases, the endonuclease target sequence being between the backbone sequence and the cleavage sites for the first recombinase and the one or more additional recombinases; (iii) a cytomegalovirus (CMV) enhancer integrated between the 3' end of a first target sequence for a first recombinase and the 5' end of a promoter in an expression cassette; (iv) a 5' untranslated region (5'UTR) comprising an intron, wherein the 5'UTR is integrated between the promoter and a nucleic acid sequence of interest in an expression cassette; (v) a vertebrate chromatin insulator integrated into the expression cassette between the nucleic acid of interest and a polyadenylation signal; (vi) a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) integrated into the expression cassette between the nucleic acid of interest and a polyadenylation signal; (vii) a scaffold / matrix attachment region (S / MAR) integrated into the expression cassette between the nucleic acid of interest and a polyadenylation signal;or (viii) an expression vector comprising one or more of the following: a DNA nuclear targeting sequence (DTS) integrated within the first and / or second target sequence for the first recombinase in a non-binding region for the first recombinase and one or more additional recombinases, the DTS being between the expression cassette and the cleavage sites for the first recombinase and one or more additional recombinases;
[0008] In one aspect, the expression vector comprises an endonuclease target sequence integrated within the first and / or second target sequence for the first recombinase in the non-binding region for the first recombinase and one or more additional recombinases, wherein the endonuclease target sequence is between the backbone sequence and the cleavage sites for the first recombinase and one or more additional recombinases. In one aspect, the endonuclease target sequence is integrated within the first and second target sequences for the first recombinase. In one aspect, the endonuclease target sequence is for a homing endonuclease. In certain aspects, the endonuclease target sequence is I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H -DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI, I-PpoI, PI-PspI, I- ScaI, I-SceI, PI-SceI, I-SceII, I-SecIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I or I-Vdi141I. In some aspects, the endonuclease target sequence is for I-SceI. In some aspects, the endonuclease target sequence is for PI-SceI. In some aspects, the endonuclease target sequence is for a Cas endonuclease. In some aspects, the Cas endonuclease is Cas9.
[0009] In one aspect, the expression vector includes a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 integrated between the 3' end of a first target sequence for a first recombinase and the 5' end of another enhancer or promoter in the expression cassette. In one aspect, the synthetic enhancer comprises multiple contiguous copies of a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12. In one aspect, the synthetic enhancer comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO: 46. In one aspect, the synthetic enhancer is integrated at the 5' end of a chicken β-actin promoter. In one aspect, a chimeric intron comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 47 is integrated at the 3' end of the chicken β-actin promoter and the 5' end of the nucleic acid sequence of interest.
[0010] In one aspect, the expression vector comprises a CMV enhancer integrated between the 3' end of the first target sequence for the first recombinase in the expression cassette and the 5' end of the promoter. In one aspect, the CMV enhancer is integrated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 or SEQ ID NO: 46. In one aspect, a CMV promoter is integrated at the 3' end of the CMV enhancer and the 5' end of the nucleic acid sequence of interest.
[0011] In one aspect, the expression vector comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39 integrated between a first target sequence for a first recombinase and a nucleic acid sequence of interest.
[0012] In one aspect, the expression vector comprises a 5'UTR containing an intron, and the 5'UTR is incorporated in the expression cassette between the promoter and the nucleic acid sequence of interest. In one aspect, the intron comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:1. In one aspect, the 5'UTR further comprises a non-coding sequence incorporated within the intron. In one aspect, the 5'UTR comprises a non-coding sequence incorporated between two nucleotides within the intron corresponding to any two nucleotides from positions 25 to 55 of SEQ ID NO:1. In one aspect, the non-coding sequence is an S / MAR. In one aspect, the S / MAR is MAR-5. In one aspect, the 5'UTR comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:3. In one aspect, the 5'UTR comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:5. In one aspect, the promoter is a chicken β-actin promoter. In one aspect, the promoter is a CMV promoter. In one aspect, the promoter is integrated at the 3' end of a CMV enhancer. In one aspect, the CMV enhancer is integrated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 or SEQ ID NO: 46.
[0013] In one aspect, the expression vector comprises a polyadenylation signal incorporated into the 3' end of the nucleic acid sequence of interest. In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:13, SEQ ID NO:14, or SEQ ID NO:15.
[0014] In one aspect, the expression vector comprises a vertebrate chromatin insulator integrated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal. In one aspect, the vertebrate chromatin insulator is the 5'-HS4 chicken β-globulin insulator (cHS4). In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15.
[0015] In one aspect, the expression vector comprises a WPRE incorporated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal. In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:13, SEQ ID NO:14, or SEQ ID NO:15.
[0016] In one aspect, the expression vector comprises an S / MAR integrated in the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal. In one aspect, the S / MAR is MAR-5. In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15.
[0017] In one aspect, the expression vector comprises an enhancer sequence flanking each side of the first and second target sequences for the first recombinase. In one aspect, the expression vector comprises at least two enhancer sequences flanking each side of the first and second target sequences for the first recombinase. In one aspect, the enhancer sequence is an SV40 enhancer sequence.
[0018] In one aspect, the expression vector comprises a DTS integrated into the first and / or second target sequence for the first recombinase in the non-binding region for the first recombinase and one or more additional recombinases, the DTS being located between the expression cassette and the cleavage sites for the first recombinase and one or more additional recombinases. In one aspect, the DTS is an SV40 enhancer sequence. In one aspect, the DTS is cell-specific.
[0019] In one aspect, the first and second target sequences and one or more additional target sequences are selected from the group consisting of a PY54 pal site, an N15 telRL site, a loxP site, a φK02 telRL site, an FRT site, a phiC31 attP site, and a λ attP site. In one aspect, an expression vector comprises each of the target sequences. In one aspect, an expression vector comprises a pal site, and a telRL recombinase target binding sequence, a loxP recombinase target binding sequence, and an FRT recombinase target binding sequence integrated within the pal site. In one aspect, the first and second target sequences for the first recombinase each comprise the nucleic acid sequence of SEQ ID NO: 33.
[0020] In some aspects, the expression vector is for producing a vector that does not contain any bacterial sequences. In some aspects, the vector that does not contain any bacterial sequences is a covalently closed circular vector. In some aspects, the vector that does not contain any bacterial sequences is a covalently closed linear vector.
[0021] The present invention relates to a vector production system comprising a recombinant cell encoding a recombinase under the control of an inducible promoter, wherein the recombinant cell comprises any of the expression vectors described above, and the recombinase targets one of first and second target sequences for a first recombinase or one or more additional target sequences for one or more additional recombinases in the expression vector. In one aspect, the recombinase is TelN, Tel, Cre, or Flp.
[0022] In one aspect, the recombinant cell further encodes an endonuclease under the control of an inducible promoter, wherein the endonuclease targets an endonuclease target sequence in an expression vector comprising the endonuclease target sequence. In one aspect, the endonuclease is a homing endonuclease. In certain aspects, the homing endonuclease is I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI, I-PpoI, PI-PspI. , I-ScaI, I-SceI, PI-SceI, I-SceII, I-SecIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I or I-Vdi141I. In some aspects, the endonuclease is I-SceI. In some aspects, the endonuclease is PI-SceI. In some aspects, the recombinant cell encodes a nuclease genome editing system comprising the endonuclease. In some aspects, the nuclease genome editing system is a clustered regularly interspaced short palindromic repeats (CRISPR) nuclease system comprising a guide RNA and a Cas endonuclease. In some aspects, the Cas endonuclease is Cas9. In some aspects, the inducible promoter is thermally regulated, chemically regulated, IPTG regulated, glucose regulated, arabinose induced, T7 polymerase regulated, cold shock induced, pH induced, or a combination thereof.
[0023] The present invention relates to a method for producing a bacterial sequence-free vector, comprising incubating the vector production system described above under conditions suitable for expression of a recombinase. In one aspect, the method further comprises incubating the vector production system described above encoding an endonuclease under conditions suitable for expression of an endonuclease. In one aspect, the method further comprises incubating any of the vector production systems described above under conditions suitable for expression of a nuclease genome editing system. In one aspect, the method further comprises recovering the bacterial sequence-free vector.
[0024] The present invention relates to a vector that does not contain bacterial sequences produced by any of the above methods for producing a vector that does not contain bacterial sequences.
[0025] The present invention relates to a bacterial sequence-free vector comprising (a) an expression cassette comprising a nucleic acid sequence of interest, and (b) one or more of the following: (i) a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 located 5' to another enhancer or promoter in the expression cassette; (ii) a CMV enhancer located 5' to the promoter in the expression cassette; (iii) a 5'UTR comprising an intron, wherein the 5'UTR is incorporated between the promoter and the nucleic acid sequence of interest in the expression cassette; (iv) a vertebrate chromatin insulator incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette; (v) a WPRE incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette; (vi) an S / MAR incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette; or (vii) a DTS located 5' to the expression cassette.
[0026] In one aspect, the bacterial sequence-free vector comprises a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 located 5' of another enhancer or promoter in the expression cassette. In one aspect, the synthetic enhancer comprises multiple contiguous copies of a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12. In one aspect, the synthetic enhancer comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO: 46. In one aspect, the synthetic enhancer is incorporated into the 5' end of a chicken β-actin promoter. In one aspect, a chimeric intron comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 47 is incorporated into the 3' end of the chicken β-actin promoter and the 5' end of the nucleic acid sequence of interest.
[0027] In one aspect, the bacterial sequence-free vector comprises a CMV enhancer located 5' of the promoter in the expression cassette. In one aspect, the CMV enhancer is incorporated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 or SEQ ID NO: 46. In one aspect, the CMV promoter is incorporated at the 3' end of the CMV enhancer and 5' of the nucleic acid sequence of interest.
[0028] In one aspect, the bacterial sequence-free vector comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39 located 5' to the nucleic acid sequence of interest.
[0029] In one aspect, the bacterial sequence-free vector comprises a 5'UTR containing an intron, the 5'UTR being incorporated between the promoter and the nucleic acid sequence of interest in the expression cassette. In one aspect, the intron comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:1. In one aspect, the 5'UTR further comprises a non-coding sequence incorporated within the intron. In one aspect, the 5'UTR further comprises a non-coding sequence incorporated between two nucleotides in the intron corresponding to any two nucleotides from nucleotide positions 25 to 55 of SEQ ID NO:1. In one aspect, the non-coding sequence is an S / MAR. In one aspect, the S / MAR is MAR-5. In one aspect, the 5'UTR comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:3. In one aspect, the 5'UTR comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:5. In one aspect, the promoter is a chicken β-actin promoter. In one aspect, the promoter is a CMV promoter. In one aspect, the promoter is integrated at the 3' end of a CMV enhancer. In one aspect, the CMV enhancer is integrated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 or SEQ ID NO: 46.
[0030] In one aspect, the bacterial sequence-free vector comprises a polyadenylation signal incorporated into the 3' end of the nucleic acid sequence of interest. In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:13, SEQ ID NO:14, or SEQ ID NO:15.
[0031] In one aspect, the bacterial sequence-free vector comprises a vertebrate chromatin insulator integrated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal. In one aspect, the vertebrate chromatin insulator is cHS4. In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15.
[0032] In one aspect, the bacterial sequence-free vector comprises a WPRE integrated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal. In one aspect, the polyadenylation signal comprises a nucleic acid sequence at least about 90% identical to SEQ ID NO:13, SEQ ID NO:14, or SEQ ID NO:15.
[0033] In one aspect, the bacterial sequence-free vector comprises an S / MAR incorporated in the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal. In one aspect, the S / MAR is MAR-5.
[0034] In some aspects, the vector without bacterial sequences comprises enhancer sequences flanking both sides of the expression cassette. In some aspects, the vector without bacterial sequences comprises at least two enhancer sequences flanking both sides of the expression cassette. In some aspects, the enhancer sequences are SV40 enhancer sequences.
[0035] In some aspects, the bacterial sequence-free vector comprises a DTS located 5' to the expression cassette. In some aspects, the DTS is an SV40 enhancer sequence. In some aspects, the DTS is cell-specific.
[0036] In one aspect, the bacterial sequence-free vector is a covalently closed circle vector.
[0037] In one aspect, the bacterial sequence-free vector is a covalently closed linear vector.
[0038] The present invention relates to a recombinant cell comprising any of the above expression vectors or any of the above bacterial sequence-free vectors.
[0039] The present invention relates to a composition comprising any of the above expression vectors or any of the above bacterial sequence-free vectors. In some aspects, the composition further comprises a delivery agent. In some aspects, the delivery agent is a nanoparticle. In some aspects, the delivery agent comprises a targeting ligand. In some aspects, the composition is a pharmaceutical composition further comprising a pharmaceutically acceptable carrier.
[0040] The present invention relates to a method of treating a disease or disorder in a subject in need thereof, comprising administering to the subject any of the above-described expression vectors, any of the above-described bacterial sequence-free vectors, or any of the above-described pharmaceutical compositions.
[0041] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:1.
[0042] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:2.
[0043] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:3.
[0044] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:5.
[0045] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:12.
[0046] The present invention relates to a polynucleotide comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:46.
[0047] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:13.
[0048] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:14.
[0049] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:15.
[0050] In one aspect, any of the above polynucleotides comprising a nucleic acid sequence at least about 90% identical to any of SEQ ID NOs: 13 to 15 further comprises 100 to 120 adenine nucleotides at the 3' end of the nucleic acid sequence.
[0051] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:16.
[0052] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:17.
[0053] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:18.
[0054] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:35.
[0055] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:36.
[0056] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:37.
[0057] The present invention relates to polynucleotides comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:38.
[0058] The present invention relates to a polynucleotide comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:39.
[0059] The present invention relates to an expression vector comprising any of the above polynucleotides.
[0060] The present invention relates to a polynucleotide comprising a nucleic acid sequence at least about 90% identical to any one of SEQ ID NOs: 2, 3, or 5, and an expression vector comprising (i) a polynucleotide comprising a nucleic acid sequence at least about 90% identical to any one of SEQ ID NOs: 13-18, or (ii) a polynucleotide comprising a nucleic acid sequence at least about 90% identical to any one of SEQ ID NOs: 13-15, and 100 to 120 adenine nucleotides at the 3' end of the nucleic acid sequence.
[0061] The present invention relates to a gene editing method comprising inserting a nucleic acid sequence of interest from any of the above expression vectors, any of the above bacterial sequence-free vectors, or any of the above pharmaceutical compositions into a target site for gene editing. In one aspect, the gene editing is by non-homologous end joining. In one aspect, the gene editing is by homology-directed repair. [Brief explanation of the drawings]
[0062] [Figure 1] FIG. 1 shows the vector map of the expression vector pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*. [Figure 2] FIG. 2 shows the vector map of the expression vector pcDNA-CMV-5′UTR-SecNLuc-P2A-eGFP-bGHpA. [Figure 3] Figure 3 shows photographs of HEK-293 cell fluorescence assessed by live imaging. (A) shows negative control cells exposed to Lipofectamine without any plasmid, (B) shows cells transfected with the expression vector shown in Figure 1, (C) shows cells transfected with the expression vector shown in Figure 2, and (D) shows positive control cells transfected with the parent expression vector pGL2-SS*-CAG-eGFP-BGpA-SS* (PP-CAG-GFP), which expresses eGFP under the control of the CAG promoter. [Figure 4]Figure 4 shows bar graphs of the relative fluorescence intensity of cells divided according to Figures 3(A) to (D). In Figures 4 and 5, "pGL2-SecNLuc-eGFP" refers to cells transfected with the pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* expression vector of Figure 1. In Figures 4 and 5, "pcDNA-SecNLuc-eGFP" refers to cells transfected with the pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA expression vector of Figure 2. [Figure 5] Figure 5 shows a bar graph of relative luciferase intensity in the medium of cells transfected according to Figures 3(A)-(C). [Figure 6] FIG. 6 shows a vector map of the expression vector pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS*. [Figure 7] FIG. 7 shows the vector map of the expression vector pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*. [Figure 8] FIG. 8 shows a vector map of the expression vector pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS*. [Figure 9] Figure 9 shows a line graph of long-term luciferase activity, expressed in relative luminometer units (also called relative light units, RLU), in the culture medium of HEK-293 cells 2, 6, 10, 14, 17, 20, 27, and 34 days after electroporation of cells with the expression vectors shown in Figure 1 (pGL2-SecNLuc-eGFP), Figure 6 (WPRE), Figure 7 (5'UTR1+WPRE), and Figure 8 (5'UTR2+WPRE), compared to a negative control (Neg. Ctl. (No Plasmid)) in which cells were electroporated with a puc57 plasmid lacking a mammalian expression cassette. *=p<0.05, **=p<0.01, ***=p<0.001, and ****=p<0.0001. [Figure 10]FIG. 10 shows a bar graph of luciferase activity, expressed as luminescence in RLU, in the culture medium of transfected cells described in FIG. 9 and negative control HEK-293 cells at passages 1, 2, 3 and 5. [Figure 11] FIG. 11 shows a line graph of relative luciferase intensity in HEK-293 cells at passages 1, 2, 3, 4, 5, 6, and 7, corresponding to days 8, 15, 24, 31, 38, 45, and 52, respectively, after electroporation of the cells with the expression vector of FIG. 7 (2nd gen pDNA (CMV+U1+W)), msDNA produced from the expression vector of FIG. 7 (2nd gen msDNA (CMV+U1+W)), or the expression vector of FIG. 2 (conventional pcDNA) containing a luciferase transgene. [Figure 12] Figure 12 shows fluorescence in cells transfected with the expression vectors shown in Figure 1 or Figure 7, as described in Figure 9. (A) Photographs show fluorescence in HEK-293 cells assessed by live imaging at passages 1, 2, 3, and 5. (B) Line graphs show the number of eGFP-positive (GFP+) cells observed within a field of view from three live fluorescent images at each passage; not significant (ns) = p>0.05. (C) Dot plots show the mean fluorescence intensity (MFI) of GFP+ cells measured from three live fluorescent images at passage 5. The bar graph below each construct shows the average MFI value from all GFP+ cells measured. **** = p<0.0001.
[0063] [Figure 13]Figure 13 shows a line graph of RLU / mg protein in plasma collected from wild-type mice on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after a single hydrodynamic tail vein injection of 50 μg of pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, PSNLuc), pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (pCAGLuc), or pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (pGSNLuc-WPRE). [Figure 14] FIG. 14 shows a line graph of RLU / mg protein in plasma collected from wild-type mice on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after a single hydrodynamic tail vein injection of 50 μg of pcDNA-CMV-5′UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, PSNLuc), pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (pCAGLuc), or pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (pCAGLucWPRE). [Figure 15] FIG. 15 shows line graphs of RLU / mg protein in plasma collected from wild-type mice on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after a single hydrodynamic tail vein injection of 5 μg of pcDNA-CMV-5′UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, pDNA CMV-U (no SSeq)), pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (SSeq pDNA CAG), or pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (SSeq pDNA CAG-W). [Figure 16]Figure 16 shows a line graph of RLU / mg protein in plasma collected from wild-type mice on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after a single hydrodynamic tail vein injection of 5 μg of pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, No-SSeq pDNA CMV-U), pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (2XSSeq pDNA CAG-W), or msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (2XSSeq msDNA CAG-W). [Figure 17] Figures 17A-D show bar graphs of GFP expression determined by ELISA in the livers of wild-type mice and negative control mice not injected with vector 56 days after a single hydrodynamic tail vein injection of 5 μg of the vector described in Figure 16. (A) GFP concentration in μg / mL. (B) μg of GFP normalized to μg of total protein. (C) μg of GFP normalized to g of total tissue. (D) GFP expression levels relative to control. *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. [Figure 18] Figure 18 is a bar graph showing cytoplasmic GFP protein concentrations (pg / mL) in the livers of wild-type mice and negative control mice not injected with vector 56 days after a single hydrodynamic tail vein injection of 5 μg of the vector described in Figure 16. [Figure 19]Figure 19 shows the results of the lipid nanoparticle (LNP) carrier (vehicle (control)) or LNP and msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (LNP-2G msDNA-CAG-SecretedNanoLuc), pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (LNP-2G ppDNA-CAG-SecretedNanoLuc), msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (LNP-2G msDNA-CMV-SecretedNanoLuc), pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (LNP-2G Bar graphs of total flux (flux) in photons / second from in vivo whole-body bioluminescence imaging at days 1, 3, 10, 30, 58, 92, 119, and 174 after a single hydrodynamic tail vein injection of mice with lipoplexes of either ppDNA-CMV-SecretedNanoLuc or pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (LNP-Conv.pDNA-CMV-SecretedNanoLuc). The bar graph for LNP-2G msDNA-CAG-SecretedNanoLuc injection is enclosed within a dotted line. [Figure 20] Figure 20 shows photomicrographs of green fluorescent protein (GFP) expression in sagittal brain sections from the cerebral cortex, thalamus, brainstem, and cerebellum of mice injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA. White arrows indicate transgene expression. Nuclei are indicated by staining with diamidino-2-phenylindole (DAPI). [Figure 21] Figure 21 shows photomicrographs of GFP expression in sagittal brain sections from the cerebral cortex and thalamus of mice injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA. Neurons are indicated by the neuronal marker NeuN. [Figure 22]Figure 22 shows photomicrographs of GFP expression in sagittal brain sections from the cerebellum and brainstem of mice injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA. Neurons are indicated by the neuronal marker NeuN. [Figure 23] Figure 23 shows photomicrographs of GFP expression in sagittal brain sections from the cerebral cortex and thalamus of mice injected with msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA. Neurons are indicated by the neuronal marker NeuN. [Figure 24] Figure 24 shows photomicrographs of GFP expression in sagittal brain sections from the cerebellum and brainstem of mice injected with msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA. Neurons are indicated by the neuronal marker NeuN. [Figure 25] Figure 25 shows lipid nanoparticle carriers (LNPs) and pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, LNP-Conv.pDNA-CMV-SecretedNLuc), pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (LNP-ppDNA-CMV-SecretedNLuc), msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (LNP-msDNA-CMV-SecretedNLuc), pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (LNP-ppDNA-CAG-SecretedNLuc), or msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (LNP-msDNA-CAG-SecretedNLuc) 25 shows a bar graph of luminescence associated with luciferase expression in human T cells (Pan-T(TA+) cells, FIG. 25) 3 and 5 days after transfection with lipoplexes containing α- and β-glucan-1, PBS control, or untreated cells. [Figure 26]Figure 26 shows lipid nanoparticle carriers (LNPs) and pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, LNP-Conv.pDNA-CMV-SecretedNLuc), pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (LNP-ppDNA-CMV-SecretedNLuc), msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (LNP-msDNA-CMV-SecretedNLuc), pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (LNP-ppDNA-CAG-SecretedNLuc), or msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (LNP-msDNA-CAG-SecretedNLuc) 26 shows a bar graph of luminescence associated with luciferase expression in human hepatocytes (Huh7 cells, FIG. 26) 3 and 5 days after transfection with lipoplexes containing α- and β-lactams, a PBS control, or untreated cells. [Figure 27] Figures 27A-C show fluorescence-activated cell sorting (FACS) scatter plots of the knock-in (KI) efficiency (Q3) of a gene of interest (GOI) on day 3 (Figures 27A-C) after transfection with a conventional plasmid or msDNA carrying a RISPR gene editing system and a gene of interest (GOI) flanked by 5' and 3' homology arms (HDR-GOI-HDR). Figure 27(A) shows FACS scatter plots of control, wild-type (WT), and induced pluripotent stem cells (iPSCs) without HDR KI of the GOI. Figure 27(B) shows FACS scatter plots of iPSCs after HDR KI of the GOI using a conventional plasmid (plasmid DNA HDR-GOI-HDR). Figure 27(C) shows FACS scatter plots of iPSCs after HDR KI of the GOI using msDNA (msDNA HDR-GOI-HDR). [Figure 28]Figures 28A-C show fluorescence-activated cell sorting (FACS) scatter plots of the knock-in (KI) efficiency (Q3) of a gene of interest (GOI) at day 7 (Figures 28A-C) after transfection with a conventional plasmid or msDNA carrying a RISPR gene editing system and a gene of interest (GOI) flanked by 5' and 3' homology arms (HDR-GOI-HDR). Figure 28(A) shows FACS scatter plots of control, wild-type (WT), and induced pluripotent stem cells (iPSCs) without HDR KI of the GOI. Figure 28(B) shows FACS scatter plots of iPSCs after HDR KI of the GOI using a conventional plasmid (plasmid DNA HDR-GOI-HDR). Figure 28(C) shows FACS scatter plots of iPSCs after HDR KI of the GOI using msDNA (msDNA HDR-GOI-HDR). [Figure 29] Figures 29A-B show fluorescence-activated cell sorting (FACS) scatter plots of the knock-in (KI) efficiency (Q3) of a gene of interest (GOI) at day 15 (Figures 29A-B) after transfection with a RISPR gene editing system and a conventional plasmid or msDNA carrying a gene of interest (GOI) flanked by 5' and 3' homology arms (HDR-GOI-HDR). Figure 29(A) shows a FACS scatter plot of iPSCs after HDR KI of a GOI using a conventional plasmid (plasmid DNA HDR-GOI-HDR). Figure 29(B) shows a FACS scatter plot of iPSCs after HDR KI of a GOI using msDNA (msDNA HDR-GOI-HDR). [Figure 30] Figure 30 shows the vector map of the expression vector SS*-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS*.
[0064] [Figure 31] Figure 31 shows the vector map of the expression vector SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS*. [Figure 32]Figure 32 shows the vector map of the expression vector SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS*. [Figure 33] Figure 33 shows the vector map of the expression vector SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS*. [Figure 34] Figure 34 shows the vector map of the expression vector SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-3'UTR[2hBGpA-A120]-SS*. [Figure 35] Figure 35 shows the vector map of the expression vector SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-3'UTR[2hBGpA-A120]-SS*. [Figure 36] Figure 36 shows the vector map of the expression vector SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS*. [Figure 37] Figure 37 shows the vector map of the expression vector SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-WPRE-3'UTR[2hBGpA-A120]-SS*. [Figure 38] Figure 38 shows the vector map of the expression vector SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-WPRE-3'UTR[2hBGpA-A120]-SS*. [Figure 39]Figure 39 shows a line graph of luciferase activity, expressed as RLU luminescence, in the medium of HEK-293 cells on days 2, 3, 7, 10, 14, 21, and 28 after electroporating the cells with the expression vectors in Figure 2 (conventional pDNA CMV-U), Figure 30 (A: CMV-U1-3'UTR), Figure 31 (B: E1-CMV-U1-3'UTR), and Figure 32 (C: E1-CMV-U1-WPRE-3'UTR). *=p<0.05 and **=p<0.01. [Figure 40] Figure 40 shows a line graph of relative luciferase intensity in HEK-293 cells at passages 1, 2, 3, 4, and 5 after passage every 7 days after electroporating the cells on day 0 with the expression vector described in Figure 39. *p<0.05 and **=p<0.01. [Figure 41] Figure 41 shows a vector map of the expression vector pGL2-CAG-SecNLuc-2A-eGFP-WPRE-bGlobin polyA. [Figure 42] Figure 42 shows a vector map of the expression vector 4-1 pGL2-SS*-CAG [CMV enhancer + CBA promoter + intron]-SecNLuc-2A-eGFP-WPRE-3'UTR (108 to 120 polyA)-SS*. [Figure 43] Figure 43 shows a vector map of the expression vector 4-2 pGL2-SS*-CAG [E1 X3 + CBA promoter + intron]-SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA)-SS*.
[0065] [Figure 44] Figure 44 shows a vector map of the expression vector 4-3 pGL2-SS*-CAG [E2(U100)+CBA promoter+intron]-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS*. [Figure 45]Figure 45 shows a vector map of the expression vector 4-4 pGL2-SS*-CAG [E1 X3+CBA promoter+UTR1]-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS*. [Figure 46] Figure 46 shows a vector map of the expression vector 4-5-pGL2-SS*-CAG [E2 (U100) + CBA promoter + UTR1]-SecNLuc-2A-eGFP-WPRE-3'UTR (108 to 120 polyA)-SS*. [Figure 47] Figure 47 shows the vector map of the expression vector 4-6-pGL2-SS*-CMV enhancer-EF1-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR(108 to 120 polyA)-SS*. [Figure 48] Figure 48 shows a bar graph of luciferase activity, as indicated by luminescence in RLU, in the medium of HEK-293 cells 3 and 6 days after transfection with the expression vectors shown in Figure 41 (conventional SSeq-less pDNA cag-W) and Figures 42 to 47 (4-1 to 4-6, respectively). [Figure 49] Figure 49 shows an exemplary sequence diagram of a self-limiting CRISPR gene editing system, including flanking supersequences (SSeq), a synthetic enhancer (El), a CMV promoter (PCMV), a synthetic 5'UTR containing an optimized internal intron with a tRNA-gRNA-PAM insert (UTR-tRNA-gRNA-PAM-1), a Casβ2 gene, and a 3'UTR containing a human β-globin polyadenylation signal and a gRNA-PAM insert (HBg3'UTR-gRNA-PAM). [Figure 50] Figure 50 shows a diagram of self-limiting Cas expression from the sequence of Figure 52 during homology-directed repair (HDR) of chromosomal DNA bearing a therapeutic GOI flanked by homology arms. [Figure 51]Figure 51 shows a diagram of two gene editing scenarios using the self-limiting CRISPR gene editing system. In scenario 1, an msDNA containing a human expression cassette (e.g., a therapeutic GOI) is first transfected for transient expression, followed by the gene editing system of Figure 42 for HDR knock-in. In scenario 2, HDR knock-in is mediated by a single msDNA containing both the self-limiting CRISPR gene editing system and the human expression cassette flanked by homology arms. DETAILED DESCRIPTION OF THE INVENTION
[0066] Detailed Description The present invention provides expression vectors, vectors free of bacterial sequences (eg, ministring DNA (msDNA)), vector production systems, methods for producing vectors free of bacterial sequences, compositions and uses thereof.
[0067] All publications cited herein, including, but not limited to, all journal articles, books, instruction manuals, patent applications, and patents, are incorporated herein by reference in their entirety to the same extent as if each individual publication was specifically and individually indicated by citation.
[0068] I. Terminology In order that the present invention may be more readily understood, certain terms are first defined. As used herein, unless otherwise stated in the specification, each of the following terms may have the meaning set forth below. Additional definitions are set forth throughout the specification.
[0069] It is noted that a noun followed by the term "a" or "an" means one or more of that noun; for example, a "nucleotide sequence" is understood to refer to one or more nucleotide sequences. Thus, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.
[0070] The term "and / or," as used herein, should be understood as the specific disclosure of each of the specified features or components with or without the other. Thus, when used in phrases such as "A and / or B," the term "and / or" is intended to include "A and B," "A or B," "A" (only), and "B" (only). Similarly, when used in phrases such as "A, B, and / or C," the term "and / or" is intended to encompass each of the following aspects: A, B, and C; A, B, or C; A, or C; A or B; B or C; A and C; A and B; B and C; A (only); B (only); and C (only).
[0071] Where an aspect is described herein by the language "comprising," it is understood that other similar aspects described in the terms "consisting of" and / or "consisting essentially of" are also provided.
[0072] The terms "about" or "substantially comprising" refer to a value or composition that is within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which may depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "substantially comprising" can mean within one standard deviation or more than one standard deviation per practice in the art. Alternatively, "about" or "substantially comprising" can mean a range of up to 10% (i.e., ±10%). Furthermore, particularly with respect to biological systems or processes, the term can mean up to an order of magnitude of the value or up to five times the value. When a particular value or composition is provided in this specification or claims, unless otherwise specified, it should be assumed that the meaning of "about" or "substantially comprising" is within an acceptable error range for that particular value or composition.
[0073] As described herein, any concentration range, percentage range, ratio range, or integer range, unless otherwise specified, should be understood to include any integer value within the recited range, as well as fractions thereof, where appropriate (e.g., tenths and hundredths of integers, etc.). Numeric ranges are inclusive of the numbers defining the range.
[0074] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as that commonly understood by those skilled in the art in the technical field to which this specification relates.For example, Concise Dictionary of Biomedicine and Molecular Biology, Juo, Pei-Show, 2nd ed., 2002, CRC Press; The Dictionary of Cell and Molecular Biology, 5th ed., 2013, Academic Press; and Oxford Dictionary of Biochemistry and Molecular Biology, 2006, Oxford University Press provide one of the techniques in many general dictionary of terms used herein.
[0075] Units, prefixes, and symbols are shown in their International System of Units (SI) accepted form.
[0076] Unless otherwise specified, nucleotide sequences are written left to right in 5' to 3' direction. Amino acid sequences are written left to right in amino to carboxy direction.
[0077] The headings provided herein are not limitations of the various aspects of the specification, which can be obtained by reference to the specification as a whole. Accordingly, the terms defined immediately below are more fully defined by reference to the specification.
[0078] An "amino acid" is a molecule having a central carbon atom (the alpha-carbon atom) bound to a hydrogen atom, a carboxylic acid group (the carbon atom being referred to herein as the "carboxyl carbon atom"), an amino group (the nitrogen atom being referred to herein as the "amino nitrogen atom"), and a side group R. When incorporated into a peptide, polypeptide, or protein, an amino acid loses one or more atoms of its amino acid carboxyl group in a dehydrogenation reaction that links one amino acid to another. As a result, when incorporated into a protein, an amino acid is referred to as an "amino acid residue."
[0079] "Protein" or "polypeptide" means any polymer (whether naturally occurring or not) of two or more individual amino acids linked via peptide bonds, which occurs when the carboxyl carbon atom of the carboxylic acid group attached to the alpha-carbon of one amino acid (or amino acid residue) becomes covalently bonded to the amino nitrogen atom of the amino group attached to the non-alpha-carbon of an adjacent amino acid. The terms "protein" and "polypeptide" may be used interchangeably herein. Similarly, fragments of proteins and polypeptides are also within the scope of the present invention and may be referred to herein as "proteins" or "polypeptides." In one aspect of the present invention, a polypeptide comprises a chimera of two or more parent peptide segments or proteins. The term "polypeptide" is also intended to mean and encompass the product of a post-translational modification ("PTM") of a polypeptide, such as, but not limited to, disulfide bond formation, glycosylation, carbamylation, lipidation, acetylation, phosphorylation, amidation, derivatization with known protecting / blocking groups, proteolytic cleavage, modification with non-naturally occurring amino acids, or any other manipulation or modification, such as conjugation with a labeling moiety. A polypeptide can be obtained from a natural biological source or produced by recombinant technology, but is not necessarily translated from a designated nucleic acid sequence. It can be produced in any manner, including chemical synthesis. An "isolated" polypeptide, or a fragment, variant, or derivative thereof, means a polypeptide that is not in its natural environment. No particular level of purification is required. For example, an isolated polypeptide can simply be removed from its native or natural environment. Recombinantly produced polypeptides and proteins expressed in host cells are considered isolated for the purposes of the present invention when they are natural or recombinant polypeptides that have been separated, fractionated, or partially or substantially purified by any suitable technique.
[0080] A recombinant polypeptide comprising two or more proteins disclosed herein (i.e., a recombinant protein) can be encoded by a single coding sequence comprising a polynucleotide sequence encoding each protein. Unless otherwise specified, the polynucleotide sequences encoding each domain and / or protein are "in frame," such that translation of a single mRNA comprising the polynucleotide sequences results in a single polypeptide comprising each protein. Typically, the proteins in a recombinant polypeptide described herein will be fused directly to each other or separated by a peptide linker. A variety of polypeptide sequences encoding peptide linkers are known in the art, including, for example, self-cleaving peptides.
[0081] As used herein, "polynucleotide" or "nucleic acid" refers to a polymeric form of nucleotides. In some cases, a polynucleotide contains sequences that are not immediately adjacent to or immediately adjacent (at the 5' or 3' end) to a coding sequence, in which case the coding sequence is in the naturally occurring genome of the organism from which it is derived. Thus, the term includes recombinant DNA, whether incorporated into a vector, autonomously replicating plasmid, or virus, or into the genomic DNA of a prokaryotic or eukaryotic cell, or existing as a separate molecule (e.g., cDNA) independent of other sequences. The nucleotides of the invention can be ribonucleotides, deoxyribonucleotides, or modified forms of either nucleotide. As used herein, polynucleotide refers, among other things, to single- and double-stranded DNA, DNA that is a mixture of single- and double-stranded regions, single- and double-stranded RNA, and RNA that is a mixture of single- and double-stranded regions, hybrid molecules containing DNA and RNA that can be single-stranded regions, more typically double-stranded regions, or a mixture of single- and double-stranded regions. The term polynucleotide encompasses genomic DNA or RNA (depending on the organism, i.e., viral RNA genomes), as well as mRNA encoded by genomic DNA, and cDNA. In certain aspects, polynucleotides contain conventional phosphodiester bonds or non-conventional bonds (e.g., amide bonds, such as those found in peptide nucleic acids (PNAs)). By "isolated" nucleic acid or polynucleotide is intended a nucleic acid molecule, e.g., DNA or RNA, that has been removed from its natural environment. For example, a nucleic acid molecule comprising a polynucleotide encoding a recombinant polypeptide contained in a vector is considered "isolated" for purposes of the present invention. Further examples of isolated polynucleotides include recombinant polynucleotides maintained in heterologous host cells or purified (partially or substantially) from other polynucleotides in solution. Isolated RNA molecules include in vivo or in vitro RNA transcripts of the polynucleotides of the present disclosure.Isolated polynucleotides or nucleic acids according to the present invention further include synthetically produced polynucleotides and nucleic acids (eg, nucleic acid molecules).
[0082] As used herein, a "coding region" or "coding sequence" is a portion of a polynucleotide consisting of codons translatable into amino acids. A "stop codon" (TAG, TGA, or TAA) typically is considered to be part of a coding region, although any adjacent sequences, such as promoters, ribosome binding sites, transcription terminators, introns, and the like, are not part of the coding region. The boundaries of a coding region are typically determined by a start codon at the 5'-terminus, which encodes the amino terminus of the resulting polypeptide, and a translation stop codon at the 3'-terminus, which encodes the carboxyl terminus of the resulting polypeptide.
[0083] As used herein, an "expression cassette" comprises a nucleic acid sequence of interest (eg, a nucleic acid sequence, DNA or RNA, for expression of a polypeptide) and an expression control region.
[0084] As used herein, "transgene" is used interchangeably with "gene of interest (GOI)" and refers to a portion of a polynucleotide that contains codons translatable into amino acids. A "stop codon" (TAG, TGA, or TAA) is not usually translated into an amino acid but can still be considered part of the transgene; however, adjacent sequences, such as promoters, ribosome binding sites, transcription terminators, and introns, are not part of the transgene. The boundaries of a transgene are generally determined by a start codon at the 5'-terminus, encoding the amino-terminus of the resulting polypeptide, and by a translation stop codon at the 3'-terminus, encoding the carboxyl-terminus of the resulting polypeptide.
[0085] As used herein, the term "expression control region" refers to a transcriptional control element operably linked to a coding region to direct or regulate the expression of a product encoded by the coding region, including, for example, cis-regulatory modules (CRMs), promoters (e.g., tissue-specific and / or inducible promoters), enhancers, operators, repressors, ribosome binding sites, translation leader sequences, introns, post-transcriptional elements, polyadenylation recognition sequences, RNA processing sites, effector binding sites, stem-loop structures, and transcription termination signals, miRNA binding sites, and combinations thereof. Expression control regions include nucleotide sequences located upstream (5'), within, or downstream (3') of a nucleic acid sequence of interest that affect the transcription, RNA processing, stability, or translation of the associated nucleic acid sequence of interest. When a transgene is intended for expression in eukaryotic cells, polyadenylation signals and transcription termination sequences are usually located 3' of the transgene.
[0086] A coding region and promoter are "operably associated" (i.e., "operably linked") if introduction of promoter function results in the transcription of an mRNA comprising the coding region that encodes a product, and if the nature of the linkage between the promoter and coding region does not interfere with the ability of the promoter to direct expression of the product encoded by the coding region or the ability of the DNA template to be transcribed. Expression control regions comprise nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding region that affect the transcription, RNA processing, stability, or translation of the associated coding region. If a coding region is intended for expression in a eukaryotic cell, a polyadenylation signal and transcription termination sequence will usually be located 3' to the coding region.
[0087] As used herein, the terms "host cell" and "cell" are used interchangeably and can refer to any type of cell or population of cells, such as primary cells, cells in culture, or cells from a cell line, that carry or are capable of carrying a nucleic acid molecule (e.g., a recombinant nucleic acid molecule). Host cells can be prokaryotic cells, or they can be eukaryotic cells, such as fungal cells, such as yeast cells, and various animal cells, such as insect cells or mammalian cells.
[0088] As used herein, "culture," "culturing," and "culturing" refer to the incubation of cells or maintaining cells in a living state under in vitro conditions that permit cell growth or cell division. As used herein, "cultured cells" refers to cells that are propagated in vitro.
[0089] A "subject" includes any human or non-human animal. The term "non-human animal" includes, for example, but is not limited to, vertebrates, such as mammals, birds, pets, livestock, non-human primates, sheep, cows, goats, pigs, chickens, dogs, cats, and rodents, such as mice, rats, and guinea pigs. In a preferred aspect, the subject is a human. The terms "subject" and "patient" are used interchangeably herein.
[0090] "Administering" refers to the physical introduction of a therapeutic agent into a subject using any of a variety of methods and delivery systems known to those skilled in the art.
[0091] The terms "treating," "treating," "treatment" of a subject, or "therapy," as used herein, refer to any type of intervention or process performed on a subject, or the administration of an active agent to a subject, for the purpose of reversing, alleviating, ameliorating, inhibiting, or otherwise delaying or preventing the progression, onset, severity, or recurrence of symptoms, complications, medical conditions, or biochemical manifestations associated with a disease, or prolonging overall survival. Treatment can be administered to subjects with a disease or to subjects without a disease (e.g., for prophylaxis, such as vaccination).
[0092] The terms "effective dose," "effective dosage," or "effective amount" are defined as an amount of an agent sufficient to achieve or at least partially achieve a desired effect. A "therapeutically effective amount" or "therapeutically effective dosage" of a drug or therapeutic agent is any amount of drug that, when used alone or in combination with another therapeutic agent, promotes disease reversal as manifested by a decrease in the severity of disease symptoms, an increase in the frequency and duration of disease-free periods, an increase in overall survival (the length of time from the date of diagnosis or initiation of treatment for the disease that a patient diagnosed with the disease remains alive), or prevention of disability or disabling illness due to the disease. A therapeutically effective amount or dosage of a drug encompasses a "prophylactically effective amount" or "prophylactically effective dosage," which is any amount of drug that prevents the onset or recurrence of disease when administered alone or in combination with another therapeutic agent to a subject at risk of developing the disease or suffering from a recurrence of the disease. The ability of a therapeutic agent to promote disease improvement or prevent the onset or recurrence of disease can be assessed using a variety of methods known to those of skill in the art, such as, for example, by testing the agent's activity in human subjects during clinical trials, in animal model systems predictive of efficacy in humans, or in in vitro assays.
[0093] Various aspects of the disclosure are described in more detail in the following subsections.
[0094] II. Expression Vectors and Vector Construction Systems for Producing Vectors Free of Bacterial Sequences Bacterial sequence-free vectors and their production are described in U.S. Patent Nos. 9,290,778 and 9,862,954; Nafissi and Slavcev, Microbial Cell Factories 11:154 (2012); and Nafissi et al., Nucleic Acids 3(6):e165 (2014), which are incorporated by reference in their entireties. These bacterial sequence-free vectors are produced from expression vectors (e.g., plasmids) that contain specialized "super sequence" ("SS" or "SSeq") sites containing target sequences for a recombinase flanking each side (i.e., 5' and 3') of an expression cassette containing a nucleic acid sequence of interest. Specifically, each SS contains a target sequence for a first recombinase, with further target sequences for one or more additional recombinases incorporated within the non-binding region for the first recombinase. When the expression vector is present in a recombinant cell expressing a suitable recombinase, the bacterial sequence-free vector containing the expression cassette is separated from the backbone DNA of the expression vector.To create a circular covalently closed (CCC) bacterial sequence-free vector, also referred to herein as ministring DNA (msDNA), the expression vector is placed into a recombinant cell expressing a recombinase such as TelN or Tel, which acts through its target sequence in the SS.The bacterial sequence-free vector obtained from the recombination can then be purified from the cells and used directly as a delivery vector.See U.S. Patent Nos. 9,290,778 and 9,862,954, Nafissi and Slavcev, and Nafissi et al.
[0095] msDNA with LCC ends is not twisted and does not undergo gyrase-directed negative supercoiling during its production in E. coli. Furthermore, due to its double-stranded LCC topology, insertion of msDNA into a cell's chromosome results in chromosomal breakage, thereby eliminating the cell from the population. Therefore, msDNA eliminates all risk of insertional mutagenesis and protects patients administered msDNA from potential genotoxicity and cancer (Nafissi et al.).
[0096] The present invention provides improved methods for producing vectors that do not contain bacterial sequences, and improved vectors that do not contain bacterial sequences. In some aspects, the production of vectors that do not contain bacterial sequences is improved by removing contaminating expression vector sequences. In some aspects, vectors that do not contain bacterial sequences are improved by their ability to establish in cells (i.e., transfection efficiency), improved transgene expression (e.g., mediated by a combination of enhanced transcription and translation), and improved spread in cells (e.g., replication and partitioning of the vector into daughter cells).
[0097] In some aspects, the improvements described herein may be adapted to CCC or LCC vectors made according to other methods known in the art.
[0098] A. Expression Vectors The present invention provides a method for producing a nucleic acid sequence comprising: (a) a backbone sequence; (b) an expression cassette comprising: (i) a nucleic acid sequence of interest; (ii) a first target sequence for a first recombinase flanking the 5' side of the expression cassette; (iii) a second target sequence for the first recombinase flanking the 3' side of the expression cassette; and (iv) one or more additional target sequences for one or more additional recombinases integrated within the non-binding regions of the first and second target sequences for the first recombinase; and (c) a sequence comprising: (i) a first recombinase and one or more additional recombinases; (ii) an endonuclease target sequence integrated within the first and / or second target sequence for the first recombinase in the non-binding region for the recombinase, the endonuclease target sequence being between the backbone sequence and the cleavage sites for the first recombinase and one or more additional recombinases; (iii) an endonuclease target sequence integrated between the 3' end of the first target sequence for the first recombinase and the 5' end of another enhancer or promoter in the expression cassette and a sequence encoding SEQ ID NO: 12 and at least about 90 (iii) a cytomegalovirus (CMV) enhancer integrated between the 3' end of the first target sequence for the first recombinase and the 5' end of the promoter in the expression cassette; (iv) a 5' untranslated region (5'UTR) including an intron, wherein the 5'UTR is integrated between the promoter and the nucleic acid sequence of interest in the expression cassette; (v) a vertebrate chromatin insulator integrated into the expression cassette between the nucleic acid of interest and a polyadenylation signal; (vi) a chromatin insulator of a vertebrate (vii) a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal; (vii) a scaffold / matrix attachment region (S / MAR) incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal; or (viii) a DNA nuclear targeting sequence (DTS) incorporated within the first and / or second target sequence for the first recombinase in the non-binding region for the first recombinase and one or more additional recombinases, wherein the DTS isand a DTS located between the expression cassette and the cleavage sites for the first recombinase and one or more additional recombinases.
[0099] As used herein, a "backbone" sequence is the sequence of an expression vector excluding the sequence of the expression cassette and the flanking SS sites containing the first and second target sequences of a first recombinase. The backbone sequence can include, for example, sequences for amplification of the expression vector in a host cell (e.g., E. coli) and antibiotic selection as described herein.
[0100] A "non-binding" region for a recombinase is a region within a target sequence for a first recombinase that is not acted upon by the recombinase (e.g., not bound and / or cleaved by the recombinase), as described herein.
[0101] A "cleavage site" for a recombinase is the site at which the recombinase initiates a double-stranded break or single-stranded nick in DNA that accompanies recombination.
[0102] In some aspects, the expression vector comprises an endonuclease target sequence integrated into the first and / or second target sequence for the first recombinase in the non-binding region for the first recombinase and one or more additional recombinases, wherein the endonuclease target sequence is between the backbone sequence and the cleavage site for the first recombinase and one or more additional recombinases. In some aspects, the endonuclease target sequence is integrated into the first target sequence for the first recombinase. In some aspects, the endonuclease target sequence is integrated into the second target sequence for the first recombinase. In some aspects, the endonuclease target sequence is integrated into both the first and second target sequences for the first recombinase. In some aspects, the same endonuclease target sequence is integrated into the first and second target sequences for the first recombinase. In some aspects, the endonuclease target sequences integrated into the first and second target sequences for the first recombinase are for the same endonuclease. In some aspects, the endonuclease target sequence integrated into the first target sequence for the first recombinase is different from the endonuclease target sequence integrated into the second target sequence for the first recombinase. In some aspects, the endonuclease target sequence integrated into the first target sequence for the first recombinase is for a different endonuclease than the endonuclease target sequence integrated into the second target sequence for the first recombinase.
[0103] By locating the endonuclease target sequence between the backbone sequence and the cleavage site of the recombinase in the expression vector, it is ensured that after recombination as described herein, the endonuclease target sequence remains attached to the backbone sequence, rather than to the vector containing no bacterial sequences. Thus, after recombination, the backbone sequence and the sequence containing the endonuclease target site can be removed from a preparation containing the vector containing no bacterial sequences by exposure to an endonuclease, reducing or avoiding the need for a purification step to remove the backbone sequence in methods for producing a vector containing no bacterial sequences. In one aspect, the endonuclease is expressed after recombination in a host cell of the vector production system described herein, where the endonuclease cleaves DNA at the endonuclease target site, and the backbone sequence and the sequence containing the endonuclease target site are degraded by an exonuclease (e.g., exonuclease V).
[0104] In some aspects, the expression vector comprises an endonuclease target sequence for a homing endonuclease. In certain aspects, the endonuclease target sequence is I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H -DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI, I-PpoI, PI-PspI, I- ScaI, I-SceI, PI-SceI, I-SceII, I-SecIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I or I-Vdi141I. In one aspect, the endonuclease target sequence is for I-SceI. In one aspect, the endonuclease target sequence is for PI-SceI. Target sequences for homing endonucleases are well known in the art.
[0105] In some aspects, the expression vector comprises an endonuclease target sequence for an endonuclease used in genome editing, including an endonuclease that is part of a nuclease genome editing system. In some aspects, the nuclease genome editing system is a clustered regularly interspaced short palindromic repeats (CRISPR) system, a transcription activator-like effector nuclease (TALEN) system, a zinc finger nuclease (ZFN) system, or a meganuclease system.
[0106] In some aspects, the expression vector comprises an endonuclease target sequence for a Cas endonuclease. In some aspects, the Cas endonuclease is Cas9 (e.g., Streptococcus pyogenes Cas9 (SpCas9), Staphlococcus aureus Cas9 (SaCas9), Francisella novicida Cas9 (FnCas9), or Neisseria meningitidis Cas9 (FnCas9). The Cas endonuclease may be selected from the group consisting of Cas9 (NmCas9), Cas9 variants (e.g., Cas9β2, xCas9, SpCas9-NG, SpCas9-NRRH, SpCas9-NRCH, SpCas9-NRTH, SpG, and SpRY), Cas3, Cas12 (e.g., Cas12a, Cas12b, Cas12c, Cas12d, or Cas12e), Cas13 (e.g., Cas13a, Cas13b, Cas13c, or Cas13d), or Cas14. In some aspects, the endonuclease target sequence for a Cas endonuclease used herein is homologous to a guide RNA (gRNA) targeting sequence and includes a protospacer adjacent motif (PAM) recognized by the Cas endonuclease. The sequence homologous to the gRNA targeting sequence with PAM site can be routinely designed based on the well-known CRISPR system. The gRNA comprises a fusion of a targeting RNA (crRNA) sequence and a transactivating RNA (tracrRNA) sequence, which interact and function to guide the Cas endonuclease to the endonuclease target site and catalyze cleavage.
[0107] In one aspect, an expression vector includes a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12, incorporated between the 3' end of a first target sequence for a first recombinase and the 5' end of another enhancer or promoter in the expression cassette. In one aspect, an expression vector includes a synthetic enhancer comprising the nucleic acid sequence of SEQ ID NO: 12, incorporated between the 3' end of a first target sequence for a first recombinase and the 5' end of another enhancer or promoter in the expression cassette. In one aspect, the synthetic enhancer includes multiple contiguous copies of the nucleic acid sequence, such as, for example, 1, 2, 3, 4, 5, or more contiguous copies. In one aspect, the synthetic enhancer includes three contiguous copies of the nucleic acid sequence. In one aspect, the synthetic enhancer comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 46. In one aspect, the synthetic enhancer comprises the nucleic acid sequence of SEQ ID NO: 46. In one aspect, the synthetic enhancer is incorporated into the 5' end of a chicken β-actin promoter. In one aspect, a chimeric intron comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 47 is incorporated into the 3' end of the chicken β-actin promoter and the 5' end of the nucleic acid sequence of interest. In one aspect, a chimeric intron comprising the nucleic acid sequence of SEQ ID NO: 47 is incorporated at the 3' end of the chicken β-actin promoter and at the 5' end of the nucleic acid sequence of interest.
[0108] In one aspect, the expression vector includes a CMV enhancer integrated between the 3' end of the first target sequence for the first recombinase and the 5' end of the promoter in the expression cassette. In one aspect, the CMV enhancer is integrated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In one aspect, the CMV enhancer is integrated at the 3' end of multiple contiguous copies of the synthetic enhancer, such as the 3' end of 1, 2, 3, 4, 5, or more contiguous copies of the synthetic enhancer. In one aspect, the CMV enhancer is incorporated at the 3' end of three consecutive copies of the synthetic enhancer. In one aspect, the CMV enhancer is incorporated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 46. In one aspect, the CMV enhancer is incorporated at the 3' end of the nucleic acid sequence of SEQ ID NO: 46. In one aspect, the CMV promoter is incorporated at the 3' end of the CMV enhancer and the 5' end of the nucleic acid sequence of interest.
[0109] In one aspect, an expression vector comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39 integrated between a first target sequence for a first recombinase and a nucleic acid sequence of interest. In one aspect, an expression vector comprises a nucleic acid sequence of SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39 integrated between a first target sequence for a first recombinase and a nucleic acid sequence of interest. In one aspect, a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39, or the nucleic acid sequence of SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39, includes all regulatory elements in an expression cassette located at the 5' end of the nucleic acid sequence of interest.
[0110] In one aspect, the expression vector comprises a 5'UTR that includes an intron, wherein the 5'UTR (i.e., the 5'UTR that includes an intron) is incorporated between the promoter and the nucleic acid sequence of interest in the expression cassette.
[0111] In one aspect, the 5'UTR is intended to improve splicing and translation of the transgene transcript from the expression vector or a vector not containing the bacterial sequence produced from the expression vector, compared to the same expression vector or vector not containing the bacterial sequence, respectively, lacking the 5'UTR.
[0112] In some aspects, the intron comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 1. In some aspects, the intron comprises the nucleic acid sequence of SEQ ID NO: 1.
[0113] In one aspect, the 5'UTR comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO:2, which is an optimized 5'UTR with an internal minimal intron, also referred to herein as "5'UTR1." In one aspect, the 5'UTR comprises the nucleic acid sequence of SEQ ID NO:2.
[0114] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 4. In some aspects, the 5'UTR comprises the nucleic acid sequence of SEQ ID NO: 4.
[0115] In some aspects, the 5'UTR further comprises non-coding sequences embedded within the intron.
[0116] In one aspect, the intron is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to or comprises SEQ ID NO:1, and the non-coding sequence is incorporated between two of the nucleotides in the intron corresponding to any two of nucleotide positions 25 to 55 of SEQ ID NO:1.
[0117] In some aspects, the non-coding sequence is non-prokaryotic and non-viral. In some aspects, the non-coding sequence is a eukaryotic sequence. In some aspects, the non-coding sequence comprises an intron, a ubiquitous chromatin opening element (UCOE), an S / MAR, an SV40 enhancer sequence (e.g., one or more SV40 enhancer sequences, such as two, three, four, five, or more SV40 enhancer sequences), a vertebrate chromatin insulator (e.g., cHS4), a WPRE, or any combination thereof.
[0118] In one aspect, the non-coding sequence comprises an S / MAR, hi one aspect, the S / MAR is MAR-5, set forth herein in SEQ ID NO:9.
[0119] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 3. In some aspects, the 5'UTR comprises SEQ ID NO: 3.
[0120] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 5. In some aspects, the 5'UTR comprises SEQ ID NO:5.
[0121] In one aspect, a 5'UTR is incorporated into the expression cassette between the chicken β-actin promoter and the nucleic acid sequence of interest.
[0122] In one aspect, a 5'UTR is incorporated into the expression cassette between the CMV promoter and the nucleic acid sequence of interest.
[0123] In one aspect, a 5'UTR is incorporated between the promoter and the nucleic acid sequence of interest in an expression cassette, wherein the promoter is incorporated at the 3' end of the CMV enhancer. In one aspect, the CMV enhancer is incorporated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In one aspect, the CMV enhancer is incorporated at the 3' end of multiple consecutive copies of the synthetic enhancer, such as the 3' end of 1, 2, 3, 4, 5, or more consecutive copies of the synthetic enhancer. In some aspects, the CMV enhancer is incorporated at the 3' end of three consecutive copies of the synthetic enhancer. In some aspects, the CMV enhancer is incorporated at the 3' end of a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO:46. In some aspects, the CMV enhancer is incorporated at the 3' end of SEQ ID NO:46.
[0124] In some aspects, the expression vector comprises a polyadenylation signal incorporated into the 3' end of the nucleic acid sequence of interest. In some aspects, the polyadenylation signal comprises a Xenopus laevis beta globin polyadenylation signal, a human beta globin polyadenylation signal, or a hybrid Xenopus laevis and human beta globin polyadenylation signal. In some aspects, the polyadenylation signal comprises multiple copies, such as 1 copy, 2 copies, 3 copies, 4 copies, or 5 copies, of the Xenopus laevis beta globin polyadenylation signal, the human beta globin polyadenylation signal, or the hybrid Xenopus laevis and human beta globin polyadenylation signal. In certain aspects, the polyadenylation signal comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15. In certain aspects, the polyadenylation signal comprises the nucleic acid sequence of SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15. In certain aspects, a polyadenylate tail (i.e., a poly(A) tail) is located at the 3' end of the polyadenylation signal. In some aspects, the poly(A) tail is 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120 or more residues in length. In some aspects, the sequence comprising the polyadenylation signal and poly(A) tail is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO:16, SEQ ID NO:17, or SEQ ID NO:18.In one aspect, the sequence comprising the polyadenylation signal and poly(A) tail comprises SEQ ID NO:16, SEQ ID NO:17, or SEQ ID NO:18.
[0125] In some aspects, the expression vector comprises a vertebrate chromatin insulator in the expression cassette. In some aspects, the vertebrate chromatin insulator is the 5'-HS4 chicken-β-globulin insulator (cHS4). See, for example, Benabdellah et al., PLoS ONE 9(1): e84268 (2014); Lu et al., FEBS Open Bio 10: 644-656 (2020); Hanawa et al., Mol. Ther. 17(4): 667-674 (2009); Walters et al., Mol. Cell. Biol. 19(5): 3714-3726 (1999). In some aspects, the vertebrate chromatin insulator is incorporated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal, as described herein. In one aspect, a vertebrate chromatin insulator, as described herein, is integrated within an intron of the 5'UTR.
[0126] In some aspects, the vertebrate chromatin insulator comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 8. In some aspects, the vertebrate chromatin insulator comprises SEQ ID NO:8.
[0127] In one aspect, the vertebrate chromatin insulator is intended to improve colonization (i.e., transfection efficiency) of an expression vector or a bacterial sequence-free vector produced from the expression vector, compared to the same expression vector or bacterial sequence-free vector, respectively, that does not comprise the vertebrate chromatin insulator.
[0128] In some aspects, the expression vector comprises a WPRE in the expression cassette.See, for example, Higashimoto et al., Gene Therapy 14: 1298-1304 (2007).In some aspects, the WPRE is incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette as described herein.In some aspects, the WPRE is incorporated at the 3' end of the S / MAR and the 5' end of the polyadenylation signal in the expression cassette as described herein.In some aspects, the WPRE is incorporated within the intron of the 5'UTR as described herein.
[0129] In some aspects, the WPRE comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 11. In some aspects, the WPRE comprises SEQ ID NO: 11.
[0130] In one aspect, the WPRE improves expression of a transgene from an expression vector or a bacterial sequence-free vector produced from the expression vector, compared to the same expression vector or bacterial sequence-free vector, respectively, lacking the WPRE.
[0131] In some aspects, the expression vector contains an S / MAR in the expression cassette. See, for example, Martens et al., Mol. Cell. Biol. 22(8): 2598-2606 (2002); Narwade et al., Nucleic Acids Res. 47(14): 7247-7261 (2019). In some aspects, the S / MAR is incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette. In some aspects, the S / MAR is incorporated at the 3' end of the nucleic acid sequence of interest and at the 5' end of the WPRE in the expression cassette, as described herein. In some aspects, the S / MAR is incorporated within the intron of the 5'UTR, as described herein.
[0132] In some aspects, the S / MAR is MAR-3, MAR-4, or MAR-5, which are fragments of human β-interferon MAR. See, for example, Wang et al., Mol. Biol. Cell 30: 2761-2770 (2019). In some aspects, the S / MAR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 9. In some aspects, the S / MAR comprises SEQ ID NO: 9.
[0133] In some aspects, the S / MAR is a human cytotoxic serine protease-B (CSP-B) MAR or a CSP-C MAR. See, e.g., Hanson and Ley, Blood 79(3):610-618 (1992); Klein et al., Tissue Antigens 35(5):220-228 (1990). In some aspects, the S / MAR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 10. In some aspects, the S / MAR comprises SEQ ID NO: 10.
[0134] In one aspect, the S / MAR is intended to improve the expression level, stability and / or durability (e.g., by episomal maintenance and replication, e.g., propagation and partitioning of the vector to daughter cells, and / or prevention of epigenetic silencing) of the expression vector or bacterial sequence-free vector (produced from the expression vector) compared to the same expression vector or bacterial sequence-free vector, respectively, lacking the S / MAR.
[0135] In one aspect, an expression vector comprising any one or more of (c)(i)-(c)(vii) above (i.e., not comprising a DTS) further comprises an enhancer sequence flanking each strand of the first and second target sequences for the first recombinase. In one aspect, the enhancer sequences flanking each strand of the first and second target sequences for the first recombinase are at least two enhancer sequences flanking each strand of the first and second target sequences for the first recombinase. In one aspect, the enhancer sequences are SV40 enhancer sequences.
[0136] In one aspect, the expression vector comprises a DTS. In one aspect, the DTS is integrated into the first and / or second target sequence for the first recombinase in the non-binding region for the first recombinase and one or more additional recombinases, wherein the DTS is located between the expression cassette and the cleavage sites for the first recombinase and one or more additional recombinases. In one aspect, the DTS is an SV40 enhancer sequence. In one aspect, the DTS is cell-specific. In one aspect, the DTS is specific for smooth muscle cells, embryonic stem cells, type II alveolar epithelial cells, endothelial cells, or osteoblasts.
[0137] The DTS, located between the expression cassette and the cleavage site for the recombinase in the expression vector, ensures that after recombination, the DTS remains attached to the bacterial sequence-free vector and not to the backbone sequences, as described herein.
[0138] In some aspects, the expression vector contains a UCOE in the expression cassette. See, for example, Mueller-Kuller et al., Nucleic Acids Res. 43(3): 1577-1592 (2015); Skipper et al., BMC Biotechnol. 19:75 (2019); Rudina et al., bioRxiv, doi.org / 10.1101 / 626713 (2019); Neville et al., Biotechnol. Adv. 35(5): 557-564 (2017). In some aspects, the UCOE is located between the 3' end of the first target sequence for the first recombinase and the 5' end of the promoter or any enhancer in the expression cassette. In some aspects, the UCOE is integrated into an intron of the 5'UTR as described herein.
[0139] In some aspects, the UCOE is an A2UCOE. In some aspects, the UCOE comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 6. In some aspects, the UCOE is SEQ ID NO: 6.
[0140] In some aspects, the UCOE is an SRF-UCOE. See, e.g., International Patent Application No. WO2020223160. In some aspects, the UCOE comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 7. In some aspects, the UCOE is SEQ ID NO: 7.
[0141] In one aspect, the UCOE improves expression of a transgene from an expression vector or a vector not containing bacterial sequences produced from the expression vector, compared to the same expression vector or vector not containing bacterial sequences, respectively, lacking the UCOE.
[0142] In one aspect, the expression vector comprises enhancer-1 in the expression cassette. In one aspect, enhancer-1 is integrated between the 3' end of the first target sequence for the first recombinase and the 5' end of the promoter or any enhancer in the expression cassette. In one aspect, enhancer-1 is integrated between the 3' end of the UCOE and the 5' end of the CMV enhancer. In one aspect, enhancer-1 comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In one aspect, enhancer-1 is SEQ ID NO: 12.
[0143] In some aspects, the expression vector comprises a CMV, EF1, SV40, cag, Rho, VDM2, HCR, or HLP promoter, or a variant thereof, in the expression cassette. In some aspects, the expression vector comprises a CMV promoter variant in the expression cassette. See, for example, International Publication WO2012099540; Xu et al., Bioengineered 10(1): 548-560, DOI: 10.1080 / 21655979.2019.1684863 (2019).
[0144] In one aspect, the expression vector comprises an EF1-alpha promoter in the expression cassette. In one aspect, the expression vector comprises a CMV enhancer and an EF1-alpha promoter in the expression cassette.
[0145] In one aspect, the expression vector contains a 3'UTR in the expression cassette that contains two copies of the beta-globin polyadenylation signal. In one aspect, the 3'UTR is integrated between the nucleic acid sequence of interest and the 5' end of the second target sequence for the first recombinase.
[0146] In one aspect, the 3'UTR comprises two copies of a Xenopus beta-globin polyadenylation signal. In one aspect, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 13. In one aspect, the 3'UTR is SEQ ID NO: 13.
[0147] In one aspect, the 3'UTR comprises two copies of the human beta-globin polyadenylation signal. In one aspect, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 14. In one aspect, the 3'UTR is SEQ ID NO: 14.
[0148] In one aspect, the 3'UTR comprises one copy of a Xenopus beta-globin polyadenylation signal and one copy of a human beta-globin polyadenylation signal. In one aspect, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 15. In one aspect, the 3'UTR is SEQ ID NO: 15.
[0149] In certain aspects, the 3'UTR further comprises a poly(A) tail (i.e., at the 3' end of the 3'UTR) comprising 100 to 120 adenine nucleotides, i.e., 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 adenine nucleotides.
[0150] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 16. In some aspects, the 3'UTR is SEQ ID NO:16.
[0151] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 17. In some aspects, the 3'UTR is SEQ ID NO:17.
[0152] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 18. In some aspects, the 3'UTR is SEQ ID NO:18.
[0153] The expression vector can include any combination of the above modifications to the first and / or second target sequences and / or expression cassettes as described herein. In some aspects, the combination provides a synergistic effect.
[0154] In one aspect, the first and second target sequences for the first recombinase and the one or more additional target sequences for the one or more additional recombinases are selected from the group consisting of a PY54 pal site, an N15 telRL site, a loxP site, a φK02 telRL site, an FRT site, a phiC31 attP site, and a λ attP site. In one aspect, an expression vector comprises each of the target sequences. In one aspect, the expression vector comprises a pal site and telRL, loxP, and FRT recombinase target binding sequences integrated within the pal site. In one aspect, the first and second target sequences for the first recombinase each comprise the nucleic acid sequence of SEQ ID NO: 33.
[0155] In certain aspects, the nucleic acid sequence of interest in any of the expression cassettes described herein comprises a sequence encoding a polypeptide, RNA (messenger RNA (mRNA), micro-RNA (miRNA), small interfering RNA (siRNA), small hairpin RNA (shRNA), ribosomal or antisense RNA), or non-coding DNA (e.g., antisense oligonucleotide). In certain aspects, the nucleic acid sequence of interest is a genomic DNA sequence containing introns and / or exons. In some aspects, the nucleic acid sequence of interest comprises a sequence encoding an anti-cancer agent, a tumor suppressor, an apoptotic agent, an anti-angiogenic agent, an enzymatic agent, a cytotoxic agent, a suicide gene, a cytokine, an interferon, an interleukin, an immunomodulator, an immunostimulator, an immunosuppressant, a chemokine, an antigen for stimulating antigen-presenting cells, an antibody (e.g., a heavy and / or light chain of an antibody such as a monoclonal antibody, a chimeric antibody, a humanized anti- or human antibody, or an antigen-binding fragment thereof), a genome editing system or portion thereof (e.g., a CRISPR-Cas, TALEN, ZFN, or meganuclease system or portion thereof (e.g., a Cas endonuclease or gRNA)), or an immunogenic substance (e.g., as a VLP and / or a vaccine). In some aspects, the nucleic acid sequence of interest comprises a sequence encoding a polypeptide capable of forming a VLP when the nucleic acid sequence is expressed in a cell.
[0156] Exemplary therapeutic targets and indications include, for example, genes associated with monogenic diseases including liver, blood, or eye diseases, galactosidase alpha (GLA, e.g., for treating Fabry disease), sodium voltage-dependent channel alpha subunit 1 (SCN1A, e.g., for treating Dravet syndrome), ATP-binding cassette subfamily A member 4 (ABCA4, e.g., for treating Stargardt disease), surfactant protein B (SP-B, e.g., for treating surfactant dysfunction diseases ... protein B), surfactant protein B (SP-B, e.g., for treating surfactant protein B), surfactant protein B (SP-B, e.g., for treating surfactant protein B), surfactant protein B (SP-B, e.g., for treating surfactant protein B), surfactant protein B (SP-B, e.g., for treating surfactant protein B), surfactant protein B (SP-B, e.g., for treating surfactant protein B), surfactant protein surfactant protein C (SP-C, e.g., for treating surfactant dysfunction diseases), ATP-binding cassette subfamily A member 3 (ABCA3, e.g., for treating surfactant dysfunction diseases), solute carrier family 34 member 2 (SLC34A2, e.g., for treating pulmonary alveolar microlithiasis and / or testicular microlithiasis), cystic fibrosis transmembrane conductance regulator (CFTR, e.g., for treating cystic fibrosis), glutamic acid decarboxylase (GAD, e.g., GAD65 or GAD6 7, e.g., for treating Parkinson's disease), aspartoacylase gene (ASPA, also known as aminoacylase (AAC), e.g., for treating Canavan disease), aromatic L-amino acid decarboxylase (AADC, e.g., for treating Parkinson's disease and / or for treating AADC deficiency), neurturin (NRTN, e.g., for treating Parkinson's disease), glial cell line-derived neurotrophic factor (GDNF, e.g., for treating Parkinson's disease), nerve growth factor (NGF, e.g., for treating Alzheimer's disease), tripeptidyl peptidase I (TPP1, for example, for treating Batten disease, e.g., CLN2 disease), arylsulfatase A (ARSA, for example, for treating metachromatic leukodystrophy), N-sulfoglucosamine sulfohydrolase (SGSH, for example, for treating Sanfilippo syndrome type A), sulfatase modifying factor 1 (SUMF1, for example, for treating Sanfilippo syndrome type A), N-acetyl-α-glucosaminidase (NAGLU, for example,Sanfilippo syndrome type B), survival factor of motor neuron 1 (SMN1, e.g., for treating spinal muscular atrophy 1), retinal pigment epithelium specific 65 kDa protein (RPE65, also known as retinoid isomerohydrolase, e.g., for treating Leber's congenital amaurosis), Rab escort protein 1 (REP1, e.g., for treating choroideremia), retinoschisin 1 (RS1, e.g., for treating X-linked juvenile retinoschisis), alpha-1 antitrypsin (AAT, e.g., for treating hereditary emphysema or AAT deficiency), mini-dystrophin (e.g., for treating Duchenne muscular dystrophy), α-sarcoglycan (αSG, e.g., for treating Duchenne muscular dystrophy or limb-girdle muscular dystrophy type 2), β-sarcoglycan (βSG), γ-sarcoglycan (γSG, e.g., for treating limb-girdle muscular dystrophy type 2), δ-sarcoglycan (δSG), lipoprotein lipase (LPL, e.g., for treating familial LPL deficiency), acid α-glucosidase (GAA, e.g., for treating Pompe disease), tumor necrosis factor receptor:Fc (TNFR:Fc, e.g., for treating arthritis, e.g., inflammatory arthritis), sarcoplasmic / endoplasmic reticulum Ca(2+)ATPase 2a (SERCA2a, e.g., for treating congestive heart failure), Factor VIII or Factor IX (FVIII or FIX, e.g., for treating hemophilia B), porphobilinogen deaminase gene (PBGD, e.g., for treating acute intermittent porphyria), soluble fms-like tyrosine kinase-1 (sFLT1, e.g., for treating age-related macular degeneration or cancer, e.g., ovarian cancer), VEGFR-1 and and VEGF-R2 domains (e.g., for treating cancer, e.g., melanoma or colon cancer), soluble VEGFR3 (e.g., for treating cancer, e.g., endometrial cancer), soluble VEGF-C decoy receptor (sVEGFR3-Fc, e.g., for treating cancer, e.g., melanoma, renal cell carcinoma, or prostate cancer), pigment epithelium-derived growth factor (PEDF, e.g., for treating cancer, e.g.,for treating Lewis lung cancer), neutralizing monoclonal antibodies against VEGFR2 (e.g., for treating cancer, e.g., melanoma or glioblastoma), endostatin (e.g., for treating cancer, e.g., bladder cancer or pancreatic cancer), angiostatin (e.g., for treating cancer, e.g., liver cancer), both endostatin and angiostatin (i.e., as a bicistronic sequence, e.g., for treating cancer, e.g., ovarian cancer), endostatin mutants (i.e., P1254A-endostatin), tin, e.g., for treating cancer, e.g., ovarian cancer), the anti-angiogenic domain of TSP-1 (3TSR, e.g., for treating cancer, e.g., pancreatic cancer), tissue factor pathway inhibitor-2 (TFPI-2, e.g., for treating cancer, e.g., glioblastoma), a fragment of plasminogen (e.g., kringle 5, e.g., for treating cancer, e.g., ovarian cancer), plasminogen kringle 1-5 (e.g., for treating cancer, e.g., melanoma or lung cancer), siRNA against unfolded protein response protein (UPR; For example, IRE1α, XBP-1 or ATF6, e.g., for treating cancer, e.g., breast cancer), vasostatin (e.g., for treating cancer, e.g., lung cancer), herpes simplex virus type 1 thymidine kinase (HSV-TK, e.g., for treating cancer, e.g., breast cancer), sc39TK (e.g., for treating cancer, e.g., cervical cancer), diphtheria toxin A (DTA, e.g., for treating cancer, e.g., cervical cancer or myeloma), modulators of p53-upregulated apoptosis (PUMA, e.g., for treating cancer, e.g., for treating cervical cancer or myeloma), tumor necrosis factor (TNF)-related apoptosis-inducing ligand (TRAIL, for example, for treating cancer, such as lymphoma, hepatocellular carcinoma, squamous cell carcinoma of the head and neck (i.e., head and neck cancer), or glioblastoma), soluble TRAIL (for example, for treating cancer, such as liver cancer or lung adenocarcinoma), IFN-β (for example, for treating cancer, such as colorectal cancer, lung cancer, neuroblastoma, or glioblastoma multiforme), IFN-α (for example, for treating cancer, such as metastatic melanoma),CD40 ligand (CD40L) or CD40L variants (e.g., for treating cancer, e.g., lung cancer), melanoma differentiation associated gene-7 and interleukin 24 (mda-7 and IL24, e.g., for treating cancer, e.g., Ehrlich ascites tumor), apoptotin and IL24 (e.g., for treating cancer, e.g., liver cancer), IL24 (e.g., for treating cancer, e.g., metastatic Lewis lung carcinoma), IL15 (e.g., for treating cancer, e.g., metastatic hepatocellular carcinoma), secondary lymphoid tissue chemokine (SLC, e.g., for treating cancer, e.g., liver cancer), Nk4 (the N-terminal hairpin and subsequent four kringle domains of hepatocyte growth factor (HGF), e.g., for treating cancer, e.g., for treating cervical cancer), tumor necrosis factor superfamily member 14 (TNFSF14, also known as LIGHT, for example, for treating cancer, such as cervical cancer), granulocyte-macrophage colony-stimulating factor (GM-CSF, for example, for treating cancer), TNF-α (for example, for treating cancer, such as glioma), dominant negative mutants of survivin (e.g., C84A or T34A, for example, for treating cancer, such as colon or stomach cancer), C-terminal fragment of human telomerase reverse transcriptase (hTERTC27, for example, for treating cancer, such as glioblastoma multiforme), maspin (for example, for treating cancer, such as prostate cancer), nm23H1 (e.g., cancer, e.g., metastatic ovarian cancer), kringle 1 domain of human hepatocyte growth factor (HGFK1, e.g., for treating cancer, e.g., colorectal cancer), anti-calcitonin ribozyme (e.g., for treating cancer, e.g., prostate cancer), eukaryotic translation initiation factor 4E binding protein 1 (4EBP1, e.g., for treating cancer, e.g., lung cancer), C-X-C motif chemokine receptor 2 (CXCR2) C-tail sequence (e.g., for treating cancer, e.g., pancreatic cancer), alpha-tocopherol-associated protein (TAP, e.g., for treating cancer, e.g., prostate cancer), trichosanthin (e.g., for treating cancer, e.g., hepatocellular carcinoma), decorin (e.g., for treating cancer, e.g., glioblastoma multiforme), cathelicidin (e.g., cancer,for treating colon cancer), Niemann-Pick type C2 (NPC2, for treating cancer, for example hepatocellular carcinoma), Mullerian inhibitory substance (MIS, for treating cancer, for example ovarian cancer), p53 (for treating cancer, for example bronchoalveolar carcinoma), shRNA highly expressed in cancer 1 (Hec1, for treating cancer, for example glioma), shRNA against Epstein-Barr virus latent membrane protein-1 (EBV LMP-1, for treating cancer, for example nasopharyngeal carcinoma), antisense RNA against human papillomavirus 16 E7 oncogene (HPV16-E7, for treating cancer, for example cervical cancer), shRNA against androgen receptor (AR, for treating cancer, for example prostate cancer), siRNA against Snail (also known as SNA1, for treating cancer, for example pancreatic cancer), siRNA against Slug (i.e., the protein product of SNAI2, e.g., for treating cancer, e.g., cholangiocarcinoma (liver cancer)), shRNA against Four and a half LIM-only protein 2 (FHL2, e.g., for treating cancer, e.g., colon cancer), miR-26a (e.g., for treating cancer, e.g., hepatocellular carcinoma), HPV16 structural protein L1 (HPV16-L1, e.g., for treating cancer, e.g., cervical cancer), HPV16 E5, E6, and E7 oncogenes (HPV16 E5 / E6 / E7, e.g., for treating cancer, e.g., cervical cancer), B-cell leukemia / lymphoma 1 (BLC1) idiotype (e.g., for treating cancer, e.g., B-cell leukemia / lymphoma 1), EBV LMP1 and LMP2 fused to heat shock proteins (EBV LMP2 / 1-hsp, for example for treating cancer, such as nasopharyngeal carcinoma), carcinoembryonic antigen (CEA, for example for treating cancer, such as colon cancer), soluble forms of B cell and T lymphocyte attenuator in combination with heat shock proteins (BTLA and HSP70, for example for treating cancer, such as melanoma lung metastasis), HPV16-L1 / E7 (for example for treating cancer, such as cervical cancer), HPV16-L1 (for example for treating cancer, such as cervical cancer),anti-EGFR antibodies (14D1, e.g., for treating cancer, e.g., vulvar cancer), anti-death receptor 5 (DR5) antibodies (adximab, e.g., for treating cancer, e.g., liver or colon cancer), anti-enolase 1 (ENOI1) antibodies (e.g., for treating cancer, e.g., pancreatic ductal adenocarcinoma), anti-VEGFA antibodies (bevacizumab, e.g., for treating cancer, e.g., metastatic lung cancer or ovarian cancer), mucin 1 (MUC1) antigen (e.g., for treating cancer, e.g., gastric cancer), or Apollins (eg, hAQP1, eg, for treating radiation-induced parotid hyposalivation, ie, xerostomia).
[0157] In some aspects, the nucleic acid sequence of interest is for use in gene editing (e.g., gene therapy, including treatment of a genetic defect, disorder, or disease).
[0158] In some aspects, the nucleic acid sequence of interest is intended to be inserted into the target site for gene editing (i.e., the site in the DNA or RNA sequence that is the target of gene editing).The target site for gene editing includes any genetic element, such as any cis-element.In some aspects, the target site for gene editing is located in the exon of gene, the intron of gene, or the regulatory element of gene.
[0159] In some aspects, gene editing comprises endonuclease.In some aspects, endonuclease is associated with genome editing system.In some aspects, endonuclease is for example, homing endonuclease, site-specific nuclease, structure-guided nuclease or RNA-guided nuclease (for example, transposon-encoded RNA-guided nuclease).
[0160] In some aspects, gene editing involves a genome editing system that generates a double-strand break within the target site for gene editing. In some aspects, the genome editing system is a CRISPR-Cas, TALEN, ZFN, or meganuclease gene editing system.
[0161] In some aspects, the nucleic acid sequence of interest is inserted into the target site for gene editing by non-homologous end joining at double-strand break.In some aspects, the double-strand break is generated by CRISPR-Cas system.In some aspects, the expression vector described herein comprises Cas endonuclease target sequence (i.e., the sequence homologous to gRNA target sequence) located between the first and second target sequences for the first recombinase and the nucleic acid sequence of interest (i.e., between the 5' super sequence and the nucleic acid sequence of interest, and between the 3' super sequence and the nucleic acid sequence of interest), and wherein the target site for gene editing (for example, the target site in chromosome) comprises the same Cas endonuclease target sequence. For example, processing of Cas endonuclease target sequences adjacent to nucleic acid sequences in a bacterial sequence-free vector (e.g., msDNA) generated from an expression vector results in the removal of the supersequence and instead renders a linear, covalently closed bacterial sequence-free vector, such as msDNA, linear and open-ended with reactive ends that are amenable to non-homologous end joining events.
[0162] In one aspect, a nucleic acid sequence of interest is inserted into a target site for gene editing by homology-directed repair, which occurs via recombination between sequences adjacent to the double-stranded break and homologous sequences related to the nucleic acid sequence of interest.
[0163] In one aspect, the nucleic acid sequence of interest has sufficient homology to the sequences adjacent to the double-strand break to support homology-directed repair.
[0164] In one aspect, the nucleic acid sequence of interest is flanked by 5' and 3' homology arms (ie, sequences with sufficient homology to sequences adjacent to the double-strand break to mediate homology-directed repair).
[0165] In some aspects, sufficient homology to mediate homology-directed repair comprises at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100% homology between the nucleic acid sequence of interest and the sequence adjacent to the double-strand break, or between the homology arms adjacent to the nucleic acid sequence of interest and the sequence adjacent to the double-strand break. In some aspects, the sequence adjacent to the double-strand break is within about 100 bases, about 90 bases, about 80 bases, about 70 bases, about 60 bases, about 50 bases, about 45 bases, about 40 bases, about 35 bases, about 30 bases, about 25 bases, about 20 bases, about 15 bases, about 10 bases, or about 5 bases of the double-strand break, or is located immediately on both sides of the double-strand break.
[0166] In some aspects, the homology-directed repair is via a CRISPR-Cas system. In some aspects, the expression vector described herein comprises a CRISPR-Cas system. In some aspects, the expression vector comprises a tRNA-gRNA polycistron flanking each side of a sequence encoding a Cas endonuclease (e.g., immunosilenced Cas9-beta2). An exemplary aspect is shown in Figure 49. In some aspects, the expression vector comprises a 5'UTR (e.g., 5'UTR1) described herein that comprises a tRNA-gRNA polycistron in an intron. In some aspects, the expression vector comprises a chimeric intron described herein that comprises a tRNA-gRNA polycistron. In some aspects, the EF1-α promoter described herein comprises a tRNA-gRNA polycistron in a unique intron. In some aspects, the polyadenylation signal or 3'UTR described herein comprises a tRNA-gRNA polycistron. For example, upon expression of the Cas endonuclease from a vector containing an adjacent tRNA-gRNA polycistron (i.e., an expression vector or a bacterial sequence-free vector (e.g., msDNA)), the gRNA is excised as free RNA, targeting the Cas endonuclease to the target site for gene editing (e.g., a target site in a chromosome) and the adjacent gRNA site on the vector. This results in self-restriction of the Cas endonuclease from the vector, limiting further expression of the Cas endonuclease. A schematic of this process is shown in Figure 50, which also illustrates mediation of homology-directed repair by a nucleic acid of interest (i.e., a gene of interest, GOI) flanked by homology arms on another vector. In certain aspects, the expression vectors described herein contain a nucleic acid sequence of interest flanked by homology arms, for example, as shown in scenario 1 of Figure 51. In certain aspects, the nucleic acid sequence of interest and the self-limiting CRISPR-Cas system described herein are located on a single expression vector described herein, as shown in scenario 2 of Figure 51. In the latter aspect, the sequence containing the self-limiting CRISPR-Cas system is located 5' to the sequence containing the nucleic acid sequence of interest, flanked by arms of homology.
[0167] In some aspects, the nucleic acid sequence of interest is homologous to the target site for gene editing and contains one or more nucleotide insertions, deletions, inversions, or rearrangements compared to the target site. In some aspects, the nucleic acid sequence of interest is a genomic sequence, coding region, exon, intron, or any portion thereof that replaces the homologous sequence at the target site.
[0168] In some aspects, the nucleic acid sequence of interest is heterologous to the target site for gene editing.
[0169] In some aspects, the nucleic acid sequence of interest restores a lost function, corrects an abnormal function, or provides an additional function associated with the target site for gene editing.
[0170] In one aspect, the nucleic acid sequence of interest is for knockout of gene expression (i.e., gene silencing) associated with the target site for gene editing.
[0171] In one aspect, the nucleic acid sequence of interest is for in vivo gene editing.
[0172] In one aspect, the nucleic acid sequence of interest is for in vitro gene editing.
[0173] In some aspects, the nucleic acid sequence of interest is for ex vivo gene editing (e.g., cell therapy such as chimeric antigen receptor (CAR) T-cell therapy).
[0174] In some aspects, gene editing involves epigenetic modification, and the expression vector described herein includes an epigenetic effector molecule as a nucleic acid sequence of interest. In some aspects, the epigenetic effector molecule mediates, for example, acetylation or deacetylation, methylation or demethylation, or phosphorylation or dephosphorylation. In some aspects, the epigenetic effector molecule inhibits acetylation or deacetylation, methylation or demethylation, or phosphorylation or dephosphorylation. In some aspects, the epigenetic modification is a histone modification. In some aspects, the histone modification is histone acetylation, and the nucleic acid sequence of interest is a histone acetyltransferase. In some aspects, the epigenetic modification is a DNA modification. In some aspects, the DNA modification is DNA methylation, and the nucleic acid sequence of interest is a DNA methylase. In some aspects, the DNA modification is DNA demethylation, and the nucleic acid sequence of interest is a DNA demethylase. In some aspects, the epigenetic effector molecule is fused to a targeting molecule, such as a DNA-binding molecule, to target the effector to a chromosomal location.
[0175] In one aspect, the expression cassette is polygenic, ie, the expression cassette comprises two or more nucleic acid sequences of interest, each encoding two or more polypeptides.
[0176] In one aspect, the expression cassette comprises a single open reading frame containing a nucleic acid sequence encoding a self-cleaving peptide between each nucleic acid sequence encoding a polypeptide, such that the translation product of the expression cassette is cleaved into two or more polypeptides within the cell. In one aspect, the self-cleaving peptide is a 2A self-cleaving peptide. In one aspect, the 2A self-cleaving peptide is P2A derived from porcine teschovirus-1. In one aspect, the 2A self-cleaving peptide is T2A derived from Thesa signavirus 2A. In one aspect, the self-cleaving peptide comprises any one or more of 2A, P2A, and T2A. In one aspect, the self-cleaving peptide comprises P2A and T2A.
[0177] In some aspects, the expression cassette further comprises a nucleic acid sequence encoding a marker for gene expression. In some aspects, the marker for gene expression is a fluorescent reporter gene, such as a green fluorescent protein (GFP, e.g., enhanced GFP (eGFP)), a red fluorescent protein (RFP), a yellow fluorescent protein (YFP), or a near-infrared fluorescent protein (iRFP); a bioluminescent reporter gene such as a luciferase (e.g., nanoluciferase, i.e., NanoLuc) (登録商標) (NLuc), England et al., Bioconjug. Chem. 27(5):1175-1187 (2016), Promega Corporation); a selectable antibiotic marker; or LacZ. In one aspect, the expression cassette includes a nucleic acid sequence encoding a self-cleaving peptide between a nucleic acid sequence encoding a marker for gene expression and any other nucleic acid sequence encoding a polypeptide.
[0178] An expression cassette can include any expression control region known to those skilled in the art operably linked to a nucleic acid sequence of interest. In some aspects, the expression control region is a promoter, enhancer, operator, repressor, ribosome binding site, translation leader sequence, intron, polyadenylation recognition sequence, RNA processing site, effector binding site, stem-loop structure, transcription termination signal, or a combination thereof.
[0179] In some aspects, the expression vector is for producing a vector that does not contain any bacterial sequences. In some aspects, the vector that does not contain any bacterial sequences is a covalently closed circular vector. In some aspects, the vector that does not contain any bacterial sequences is a covalently closed linear vector.
[0180] B. Vector Production System The present invention provides a vector production system comprising a recombinant cell encoding a recombinase under the control of an inducible promoter, wherein the recombinant cell comprises an expression vector described herein comprising first and second target sequences for a first recombinase and one or more additional target sequences for one or more additional recombinases, and wherein the recombinase targets one of the first and second target sequences for the first recombinase or the one or more additional target sequences for the one or more additional recombinases.
[0181] Suitable host cells for use in the vector production system include microbial cells, for example bacterial cells such as E. coli cells, and yeast cells such as S. cerevisiae cells. Mammalian host cells can also be used, including Chinese hamster ovary (CHO) cells (e.g., K1 line (ATCC CCL 61) or Pro5 mutant (ATCC CRL 1281)); SV40-transformed fibroblast-like cells from African green monkey kidney of the CV-1 line (ATCC CCL 70), COS-1 line (ATCC CRL 1650), or COS-7 line (ATCC CRL 1651); mouse L cells; mouse 3T3 cells (ATCC CRL 1658); mouse C127 cells; human embryonic kidney cells of the 293 line (ATCC CRL 1573); human carcinoma cells, including those of the HeLa line (ATCC CCL 2); and neuroblastoma cells of the IMR-32 (ATCC CCL 127), SK-N-MC (ATCC HTB 10), or SK-N-SH (ATCC HTB 11) lines.
[0182] Suitable recombinase enzymes catalyze DNA replacement at the target sequence of the recombinases described herein, including, but not limited to, TelN, Tel, Tel (gp26 K02 phage), Cre, Flp, phiC31, Int, and other λ phage integrases, such as phi 80, HK022, and HP1 recombinase enzymes. In one aspect, the recombinase is TelN, Tel, Cre, or Flp.
[0183] In one aspect, the recombinant cell further encodes an endonuclease under the control of an inducible promoter, wherein the endonuclease targets an endonuclease target sequence in the expression vector.
[0184] A suitable endonuclease cleaves the polynucleotide at the endonuclease target sequence. In one aspect, the endonuclease is a homing endonuclease. In certain aspects, the homing endonuclease is I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI, I-PpoI, PI-PspI. , I-ScaI, I-SceI, PI-SceI, I-SceII, I-SecIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliII, I-Tsp061I or I-Vdi141I. In some aspects, the endonuclease is I-SceI. In some aspects, the endonuclease is PI-SceI. In some aspects, the recombinant cell encodes a nuclease genome editing system comprising the endonuclease. In some aspects, the genome editing system is a CRISPR-Cas, TALEN, ZFN, or meganuclease system. In some aspects, the nuclease genome editing system is a Class 1 or Class 2 CRISPR-Cas system. In some aspects, the nuclease genome editing system is a Type I, II, III, IV, V, or VI CRISPR-Cas system. In some aspects, the Cas endonuclease in the CRISPR-Cas system is Cas9 (e.g., SpCas9, SaCas9, FnCas9, or NmCas9), a Cas9 mutant (e.g., Casβ9, xCas9, SpCas9-NG, SpCas9-NRRH, SpCas9-NRCH, SpCas9-NRTH, SpG, SpRY), Cas3, Cas12 (e.g., Cas12a, Cas12b, Cas12c, Cas12d, or Cas12e), Cas13 (e.g., Cas13a, Cas13b, Cas13c, or Cas13d), or Cas14.
[0185] Recombinant host cells encoding recombinases, or recombinases and endonucleases, are prepared using well-known techniques. For example, a nucleic acid sequence encoding a selected recombinase or endonuclease is introduced into a cell using an appropriate vector under appropriate conditions for cell transformation. Recombinant host cells can be transformed via an expression vector or by integrating a nucleic acid sequence encoding a recombinase and / or an endonuclease into the host cell genome. In embodiments where an endonuclease is associated with a nuclease genome editing system, the host cell can be designed to encode all of the components of the nuclease genome editing system, either by transforming the host cell with one or more expression vectors containing all of the components, by integrating all of the components into the host cell genome, or by a combination of transformation and integration of the components. In some embodiments, the host cell encodes a Cas or Cas-like endonuclease and a gRNA.
[0186] The expression of the recombinase or endonuclease, including the endonuclease of the nuclease genome editing system, is under the control of an inducible promoter, i.e., a promoter that is activated under specific physical or chemical conditions or stimuli. In some aspects, the inducible promoter is thermally regulated, chemically regulated, IPTG regulated, glucose regulated, arabinose induced, T7 polymerase regulated, cold shock induced, pH induced, or a combination thereof.
[0187] The present invention provides a recombinant cell comprising an expression vector described herein, which comprises a first and second target sequence for a first recombinase and one or more additional target sequences for one or more additional recombinases. In some aspects, the recombinant cell encodes the first recombinase and / or one or more additional recombinases described herein. In some aspects, the recombinant cell encodes one or more endonucleases described herein. In some aspects, the recombinant cell encodes the nuclease genome editing system described herein.
[0188] The present invention provides a method for producing a bacterial sequence-free vector, comprising incubating the vector production system described herein under conditions suitable for expression of a recombinase. In some aspects, the method further comprises incubating the vector production system under conditions suitable for expression of an endonuclease encoded by the recombinant cell. In some aspects, the method further comprises incubating the vector production system under conditions suitable for expression of a nuclease genome editing system encoded by the recombinant cell. In some aspects, the method further comprises recovering the bacterial sequence-free vector.
[0189] The present invention provides vectors that do not contain bacterial sequences produced by the methods for producing vectors that do not contain bacterial sequences described herein.
[0190] III. Vectors without bacterial sequences The present invention provides bacterial sequence-free vectors comprising (a) an expression cassette comprising a nucleic acid sequence of interest, and (b) one or more of the following: (i) a synthetic enhancer comprising a nucleic acid sequence at least about 90% identical to SEQ ID NO: 12 located 5' to another enhancer or promoter in the expression cassette; (ii) a CMV enhancer located 5' to the promoter in the expression cassette; (iii) a 5'UTR comprising an intron, wherein the 5'UTR is incorporated between the promoter and the nucleic acid sequence of interest in the expression cassette; (iv) a vertebrate chromatin insulator incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette; (v) a WPRE incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette; (vi) an S / MAR incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette; or (vii) a DTS located 5' to the expression cassette.
[0191] In some aspects, a bacterial sequence-free vector comprises a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12 located 5' to another enhancer or promoter in the expression cassette. In some aspects, a bacterial sequence-free vector comprises a synthetic enhancer comprising the nucleic acid sequence of SEQ ID NO: 12 located 5' to another enhancer or promoter in the expression cassette. In some aspects, the synthetic enhancer comprises multiple contiguous copies of the nucleic acid sequence, e.g., 1, 2, 3, 4, 5, or more contiguous copies. In some aspects, the synthetic enhancer comprises three contiguous copies of the nucleic acid sequence. In one aspect, the synthetic enhancer comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 46. In one aspect, the synthetic enhancer comprises the nucleic acid sequence of SEQ ID NO: 46. In one aspect, the synthetic enhancer is incorporated into the 5' end of a chicken β-actin promoter. In one aspect, a chimeric intron comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 47 is incorporated into the 3' end of the chicken β-actin promoter and 5' of the nucleic acid sequence of interest. In one aspect, a chimeric intron comprising the nucleic acid sequence of SEQ ID NO: 47 is incorporated at the 3' end of the chicken β-actin promoter and 5' of the nucleic acid sequence of interest.
[0192] In one aspect, the bacterial sequence-free vector comprises a CMV enhancer located 5' to the promoter in the expression cassette. In one aspect, the CMV enhancer is incorporated at the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In one aspect, the CMV enhancer is incorporated at the 3' end of multiple contiguous copies of the synthetic enhancer, for example, at the 3' end of one, two, three, four, five, or more contiguous copies of the synthetic enhancer. In one aspect, the CMV enhancer is incorporated at the 3' end of three contiguous copies of the synthetic enhancer. In one aspect, a CMV enhancer is incorporated into the 3' end of a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 46. In one aspect, a CMV enhancer is incorporated into the 3' end of the nucleic acid sequence of SEQ ID NO: 46. In one aspect, a CMV promoter is incorporated into the 3' end of the CMV enhancer and the 5' end of the nucleic acid sequence of interest.
[0193] In one aspect, the bacterial sequence-free vector comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39 located 5' to the nucleic acid sequence of interest. In one aspect, the bacterial sequence-free vector comprises the nucleic acid sequence of SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39 located 5' to the nucleic acid sequence of interest. In one aspect, a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39, or the nucleic acid sequence of SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, or SEQ ID NO:39, includes all regulatory elements in an expression cassette located 5' to the nucleic acid sequence of interest.
[0194] In one aspect, the bacterial sequence-free vector comprises a 5'UTR containing an intron, wherein the 5'UTR (i.e., the 5'UTR containing the intron) is incorporated between the promoter and the nucleic acid sequence of interest in the expression cassette.
[0195] In one aspect, the 5'UTR is for improving transgene transcript splicing and translation from a bacterial sequence-free vector compared to the same bacterial sequence-free vector lacking the 5'UTR.
[0196] In some aspects, the intron comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 1. In some aspects, the intron comprises the nucleic acid sequence of SEQ ID NO: 1.
[0197] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 2. In some aspects, the 5'UTR comprises the nucleic acid sequence of SEQ ID NO: 2.
[0198] In some aspects, the 5' UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 4. In some aspects, the 5' UTR comprises the nucleic acid sequence of SEQ ID NO: 4.
[0199] In some aspects, the 5'UTR further comprises non-coding sequences embedded within the intron.
[0200] In one aspect, the intron is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to or comprises SEQ ID NO:1, and the non-coding sequence is incorporated between two of the nucleotides in the intron corresponding to any two nucleotides at positions 25 to 55 of SEQ ID NO:1.
[0201] In some aspects, the non-coding sequence is non-prokaryotic and non-viral. In some aspects, the non-coding sequence is eukaryotic. In some aspects, the non-coding sequence comprises an intron, a UCOE, an S / MAR, an SV40 enhancer sequence (e.g., one or more SV40 enhancer sequences, e.g., two, three, four, five or more SV40 enhancer sequences), a vertebrate chromatin insulator (e.g., cHS4), a WPRE, or any combination thereof.
[0202] In one aspect, the non-coding sequence comprises an S / MAR. In one aspect, the S / MAR is MAR-5, which is shown in SEQ ID NO: 9 in the present invention.
[0203] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 3. In some aspects, the 5'UTR comprises SEQ ID NO: 3.
[0204] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 5. In some aspects, the 5'UTR comprises SEQ ID NO:5.
[0205] In one aspect, a 5'UTR is incorporated into the expression cassette between the chicken β-actin promoter and the nucleic acid sequence of interest.
[0206] In one aspect, a 5'UTR is incorporated into the expression cassette between the CMV promoter and the nucleic acid sequence of interest.
[0207] In one aspect, a 5'UTR is incorporated into an expression cassette between a promoter and a nucleic acid sequence of interest, wherein the promoter is incorporated into the 3' end of a CMV enhancer. In one aspect, the CMV enhancer is incorporated into the 3' end of a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In one aspect, the CMV enhancer is incorporated into the 3' end of a synthetic enhancer comprising the nucleic acid sequence of SEQ ID NO: 12. In one aspect, the CMV enhancer is incorporated into the 3' end of multiple consecutive copies of the synthetic enhancer, for example, into the 3' end of one, two, three, four, five, or more consecutive copies of the synthetic enhancer. In one aspect, the CMV enhancer is incorporated at the 3' end of three consecutive copies of the synthetic enhancer. In one aspect, the CMV enhancer is incorporated at the 3' end of a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 46. In one aspect, the CMV enhancer is incorporated at the 3' end of the nucleic acid sequence of SEQ ID NO: 46.
[0208] In one aspect, the bacterial sequence-free vector comprises a polyadenylation signal incorporated into the 3' end of the nucleic acid sequence of interest. In one aspect, the polyadenylation signal comprises a Xenopus laevis beta globin polyadenylation signal, a human beta globin polyadenylation signal, or a hybrid Xenopus laevis and human beta globin polyadenylation signal. In one aspect, the polyadenylation signal comprises multiple copies, such as 1 copy, 2 copies, 3 copies, 4 copies, or 5 copies, of the Xenopus laevis beta globin polyadenylation signal, the human beta globin polyadenylation signal, or the hybrid Xenopus laevis and human beta globin polyadenylation signal. In certain aspects, the polyadenylation signal comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15. In certain aspects, the polyadenylation signal comprises the nucleic acid sequence of SEQ ID NO: 13, SEQ ID NO: 14, or SEQ ID NO: 15. In certain aspects, a polyadenylate tail (i.e., a poly(A) tail) is located at the 3' end of the polyadenylation signal. In some aspects, the poly(A) tail is 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120 or more residues in length. In some aspects, the sequence comprising the polyadenylation signal and poly(A) tail is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO:16, SEQ ID NO:17, or SEQ ID NO:18.In one aspect, the sequence comprising the polyadenylation signal and poly(A) tail comprises SEQ ID NO:16, SEQ ID NO:17, or SEQ ID NO:18.
[0209] In one aspect, the bacterial sequence-free vector comprises a vertebrate chromatin insulator in the expression cassette. In one aspect, the vertebrate chromatin insulator is cHS4. In one aspect, the vertebrate chromatin insulator is incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette as described herein. In one aspect, the vertebrate chromatin insulator is incorporated within an intron of the 5'UTR as described herein.
[0210] In some aspects, the vertebrate chromatin insulator comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 8. In some aspects, the vertebrate chromatin insulator comprises SEQ ID NO:8.
[0211] In one aspect, the vertebrate chromatin insulator is for improving colonization (i.e., transfection efficiency) of a vector that does not contain bacterial sequences compared to the same vector that does not contain the vertebrate chromatin insulator.
[0212] In some aspects, the bacterial sequence-free vector comprises a WPRE in the expression cassette. In some aspects, the WPRE is incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette as described herein. In some aspects, the WPRE is incorporated at the 3' end of the S / MAR and the 5' end of the polyadenylation signal in the expression cassette as described herein. In some aspects, the WPRE is incorporated within the intron of the 5'UTR as described herein.
[0213] In one aspect, the WPRE comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to SEQ ID NO:11.
[0214] In one aspect, the WPRE improves expression of a transgene from a vector without bacterial sequences compared to the same vector without bacterial sequences lacking the WPRE.
[0215] In some aspects, the bacterial sequence-free vector contains an S / MAR in the expression cassette. In some aspects, the S / MAR is incorporated between the nucleic acid sequence of interest and the polyadenylation signal in the expression cassette. In some aspects, the S / MAR is incorporated at the 3' end of the nucleic acid sequence of interest and at the 5' end of the WPRE in the expression cassette as described herein. In some aspects, the S / MAR is incorporated within the intron of the 5'UTR as described herein.
[0216] In some aspects, the S / MAR is MAR-3, MAR-4, or MAR-5. In some aspects, the S / MAR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 9. In some aspects, the S / MAR comprises SEQ ID NO: 9.
[0217] In some aspects, the S / MAR is a human CSP-B MAR or CSP-C MAR. In some aspects, the S / MAR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 10. In some aspects, the S / MAR comprises SEQ ID NO: 10.
[0218] In one aspect, the S / MAR is intended to improve the expression level, stability and / or durability (e.g., by episomal maintenance and replication, e.g., propagation and partitioning of the vector to daughter cells, and / or prevention of epigenetic silencing) of the bacterial sequence-free vector compared to the same bacterial sequence-free vector lacking the S / MAR.
[0219] In some aspects, a bacterial sequence-free vector comprising any one or more of (b)(i)-(b)(v) above (i.e., not comprising a DTS) further comprises enhancer sequences flanking the expression cassette. In some aspects, the enhancer sequences flanking the expression cassette are at least two enhancer sequences flanking each strand of the expression cassette. In some aspects, the enhancer sequences are SV40 enhancer sequences.
[0220] In some aspects, the bacterial sequence-free vector comprises a DTS. In some aspects, the DTS is located 5' of the expression cassette. In some aspects, the DTS is an SV40 enhancer sequence. In some aspects, the DTS is cell-specific. In some aspects, the DTS is specific to smooth muscle cells, embryonic stem cells, type II alveolar epithelial cells, endothelial cells, or osteoblasts.
[0221] In some aspects, the bacterial sequence-free vector described herein further comprises a UCOE in the expression cassette. In some aspects, the UCOE is located 5' to the promoter or any enhancer in the expression cassette. In some aspects, the UCOE is integrated into an intron of the 5'UTR described herein.
[0222] In some aspects, the UCOE is an A2UCOE. In some aspects, the UCOE comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 6. In some aspects, the UCOE is SEQ ID NO: 6.
[0223] In some aspects, the UCOE is an SRF-UCOE. In some aspects, the UCOE comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 7. In some embodiments, the UCOE is SEQ ID NO: 7.
[0224] In one aspect, the UCOE improves expression of a transgene from a vector without bacterial sequences compared to the same vector without bacterial sequences lacking the UCOE.
[0225] In some aspects, the bacterial sequence-free vector comprises enhancer-1 in the expression cassette. In some aspects, enhancer-1 is integrated 5' to the promoter or any other enhancer in the expression cassette. In some aspects, enhancer-1 is integrated between the 3' end of the UCOE and the 5' end of the CMV enhancer. In some aspects, enhancer-1 comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In some embodiments, enhancer-1 is SEQ ID NO: 12.
[0226] In some aspects, the vector without bacterial sequences comprises a CMV, EF1, SV40, cag, Rho, VDM2, HCR or HLP promoter, or a variant thereof, in the expression cassette, hi some aspects, the vector without bacterial sequences comprises a CMV promoter variant in the expression cassette.
[0227] In one aspect, the vector without bacterial sequences comprises an EF1-α promoter in the expression cassette. In one aspect, the vector without bacterial sequences comprises a CMV enhancer and an EF1-α promoter in the expression cassette.
[0228] In one aspect, the bacterial sequence-free vector contains a 3'UTR in the expression cassette that contains two copies of the β-globin polyadenylation signal. In one aspect, the 3'UTR is integrated 3' to the nucleic acid sequence of interest.
[0229] In some aspects, the 3'UTR contains two copies of the Xenopus beta-globin polyadenylation signal. In some aspects, the 3'UTR contains a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 13. In some aspects, the 3'UTR is SEQ ID NO: 13.
[0230] In one aspect, the 3'UTR comprises two copies of the human beta-globin polyadenylation signal. In one aspect, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 14. In one aspect, the 3'UTR is SEQ ID NO: 14.
[0231] In some aspects, the 3'UTR comprises one copy of a Xenopus beta-globin polyadenylation signal and one copy of a human beta-globin polyadenylation signal. In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 15. In some embodiments, the 3'UTR is SEQ ID NO: 15.
[0232] In one aspect, the 3'UTR further comprises a poly(A) tail comprising 100 to 120 adenine nucleotides, i.e., 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 adenine nucleotides.
[0233] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 16. In some aspects, the 3'UTR is SEQ ID NO:16.
[0234] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 17. In some aspects, the 3'UTR is SEQ ID NO:17.
[0235] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 18. In some aspects, the 3'UTR is SEQ ID NO:18.
[0236] Nucleic acid sequences of interest for the bacterial sequence-free vectors described herein include any of the nucleic acid sequences described herein with respect to expression vectors for generating bacterial sequence-free vectors.
[0237] In certain aspects, the bacterial sequence-free vectors described herein comprise Cas endonuclease target sequences (i.e., sequences homologous to a gRNA targeting sequence) located 5' and 3' to a nucleic acid sequence of interest, where the target site for gene editing (e.g., a target site in a chromosome) comprises the same Cas endonuclease target sequence.
[0238] In some aspects, the vectors described herein that do not contain bacterial sequences comprise a CRISPR-Cas system. In some aspects, the vectors described herein that do not contain bacterial sequences comprise a tRNA-gRNA polycistron flanked on each side by sequences encoding a Cas endonuclease (e.g., immunosilenced Cas9-β2). In some aspects, the vectors described herein that do not contain bacterial sequences comprise a 5'UTR (e.g., 5'UTR1) that comprises a tRNA-gRNA polycistron in an intron. In some aspects, the vectors described herein that do not contain bacterial sequences comprise a chimeric intron that comprises a tRNA-gRNA polycistron. In some aspects, the EF1-alpha promoter described herein comprises a tRNA-gRNA polycistron in a unique intron. In some aspects, the polyadenylation signal or 3'UTR described herein comprises a tRNA-gRNA polycistron. In some aspects, the nucleic acid sequence of interest and the self-limiting CRISPR-Cas system described herein are located on a single expression vector described herein. In the latter aspect, the sequence containing the self-limiting CRISPR-Cas system is located 5' to the sequence containing the nucleic acid sequence of interest, flanked by arms of homology.
[0239] The bacterial sequence-free vectors described herein can include any combination of the above modifications. In some aspects, the combination provides a synergistic effect.
[0240] In one aspect, the bacterial sequence-free vector is a covalently closed circle vector.
[0241] In one aspect, the bacterial sequence-free vector is a covalently closed linear vector.
[0242] The bacterial sequence-free vector is a circular, covalently closed vector.
[0243] The present invention provides recombinant cells comprising the bacterial sequence-free vectors described herein.
[0244] IV. Other Expression Vectors The above improvements and modifications can also be applied to other expression vectors, such as, but not limited to, expression vectors used for direct gene expression, rather than producing vectors that do not contain bacterial sequences. In some aspects, the nucleic acid sequences described herein are provided as DNA sequences, and the expression vectors are DNA expression vectors. In some aspects, the nucleic acid sequences described herein are provided as RNA sequences, and the expression vectors are RNA expression vectors. The RNA sequences can correspond to any DNA sequence provided herein as a SEQ ID NO:, or can correspond to a DNA sequence complementary to any DNA sequence provided herein as a SEQ ID NO:.
[0245] The present invention provides polynucleotides comprising any combination of the nucleic acid sequences described herein.
[0246] The present invention provides polynucleotides comprising the nucleic acid sequences of the introns, 5'UTRs containing introns, and / or 3'UTRs described herein.
[0247] The present invention provides polynucleotides comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of SEQ ID NOs: 1, 2, 3, 5, 13, 14, 15, 16, 17, or 18. In some aspects, the polynucleotide comprises 100-120 adenine nucleotides at the 3' end of the nucleic acid sequence. In some aspects, the polynucleotide comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of SEQ ID NOs: 13, 14, or 15, as well as 100-120 adenine nucleotides at the 3' end of the nucleic acid sequence. In one aspect, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 1, 2, 3, 5, 13, 14, 15, 16, 17 or 18.
[0248] The present invention provides expression vectors comprising one or more polynucleotides described herein. In one aspect, the expression vector comprises a polynucleotide comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of SEQ ID NOs: 1, 2, 3, 5, 13, 14, 15, 16, 17, or 18. In one aspect, the polynucleotide comprises 100 to 120 adenine nucleotides at the 3' end of the nucleic acid sequence. In one aspect, the polynucleotide comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to any one of SEQ ID NOs: 13, 14, or 15, and 100 to 120 adenine nucleotides at the 3' end of the nucleic acid sequence. In one aspect, the expression vector comprises a polynucleotide comprising the nucleic acid sequence of any one of SEQ ID NOs: 1, 2, 3, 5, 13, 14, 15, 16, 17, or 18. In one aspect, the expression vector comprises a polynucleotide comprising the nucleic acid sequence of any one of SEQ ID NOs: 2, 3, or 5, and (a) a polynucleotide comprising the nucleic acid sequence of any one of SEQ ID NOs: 13, 14, 15, 16, 17, or 18, or (b) a polynucleotide comprising the nucleic acid sequence of any one of SEQ ID NOs: 13, 14, or 15 and 100 to 120 adenine nucleotides at the 3' end of the nucleic acid sequence.
[0249] The present invention provides an expression vector comprising a 5'UTR containing an intron, the 5'UTR being incorporated between a promoter and a nucleic acid sequence of interest in an expression cassette, and / or a 3'UTR containing two copies of a beta-globin polyadenylation signal incorporated 3' to the nucleic acid sequence of interest in the expression cassette.
[0250] In some aspects, the 5' UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 2. In some aspects, the 5' UTR comprises the nucleic acid sequence of SEQ ID NO: 2.
[0251] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 4. In some aspects, the 5'UTR comprises the nucleic acid sequence of SEQ ID NO:4.
[0252] In one aspect, the 5'UTR further comprises non-coding sequences embedded within the intron.
[0253] In one aspect, the intron is at least about 90%, at least about 91%, at least about 92%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to or comprises SEQ ID NO:1, and the non-coding sequence is incorporated between two nucleotides within the intron corresponding to any two nucleotides at positions 25 to 55 of SEQ ID NO:1.
[0254] In some aspects, the non-coding sequence is non-prokaryotic and non-viral. In some aspects, the non-coding sequence is a eukaryotic sequence. In some aspects, the non-coding sequence includes an intron, a UCOE, an S / MAR, an SV40 enhancer sequence (e.g., one or more SV40 enhancer sequences, e.g., two, three, four, five or more SV40 enhancer sequences), a vertebrate chromatin insulator (e.g., cHS4), a WPRE, or any combination thereof.
[0255] In one aspect, the non-coding sequence comprises an S / MAR. In one aspect, the S / MAR is MAR-5, which is shown in SEQ ID NO: 9 in the present invention.
[0256] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 3. In some aspects, the 5'UTR comprises SEQ ID NO: 3.
[0257] In some aspects, the 5'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 5. In some aspects, the 5'UTR comprises SEQ ID NO:5.
[0258] In one aspect, the 3'UTR comprises two copies of a Xenopus beta-globin polyadenylation signal. In one aspect, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 13. In one aspect, the 3'UTR is SEQ ID NO: 13.
[0259] In some aspects, the 3'UTR contains two copies of the human beta-globin polyadenylation signal. In some aspects, the 3'UTR contains a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 14. In some aspects, the 3'UTR is SEQ ID NO: 14.
[0260] In one aspect, the 3'UTR comprises one copy of a Xenopus beta-globin polyadenylation signal and one copy of a human beta-globin polyadenylation signal. In one aspect, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 15. In one aspect, the 3'UTR is SEQ ID NO: 15.
[0261] In one aspect, the 3'UTR further comprises a poly(A) tail comprising 100 to 120 adenine nucleotides, i.e., 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, or 120 adenine nucleotides.
[0262] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 16. In some aspects, the 3'UTR is SEQ ID NO:16.
[0263] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 17. In some aspects, the 3'UTR is SEQ ID NO:17.
[0264] In some aspects, the 3'UTR comprises a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 18. In some aspects, the 3'UTR is SEQ ID NO:18.
[0265] The present invention provides an expression vector comprising a synthetic enhancer comprising a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 12. In some aspects, the expression vector comprises a synthetic enhancer comprising the nucleic acid sequence of SEQ ID NO: 12. In some aspects, the synthetic enhancer comprises multiple contiguous copies of the nucleic acid sequence, e.g., 1, 2, 3, 4, 5, or more contiguous copies. In some aspects, the synthetic enhancer comprises three contiguous copies of the nucleic acid sequence. In some aspects, the synthetic enhancer comprises a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 46. In one aspect, the synthetic enhancer comprises the nucleic acid sequence of SEQ ID NO:46.
[0266] The present invention provides an expression vector comprising a nucleic acid sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39. In one aspect, the expression vector comprises the nucleic acid sequence of SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39. In one aspect, a nucleic acid sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% identical to SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38 or SEQ ID NO:39, or the nucleic acid sequence of SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38 or SEQ ID NO:39, comprises all regulatory elements in an expression cassette located 5' to the nucleic acid sequence of interest in an expression vector.
[0267] V. Composition The present invention provides compositions comprising the expression vectors or bacterial sequence-free vectors described herein.
[0268] Various methods are known in the art and suitable for introducing nucleic acids into cells, including, but not limited to, electroporation, calcium phosphate-mediated transfection, nucleofection, sonication, heat shock, magnetofection, liposome-mediated transfection, microinjection, microprojectile-mediated transfection (nanoparticles), cationic polymer-mediated transfection (DEAE-dextran, polyethyleneimine, polyethylene glycol (PEG), etc.), or cell fusion.
[0269] Nanoparticle carriers, such as liposomes, micelles, and polymeric nanoparticles, have been explored to improve the bioavailability and pharmacokinetic properties of therapeutics through various mechanisms, such as the enhanced vascular permeability and retention (EPR) effect.
[0270] Further improvements can be achieved by conjugating targeting ligands onto nanoparticles to achieve selective delivery to target cells. For example, receptor-targeted nanoparticle delivery has been shown to improve therapeutic responses both in vitro and in vivo. Targeting ligands that have been investigated include folate, transferrin, antibodies, peptides, and aptamers. Furthermore, multiple functionalities can be incorporated into the design of nanoparticles, for example, to enable imaging and to trigger intracellular drug release.
[0271] In some aspects, the composition further comprises a delivery agent. In some aspects, the delivery agent is a nanoparticle. In some aspects, the delivery agent is selected from the group consisting of a liposome, a non-lipid polymer, an endosome, and any combination thereof.
[0272] In some aspects, the delivery agent (eg, nanoparticle) comprises a targeting ligand.
[0273] In some aspects, the composition further comprises a physiologically acceptable carrier, excipient, or stabilizer. See, e.g., Remington: The Science and Practice of Pharmacy, 22 nd ed. (2013). Acceptable carriers, excipients, or stabilizers may include those that are non-toxic to the subject. In some aspects, the composition or one or more components of the composition are sterile. Sterile components can be prepared, for example, by filtration (e.g., through sterile filtration membranes) or by irradiation (e.g., by gamma irradiation).
[0274] In some aspects, a composition comprising an expression vector or a bacterial sequence-free vector described herein is a pharmaceutical composition that further comprises a pharmaceutically acceptable carrier.
[0275] The excipients of the present invention, when added to pharmaceutical compositions, can be described as "pharmaceutically acceptable" excipients, meaning that the excipients are compounds, materials, compositions, salts, and / or dosage forms that, within the scope of sound medical judgment, are suitable for contact with human or animal tissues without undue toxicity, irritation, allergic response, or other undesirable complications for the desired period of contact commensurate with a reasonable benefit / risk ratio. In some aspects, the term "pharmaceutically acceptable" means approved by a federal or state regulatory agency or listed in the United States Pharmacopoeia or other generally recognized international pharmacopoeias for use in animals, more particularly humans. A variety of excipients can be used. In some aspects, the excipient can be, but is not limited to, an alkalizing agent, a stabilizer, an antioxidant, an adhesive agent, a separating agent, a coating agent, an external phase component, a controlled-release component, a solvent, a surfactant, a wetting agent, a buffering agent, a bulking agent, an emollient, or a combination thereof. Excipients, in addition to those discussed herein, include, but are not limited to, those listed in Remington: The Science and Practice of Pharmacy, 22 nd ed. (2013). The inclusion of an excipient in a particular category (e.g., "solvent") herein is intended to illustrate, rather than limit, the role of the excipient. A particular excipient may be included in multiple categories.
[0276] The pharmaceutical compositions of the present invention are formulated to be compatible with their intended route of administration. Examples of routes of administration include enteral, topical, parenteral, oral, pulmonary, intranasal, intravenous, epidermal, transdermal, subcutaneous, intramuscular, or intraperitoneal administration, or inhalation. As used herein, "parenteral administration" refers to modes of administration other than enteral and topical administration, usually by injection or infusion, including, but not limited to, intravenous, intramuscular, intraarterial, intrathecal, intralymphatic, intralesional, intraarticular, intraorbital, intracardiac, intradermal, intraperitoneal, transtracheal, subcutaneous, subcuticular, interarticular, subcapsular, subarachnoid, intrathecal, epidural, intrapleural, and intrasternal injection and infusion, and in vivo electroporation. In some aspects, the formulation is administered enterally, and in some aspects, orally. Other enteral routes include topical, epidermal, or mucosal routes of administration, such as intranasal, vaginal, rectal, sublingual, or topical.
[0277] In some aspects, the pharmaceutical composition is lyophilized.
[0278] VI. Therapeutic Uses and Methods The present invention provides methods of treating a disease or disorder in a subject in need thereof, comprising administering to the subject an expression vector, bacterial sequence-free vector, or pharmaceutical composition described herein.
[0279] The expression vector, bacterial sequence-free vector or composition can be administered to a subject by any route of administration that is effective in treating the disease or disorder.
[0280] In certain aspects, the administering step is enteral, topical, parenteral, oral, pulmonary, intranasal, intravenous, epidermal, transdermal, subcutaneous, intramuscular, intrathecal or intraperitoneal administration, or by inhalation, or by cerebrospinal fluid (CSF)-based delivery via intracerebroventricular (ICV) injection, cisterna magna administration (ICM) or lumbar puncture (LIT).
[0281] In some aspects, the administering step is performed by parenteral administration or enteral administration.
[0282] In one aspect, parenteral administration is by injection or infusion.
[0283] In certain aspects, parenteral administration is by intravenous, intramuscular, intraarterial, intrathecal, intralymphatic, intralesional, intraarticular, intraorbital, intracardiac, intradermal, intraperitoneal, transtracheal, subcutaneous, subcuticular, interarticular, subcapsular, subarachnoid, intrathecal, epidural, intrapleural, or intrasternal injection or infusion, or by in vivo electroporation, nucleofection, microbubbles, or ultrasound.
[0284] In some aspects, enteral administration is oral, topical, epidermal, transmucosal, intranasal, vaginal, rectal, or sublingual administration.
[0285] In some aspects, the administering step is performed by oral, pulmonary, intranasal, intravenous, epidermal, transdermal, subcutaneous, intramuscular, or intraperitoneal injection, or by inhalation.
[0286] In some aspects, the administering step is performed by oral, nasal, or pulmonary administration. In some aspects, the administering step is performed by intranasal administration.
[0287] The administration step can be carried out, for example, once, multiple times, and / or over one or more extended periods of time. In some aspects, the administration step is once, twice (e.g., a first administration followed by a second administration about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, or later), about once every week, about once every month, about once every 2 months, about once every 3 months, about once every 4 months, about once every 6 months, about once every year, or about once every 10 years.
[0288] The present invention provides methods for gene editing, comprising inserting a nucleic acid sequence of interest from an expression vector, bacterial sequence-free vector, or pharmaceutical composition described herein into a target site for gene editing.
[0289] In one aspect, the inserting step is performed by non-homologous end joining.
[0290] In some aspects, the insertion step is performed by homology directed repair. In some aspects, the nucleic acid sequence of interest is flanked by 5' and 3' homology arms as described herein.
[0291] In one aspect, the nucleic acid sequence of interest is homologous to a target site for gene editing and contains one or more nucleotide insertions, deletions, inversions, or rearrangements compared to the target site.
[0292] In some aspects, the nucleic acid sequence of interest is heterologous to the target site for gene editing.
[0293] In some aspects, the nucleic acid sequence of interest restores a lost function, corrects an abnormal function, or provides an additional function associated with the target site for gene editing.
[0294] In one aspect, the nucleic acid sequence of interest is for knockout of gene expression associated with the target site for gene editing.
[0295] In one aspect, the method of gene editing is a method of treating a disease or disorder in a subject in need thereof.
[0296] In one aspect, the nucleic acid sequence of interest is for in vivo gene editing.
[0297] In one aspect, the nucleic acid sequence of interest is for in vitro gene editing.
[0298] In some aspects, the nucleic acid sequence of interest is for ex vivo gene editing (e.g., cell therapy, e.g., CAR T cell therapy).
[0299] In some aspects, this method is an in vitro method.In some aspects, the in vitro method further comprises administering to cells an expression vector, a vector that does not contain bacterial sequences, or a pharmaceutical composition (for example, for in vitro or ex vivo gene editing).In some aspects, the in vitro method further comprises administering to cells an endonuclease for gene editing, or a genome editing system or its components (for example, the Cas endonuclease and gRNA for CRISPR-Cas system).In some aspects, the genome editing system is CRISPR-Cas, TALEN, ZFN, or meganuclease gene editing system.
[0300] In some aspects, this method is an in vivo method.In some aspects, the in vivo method further comprises administering to the subject an expression vector, a vector that does not contain bacterial sequence, or a pharmaceutical composition.In some aspects, the in vivo method further comprises administering to the subject an endonuclease for gene editing, or a gene editing system or its components (for example, the Cas endonuclease and gRNA for the CRISPR-Cas system).In some aspects, the gene editing system is CRISPR-Cas, TALEN, ZFN or meganuclease gene editing system.
[0301] The gene editing endonuclease or genome editing system or its components can be administered by any method described herein or any method known in the art for administering nucleic acid sequences and / or polypeptides to cells or subjects, including via electroporation or a vector adapted for administration. For example, in aspects involving a CRISPR-Cas system, RNA encoding the Cas and / or gRNA can be administered, the Cas and / or gRNA can be administered directly, a bacterial sequence-free vector or expression vector described herein encoding the Cas and / or gRNA can be administered, or any other suitable vector known in the art encoding the Cas and / or gRNA can be administered.
[0302] In some aspects, the nucleic acid sequence of interest is provided in a linear, covalently closed, bacterial sequence-free vector (i.e., msDNA) as described herein. In some aspects, the use of a linear, covalently closed, bacterial sequence-free vector in gene editing avoids undesired non-homologous end joining because the ends of the bacterial sequence-free vector are closed and non-reactive with double-strand breaks. In some aspects, the use of a linear, covalently closed, bacterial sequence-free vector in gene editing enhances homology-directed repair. In some aspects, the recombination rate for homology-directed repair is higher when the nucleic acid sequence of interest is provided by a linear, covalently closed, bacterial sequence-free vector as described herein than when it is provided by a circular supercoiled vector.
[0303] The following examples are offered to illustrate, but not to limit, the present invention. [Example]
[0304] Example Example 1 - Expression vectors containing chimeric introns or 5'UTRs A. Expression Vectors The multigene expression vector was the parent ministring expression vector (Mediphage Bioceuticals, Inc., Toronto, CA, U.S. Pat. Nos. 9,290,778 and 9,862,954), which was modified with the eGFP coding sequence of pGL2-SS*-CAG-eGFP-BGpA-SS* and the NanoLuc vector modified with an enhanced green fluorescent protein (eGFP) and a secretion sequence for extracellular expression. (登録商標) An expression cassette encoding a luciferase reporter (NLuc, Promega Corporation) was prepared by substituting between two specialized supersequence (SS*) sites in the parent vector.
[0305] The expression cassettes of the parental and multigene vectors contained a cytomegalovirus (CMV) enhancer, a promoter derived from chicken β-actin, and the CAG promoter, a synthetic promoter containing a chimeric intron.
[0306] A map of the multigene expression vector is shown in Figure 1 (pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*), which contains recombinase target sequences (telL, FRT(minimal), and loxP) flanking a multigene expression cassette containing a CAG promoter, a sequence encoding enhanced green fluorescent protein (eGFP) and secreted nanoluciferase (SecNLuc) bound by P2A and T2A self-cleaving peptides (SecNLuc-2A-eGFP), and a specific supersequence site (SS*) with a rabbit β-globin polyadenylation signal (BGpA). The nucleic acid sequence of the vector is provided as SEQ ID NO: 19.
[0307] A second multigene expression vector was prepared by cloning the same eGFP and Nluc sequences along with the 5'UTR into the pcDNA3.1 vector (Thermo Fisher Scientific). A map of the expression vector, containing a multigene expression cassette containing a CMV enhancer / promoter, sequences encoding eGFP and SecNLuc linked by a P2A self-cleaving peptide (SecNLuc-P2A-eGFP), and a bovine growth hormone polyadenylation signal (bGHpA), is shown in Figure 2 (vector pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA). The nucleic acid sequence of the vector is provided as SEQ ID NO: 20.
[0308] B. Transfection of HEK293 cells Adherent human embryonic kidney 293 (HEK293) cells were cultured at 1 × 10 5 Cells / well were seeded into 24-well plates.
[0309] Complexes of expression vector (1 μg) and Lipofectamine (3 μL) were prepared and incubated using standard operating procedures for (1) pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*, (2) pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA, and (3) pGL2-SS*-CAG-eGFP-BGpA-SS*, respectively.
[0310] HEK293 cells were transfected separately in individual wells with the three complexes by electroporation and then incubated for 48 hours. HEK293 cells in other wells were treated with 3 μL of Lipofectamine without plasmid as a negative control.
[0311] Cells were assessed for cytoplasmic GFP and luciferase expression 48 hours after transfection.
[0312] C. Cytoplasmic GFP expression Cytoplasmic GFP expression was used as a measure of transfection efficiency and the level of gene expression by the multigene expression vector. Expression was assessed by fluorescence microscopy, and the average GFP expression / intensity of the experimental expression vectors (pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* and pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA) was measured relative to the negative control (cells treated with Lipofectamine but without the plasmid) and the positive control (pGL2-SS*-CAG-eGFP-BGpA-SS*) (also referred to herein as the parental plasmid CAG-GFP, i.e., PP-CAG-GFP).
[0313] Live imaging of fluorescent cells under autoexposure mode showed that the chimeric introns of pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* and pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA had similar expression, indicating that the experimental expression vectors produced GFP. See Figures 3 and 4. Multigene expression by the experimental expression vectors did not affect GFP expression based on the mean relative fluorescence intensity compared to the positive control (Id). The mean fluorescence intensity in cells transfected with the experimental expression vectors was at least three-fold higher than that in negative control cells. See Figure 4.
[0314] D. Luciferase Expression Luciferase expression was measured using Nano-Glo (登録商標) The intensity of secreted luciferase in the medium of transfected and negative control cells was assessed using a luciferase assay system (Promega) according to the manufacturer's protocol. Both experimental expression vectors expressed luciferase. See Figure 5. The average relative luciferase intensity in the medium of cells transfected with the experimental expression vectors was at least 300-fold higher than that in the medium of negative control cells (Id).
[0315] Example 2 - Expression vector containing WPRE and engineered 5'UTR A. Expression Vectors A multigene expression vector was prepared by cloning the woodchuck hepatitis virus posttranscriptional regulatory element (WPRE) between the sequences encoding eGFP and BGpA in the expression vector of Figure 1. A map of the resulting expression vector is shown in Figure 6 (pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS*). The nucleic acid sequence of the vector is provided as SEQ ID NO:21.
[0316] Another multigene expression vector was prepared containing an engineered 5'UTR containing a CMV enhancer / promoter and internal minimal intron sequence (i.e., 5'UTR1, SEQ ID NO: 2) in place of the CAG promoter in Figure 6. A map of the resulting expression vector is shown in Figure 7 (pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*). The nucleic acid sequence of the vector is provided as SEQ ID NO: 22.
[0317] An additional multigene expression vector was prepared containing an engineered 5'UTR (i.e., 5'UTR2, SEQ ID NO: 5) containing a CMV enhancer / promoter and an intron with an integrated MAR-5 instead of the CAG promoter of Figure 6. A map of the resulting expression vector is shown in Figure 8 (pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS*). The nucleic acid sequence of the vector is provided as SEQ ID NO: 23.
[0318] B. Luciferase Expression Levels and Persistence Adherent HEK293 cells were detached and analyzed in electroporation medium. 1 x 10 6 Cells / tube were counted.
[0319] Expression vectors (1 μg) were prepared and incubated with cells using standard operating procedures for each of (1) pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (see Example 1), (2) pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, (3) pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, and (4) pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS*.
[0320] HEK293 cells electroporated with the puc57 plasmid lacking the mammalian expression cassette served as a negative control.
[0321] After electroporation, HEK293 cells were seeded at 3 x 10 5 Cells / well were allowed to attach to the wells.
[0322] Luciferase expression was measured 2, 6, 10, 14, 17, 20, 27, and 34 days after electroporation using Nano-Glo™ according to the manufacturer's protocol. (登録商標) The luciferase activity was assessed by measuring the intensity of secreted luciferase in 20 μL of cell culture medium in triplicate for each of the four transfections and the negative control using a luciferase assay system (Promega). (登録商標) Luciferase activity was measured using a plate reader and expressed as relative luminometer units (RLU). Statistical analysis of luciferase activity was performed by Student's T-test. See Figure 9, which shows the expression levels in the medium from cells transfected with pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (pGL2-SecNLuc-eGFP), pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (WPRE), pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (5'UTR1+WPRE), or pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (5'UTR2+WPRE) compared with the negative control (Neg. Ctl. (no plasmid)). *=p<0.05, **=p<0.01, ***=p<0.001 and ****=p<0.0001.
[0323] Luciferase expression was detected throughout the experiment in cells transfected with any of the four expression vectors. pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, and pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS* all showed significantly higher luciferase expression than pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*, with pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* showing the highest expression enhancement.
[0324] C. Vector propagation to daughter cells and luciferase expression HEK293 cells were transfected with the four expression vectors or the puc57 plasmid as a negative control, as described in Part B of this Example. Cells were passaged five times weekly. At the time of cell passage, cells were replated at 1 / 6 of the original cell density for passages 1-3 and at 1 / 10 of the original cell density for passages 4-5. For each cell passage, secreted luciferase expression was measured 6-8 days after cell replated, as described in Part B of this Example. See Figure 10, which shows the expression levels in the medium from vector-transfected cells compared to the negative control at each passage number. Statistical analysis and p-values were as described in Part B of this Example.
[0325] Luciferase expression was detected in cells transfected with any of the four expression vectors at each passage, indicating that the vector was inherited by daughter cells with sustained expression of luciferase. pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, and pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS* all showed significantly higher luciferase expression at each passage compared with pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*, with pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* showing the highest expression enhancement.
[0326] Subsequent studies also observed propagation of msDNA to daughter cells accompanied by persistent expression of luciferase.
[0327] Briefly, msDNA was produced from pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* in an inducible E. coli vector production system using the methods described herein and in U.S. Patent Nos. 9,290,778 and 9,862,954. Separate complexes with Lipofectamine were prepared using (1) msDNA (i.e., msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA), (2) the parental plasmid (i.e., pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*), and (3) a conventional plasmid carrying a luciferase expression cassette (i.e., pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA). HEK293 cells were transfected separately in individual wells by electroporation with a total of 0.25 pmol of vector / well. Cells were passaged seven times, with a 10-fold cell dilution at each passage. Relative luciferase intensities were measured on days 8, 15, 24, 31, 38, 45, and 52 for passages 1, 2, 3, 4, 5, 6, and 7, respectively.
[0328] As shown in Figure 11, pGL2-SecNLuc*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (i.e., pDNA(CMV+U1+W)) and msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (i.e., msDNA(CMV+U1+W)) demonstrated sustained transgene expression at much higher levels across all passage numbers than pcDNA-CMV-5'UTR-SS-P2A-eGFP-bGhPa (i.e., conventional plasmid without a supergene (conventional pcDNA)). **=p<0.01, ***=p<0.001, and ****=p<0.0001 compared to the conventional plasmid dataset.
[0329] D. Vector propagation to daughter cells and eGFP expression Cells transfected with pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (pGL2-SecNLuc-eGFP) or pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (5'UTR1+WPRE) and passaged according to part C of this example were analyzed for eGFP expression.
[0330] Imaging was performed 6–8 days after each subculture. BioTek (登録商標) Cytation (商標) Live-cell imaging was performed using a plate reader. See Figure 12A, which shows representative micrographs of fluorescence in HEK-293 cells at passages 1, 2, 3, and 5. eGFP expression was detected in cells transfected with either expression vector at each passage, indicating that the vector was inherited by daughter cells with persistent expression of eGFP. pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* showed stronger fluorescent signals at each passage compared with pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*, indicating higher transfection efficiency.
[0331] One representative image was taken from each triplicate well for each expression vector, and eGFP-expressing cells (GFP+) were quantified by manual cell counting using ImageJ computer software. + Statistical analysis of cells was performed using Student's t-test. See Figure 12B, which shows the GFP counts observed within the field of view from triplicate raw fluorescent images at each passage number. + Line graphs of cells are shown (not significant (ns) = p>0.05 due to variance of triplicate images).
[0332] Using ImageJ software, each GFP +Cells were manually selected, and the mean fluorescence intensity (MFI) of each cell was measured based on pixel intensity. The following formula was used to calculate the final MFI value for each cell: final MFI = cell MFI - background MFI. MFI measurements were obtained for at least 50 cells from each of the three images taken for each treatment group. All MFI measurements were then pooled and used to generate dot plots. Statistical analysis was performed using Student's t-test. See Figure 12C, which shows a dot plot of MFI at passage 5; n = 257 cells for pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* and n = 414 cells for pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*; **** = p < 0.0001. The bar graph at the bottom of Figure 12C represents the percentage of all GFP cells measured. + The average MFI values for the cells are shown. The MFI of pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* was measured to be approximately three times higher than that of pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*.
[0333] Example 4 - Non-viral delivery of msDNA in animal models A study was conducted to evaluate targeted delivery of msDNA to the liver, retina, and brain. For each target tissue, different routes of administration (ROA), doses, dosing regimens, and delivery techniques were evaluated. Secreted luciferase expression kinetics, cytoplasmic eGFP expression levels, and transfection efficiency (TE) were assessed. Furthermore, tolerability to msDNA was assessed after single and multiple injections by physiological assessment, histomorphometric analysis, plasma cytokine assay, and hepatotoxicity analysis.
[0334] Across all delivery techniques, msDNA demonstrated a strong efficacy and tolerability profile in brain and liver tissues following multiple intracerebroventricular (ICV) or hydrodynamic injection (HDI) and intravenous (IV) injections, respectively. Adult mice treated with msDNA demonstrated sustained secreted luciferase levels (>10 s) after a single IV injection. 8 The results showed a significant increase in RLU / mg protein. msDNA showed sustained (>100 days) expression in liver tissue after a single IV injection. Significant biodistribution to deep tissue regions was also demonstrated, with 80%-97% TE observed in the brainstem, cerebellum, cortex, and thalamus. Triple ICV injections using nanocarrier-msDNA complexes showed no side effects.
[0335] A. Liver Luciferase expression by hydrodynamic injection of a single high dose of 1.2 mg / kg (50 μg) of carrier-free intact plasmid Eight- to 12-week-old C57BL / 6J male wild-type adult mice were administered a high dose of 2 mg / kg (50 μg) of pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (a positive control without a supersequence), pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*, pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, or pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* via hydrodynamic tail vein injection (HDI) via the tail vein. Plasma samples were collected from treated mice on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after HDI to examine luciferase gene expression. One day after vector administration, all mice showed high levels of luciferase expression (10 8 ~10 9 On day 7 after vector administration, pCAGLuc and pCAGLuc-WPRE treated mice showed 10 7pGSNLuc-WPRE-treated mice showed low levels of luciferase expression (approximately 10 6 Eight weeks after vector administration, all mice showed low levels of luciferase expression (approximately 10 RLU / mg protein). 5 The rapid decline in luciferase levels may be due to the humoral or cellular immune response induced in the plasmid-treated mice (see Figures 13-14).
[0336] 2. Expression of luciferase by hydrodynamic injection of a single low dose of 0.2 mg / kg (5 μg) of carrier-free intact plasmid. To test the dose response of plasmid DNA after non-viral gene delivery in an animal model, C57BL / 6J male wild-type adult mice, 8–12 weeks old, were administered a low dose of 0.2 mg / kg (5 μg) of carrier-free pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control without supersequence, 2 mice), pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* (2 mice), or pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* (2 mice) via the tail vein via HDI. Two additional mice were not injected and served as negative controls. Plasma was collected from the mice on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after HDI administration, and luciferase gene expression was examined. Mice treated with pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS* and pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* showed sustained high-level luciferase expression (10 ) for more than 8 weeks after vector administration. 7 ~10 8 RLU / mg protein), demonstrating over 100-fold higher expression than a conventional control plasmid carrying the isogenic expression cassette but without the supersequence (SS). See Figure 15.
[0337] In vivo whole-body bioluminescence imaging (BLI) using IVIS was performed by intraperitoneal injection of a 1:5 dilution of fluorofurimazine (FFz) 24 h after HDI of the vector. BLI was shown to correlate with the level of luceriferase in plasma samples (data not shown).
[0338] 3. Luciferase expression by a single low-dose hydrodynamic injection of 0.2 mg / kg (5 μg) of carrier-free intact msDNA MSDNA was generated from pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* and pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* in an inducible E. coli vector production system using methods described herein and in U.S. Pat. Nos. 9,290,778 and 9,862,954.
[0339] Eight to 12-week-old C57BL / 6J male wild-type adult mice were administered a single 0.2 mg / kg (5 μg) low dose of carrier-free pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control, 5 mice), msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (5 mice), or msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (5 mice) via hydrodynamic tail vein injection (HDI). Two additional mice were not injected and served as negative controls. Plasma from treated mice was collected on days 1, 3, 7, 10, 15, 22, 28, 42, and 56 after HDI administration, and luciferase gene expression was examined.
[0340] Similar to the plasmid-treated mice, the msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-treated mice showed sustained high levels of luciferase expression (10 s) for more than 8 weeks after vector administration. 7 ~10 8Although luciferase expression in msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-treated mice yielded 10 RLU / mg protein, it remained at a low level (approximately 10 RLU / mg protein) for less than 1 month. 6 The rapid decrease in luciferase expression in msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-treated mice was attributed to silencing of the CMV promoter in hepatocytes.
[0341] Luciferase gene expression was confirmed by whole-body live imaging using IVIS.
[0342] Table 1 below provides data from individual mice on days 1, 7 and 28 as detected by luciferase expression (RLU / mg protein) and BLI (photons / sec) in plasma samples.
[0343] [Table 1]
[0344] As shown in Figure 16, mice treated with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA showed a 10-fold increase in luciferase expression compared to the parental plasmid pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* at 56 days after HDI.
[0345] This data demonstrates that non-viral delivery of msDNA in mice is highly efficient and the resulting gene expression was stable for more than two months.
[0346] 4. Expression of eGFP from a single low dose of 0.2 mg / kg (5 μg), hydrodynamic injection of carrier-free, intact msDNA Intracellular cytoplasmic eGFP expression levels were assessed by ELISA. Briefly, liver samples were collected from mice 56 days after HDI with a single low dose of 0.2 mg / kg (5 μg) of vector, as described in Part 3, and homogenized for protein extraction. Total protein concentrations were measured from liver lysates, and GFP protein levels were analyzed by ELISA.
[0347] As is evident by comparing the data in Figure 17 with the luciferase data, the level of cytoplasmic GFP expression directly correlated with the level of luciferase secretion from the same construct.
[0348] As shown in Figure 18, a single HDI tail vein injection of 5 μg of carrier-free msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BgpA demonstrated strong and sustained cytoplasmic eGFP expression in liver tissue for at least 56 days after HDI.
[0349] 5. Expression and tolerance of msDNA in the liver following a single low-dose intravenous administration C57BL / 6J male wild-type adult 8-12 week-old mice were administered a single intravenous tail vein injection of 0.3 mg / kg of pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control without supersequence), msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA, pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA, or pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* lipoplexed with lipid nanoparticle carriers. The carrier was also used as a negative vehicle control.
[0350] In vivo whole-body bioluminescence imaging (BLI) using IVIS was performed as described above on days 1, 3, 10, 30, 58, 92, 119, and 174 after a single IV injection of the vector. As shown in Figure 19, the msDNA constructs exhibited high and sustained luciferase secretion, superior to the precursor plasmid and conventional plasmids.
[0351] Serum alanine aminotransferase (ALT) levels, hepatotoxicity, and cytokine responses were also assessed after vector injection. The precursor plasmid and msDNA containing the cag promoter demonstrated a well-tolerated profile compared to constructs containing the CMV promoter. However, msDNA containing the CMV promoter demonstrated significantly lower cytokine and hepatotoxic responses compared to the CMV precursor parental plasmid and conventional plasmids. See Table 2 below, which shows cytokine concentrations (pg / mL) and enzyme concentrations (U / L) of liver function markers at 4 hours and 14 days post-injection.
[0352] [Table 2]
[0353] B. Brain Using the methods described herein and in U.S. Pat. Nos. 9,290,778 and 9,862,954, msDNA was produced from pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* and pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS* in an inducible E. coli vector production system.
[0354] Adult wild-type mice were administered 1 μg of DNA via intracerebroventricular (ICV) injection three times via the implanted cannula with either msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (3 mice) or msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (3 mice) prepared with nanocarriers on days 0, 14, and 28 after transplantation. The animals were euthanized on day 42 after transplantation, and sagittal brain sections were collected from the cortex, thalamus, brainstem, and cerebellum.
[0355] Figure 20 shows sections of the cortex, thalamus, brainstem, and cerebellum from mouse #1 in the treatment group injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA. The transfection efficiency of msDNA in the sections was measured to be 81.9%, 73.0%, 69.2%, and 96.0% in the cortex, thalamus, brainstem, and cerebellum (Purkinje cells), respectively. Transfection efficiency was calculated as the percentage of cells positive for both GFP and DAPI among all DAPI-positive cells.
[0356] Comparison of GFP expression in cortex, thalamus, and brainstem sections from mice injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA and mice injected with the control plasmid pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA showed that transfection efficiency and resulting GFP expression were higher with msDNA compared with conventional plasmids (data not shown).
[0357] Figure 21 shows sections from the cortex and thalamus, and Figure 22 shows sections from the brainstem and cerebellum from mouse #2 in the treatment group injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA. Neurons were labeled with the neuronal marker NeuN, and transfected cells were shown to express GFP. Transfection efficiencies were determined to be 99.6%, 98.8%, 98.5%, and 80.8% in the cortex, thalamus, brainstem, and cerebellum (Purkinje cells), respectively. Transfection efficiency was calculated as the percentage of cells positive for both GFP and NeuN among all NeuN-positive cells.
[0358] Figure 23 shows sections from the cortex and thalamus, and Figure 24 shows sections from the brainstem and cerebellum from mouse #1 in the treatment group injected with msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA. Neurons were labeled with the neuronal marker NeuN, and transfected cells were shown to express GFP. Transfection efficiencies were measured to be 91.1%, 88.8%, 73.7%, and 92.1% in the cortex, thalamus, brainstem, and cerebellum (Purkinje cells), respectively. Transfection efficiency was calculated as the percentage of cells positive for both GFP and NeuN among all NeuN-positive cells.
[0359] Table 3 below summarizes the transfection efficiencies discussed above for mice injected with msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA (“CAG-WPRE”) or msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA (“CMV-WPRE”).
[0360] [Table 3]
[0361] Repeated injections of ICV through the implanted cannula resulted in good overall tissue integrity without signs of cytotoxicity or neurodegeneration.
[0362] The data demonstrate that msDNA is redosable and resulted in high transfection efficiency, biodistribution, and transgene expression in multiple brain regions without adverse morphological effects.
[0363] Example 5 - Efficacy and safety in human cells pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (positive control), msDNA-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA, pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*, msDNA-CAG-SecNLuc-2A-eGFP-WPRE-BGpA, or pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* was lipoplexed with lipid nanoparticle carriers.
[0364] Human T cells (Pan-T(TA+)) and hepatocytes (Huh7) were transfected with lipoplexed vectors at a dose of 0.3 μg / mL or 1 μg / mL.
[0365] The lipoplexed msDNA vectors showed higher expression in both cell types compared to the parental and conventional plasmids at days 3 and 5 post-transfection. See Figures 25 and 26.
[0366] Lipoplexed msDNA was also well tolerated in human peripheral blood mononuclear cells (PBMCs) ex vivo. Notably, msDNA exhibited significantly lower cytokine profile levels in human PBMCs compared with conventional plasmids (data not shown).
[0367] Example 6 - Homology-directed repair in msDNA Studies were performed to evaluate homology-directed repair mediated by msDNA compared with conventional plasmid DNA.
[0368] A conventional plasmid was generated with an expression cassette containing the gene of interest (GOI) flanked by 5' and 3' homology arms (plasmid DNA HDR-GOI-HDR).
[0369] The msDNA expression vector was constructed using the same HDR-GOI-HDR sequence as used in the conventional plasmid, flanked by two supersequence sites. msDNA containing HDR-GOI-HDR (msDNA HDR-GOI-HDR) was then produced in an inducible E. coli vector production system using the methods described herein and in U.S. Patent Nos. 9,290,778 and 9,862,954.
[0370] Induced pluripotent stem cells (iPSCs) were transfected with equimolar concentrations of either plasmid DNA HDR-GOI-HDR or msDNA HDR-GOI-HDR together with the CRISPR gene editing system to mediate homology-directed repair knock-in (HDR KI) of the GOI.
[0371] Homology-directed repair knock-in (HDR KI) efficiency of the GOI was assessed by fluorescence-activated cell sorting (FACS) by counting the total number of integrated, healthy iPSCs expressing the GOI on their surface compared to the total number of transfected cells at days 3, 7, and 15 after transfection. As shown in Figures 27B, 28B, and Q3 of 29A, HDR KI efficiencies were 8.60%, 7.76%, and 8.05%, respectively, for conventional plasmids at days 3, 7, and 15 after transfection. Higher HDR KI efficiencies of 15.4%, 15.4%, and 15.7%, respectively, were observed for msDNA at days 3, 7, and 15 after transfection, as shown in Figures 27C, 28C, and Q3 of 29B.
[0372] Example 7 - Expression vectors containing regulatory sequence modifications A. Expression Vectors Expression vectors containing two supersequence sites, a CMV enhancer / promoter, an engineered 5' UTR containing an internal minimal intron sequence, and a multigenic expression cassette encoding eGFP and Nluc as described in Examples 1 and 2 were also designed to contain a 3' UTR containing a human β-globin polyadenylation signal and two copies of 120 adenine nucleotides (i.e., 2huBGpA-A120, SEQ ID NO: 17), and one or more of: (1) a synthetic enhancer (i.e., enhancer-1 (E1), SEQ ID NO: 12) located at the 5' end of the CMV enhancer; (2) a WPRE located at the 5' end of the 3' UTR; (3) an SRF-UCOE located at the 3' end of the 5' supersequence; and (4) a human CSP-B MAR (huMAR) located at the 3' end of eGFP. Maps of the designed vectors are shown in Figures 30 through 38. Figure 30 shows a map of an expression vector containing the 3'UTR (SS*-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 24). Figure 31 shows a map of an expression vector containing E1 and the 3'UTR (SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 25). Figure 32 shows a map of an expression vector containing E1, WPRE, and the 3'UTR (SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 26). Figure 33 shows a map of an expression vector containing UCOE, E1, WPRE, and 3'UTR (SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 27). Figure 34 shows a map of an expression vector containing E1, huMAR, and 3'UTR (SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 28).Figure 35 shows a map of an expression vector containing UCOE, E1, huMAR, and a 3'UTR (SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 29). Figure 36 shows a map of an expression vector containing UCOE, E1, WPRE, and a 3'UTR (SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 30). Figure 37 shows a map of an expression vector containing E1, huMAR, WPRE, and 3'UTR (SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-WPRE-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 31). Figure 38 shows a map of an expression vector containing UCOE, E1, huMAR, WPRE, and 3'UTR (SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-WPRE-3'UTR[2hBGpA-A120]-SS*, SEQ ID NO: 32).
[0373] B. Luciferase expression levels HEK293 cells were transfected separately with (1) the conventional plasmid pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA as shown in Figure 2, (2) SS*-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS*, (3) S*-E1-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS*, and (4) SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS* using standard operating procedures.
[0374] On days 2, 3, 7, 10, 14, 21, and 28 after electroporation, luciferase expression was assessed by measuring the intensity of luciferase secreted from the culture medium of the cultured cells as described in Example 2B. See Figure 39, which shows expression levels in medium from cells transfected with pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA (conventional pDNA CMV-U), SS*-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS* (A: CMV-U1-3'UTR), SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2hBGpA-A120]-SS* (B: E1-CMV-U1-3'UTR), and SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2hBGpA-A120]-SS* (C: E1-CMV-U1-WPRE-3'UTR).
[0375] As shown in Figure 39, luciferase expression was increased and sustained in the msDNA expression vector containing the 3'UTR compared to the conventional plasmid with the same promoter and multigene expression cassette. Expression was further increased upon addition of the E1 (A vs. B) and WPRE (B vs. C) gene elements to the msDNA expression vector, with the additive effect of the E1 and WPRE gene elements resulting in the highest luciferase expression associated with construct C. Statistical analysis was performed using a t-test. *=p<0.05 and **=p<0.01.
[0376] C. Vector propagation to daughter cells and luciferase expression HEK293 cells were transfected with the four expression vectors described in Part B of this Example. Cells were passaged five times every seven days. At cell passage, cells were replated at 1 / 10 of the original cell density. For each cell passage, secreted luciferase expression was measured as described in Example 2B. See Figure 40, which shows the expression level in the medium from cells transfected with the msDNA expression vector compared to conventional plasmids. Statistical analysis and p-values were as described in Part B of this Example.
[0377] Luciferase expression was detected from cells transfected with either of the msDNA expression vectors at each passage number, indicating that the vector was inherited by daughter cells with persistent expression of luciferase.
[0378] As shown in Figure 40, the additive effect of the E1 and WPRE gene elements resulted in the highest luciferase expression associated with construct C at each passage number and the most sustained expression after multiple passages.
[0379] Example 8 - Expression vectors containing synthetic promoter sequences A. Expression Vectors Five synthetic promoter sequences were generated: (1) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a CAG containing a chimeric intron [E1X3 + CBA promoter + intron] (SEQ ID NO: 35), (2) E2 (U100), a chicken β-actin promoter, and a CAG containing a chimeric intron [E2 + CBA promoter + intron] (SEQ ID NO: 36), (3) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a chimeric intron [E2 + CBA promoter + intron] (SEQ ID NO: 36), (4) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a chimeric intron [E2 + CBA promoter + intron] (SEQ ID NO: 36), (5) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a chimeric intron [E2 + CBA promoter + intron] (SEQ ID NO: 36), (6) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a chimeric intron [E2 + CBA promoter + intron] (SEQ ID NO: 36), (7) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a chimeric intron [E2 + CBA promoter + intron] (SEQ ID NO: 36), (8) three copies of the synthetic enhancer E1 (i.e., three copies of SEQ ID NO: 12), a chicken β-actin promoter, and a chimeric intron [E2 + C - E1, CAG [E1X3 + CBA promoter + UTR1] (SEQ ID NO: 37) including chicken β-actin promoter and 5'UTR1, (4) E2 (U100), CAG [E2 (U100) + CBA promoter + UTR1] (SEQ ID NO: 38) including chicken β-actin promoter and 5'UTR1, and (5) CMV enhancer-EF1-UTR1 (SEQ ID NO: 39) including CMV enhancer, EF1a short promoter, and 5'UTR1.
[0380] A conventional plasmid containing a CMV enhancer, chicken β-actin promoter, and chimeric intron, and a multigene expression cassette encoding eGFP and Nluc was constructed as described in Examples 1 and 2. A map of the conventional plasmid is shown in Figure 41 (pGL2-CAG-SecNLuc-2A-eGFP-WPRE-bGlobin polyA, SEQ ID NO: 34).
[0381] We constructed an msDNA expression vector containing two supersequence sites, a CMV enhancer, a chicken β-actin promoter, a chimeric intron, a multigene expression cassette encoding eGFP and Nluc, a WPRE, and a 3'UTR. The vector map is shown in Figure 42 (4-1 pGL2-SS*-CAG[CMV enhancer + CBA promoter + intron]-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS*, SEQ ID NO: 40).
[0382] Figure 43 (4-2 pGL2-SS*-CAG [E1 X3 + CBA promoter + intron] -SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA) -SS*, SEQ ID NO: 41), Figure 44 (4-3 pGL2-SS*-CAG [E2 (U100) + CBA promoter + intron] -SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA) -SS*, SEQ ID NO: 42), Figure 45 (4-4 pGL2-SS*-CAG[E2(U100)+CBA promoter+UTR1]-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS*, SEQ ID NO: 43), Figure 46 (4-5-pGL2-SS*-CAG[E2(U100)+CBA promoter+UTR1]-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS*, SEQ ID NO: 44), and Figure 47 (4-6-pGL2-SS*-CMV enhancer-EF1-UTR1-SecNLuc-2 Using the respective vector maps shown in Figure 1, five msDNA expression vectors were constructed containing two supersequence sites, a multigenic expression cassette encoding eGFP and Nluc, a WPRE, and a 3'UTR, along with one of the synthetic promoters (1) to (5) described above.
[0383] B. Luciferase expression levels 1 x 10 HEK293 cells in a 24-well plate 5 Cells / well were seeded and separately transfected with complexes of Lipofectamine and the vector described in Part A of this Example at 0.25 pmol DNA / well. Secreted luciferase expression was measured 3 and 6 days after transfection as described in Example 2B. See Figure 48, which shows the expression levels in the medium from cells transfected with the msDNA expression vector compared to the conventional plasmid. Statistical analysis was performed using two-way ANOVA compared to the conventional plasmid. *=p<0.05, **=p<0.01, and ****=p<0.0001.
[0384] Luciferase expression levels were higher in all msDNA expression vectors compared to conventional plasmids. The highest expression was observed with 4-6-pGL2-SS*-CMV enhancer-EF1-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA (4-6:CMV-EF1-UTR1-W-3'UTR)), which contains the EF-1 promoter element in combination with the CMV enhancer and 5'UTR1.
[0385] Example 9 - SS and Expression Cassette Modifications The impact of modifications to the supersequences (SS) and expression cassettes of the expression vectors described herein is evaluated with respect to transfection efficiency, expression of nucleic acid sequences of interest (including reporter genes such as the multigene GFP and luciferase expression cassettes described in Examples 1 and 2), and vector durability / propagation in dividing cells (including rapidly and slowly dividing cells). Modifications to the SS are also evaluated for restriction enzyme activity at these sites.
[0386] Modifications include, but are not limited to, an endonuclease target sequence incorporated into the non-binding region of the recombinase in the SS between the vector backbone and the recombinase cleavage site, a CAG promoter incorporated between the 3' end of the first target sequence for the first recombinase (i.e., the 3' end of the 5' SS) and the 5' of the promoter in the expression cassette, a CMV enhancer incorporated between the 3' end of the first target sequence for the first recombinase (i.e., the 3' end of the 5' SS) and the 5' of the promoter in the expression cassette, an enhancer-1 sequence located 5' of the CMV enhancer and / or 3' of the UCOE, a CMV, EF1, SV40, CAG, Rho, VDM2, HCR or HLP promoter or variants thereof, a CMV promoter variant, an EF1-α promoter, a synthetic promoter, an intron incorporated into the expression cassette between the promoter and the nucleic acid sequence of interest with or without a non-coding sequence incorporated within the intron. These include individual modifications and combinations such as a 5' UTR containing a 5' UTR (e.g., a 5' UTR containing any one of the nucleic acid sequences of SEQ ID NOs: 2-5), a vertebrate chromatin insulator incorporated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal, a woodchuck hepatitis virus post-transcriptional regulatory element incorporated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal, a scaffold / matrix attachment region incorporated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal, a ubiquitous chromatin opening element located 5' of the promoter in the expression cassette (e.g., 3' of the 5' SS and before any other sequences in the expression cassette), a 3' UTR incorporated into the expression cassette between the nucleic acid of interest and the 3' SS, such as immediately after a stop codon (e.g., a 3' UTR containing any one of the nucleic acid sequences of SEQ ID NOs: 13-16), and / or a poly(A) tail containing 100-120 adenine nucleotides (e.g., at the 3' end of the 3' UTR).
[0387] array Sequence number 1: Artificial intron gtaagtcgacgggccgggcctgggccgggtccgggccgggtcgttggatccccactacagcccgatactcaagcttgacgaattcgagtatccaaggtagtggactagtgtgacgctgctgacccctttcttttcccttctgcag SEQ ID NO: 2 5'UTR1 ctgccttctccctcctgtgagtttggtaagtcgacggggccgggcctgggccgggtccgggccgggtcgttggatccccactacagcccgatactc aagcttgacgaattcgagtatccaaggtagtggactagtgtgacgctgctgacccctttctttcccttctgcaggttggtgtacagtagcttcca SEQ ID NO: 3 5'UTR1 with MAR-5 insertion ctgccttctccctcctgtgagtttggtaagtcgacggggccgggcctgggccgggtccgggccgggtatccatagctgattggtctaaaatgagata catcaacgctcctccatgttttttgttttctttttaaatgaaaaactttattttttaagaggagtttcaggttcatagcaaaattgagaggaaggt acattcaagctgaggaagttttcctctattcctagtttactgagagattgcatcatgaatgggtgttaaattttgtcaaatgctttttctgtgtctatcaatatgaccatgtgattttcttctttaacctgttgatgggacaaattacgttaattgattttcaaacgttgaaccacccttacatatctggaat aaattctacttggttgtggtgtatattttttgatacattcttggattctttttgctaatattttgttgaaaatgtttgtatctttgttcatgagag atattggtctgttgttttcttttcttgtaatgtcattttctagttccggtattaaggtaatgctggcctagttgaatgatttaggaagtattccctc tgcttctgtcttctgaaagagattgtagaaagttgatacaatttttttttctttaaatatcttgatagtcgttggatccccactacagcccgatac tcaagcttgacgaattcgagtatccaaggtagtggactagtgtgacgctgctgacccctttctttcccttctgcaggttggtgtacagtagcttcca SEQ ID NO: 4 5'UTR attgggatcttcacacagcaggtaaggttgcgggccgggcctgggccgggtccgggccgggccgcactgaccctggtgttgcttttttttttaggccgcaagctgaagcgtgtcc SEQ ID NO:5 5'UTR2 (5'UTR of SEQ ID NO:4 with MAR-5 insertion) attgggatcttcacacagcaggtaaggttgcgggccggggcctgggccgggtccgggccgggtatccatagctgattggtctaaaatgagatacatcaacgctcctccatgtttttgttttctttttaaatgaaaaactttattttttaagaggatttcaggtcatagcaaa attgagaggaaggtacattcaagctgaggaagttttcccttattcctagtttactgagagattgcatcatgaatgggtgttaaattttgtcaaatgcttttctgtgtctatcaatatgaccatgtgattttcttctttaacctgttgatgggacaaattacgttaattgatttt caaacgttgaaccacccttacatatctggaataaattctacttggttgtggtgtatattttttgatacattcttggattctttttgctaatattttgttgaaaatgtttgtatctttgttcatgagagatatggtctgttgttttcttttcttgtaatgtcattttctagttc cggtattaaggataatgctggcctagttgaatgatttaggaagtattccctctgcttctgtcttctgaaagagattgtagaaagttgatacaattttttttctttaaatatcttgatagccgcactgacccctggtgttgcttttttttttaggccgcaagctgaagcgtgtcc அக்கிய்குக்குக்கு6 A2UCOE element Array number 7 SRF-UCOE gcacacgaccacaattccactgaaagcattttaatacggaacttgtcactcccagggagcctccgctcagccggcagttggttcatttcaatccccacgacaacccttcaaagtgcagggcagacagcaggtggctctgcccaggcgcctggatcacagcccggcctgcagccctcacctgggcgcggggagaccctgaggacgctcctccaggcggcgctggccggggcctgcggacacggacgggcgggctgagctccgggacccctccccgcgccccgcaccccgcaccccgcaccccgcaccccgcacccggcgctcacccgtcccagccccgccgcccgcagccccagctgcaacgcagccaccgccgccatcgcacccggccccgcgggcgcttccgggacgcaggaggcatctgcatccggggcgccgctgagtcccgcccagagccccgcccccggctccaggttctgcgagcggcttccgccgggctgctccgcgggcgcgtcggccatgagcgagttgccgggcgacgtgcgggcgtttctgcgggagcacccgagcctgcggctccagacggacgcccgcaaggttcgcagcgcgggaggggaacggagtggcggagaagggcgcagttgggatgaggggctgaggggagggcagggga gaggagagggcaggggagaggggagaggggagagcaggagagaggggaaggcaggggagagggcgcggcgggatcaggggaggagagggaa Array number 8 cHS4 insulator ggggagctcacggggacagcccccccccaaagcccccagggatgtaattacgtccctcccccgctagggggcagcagcgaccgcccggggctccgctccggtccggcgctccccccgcatcccgagccggcagcgtgcggggacagcccgggcacggggaaggtggcacgggatcgctttcctctgaacgcttctcgctgctctttgagcctgcagacacctgggggatacggggaaaaggggagctcacggggacagcccccccccaaagcccccagggatgtaattacgtccctcccccgctagggggcagcagcgaccgcccggggctccgctccggtccggcgctccccccgcatcccgagccggcagcgtgcggggacagcccgggcacggggaaggtggcacgggatcgctttcctctgaacgcttctcgctgctctttgagcctgcagacacctgggggatacggggaaaa SEQ ID NO: 9 MAR-5 tatccatagctgattggtctaaaatgagatacatcaacgctcctccatgttttttgttttctttttaaatgaaaaactttattttttaagaggagtttcaggttcatagcaaaattgagaggaaggtacattcaagctgaggaagttttcctctattcctagtttactgagagattgcatcatgaatgggtgttaaattttgtcaaatgctttttctgtgtctatcaatatgaccatgtgattttcttctttaacctgttgatgggacaaattacgttaattgattttcaaacgttgaaccacccttacatatctggaataaattctacttggttgtggtgtatattttttgatacattcttggattctttttgctaatattttgttgaaaatgtttgtatctttgttcatgagagatattggtctgttgttttcttttcttgtaatgtcattttctagttccggtattaaggtaatgctggcctagttgaatgatttaggaagtattccctctgcttctgtcttctgaaagagattgtagaaagttgatacaatttttttttctttaaatatcttgatag
[0388] SEQ ID NO: 10 Human CSP-B MAR (huMAR) SEQ ID NO: 11 WPRE Tcgacaatcaacctctggattacaaaatttgtgaaagattgactggtattcttaactatgttgctccttttacgctatgtggatacgctgctttaatgcctttgtatcatgctattgcttcccgtatggctttcattttctcctccttg tataaatcctggttgctgtctctttatgaggagttgtggcccgttgtcaggcaacgtggcgtggtgtgcactgtgtttgctgacgcaacccccactggttggggcattgccaccacctgtcagctcctttccgggactttcgctttccc cctccctattgccacggcggaactcatcgccgcctgccttgcccgctgctggacaggggctcggctgttgggcactgacaattccgtggtgttgtcggggaagctgacgtcctttccatggctgctcgcctgtgttgccacctggattc tgcgcgggacgtccttctgctacgtcccttcggccctcaatccagcggaccttccttcccgcggcctgctgccggctctgcggcctcttccgcgtcttcgccttcgccctcagacgagtcggatctccctttgggccgcctccccgcctg SEQ ID NO: 12 Enhancer-1 gggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcac SEQ ID NO: 13 Two copies of Xenopus beta globulin polyadenylation signal (2xlBGpA) aaccagcctcaagaacacccgaatggagtctctaagctacataataccaacttacactttacaaaatgttgtcccccaaaatgtagccattcgtatctgctcctaataaaaagaaagtttcttcac aaccagcctcaagaacacccgaatggagtctctaagctacataataccaacttacactttacaaaatgttgtcccccaaaatgtagccattcgtatctgctcctaataaaaagaaagtttcttcac SEQ ID NO: 14 2 copies of human beta globulin polyadenylation signal (2huBGpA) gctcgctttcttgctgtccaatttctattaaaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttattttcattgcaa gctcgctttcttgctgtccaatttctattaaaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttattttcattgcaa SEQ ID NO: 15 Hybrid Xenopus and human beta globulin polyadenylation signal (xlhuBGpA) aaccagcctcaagaacacccgaatggagtctctaagctacataataccaacttacactttacaaaatgttgtcccccaaaatgtagccattcgtatctgctcctaataaaaagaaagtttcttcacgctc gctttcttgctgtccaatttctattaaaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttattttcattgcaa SEQ ID NO: 16 2xlBGpA-A120 aaccagcctcaagaacacccgaatggagtctctaagctacataataccaacttacactttacaaaatgttgtcccccaaaatgtagccattcgtatctgctcctaataaaaagaaagtttcttcacaaccagcctcaagaacacccgaatggagtctctaagctacataataccaacttacacttt acaaaatgttgtcccccaaaatgtagccattcgtatctgctcctaataaaaagaaagtttcttcacaaaaaaaaaaaaaaaaaaaaaaaaaaa aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa SEQ ID NO: 17 2huBGpA-A120 gctcgctttcttgctgtccaatttctattaaaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttattttcattgcaa gctcgctttcttgctgtccaatttctattaaaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttattttcattgcaa aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa SEQ ID NO: 18 xlhuBGpA-A120 aaccagcctcaagaacacccgaatggagtctctaagctacataataccaacttacactttacaaaatgttgtcccccaaaatgtagccattcgtatctgctcctaataaaaagaaagtttcttcacgctcgctttcttgctgtccaatttctattaaaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttattttcattgcaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa Accession No. 19 pGL2-SS*-CAG-SecNLuc-2A-eGFP-BGpA-SS*
[0389] SEQ ID NO: 20 pcDNA-CMV-5'UTR-SecNLuc-P2A-eGFP-bGHpA tctcatgagcggatacatatttgaatgtatttagaaaaataaacaaataggggttccgcgcacatttccccgaaaagtgccacctgacgtc
[0390] SEQ ID NO: 21 pGL2-SS*-CAG-SecNLuc-2A-eGFP-WPRE-BGpA-SS* SEQ ID NO: 22 pGL2-SS*-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-BGpA-SS*
[0391] SEQ ID NO: 23 pGL2-SS*-CMV-UTR2-SecNLuc-2A-eGFP-WPRE-BGpA-SS* SEQ ID NO: 24 SS*-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2huBGpA-A120]-SS* SEQ ID NO: 25 SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-3'UTR[2huBGpA-A120]-SS*
[0392] SEQ ID NO: 26 SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2huBGpA-A120]-SS* SEQ ID NO: 27 SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2huBGpA-A120]-SS*
[0393] SEQ ID NO: 28 SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-3'UTR[2huBGpA-A120]-SS* SEQ ID NO: 29 SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-huMAR-3'UTR[2huBGpA-A120]-SS*
[0394] SEQ ID NO: 30 SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR[2huBGpA-A120]-SS* SEQ ID NO: 31 SS*-E1-CMV-UTR1-SecNLuc-2A-eGFP-MAR-WPRE-3'UTR[2huBGpA-A120]-SS*
[0395] SEQ ID NO: 32 SS*-UCOE-E1-CMV-UTR1-SecNLuc-2A-eGFP-MAR-WPRE-3'UTR[2huBGpA-A120]-SS* SEQ ID NO: 33 (Supersequence, SS*) taaagtaacccaatcagcacacaattgccattatacgcgcgtataatggactattgtgtgctgataaacctatttcagcatactacgcgcgtagtatgctgaaataggtgactagaagttcctatactttctagagaataggaacttcataacttcgtataatgtatgct atacgaagttatgggttactttaatttggttgctgactaattgagatgcatgctttgcatacttctgcctgctggggagcctggggactttccacacctggttgctgactaattgagatgcatgctttgcatacttctgcctgctggggagcctggggactttccacacc SEQ ID NO: 34 pGL2-CAG-SecNLuc-2A-eGFP-WPRE-bGlobin polyA
[0396] SEQ ID NO: 35 CAG [E1X3+CBA promoter+intron] SEQ ID NO: 36 CAG [E2+CBA promoter+intron] SEQ ID NO: 37 CAG [E1X3 + CBA promoter + UTR1] gggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcacgggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcacgggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcacgtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggaaaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgctgccttctccctcctgtgagtttggtaagtcgacgggccgggcctgggccgggtccgggccgggtcgttggatccccactacagcccgatactcaagcttgacgaattcgagtatccaaggtagtggactagtgtgacgctgctgacccctttctttcccttctgcaggttggtgtacagtagcttccaaattgattaattcgagcgaacgcgtc SEQ ID NO: 38 CAG [E2 (U100) + CBA promoter + UTR1] Tgggactttccactagacatgacacagcaatctgatatgcttgcgtgagaagaggattcatatcctgggactttccacagattttaccggaagttgttagatgcttgcgtgagaagatctaacatgacacagcaatccttagtgggactttccaagtatgtggggcggggagtatacatgacacagcaattgatcattaccggaagtttataggtgggactttccagacctatgcttgcgtgagaagaaaggtctgggactttccagtcgaggtgagccccacgttctgcttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcggggggggggggggcgcgcgccaggcggggcggggcggggcgaggggcggggcggggcgaggcggaaaggtgcggcggcagccaatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgctgccttctccctcctgtgagtttggtaagtcgacgggccgggcctgggccgggtccgggccgggtcgttggatccccactacagcccgatactcaagcttgacgaattcgagtatccaaggtagtggactagtgtgacgctgctgacccctttctttcccttctgcaggttggtgtacagtagcttccaaattgattaattcgagcgaacgcgtc SEQ ID NO: 39 CMV enhancer - EF1 - UTR sequence 1 Gacattgattattgactagttattaatagtaatcaattacggggtcattagttcatagcccatatatggagttccgcgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaataatgacgtatgttcccatagtaacgccaatagggactttccattgacgtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggcccgcctggcattatgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggggcagagcgcacatcgcccacagtccccgagaagttggggggaggggtcggcaattgaaccggtgcctagagaaggtggcgcggggtaaactgggaaagtgatgtcgtgtactggctccgcctttttcccgagggtgggggagaaccgtatataagtgcagtagtcgccgtgaacgttctttttcgcaacgggtttgccgccagaacacagctgccttctccctcctgtgagtttggtaagtcgacgggccgggcctgggccgggtccgggccgggtcgttggatccccactacagcccgatactcaagcttgacgaattcgagtatccaaggtagtggactagtgtgacgctgctgacccctttctttcccttctgcaggttggtgtacagtagcttccaaattgattaattcgagcgaacgcgtc Accession No. 40 4-1 pGL2-SS*-CAG [CMV enhancer + CBA promoter + intron]-SecNLuc-2A-eGFP-WPRE-3’UTR(108~120 polyA)-SS* SEQ ID NO: 41 4-2 pGL2-SS*-CAG [E1 X3 + CBA promoter + intron]-SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA)-SS*
[0397] SEQ ID NO: 42 4-3 pGL2-SS*-CAG [E2(U100) + CBA promoter + intron]-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS* SEQ ID NO: 43 4-4 pGL2-SS*-CAG [E1 X3 + CBA promoter + UTR1]-SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA)-SS* SEQ ID NO: 44 4-5-pGL2-SS*-CAG [E2 (U100) + CBA promoter + UTR1]-SecNLuc-2A-eGFP-WPRE-3'UTR (108-120 polyA)-SS*
[0398] SEQ ID NO: 45 4-6-pGL2-SS*-CMV enhancer-EF1-UTR1-SecNLuc-2A-eGFP-WPRE-3'UTR(108-120 polyA)-SS* Enhancer-1 with 3 copies of Sequence No. 46 gggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcacgggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcacgggactttccggggcggggcacgtggtgcacgggactttccgtgcacgtgcacgggactttccgggactttccgggactttccgtgcaccacgtggggactttccgtgcac Chimeric intron of Sequence No. 47
[0399] The present invention is not to be limited in scope by the specific aspects described herein. Indeed, various modifications of the invention in addition to those described herein will become apparent to those skilled in the art from the foregoing description and accompanying drawings. Such modifications are intended to be included within the scope of the appended claims. Other aspects are within the scope of the following claims.
Claims
1. (a) an expression cassette comprising a nucleic acid sequence of interest, and (b) one or more of the following: (i) a 5'untranslated region (5'UTR) containing an intron, wherein the 5'UTR is incorporated between a promoter and the nucleic acid sequence of interest in the expression cassette, the 5'UTR; (ii) a synthetic enhancer comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 12 and is located 5' to another enhancer or promoter in the expression cassette; (iii) a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal; (iv) a DNA nuclear targeting sequence (DTS) located 5' of the expression cassette; (v) a cytomegalovirus (CMV) enhancer located 5' of the promoter in the expression cassette; (vi) a scaffold / matrix attachment region (S / MAR) incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal, or (vii) a vertebrate chromatin insulator incorporated into the expression cassette between the nucleic acid sequence of interest and the polyadenylation signal, A vector free of bacterial sequences.
2. (a) a backbone sequence, (b) (i) an expression cassette comprising a nucleic acid sequence of interest, (ii) a first target sequence for a first recombinase adjacent to the 5' side of the expression cassette, (iii) a second target sequence for the first recombinase adjacent to the 3' side of the expression cassette, and (iv) one or more additional target sequences for one or more additional recombinases incorporated within the non-binding regions of the first and second target sequences for the first recombinase A sequence comprising, and (c) one or more of the following: (i) a 5'UTR containing an intron, wherein the 5'UTR is incorporated between a promoter and the nucleic acid sequence of interest in the expression cassette, the 5'UTR; (ii) a synthetic enhancer comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 12 and is incorporated between the 3' end of the first target sequence for the first recombinase and the 5' end of another enhancer or promoter in the expression cassette; (iii) a WPRE incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal; (iv) a DTS incorporated within the first and / or second target sequences for the first recombinase in the unbound region for the first recombinase and one or more additional recombinases, wherein said DTS is between the expression cassette and the cleavage site for the first recombinase and one or more additional recombinases, DTS (v) a CMV enhancer incorporated between the 3' end of the first target sequence for the first recombinase and the 5' end of the promoter in the expression cassette, (vi) an S / MAR incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal, (vii) a vertebrate chromatin insulator incorporated into the expression cassette between the nucleic acid of interest and the polyadenylation signal, or (viii) an endonuclease target sequence incorporated within the first and / or second target sequences for the first recombinase in the unbound region for the first recombinase and one or more additional recombinases, wherein said endonuclease target sequence is between the backbone sequence and the cleavage site for the first recombinase and one or more additional recombinases, endonuclease target sequence An expression vector comprising.
3. The bacteria sequence-free vector according to claim 1 or the expression vector according to claim 2, wherein (a) (i) the intron of (b) (i) in the bacteria sequence-free vector or (c) (i) in the expression vector comprises a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 1 and / or a non-coding sequence incorporated within the intron, or (ii) the 5' UTR of (b) (i) in the bacteria sequence-free vector or (c) (i) in the expression vector comprises a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 3 or SEQ ID NO: 5, (b) the promoter of (b) (i) in the bacteria sequence-free vector or (c) (i) in the expression vector is a chicken β-actin promoter or a CMV promoter, and / or (c) the promoter of (b) (i) in the bacteria sequence-free vector or (c) (i) in the expression vector is incorporated at the 3' end of the CMV enhancer, or (d) The synthetic enhancer of (b)(ii) in the vector without the bacterial sequence or (c)(ii) in the expression vector does not contain multiple contiguous copies of a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 12, and / or (e) The synthetic enhancer of (b)(ii) in the vector without the bacterial sequence or (c)(ii) in the expression vector is integrated at the 5' end of the chicken β-actin promoter, or (f) The CMV enhancer of (b)(v) in the vector without the bacterial sequence or (c)(v) in the expression vector is integrated at the 3' end of a synthetic enhancer containing a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 12 or SEQ ID NO: 46, and / or the promoter of (b)(v) in the vector without the bacterial sequence or (c)(v) in the expression vector is a CMV promoter integrated at the 3' end of the CMV enhancer and the 5' end of the nucleic acid sequence of interest, or (g) The vector without the bacterial sequence or the expression vector contains a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 35, SEQ ID NO: 36, SEQ ID NO: 37, SEQ ID NO: 38, or SEQ ID NO: 39, which is located 5' to the nucleic acid sequence of interest in the vector without the bacterial sequence or integrated between the first target sequence for the first recombinase and the nucleic acid sequence of interest in the expression vector, The vector without the bacterial sequence according to claim 1 or the expression vector according to claim 2. **Claim 4** The vector without the bacterial sequence or the expression vector according to claim 3, wherein the non-coding sequence of (a)(i) is integrated between two nucleotides within an intron corresponding to any two nucleotides from position 25 to position 55 of SEQ ID NO:
1. **Claim 5** The vector without the bacterial sequence or the expression vector according to claim 3, wherein the non-coding sequence of (a)(i) is integrated between two nucleotides within an intron corresponding to any two nucleotides from position 25 to position 55 of SEQ ID NO: 1 and is an S / MAR. **Claim 6** The vector without the bacterial sequence or the expression vector according to claim 3, wherein the non-coding sequence of (a)(i) is integrated between two nucleotides within an intron corresponding to any two nucleotides from position 25 to position 55 of SEQ ID NO: 1 and is a MAR-5. **Claim 7** The bacterial sequence-free vector or expression vector according to claim 3, wherein the CMV enhancer in (c) is incorporated into the 3' end of a synthetic enhancer comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 12 or SEQ ID NO:
46.
8. The bacterial sequence-free vector or expression vector according to claim 3, wherein the synthetic enhancer in (d) comprises a nucleic acid sequence that is at least about 90% identical to SEQ ID NO:
46.
9. The bacterial sequence-free vector or expression vector according to claim 3, wherein the synthetic enhancer in (e) comprises a chimeric intron comprising a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 47, which is incorporated into the 3' end of the chicken β-actin promoter and the 5' end of the nucleic acid sequence of interest.
10. The bacterial sequence-free vector according to claim 1 or the expression vector according to claim 2, wherein (a) the polyadenylation signal is incorporated into the 3' end of the nucleic acid sequence of interest and comprises a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 13, SEQ ID NO: 14 or SEQ ID NO: 15, (b) the vertebrate chromatin insulator in (b)(vii) in the bacterial sequence-free vector or (c)(vii) in the expression vector is the 5'-HS4 chicken-β-globin insulator (cHS4), (c) the S / MAR in (b)(vi) in the bacterial sequence-free vector or (c)(vi) in the expression vector is MAR-5, (d) the polyadenylation signal in (b)(iii), (vi) or (vii) in the bacterial sequence-free vector or (c)(iii), (vi) or (vii) in the expression vector comprises a nucleic acid sequence that is at least about 90% identical to SEQ ID NO: 13, SEQ ID NO: 14 or SEQ ID NO: 15, or (e) the DTS in (b)(iv) in the bacterial sequence-free vector or (c)(iv) in the expression vector is the SV40 enhancer sequence or a cell-specific sequence. The bacterial sequence-free vector according to claim 1 or the expression vector according to claim 2.
11. The endonuclease target sequence in (c)(viii) is (a) for a homing endonuclease. (b) those for I-AniI, I-CeuI, I-ChuI, I-CpaI, I-CpaII, I-CreI, I-DmoI, H-DreI, I-HmuI, I-HmuII, I-LlaI, I-MsoI, PI-PfuI, PI-PkoII, I-PorI, I-PpoI, PI-PspI, I-ScaI, I-SceI, PI-SceI, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-Ssp6803I, I-TevI, I-TevII, I-TevIII, PI-TliI, PI-TliIII, I-Tsp061I or I-Vdi141I; (c) those for I-SceI; (d) those for PI-SceI; (e) those for a Cas endonuclease, or (f) those for Cas9; The expression vector according to claim 2.
12. (a) The first and second target sequences and one or more additional target sequences are selected from the group consisting of the PY54 pal site, N15 telRL site, loxP site, φK02 telRL site, FRT site, phiC31 attP site and λ attP site, or (b) The first and second target sequences for the first recombinase each contain the nucleic acid sequence of SEQ ID NO: 33; The expression vector according to claim 2.
13. The expression vector according to claim 2, wherein the first and second target sequences and one or more additional target sequences are selected from the group consisting of the PY54 pal site, N15 telRL site, loxP site, φK02 telRL site, FRT site, phiC31 attP site and λ attP site, and the expression vector contains each of the target sequences.
14. The expression vector according to claim 2, wherein the first and second target sequences and one or more additional target sequences are selected from the group consisting of the PY54 pal site, N15 telRL site, loxP site, φK02 telRL site, FRT site, phiC31 attP site and λ attP site, and the expression vector contains the pal site and telRL recombinase target binding sequences, loxP recombinase target binding sequences and FRT recombinase target binding sequences incorporated within the pal site.
15. A vector production system comprising a recombinant cell encoding a recombinase under the control of an inducible promoter, wherein the recombinant cell comprises the expression vector according to claim 2, and the recombinase targets one of a first and a second target sequence for a first recombinase, or one or more additional target sequences for one or more additional recombinases in the expression vector.
16. The vector production system according to claim 15, wherein the recombinant cell further encodes an endonuclease under the control of an inducible promoter, and the endonuclease targets an endonuclease target sequence in an expression vector comprising the endonuclease target sequence.
17. A method for producing a bacteria-free vector, comprising the step of incubating the vector production system according to claim 15 or 16 under conditions suitable for the expression of the recombinase.
18. The method according to claim 17, further comprising the step of recovering the bacteria-free vector.
19. The bacteria-free vector according to claim 1, which is a covalently closed circular vector or a covalently closed linear vector.
20. A recombinant cell comprising the bacteria-free vector according to claim 1 or the expression vector according to claim 2.
21. A composition comprising the bacteria-free vector according to claim 1.
22. A composition comprising the expression vector according to claim 2.
23. The bacteria-free vector according to claim 1, the expression vector according to claim 2, or the composition according to claim 21 or 22 for use in a method of treating a disease or disorder in a subject.
24. The bacteria-free vector according to claim 1, the expression vector according to claim 2, or the composition according to claim 21 or 22 for use in a gene editing method comprising inserting a nucleic acid sequence of interest into a target site for gene editing from the bacteria-free vector, the expression vector, or the composition.
25. A polynucleotide comprising a nucleic acid sequence that is at least about 90% identical to any one of SEQ ID NOs: 1, 2, 3, 5, 12 - 18, 35 - 39, and 46.
26. (a) The polynucleotide according to claim 25, or (b)a polynucleotide comprising a nucleic acid sequence that is at least about 90% identical to any one of SEQ ID NOs: 2, 3, and 5, and a polynucleotide comprising a nucleic acid sequence that is at least about 90% identical to any one of SEQ ID NOs: 13-18 An expression vector comprising the same.