Operation ITR sequence and usage

JP2024532261A5Pending Publication Date: 2025-08-26BIOVERATIV THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024512004
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-14
Filing Date
2022-08-19
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Adeno-associated virus (AAV) vectors for gene therapy face challenges due to the inhibitory effect of host cell proteins on transgene expression caused by the T-shaped hairpin structure of inverted terminal repeats (ITRs), necessitating improved ITRs for efficient and sustained expression.

Method used

The use of modified inverted terminal repeats (ITRs) derived from parvoviruses like B19 and goose parvovirus (GPV) that form alternative or shorter hairpin structures, retaining functional properties while minimizing inhibition, flanking a gene cassette containing heterologous polynucleotide sequences.

Benefits of technology

The modified ITRs enhance transgene expression by avoiding T-shaped hairpin inhibition, leading to improved stability and persistence of nucleic acid molecules within the cell nucleus, thereby facilitating efficient and sustained therapeutic protein production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000068_0000
    Figure 00000068_0000
  • Figure 00000068_0001
    Figure 00000068_0001
  • Figure 00000068_0002
    Figure 00000068_0002
Patent Text Reader

Abstract

The present disclosure provides a nucleic acid molecule comprising a first inverted terminal repeat (ITR), a second ITR, and a gene cassette encoding a target sequence.In some embodiments, the first ITR and / or the second ITR are ITRs of a virus other than adeno-associated virus (AAV).Also disclosed is a method of using the nucleic acid molecule in gene therapy applications.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Related Applications This application claims priority to International Application No. PCT / US2021 / 047207, filed August 23, 2021, and U.S. Provisional Application No. 63 / 310,042, filed February 14, 2022, the disclosures of which are incorporated herein by reference in their entireties.

[0002] REFERENCE TO ELECTRONICALLY SUBMITTED SEQUENCE LISTING The contents of the sequence listing submitted electronically as an ASCII text file (Name: 732871_SA9_478BPC_ST26.xml.xml; Size: 90.1 KB; Creation Date: August 17, 2022) are incorporated by reference in their entirety into this specification. [Background technology]

[0003] Gene therapy offers the potential for a long-term means of treating various diseases. In the past, many gene therapy treatments have typically relied on the use of viruses. There are numerous viral agents selected for this purpose, each with significantly different properties that make them more or less suitable for gene therapy. However, the undesirable properties of some viral vectors result in concerns about clinical safety, limiting their therapeutic use. Summary of the Invention [Problem to be solved by the invention]

[0004] Adeno-associated virus (AAV) is a common gene therapy vector, but without its drawbacks. The coding sequence of the AAV genome is flanked by inverted terminal repeats (ITRs), which are required for viral replication and packaging, as well as transgene expression. The T-shaped hairpin structure of the AAV ITRs is susceptible to binding by host cell proteins that inhibit transgene expression within the AAV vector. There is a need to provide efficient and sustained expression of target sequences while circumventing the limitations of existing AAV vector technology. [Means for solving the problem]

[0005] Disclosed herein is a nucleic acid molecule and its use, comprising a first modified inverted terminal repeat (ITR) and / or a second modified ITR flanking a gene cassette comprising a heterologous polynucleotide sequence. The modified ITR disclosed herein provides a short ITR and / or an alternative ITR to the wild-type ITR while retaining the functional properties of the wild-type ITR. The modified ITR disclosed herein also provides a short ITR and / or an alternative ITR to the wild-type ITR while retaining the functional properties of the protein produced by the gene cassette.

[0006] In one aspect herein, a nucleic acid molecule is provided comprising a first inverted terminal repeat (ITR) and a second ITR flanking a gene cassette comprising a heterologous polynucleotide sequence, wherein the first ITR comprises a polynucleotide sequence that is at least about 75% identical to nucleotides 1-49, 50-58, and 59-125 of SEQ ID NO:1, or nucleotides 1-27 and 50-114 of SEQ ID NO:15; and the second ITR comprises a polynucleotide sequence that is at least about 75% identical to nucleotides 1-67, 68-76, and 77-125 of SEQ ID NO:2, or nucleotides 1-65 and 88-114 of SEQ ID NO:16, or SEQ ID NO:25 or 26.

[0007] In some embodiments, the first ITR comprises a polynucleotide sequence that is at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to nucleotides 1-49, 50-58, and 59-125 of SEQ ID NO:1, or nucleotides 1-27 and 50-114 of SEQ ID NO:15, or SEQ ID NO:25. In some embodiments, the first ITR comprises nucleotides 1-49, 50-58, and 59-125 of SEQ ID NO:1, or nucleotides 1-27 and 50-114 of SEQ ID NO:15, or SEQ ID NO:25.

[0008] In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:1. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:3. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:5. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:9. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:13. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:15. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:17. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:19. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:21. In some embodiments, the second ITR comprises the nucleotide sequence of SEQ ID NO:23. In some embodiments, the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO:25.

[0009] In some embodiments, the second ITR comprises a polynucleotide sequence that is at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to nucleotides 1-67, 68-76, and 77-125 of SEQ ID NO:2, or nucleotides 1-65 and 88-114 of SEQ ID NO:16, or SEQ ID NO:26. In some embodiments, the second ITR comprises nucleotides 1-67, 68-76, and 77-125 of SEQ ID NO:2, or nucleotides 1-65 and 88-114 of SEQ ID NO:16, or SEQ ID NO:26. In some embodiments, the second ITR comprises the nucleotide sequence of SEQ ID NO:24.

[0010] In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:2. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:4. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:6. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:10. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:14. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:16. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:18. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:20. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:22. In some embodiments, the second ITR comprises the nucleotide sequence of SEQ ID NO:24. In some embodiments, the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:26.

[0011] In some embodiments, the first ITR is selected from the polynucleotide sequences set forth in SEQ ID NO: 1, 3, 5, 9, 13, 15, 17, 19, 21, 23, or 25, and the second ITR is selected from the polynucleotide sequences set forth in SEQ ID NO: 2, 4, 6, 10, 14, 16, 18, 20, 22, 24, or 26.

[0012] In some embodiments, the nucleic acid molecule further comprises a promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter drives expression of the heterologous polynucleotide sequence in an organ or tissue, the organ or tissue comprising muscle, central nervous system (CNS), eye, liver, heart, kidney, pancreas, lung, skin, bladder, urinary tract, spleen, myeloid cell lineage, and lymphoid cell lineage, or any combination thereof. In some embodiments, the promoter drives expression of the heterologous polynucleotide sequence in hepatocytes, epithelial cells, endothelial cells, cardiac myocytes, skeletal muscle cells, sinusoidal cells, afferent neurons, efferent neurons, interneurons, glial cells, astrocytes, oligodendrocytes, microglia, ependymal cells, lung epithelial cells, Schwann cells, satellite cells, photoreceptor cells, retinal ganglion cells, T cells, B cells, NK cells, macrophages, dendritic cells, or any combination thereof. In some embodiments, the promoter is located 5' to the heterologous polynucleotide sequence. In some embodiments, the promoter is a mouse transthyretin promoter (mTTR), a native human factor VIII promoter, a human alpha 1 antitrypsin promoter (hAAT), a human albumin minimal promoter, a mouse albumin promoter, a tristetraprolin (TTP) promoter, a CASI promoter, a CAG promoter, a cytomegalovirus (CMV) promoter, an alpha 1 antitrypsin (AAT) promoter, a muscle creatine kinase (MCK) promoter, a myosin heavy chain alpha (αMHC) promoter, a myoglobin (MB) promoter, a desmin (DES) promoter, a SPc5-12 promoter, a 2R5Sc5-12 promoter, a dMCK promoter, a tMCK promoter, or a phosphoglycerate kinase (PGK) promoter. In some embodiments, the promoter comprises the nucleic acid sequence of SEQ ID NO:31.

[0013] In some embodiments, the heterologous polynucleotide sequence further comprises an intron sequence. In some embodiments, the intron sequence is located 5' to the heterologous polynucleotide sequence. In some embodiments, the intron sequence is located 3' to the promoter. In some embodiments, the intron sequence comprises a synthetic intron sequence. In some embodiments, the intron sequence comprises the nucleic acid sequence of SEQ ID NO:32.

[0014] In some embodiments, the gene cassette further comprises a post-transcriptional regulatory element. In some embodiments, the post-transcriptional regulatory element is located 3' to the heterologous polynucleotide sequence. In some embodiments, the regulatory element comprises a mutant woodchuck hepatitis virus post-transcriptional regulatory element (WPRE), a microRNA binding site, a DNA nuclear targeting sequence, a TLR9 inhibitory sequence, or any combination thereof. In some embodiments, the post-transcriptional regulatory element comprises the nucleic acid sequence of SEQ ID NO: 33.

[0015] In some embodiments, the gene cassette further comprises a 3'UTR poly(A) tail sequence. In some embodiments, the 3'UTR poly(A) tail sequence is selected from the group consisting of bGH poly(A), actin poly(A), hemoglobin poly(A), and any combination thereof.

[0016] In some embodiments, the gene cassette further comprises an enhancer sequence.In some embodiments, the enhancer sequence is located between the first ITR and the second ITR.In some embodiments, the enhancer comprises the nucleic acid sequence of SEQ ID NO:30.

[0017] In some embodiments, the nucleic acid molecule comprises, from 5' to 3', a first ITR, a gene cassette, and a second ITR, where the gene cassette comprises a tissue-specific promoter sequence, an intron sequence, a heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.

[0018] In some embodiments, the gene cassette comprises, from 5' to 3', a tissue-specific promoter sequence, an intron sequence, a heterologous polynucleotide sequence, a post-transcriptional regulatory element, and a 3'UTR poly(A) tail sequence.

[0019] In some embodiments, the gene cassette comprises a single stranded nucleic acid. In some embodiments, the gene cassette comprises a double stranded nucleic acid.

[0020] In some embodiments, the heterologous polynucleotide sequence encodes a therapeutic protein.

[0021] In some embodiments, the heterologous polynucleotide sequence encodes a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or any combination thereof. In some embodiments, the heterologous polynucleotide sequence encodes a growth factor. In some embodiments, the heterologous polynucleotide sequence encodes a hormone. In some embodiments, the heterologous polynucleotide sequence encodes a cytokine.

[0022] In some embodiments, the heterologous polynucleotide sequence encodes an antibody or fragment thereof.

[0023] In some embodiments, the heterologous polynucleotide sequence encodes X-linked dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha glucosidase), GALT, OTC, CMD1A, LAMA2, or any combination thereof.

[0024] In some embodiments, the heterologous polynucleotide sequence encodes a microRNA (miRNA). In some embodiments, the miRNA downregulates the expression of target genes, including SOD1, HTT, RHO, CD38, or any combination thereof.

[0025] In some embodiments, the heterologous polynucleotide sequence encodes a coagulation factor, where the coagulation factor is Factor I (FI), Factor II (FII), Factor III (FIII), Factor IV (FIV), Factor V (FV), Factor VI (FVI), Factor VII (FVII), Factor VIII (FVIII), Factor IX (FIX), Factor X (FX), Factor XI (FXI), Factor XII (FXII), Factor XIII (FXIII), von Willebrand Factor (VWF), , prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2 antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor 1 (PAI-1), plasminogen activator inhibitor 2 (PAI2), or any combination thereof.

[0026] In some embodiments, the heterologous polynucleotide sequence is codon optimized. In some embodiments, the heterologous polynucleotide sequence is codon optimized for expression in humans.

[0027] In some embodiments, the nucleic acid molecule is formulated with a delivery agent. In some embodiments, the delivery agent comprises a lipid nanoparticle. In some embodiments, the delivery agent comprises a liposome, a non-lipid polymer molecule, an endosome, or any combination thereof.

[0028] In some embodiments, the nucleic acid molecule is formulated for intravenous, transdermal, intradermal, intraneuronal, intraocular, intrathecal, subcutaneous, pulmonary, or oral administration, or any combination thereof. In some embodiments, the nucleic acid molecule is formulated for intravenous administration. In some embodiments, the nucleic acid molecule is formulated for administration by in situ injection. In some embodiments, the nucleic acid molecule is formulated for administration by inhalation.

[0029] In another aspect of the present specification, there is provided a vector comprising a nucleic acid molecule described herein.

[0030] In another aspect of the present specification, there is provided a host cell comprising a nucleic acid molecule described herein, or a vector described herein.

[0031] In another aspect of the present specification, there is provided a pharmaceutical composition comprising a nucleic acid molecule described herein.

[0032] In another aspect of the present specification, there is provided a pharmaceutical composition comprising a vector described herein and a pharma- ceutically acceptable excipient.

[0033] In another aspect of the present specification, there is provided a pharmaceutical composition comprising a host cell described herein and a pharma- ceutically acceptable excipient.

[0034] In another aspect of the present specification, there is provided a kit comprising a nucleic acid molecule described herein and instructions for administering the nucleic acid molecule to a subject in need thereof.

[0035] In another aspect of the present specification, there is provided a baculovirus system for producing the nucleic acid molecules described herein.

[0036] In some embodiments, the nucleic acid molecule is produced in an insect cell.

[0037] In another aspect herein, there is provided a nanoparticle delivery system comprising a nucleic acid molecule as described herein.

[0038] In another aspect herein, there is provided a method of expressing a heterologous polynucleotide sequence in a subject in need thereof, the method comprising administering to the subject a nucleic acid molecule described herein, a vector described herein, or a pharmaceutical composition described herein.

[0039] In another aspect herein, there is provided a method of treating a disease or disorder in a subject in need thereof, the method comprising administering to the subject a nucleic acid molecule described herein, a vector described herein, or a pharmaceutical composition described herein.

[0040] In some embodiments, the nucleic acid molecule is administered intravenously, transdermally, intradermally, subcutaneously, orally, pulmonary, intraneuronally, intraocularly, intrathecally, or any combination thereof. In some embodiments, the nucleic acid molecule is administered intravenously. In some embodiments, the nucleic acid molecule is administered by in situ injection. In some embodiments, the nucleic acid molecule is administered by inhalation.

[0041] In some embodiments, the subject is a mammal, hi some embodiments, the subject is a human. [Brief description of the drawings]

[0042] [Figure 1]FIG. 1A is a graphical representation of the predicted structure of wild-type B19 ITR and truncated B19Δ135 ITR. The shaded region indicates the B19 Rep binding element (RBE). Gibbs free energy (ΔG) is provided for each ITR sequence. Note: The illustration does not represent the exact location of the RBE within the ITR. FIG. 1B is a graphical representation of the predicted structure of wild-type GPV ITR and truncated GPVΔ162 ITR. The shaded region indicates the GPV Rep binding element (RBE). Gibbs free energy (ΔG) is provided for each ITR sequence. Note: The illustration does not represent the exact location of the RBE within the ITR. [Figure 2A] Graphical representation of the predicted structures of wild-type B19 ITR, truncated B19Δ151 ITR, truncated B19Δ223 ITR, and truncated B19_minimal. The shaded region indicates the Rep-binding element (RBE) of B19. Gibbs free energy (ΔG) is provided for each ITR sequence. Note: The illustration does not represent the exact location of the RBE within the ITR. [Figure 2B] Graphical representation of the predicted structures of wild-type GPV ITR, truncated GPVΔ120 ITR, truncated GPVΔ186 ITR, and truncated GPV_minimal. Shaded regions indicate the Rep-binding elements (RBEs) of GPV. Gibbs free energies (ΔG) are provided for each ITR sequence. Note: The illustration does not represent the exact location of the RBE within the ITRs. [Diagram 3] Figure 1 shows FVIII activity measured in HemA mice receiving 34.7 μg of ssFVIII-DNA (circles) or dsFVIII-DNA (squares) of FVIII expression constructs flanked by modified ITRs, B19_min. Blood samples were taken 3 and 7 days after injection. [Figure 4]FIG. 1 is a graphical representation of the predicted DNA structures of modified ITRs according to SEQ ID NOs: 1 to 22. The predicted structures were generated using Geneious software (Biomatters Ltd., Auckland, New Zealand). [Diagram 5] FIG. 1 is a graphical representation of the predicted DNA structures of modified ITRs according to SEQ ID NOs: 1 to 22. The predicted structures were generated using Geneious software (Biomatters Ltd., Auckland, New Zealand). [Figure 6] FIG. 1 is a graphical representation of the predicted DNA structures of modified ITRs according to SEQ ID NOs: 1 to 22. The predicted structures were generated using Geneious software (Biomatters Ltd., Auckland, New Zealand). [Figure 7] FIG. 1 is a graphical representation of the predicted DNA structures of modified ITRs according to SEQ ID NOs: 1 to 22. The predicted structures were generated using Geneious software (Biomatters Ltd., Auckland, New Zealand). [Figure 8] FIG. 1 is a graphical representation of the predicted DNA structures of modified ITRs according to SEQ ID NOs: 1 to 22. The predicted structures were generated using Geneious software (Biomatters Ltd., Auckland, New Zealand). [Figure 9] FIG. 1 is a graphical representation of the predicted DNA structures of modified ITRs according to SEQ ID NOs: 1 to 22. The predicted structures were generated using Geneious software (Biomatters Ltd., Auckland, New Zealand). [Figure 10]Figures 10A-10D are schematic representations of modified FVIIIXTEN expression cassettes with modified parvovirus ITRs according to embodiments of the present invention. Figure 10A shows a linear schematic map for a modified FVIIIXTEN expression cassette flanked by AAV2 WT ITRs (described in USPTO Application No. 63 / 069,073). Figure 10B shows a linear schematic map for a modified FVIIIXTEN expression cassette flanked by HBoV1 WT ITRs. Figure 10C shows a linear schematic map for a modified FVIIIXTEN expression cassette flanked by B19 WT (described in USPTO Application No. 63 / 069,073) or B19 Minimal (SEQ ID NO: 1, SEQ ID NO: 2) ITRs. FIG. 10D shows a linear schematic map for the modified FVIIIXTEN expression cassette flanked by GPVΔ186 ITRs (SEQ ID NO:9, SEQ ID NO:10), GPVΔ120 ITRs (SEQ ID NO:13, SEQ ID NO:14), or GPV Minimal ITRs (SEQ ID NO:15, SEQ ID NO:16). [Figure 11] Schematic representation of the approach used to generate ssDNA in which a FVIIIXTEN expression cassette flanked by parvoviral ITRs was digested with a restriction enzyme that recognizes ITR-related sequences and results in blunt-ended DNA. The double-stranded DNA products of the digestion (FVIII expression cassette and plasmid backbone) were heat denatured at 95° C. (denaturation) followed by cooling at 4° C. (renaturation) to allow the palindromic ITR sequences to form hairpins. The resulting ss(ssDNA)FVIIIXTEN was used for systemic delivery via hydrodynamic tail vein injection in HemA mice. [Figure 12]1 is a graphical representation of plasma FVIII activity levels measured by Chromogenix Coatest® SP Factor VIII chromogenic assay. Plasma samples were collected at different intervals from hFVIIIR593C+ / + / HemA mice systemically injected with 200, 800, or 1600 μg / kg of single-stranded V2.0 ss (ssDNA) FVIIIXTEN flanked by human bocavirus (HBoV1), human erythrovirus (B19), goose parvovirus (GPV), or mutants or combinations thereof ("hybrids") as indicated, via fluid tail vein injection. Error bars represent standard deviation. [Figure 13] Figures 13A-13B show representations of purified ce(ceDNA)FVIIIXTEN obtained from the baculovirus system and their in vivo efficacy studies. Figure 13A shows agarose gel images of purified ce(ceDNA)FVIIIXTEN flanked by AAV2 WT ITR or HBoV1 WT ITR obtained from sequential elution electrophoresis. Purity is shown relative to starting material (SM) and arrows point to DNA bands corresponding to the sizes of FVIIIXTEN ceDNA vector (ceDNA), baculovirus DNA (vDNA), and Sf9 cell genomic DNA (gDNA). Figure 13B shows a graphical representation of plasma FVIII activity levels measured by Chromogenix Coatest® SP Factor VIII chromogenic assay. Plasma samples were collected at different intervals from hFVIIIR593C+ / + / HemA mice systemically injected with 80, 40, or 12 μg / kg of ce(ceDNA)FVIIIXTEN flanked by AAV2 ITRs or HBoV1 ITRs as indicated via hydrodynamic tail vein injection. Error bars represent standard deviation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0043] Disclosed herein is a nucleic acid molecule and its use, comprising a first modified inverted terminal repeat (ITR) and / or a second modified ITR flanking a gene cassette comprising a heterologous polynucleotide sequence. In some embodiments, the first ITR and / or the second ITR are derived from parvovirus B19 or goose parvovirus (GPV). The modified ITR disclosed herein provides a shortened ITR and / or alternative ITR to the wild-type ITR while retaining the functional properties of the wild-type ITR. The modified ITR disclosed herein also provides a shortened ITR and / or alternative ITR to the wild-type ITR while retaining the functional properties of the protein produced by the gene cassette.

[0044] Exemplary constructs of the present disclosure are illustrated in the accompanying figures and sequence listing. For a clear understanding of the specification and claims, the following definitions are provided below.

[0045] I. Definition It is noted that the term "a" entity or "an" entity refers to one or more of that entity: for example, "a nucleotide sequence" is understood to refer to one or more nucleotide sequences. Similarly, "a therapeutic protein" and "a miRNA" are understood to refer to one or more therapeutic proteins and one or more miRNAs, respectively. Thus, the terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein.

[0046] The term "about" is used herein to mean approximately, in the region of, roughly, or in the vicinity thereof. When used in conjunction with a numerical range, the term "about" modifies the range by extending the boundaries above and below the numerical values ​​set forth. In general, the term "about" is used herein to modify numerical values ​​above and below the stated value by a variance of 10 percent upward or downward (high or low).

[0047] As used herein, "and / or" when interpreted in the alternative ("or") also refers to and includes any / all possible combinations of one or more of the associated listed items, as well as the absence of combinations.

[0048] "Nucleic acid", "nucleic acid molecule", "nucleotide", "nucleotide(s) sequence", and "polynucleotide" are used interchangeably and refer to the phosphate polymeric form of ribonucleosides (adenosine, guanosine, uridine, or cytidine; "RNA molecule") or deoxyribonucleosides (deoxyadenosine, deoxyguanosine, deoxythymidine, or deoxycytidine; "DNA molecule"), or any of their phosphate analogs, such as phosphorothioates and thioesters, in single-stranded form or within a double-stranded helix. Single-stranded nucleic acid sequence refers to single-stranded DNA (ssDNA) or single-stranded RNA (ssRNA). DNA-DNA helices, DNA-RNA helices, and RNA-RNA helices, which are double-stranded, are also possible. The term nucleic acid molecule, particularly DNA or RNA molecule, refers only to the primary and secondary structure of the molecule and does not limit the molecule to any particular tertiary form. Thus, the term includes double-stranded DNA found, inter alia, in linear or circular DNA molecules (e.g., restriction fragments), plasmids, supercoiled DNA, and chromosomes. When discussing the structure of a particular double-stranded DNA molecule, the sequence is described herein, following the usual convention, showing only the sequence along the non-transcribed strand of DNA (i.e., the strand with sequence homology to mRNA) in the 5' to 3' direction. A "recombinant DNA molecule" is a DNA molecule that has undergone molecular biological manipulation. DNA includes, but is not limited to, cDNA, genomic DNA, plasmid DNA, synthetic DNA, and semi-synthetic DNA. A "nucleic acid composition" of the present disclosure comprises one or more nucleic acids as described herein.

[0049] As used herein, "inverted terminal repeat" (or "ITR") refers to a nucleic acid subsequence located at the 5' or 3' end of a single stranded nucleic acid sequence that comprises a set of nucleotides (initial sequence) followed downstream by its reverse complement, i.e., a palindromic sequence. The intervening nucleotide sequence between the initial sequence and the reverse complement can be of any length, including zero. In one embodiment, an ITR useful in the present disclosure comprises one or more "palindromic sequences." An ITR can have any number of functions. In some embodiments, an ITR described herein forms a hairpin structure. In some embodiments, an ITR forms a T-shaped hairpin structure. In some embodiments, an ITR forms a hairpin structure other than a T-shaped, e.g., a U-shaped hairpin structure. In some embodiments, an ITR promotes the survival of a nucleic acid molecule in a cell nucleus over an extended period of time. In some embodiments, an ITR promotes the permanent survival (e.g., for the entire lifespan of a cell) of a nucleic acid molecule in a cell nucleus. In some embodiments, an ITR promotes the stability of a nucleic acid molecule in a cell nucleus. In some embodiments, the ITRs promote the retention of the nucleic acid molecule in the cell nucleus. In some embodiments, the ITRs promote the persistence of the nucleic acid molecule in the cell nucleus. In some embodiments, the ITRs inhibit or prevent the degradation of the nucleic acid molecule in the cell nucleus.

[0050] In one embodiment, the initial sequence and / or reverse complement of the ITR comprises from about 2 to 600 nucleotides, from about 2 to 550 nucleotides, from about 2 to 500 nucleotides, from about 2 to 450 nucleotides, from about 2 to 400 nucleotides, from about 2 to 350 nucleotides, from about 2 to 300 nucleotides, or from about 2 to 250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5-600 nucleotides, about 10-600 nucleotides, about 15-600 nucleotides, about 20-600 nucleotides, about 25-600 nucleotides, about 30-600 nucleotides, about 35-600 nucleotides, about 40-600 nucleotides, about 45-600 nucleotides, about 50-600 nucleotides, about 60-600 nucleotides, about 70-600 nucleotides, about 80-600 nucleotides, about 90-600 nucleotides, about 100-600 nucleotides, about 150-600 nucleotides, about 200-600 nucleotides, about 300-600 nucleotides, about 350-600 nucleotides, about 400-600 nucleotides, about 450-600 nucleotides, about 500-600 nucleotides, or about 550-600 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 5-550 nucleotides, about 5-500 nucleotides, about 5-450 nucleotides, about 5-400 nucleotides, about 5-350 nucleotides, about 5-300 nucleotides, or about 5-250 nucleotides. In some embodiments, the initial sequence and / or reverse complement comprises about 10-550 nucleotides, about 15-500 nucleotides, about 20-450 nucleotides, about 25-400 nucleotides, about 30-350 nucleotides, about 35-300 nucleotides, or about 40-250 nucleotides. In certain embodiments, the initial sequence and / or the reverse complement comprises about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, or about 600 nucleotides.In certain embodiments, the initial sequence and / or the reverse complement comprises about 400 nucleotides.

[0051] In other embodiments, the initial sequence and / or reverse complement of the ITR comprises about 2-200 nucleotides, about 5-200 nucleotides, about 10-200 nucleotides, about 20-200 nucleotides, about 30-200 nucleotides, about 40-200 nucleotides, about 50-200 nucleotides, about 60-200 nucleotides, about 70-200 nucleotides, about 80-200 nucleotides, about 90-200 nucleotides, about 100-200 nucleotides, about 125-200 nucleotides, about 150-200 nucleotides, or about 175-200 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-150 nucleotides, about 5-150 nucleotides, about 10-150 nucleotides, about 20-150 nucleotides, about 30-150 nucleotides, about 40-150 nucleotides, about 50-150 nucleotides, about 75-150 nucleotides, about 100-150 nucleotides, or about 125-150 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-100 nucleotides, about 5-100 nucleotides, about 10-100 nucleotides, about 20-100 nucleotides, about 30-100 nucleotides, about 40-100 nucleotides, about 50-100 nucleotides, or about 75-100 nucleotides. In other embodiments, the initial sequence and / or reverse complement comprises about 2-50 nucleotides, about 10-50 nucleotides, about 20-50 nucleotides, about 30-50 nucleotides, about 40-50 nucleotides, about 3-30 nucleotides, about 4-20 nucleotides, or about 5-10 nucleotides. In another embodiment, the initial sequence and / or reverse complement consists of 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides.In other embodiments, the intervening nucleotides between the initial sequence and the reverse complement are (e.g., consist of) 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0052] Thus, an "ITR" as used herein may fold back on itself to form a double-stranded segment. For example, the sequence GATCXXXXGATC includes an initial sequence of GATC and its complement (3'CTAG5') such that when folded, they form a double helix. In some embodiments, an ITR includes a continuous palindromic sequence (e.g., GATCGATC) between the initial sequence and the reverse complement. In some embodiments, an ITR includes an interrupted palindromic sequence (e.g., GATCXXXXGATC) between the initial sequence and the reverse complement. In some embodiments, the complementary portions of the continuous or interrupted palindromic sequences interact with each other to form a "hairpin loop" structure. A "hairpin loop" structure, as used herein, occurs when at least two complementary sequences on a single-stranded nucleotide molecule base pair to form a double-stranded portion. In some embodiments, only a portion of the ITR forms a hairpin loop. In other embodiments, the entire ITR forms a hairpin loop. In some embodiments, the ITRs retain the Rep binding element (RBE) of the wild-type ITR from which they are derived. Preservation of the RBE can be important for ITR stability and manufacturing purposes.

[0053] As used herein, the term "parvovirus" encompasses the Parvoviridae family, including, but not limited to, Parvovirus and Dependovirus, which are autonomously replicating. Autonomous parvoviruses include, for example, members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Abeparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstildensovirus. Exemplary autonomous parvoviruses include, but are not limited to, porcine parvovirus, minute virus of mice, canine parvovirus, mink enteritis virus, bovine parvovirus, chicken parvovirus, feline panleukopenia virus, feline parvovirus, goose parvovirus (GPV), H1 parvovirus, Muscovy duck parvovirus, snake parvovirus, and B19 virus. Other autonomous parvoviruses are known to those of skill in the art, see, e.g., Fields et al., Virology, Vol. 2, Chapter 69 (4th ed., Lippincott-Raven Publishers).

[0054] As used herein, the term "non-AAV" encompasses nucleic acids, proteins, and viruses derived from the Parvoviridae family, excluding any adeno-associated viruses (AAVs) within the Parvoviridae family. "Non-AAV" includes, but is not limited to, autonomously replicating members of the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Abeparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstildensovirus.

[0055] As used herein, the term "adeno-associated virus" (AAV) includes, but is not limited to, AAV type 1, AAV type 2, AAV type 3 (including types 3A and 3B), AAV type 4, AAV type 5, AAV type 6, AAV type 7, AAV type 8, AAV type 9, AAV type 10, AAV type 11, AAV type 12, AAV type 13, snake AAV, avian AAV, bovine AAV, canine AAV, equine AAV, ovine AAV, caprine AAV, shrimp AAV, the AAV serotypes and clades disclosed by Gao et al. (J. Virol., 78:6381 (2004)) and Morris et al. (Virol., 33:375 (2004)), as well as any other AAV now known or subsequently discovered. See, e.g., FIELDS et al., VIROLOGY, Volume 2, Chapter 69 (4th ed., Lippincott-Raven Publishers).

[0056] As used herein, the term "derived from" refers to a component isolated from or made using a specified molecule or organism, or information (e.g., amino acid sequence or nucleic acid sequence) derived from a specified molecule or organism. For example, a nucleic acid sequence (e.g., ITR) derived from a second nucleic acid sequence (e.g., ITR) may contain a nucleotide sequence that is identical or substantially similar to the nucleotide sequence of the second nucleic acid sequence. In the case of a nucleotide or polypeptide, the derived molecular species may be obtained, for example, by natural mutagenesis, artificial directed mutagenesis, or artificial random mutagenesis. The mutagenesis used to derive a nucleotide or polypeptide may be intentionally directed, intentionally random, or a mixture of each. Mutagenesis of a nucleotide or polypeptide that creates a different nucleotide or polypeptide derived from a first nucleotide or polypeptide may be a different nucleotide or polypeptide derived from a first nucleotide or polypeptide, and the identification of the derived nucleotide or polypeptide may be made, for example, by a suitable screening method, as discussed herein. Mutagenesis of a polypeptide typically involves the manipulation of a polypeptide that encodes a polynucleotide.

[0057] A "capsid-free" or "capsid-less" vector or nucleic acid molecule refers to a vector construct that does not contain a capsid.

[0058] As used herein, a "coding region" or "coding sequence" is a portion of a polynucleotide that consists of codons that can be translated into amino acids. A "stop codon" (TAG, TGA, or TAA) is typically not translated into an amino acid but is considered part of the coding region, whereas any adjacent sequences, such as promoters, ribosome binding sites, transcription terminators, introns, etc., are not part of the coding region. The boundaries of a coding region are typically determined by a start codon at the 5' end that codes for the amino terminus of the resulting polypeptide and a translation stop codon at the 3' end that codes for the carboxyl terminus of the resulting polypeptide. Two or more coding regions can be present within a single polynucleotide construct, e.g., on a single vector, or in separate polynucleotide constructs, e.g., on separate (different) vectors. Thus, a single vector may contain only a single coding region or may include two or more coding regions.

[0059] Certain proteins secreted by mammalian cells are associated with secretory signal peptides that are cleaved from the mature protein once the growing protein chain is triggered to export across the rough endoplasmic reticulum. Those skilled in the art are aware that signal peptides are generally fused to the N-terminus of a polypeptide and are cleaved from the complete or "full-length" polypeptide to yield a secreted or "mature" form of the polypeptide. In certain embodiments, the native signal peptide or a functional derivative of this sequence retains the ability to direct the secretion of a polypeptide operatively associated therewith. Alternatively, a heterologous mammalian signal peptide, such as human tissue plasminogen activator (TPA), or mouse β-glucuronidase signal peptide, or a functional derivative thereof, is used.

[0060] The term "downstream" refers to a nucleotide sequence located 3' to a reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence relates to the sequence following the start of transcription. For example, the translation start codon of a gene is located downstream of the transcription start site.

[0061] The term "upstream" refers to a nucleotide sequence located 5' to a reference nucleotide sequence. In certain embodiments, the upstream nucleotide sequence refers to a sequence located 5' of a coding region or at the origin of transcription. For example, most promoters are located upstream of the transcription start site.

[0062] As used herein, the term "gene cassette" or "expression cassette" refers to a DNA sequence capable of directing the expression of a particular polynucleotide sequence in a suitable host cell, comprising a promoter operably linked to the polynucleotide sequence of interest. A gene cassette may be located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding region and may include nucleotide sequences that affect transcription, RNA processing, stability, or translation of the associated coding region. If the coding region is intended for expression in a eukaryotic cell, a polyadenylation signal sequence and a transcription termination sequence will typically be located 3' to the coding sequence. In some embodiments, a gene cassette comprises a polynucleotide that encodes a gene product. In some embodiments, a gene cassette comprises a polynucleotide that encodes a miRNA. In some embodiments, a gene cassette comprises a heterologous polynucleotide sequence.

[0063] A polynucleotide encoding a product, e.g., a miRNA or gene product (e.g., a polypeptide such as a therapeutic protein), may include a promoter and / or other expression (e.g., transcription or translation) control sequence operably associated with one or more coding regions. When in operably associated, a coding region of a gene product, e.g., a polypeptide, is associated with one or more regulatory regions in such a manner that expression of the gene product is under the influence or control of the regulatory region(s). For example, a coding region and a promoter are "operably associated" if induction of promoter function results in transcription of an mRNA that encodes the gene product encoded by the coding region, and the nature of the linkage between the promoter and the coding region does not interfere with the ability of the promoter to direct expression of the gene product or the ability of the DNA template to be transcribed. Other expression control sequences other than promoters, e.g., enhancers, operators, repressors, and transcription termination signals, may also be operably associated with a coding region to direct expression of the gene product.

[0064] "Expression control sequence" refers to a regulatory nucleotide sequence, such as a promoter, enhancer, or terminator, that results in the expression of a coding sequence in a host cell. Expression control sequences generally encompass any regulatory nucleotide sequence that facilitates efficient transcription and translation of an operably linked coding nucleic acid. Non-limiting examples of expression control sequences include promoters, enhancers, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, or stem-loop structures. A variety of expression control sequences are known to those skilled in the art. These include expression control sequences that function in vertebrate cells, such as, but are not limited to, promoter and enhancer segments derived from cytomegalovirus (immediate early promoter with intron A), simian virus 40 (early promoter), and retroviruses (such as Rous sarcoma virus). Other expression control sequences include expression control sequences derived from vertebrate genes, such as actin, heat shock proteins, bovine growth hormone, and rabbit β-globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable expression control sequences include tissue-specific promoters and enhancers, as well as lymphokine-inducible promoters (e.g., promoters induced by interferons or interleukins). Other expression control sequences include intron sequences, post-transcriptional regulatory elements, and polyadenylation signals. Additional exemplary expression control sequences are discussed elsewhere in this disclosure.

[0065] Likewise, various translation control elements are known to those of skill in the art, including, but not limited to, ribosome binding sites, translation initiation / termination codons, and elements derived from picornaviruses (particularly internal ribosome entry sites, or IRES).

[0066] As used herein, the term "expression" refers to the process by which a polynucleotide results in a gene product, e.g., an RNA or a polypeptide. "Expression" includes, without limitation, the transcription of a polynucleotide into messenger RNA (mRNA), transfer RNA (tRNA), small hairpin RNA (shRNA), small interfering RNA (siRNA), or any other RNA product, and the translation of an mRNA into a polypeptide. Expression results in a "gene product." As used herein, a gene product can be a nucleic acid, e.g., a messenger RNA produced by transcription of a gene, or a polypeptide translated from a transcript. Gene products as described herein further include nucleic acids with post-transcriptional modifications, e.g., polyadenylation or splicing, or polypeptides with post-translational modifications, e.g., methylation, glycosylation, addition of lipids, association with other protein subunits, or proteolytic cleavage. As used herein, the term "yield" refers to the amount of polypeptide resulting from expression of a gene.

[0067] "Vector" refers to any vehicle for cloning and / or introduction of a nucleic acid into a host cell. A vector can be a replicon to which another nucleic acid segment is attached to effect replication of the attached segment. "Replicon" refers to any genetic element (e.g., plasmid, phage, cosmid, chromosome, virus) that functions as an autonomous unit of replication in vivo, i.e., capable of replication under its own control. The term "vector" includes vehicles for introducing a nucleic acid into a cell in vitro, ex vivo, or in vivo. Numerous vectors are known and used in the art, including, for example, plasmids, modified eukaryotic viruses, or modified bacterial viruses. Insertion of a polynucleotide into a suitable vector is accomplished by ligating a suitable polynucleotide fragment into a selected vector with complementary cohesive termini.

[0068] Vectors are engineered to encode a selectable marker or reporter that allows for the selection or identification of cells that have incorporated the vector. Expression of the selectable marker or reporter allows for the identification and / or selection of host cells that incorporate and express other coding regions contained on the vector. Examples of selectable marker genes known and used in the art include genes that provide resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, the herbicide bialaphos, sulfonamides, and the like; and genes used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase genes, and the like. Examples of reporters known and used in the art include luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), and the like. A selectable marker is also considered to be a reporter.

[0069] As used herein, the term "host cell" refers to, for example, microorganisms, yeast cells, insect cells, and mammalian cells that are or have been used as recipients of ssDNA or vectors. The term "host cell" includes the progeny of the original cell that has been transduced. Thus, as used herein, "host cell" generally refers to a cell that has been transduced with an exogenous DNA sequence. It is understood that the progeny of a single parent cell are not necessarily completely identical in shape or genomic or total DNA complement to the original parent due to natural, accidental, or deliberate mutations. In some embodiments, the host cell may be an in vitro host cell.

[0070] The term "selectable marker" refers to an identifying factor, typically an antibiotic resistance gene or a chemical resistance gene, that allows selection based on the effect of the marker gene, i.e., resistance to antibiotics, resistance to herbicides, colorimetric markers, enzymes, fluorescent markers, etc., where the effect is used to trace the inheritance of the nucleic acid of interest and / or to identify cells or organisms that have inherited the nucleic acid of interest. Examples of selectable marker genes known and used in the art include genes that confer resistance to ampicillin, streptomycin, gentamicin, kanamycin, hygromycin, the herbicide bialaphos, sulfonamides, etc.; and genes used as phenotypic markers, i.e., anthocyanin regulatory genes, isopentanyl transferase genes, etc.

[0071] The term "reporter gene" refers to a nucleic acid encoding an identifying factor that allows identification based on the effect of the reporter gene, where the effect is used to trace the inheritance of the nucleic acid of interest, to identify cells or organisms that have inherited the nucleic acid of interest, and / or to measure induction of gene expression or transcription of the gene. Examples of reporter genes known and used in the art include luciferase (Luc), green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), β-galactosidase (LacZ), β-glucuronidase (Gus), and the like. Selectable marker genes are also considered reporter genes.

[0072] "Promoter" and "promoter sequence" are used interchangeably and refer to a DNA sequence capable of controlling the expression of a coding sequence or functional RNA. Generally, the coding sequence is located 3' to the promoter sequence. Promoters are composed of different elements in their entirety from a natural gene or from different promoters found in nature, or even contain synthetic DNA segments. Those skilled in the art will understand that different promoters may direct the expression of a gene in different tissues or cell types, may direct the expression of a gene at different developmental stages, or may direct the expression of a gene in response to different environmental or physiological conditions. A promoter that causes a gene to be expressed in most cell types at most times is generally referred to as a "constitutive promoter". A promoter that causes a gene to be expressed in a specific cell type is generally referred to as a "cell-specific promoter" or "tissue-specific promoter". A promoter that causes a gene to be expressed at a specific stage of development or cell differentiation is generally referred to as a "development-specific promoter" or "cell differentiation-specific promoter". A promoter that is induced to express a gene after exposure or treatment of cells with a drug, biomolecule, chemical, ligand, light, etc. that induces the promoter is generally referred to as an "inducible promoter" or "regulatable promoter". It is further recognized that in most cases, the exact boundaries of regulatory sequences are not fully defined, so that DNA fragments of different lengths may have the same promoter activity. Additional exemplary promoters are discussed elsewhere in this disclosure.

[0073] A promoter sequence is typically bounded at its 3' end by a transcription initiation site and extends upstream (5' direction) to incorporate the minimum number of bases or elements necessary to induce transcription at a detectable level above background. Within the promoter sequence will be found a transcription initiation site (conveniently defined, for example, by mapping with nuclease S1), as well as protein binding domains (consensus sequences) responsible for the binding of RNA polymerase.

[0074] In some embodiments, the nucleic acid molecule comprises a tissue-specific promoter. In certain embodiments, the tissue-specific promoter drives the expression of therapeutic protein in liver, hepatocytes, and / or endothelial cells. In a particular embodiment, the promoter comprises the TTP promoter. In a particular embodiment, the promoter comprises the mTTR promoter.

[0075] The terms "restriction endonucleases" and "restriction enzymes" are used interchangeably and refer to enzymes that bind to and cut within specific nucleotide sequences within double-stranded DNA.

[0076] The term "plasmid" refers to an extrachromosomal element that often carries genes that are not part of the central metabolism of the cell and are usually in the form of circular double-stranded DNA molecules. Such elements can be autonomously replicating sequences, genomic integration sequences, phages, or nucleotide sequences of single- or double-stranded DNA or RNA, linear, circular, or supercoiled, from any source, in which multiple nucleotide sequences are joined or recombined into unique constructs that can introduce promoter fragments and DNA sequences for selected gene products, along with appropriate 3' untranslated sequences, into cells.

[0077] Eukaryotic viral vectors that may be used include, but are not limited to, adenovirus vectors, retrovirus vectors, adeno-associated virus vectors, poxviruses, such as vaccinia virus vectors, baculovirus vectors, or herpes virus vectors. Non-viral vectors include plasmids, liposomes, electrically charged lipids (cytofectins), DNA-protein complexes, and biopolymers.

[0078] "Cloning vector" refers to a "replicon", a unit length of sequentially replicated nucleic acid, such as a plasmid, phage, or cosmid, to which another nucleic acid segment is attached to effect replication of the attached segment, and which contains an origin of replication. Certain cloning vectors are capable of replication in one cell type, e.g., bacteria, and expression in another cell, e.g., eukaryotic cells. Cloning vectors typically contain one or more sequences used for the insertion of a nucleic acid sequence of interest into the vector and / or for the selection of cells that contain one or more multiple cloning sites.

[0079] The term "expression vector" refers to a vehicle designed to allow for the expression of an inserted nucleic acid sequence after insertion into a host cell, the inserted nucleic acid sequence being placed in operable association with a regulatory region, as described above.

[0080] The vector is introduced into the host cell by methods well known in the art, such as transfection, electroporation, microinjection, transduction, cell fusion, DEAE dextran, calcium phosphate precipitation, lipofection (lysosomal fusion), using a gene gun, or a DNA vector transporter. As used herein, "culture", "to culture" and "to culture" refer to incubating cells under in vitro conditions that allow cells to grow or divide, or to maintaining cells in a viable state. As used herein, "cultured cells" refers to cells that have been propagated in vitro.

[0081] As used herein, the term "polypeptide" is intended to encompass the singular "polypeptide" as well as the plural "polypeptides" and refers to a molecule composed of monomers (amino acids) linked in a linear chain by amide bonds (also known as peptide bonds). The term "polypeptide" refers to any chain or chains of two or more amino acids and does not refer to a specific length of the product. Thus, peptide, dipeptide, tripeptide, oligopeptide, "protein", "amino acid chain", or any other term used to refer to one or more chains of two or more amino acids are included within the definition of "polypeptide", and the term "polypeptide" may be used in place of or interchangeably with any of these terms. The term "polypeptide" is also intended to refer to post-expression modified products of a polypeptide, including, without limitation, glycosylation, acetylation, phosphorylation, amidation, derivatization with known protecting / blocking groups, proteolytic cleavage, or modification with non-naturally occurring amino acids. A polypeptide may be derived from a natural biological source or may be produced by recombinant technology, but is not necessarily translated from a designated nucleic acid sequence. The polypeptides may be produced in any manner, including by chemical synthesis.

[0082] The term "amino acid" includes alanine (Ala or A); arginine (Arg or R); asparagine (Asn or N); aspartic acid (Asp or D); cysteine ​​(Cys or C); glutamine (Gln or Q); glutamic acid (Glu or E); glycine (Gly or G); histidine (His or H); isoleucine (Ile or I); leucine (Leu or L); lysine (Lys or K); methionine (Met or M); phenylalanine (Phe or F); proline (Pro or P); serine (Ser or S); threonine (Thr or T); tryptophan (Trp or W); tyrosine (Tyr or Y); and valine (Val or V). Unconventional amino acids are also within the scope of the present disclosure, including norleucine, ornithine, norvaline, homoserine, and other amino acid residue analogs, such as those described in Ellman et al., Meth. Enzym., 202:301-336 (1991). To generate such unnatural amino acid residues, the procedures of Noren et al., Science, 244:182 (1989); and Ellman et al., supra, are used. Briefly, these procedures involve chemical activation of a suppressor tRNA with the unnatural amino acid residue, followed by in vitro transcription and translation of the RNA. Introduction of unconventional amino acids can also be accomplished using peptide chemistry reactions known in the art. As used herein, the term "polar amino acid" includes amino acids that have a net charge of zero, but have nonzero partial charges at different portions of their side chains (e.g., M, F, W, S, Y, N, Q, C). These amino acids may participate in hydrophobic and electrostatic interactions. As used herein, the term "charged amino acids" includes amino acids that may have a non-zero net charge on their side chains (e.g., R, K, H, E, D). These amino acids may participate in hydrophobic and electrostatic interactions.

[0083] The present disclosure also includes fragments or variants of the polypeptides, and any combination thereof. The term "fragment" or "variant" when referring to the polypeptide-binding domains or polypeptide-binding molecules of the present disclosure includes any polypeptide that retains at least some of the properties of the reference polypeptide (e.g., FcRn binding affinity for FcRn binding domains or Fc variants, coagulation activity for FVIII variants, or FVIII binding activity for VWF fragments). Polypeptide fragments include proteolytic fragments as well as deletion fragments, in addition to specific antibody fragments discussed elsewhere herein, but do not include naturally occurring full-length polypeptides (or mature polypeptides). Variants of the polypeptide-binding domains or polypeptide-binding molecules of the present disclosure include the fragments described above, and also include polypeptides in which the amino acid sequence is altered due to amino acid substitution, deletion, or insertion. Variants may be naturally occurring variants or non-naturally occurring variants. Non-naturally occurring variants are generated using mutagenesis methods known in the art. Variant polypeptides can contain conservative amino acid substitutions, deletions, or additions, or can contain non-conservative amino acid substitutions, deletions, or additions.

[0084] "Conservative amino acid substitution" refers to the amino acid substitution in which an amino acid residue is replaced with an amino acid residue having a similar side chain. In the art, a family of amino acid residues with similar side chains is defined, including basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine), non-polar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), beta-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Thus, if an amino acid in a polypeptide is replaced with another amino acid from the same side chain family, the substitution is considered to be conservative. In another embodiment, a stretch of amino acids is conservatively replaced with a structurally similar stretch of side chain family members that differs in order and / or composition.

[0085] The term "percent identity," as known in the art, is a relationship between two or more polypeptide sequences, or two or more polynucleotide sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of sequence relatedness between polypeptide or polynucleotide sequences, as the case may be, as determined by the match between strings of such sequences. "Identity" is readily calculated by known methods, including but not limited to those described in "Computational Molecular Biology" (Lesk, AM, ed.), Oxford University Press, New York (1988); "Biocomputing: Informatics and Genome Projects" (Smith, DW, ed.), Academic Press, New York (1993); "Computer Analysis of Sequence Data", Part I (Griffin, AM and Griffin, HG, eds.), Humana Press, New Jersey (1994); "Sequence Analysis in Molecular Biology" (Von Heijne, G., ed.), Academic Press (1987); and "Sequence Analysis Primer" (Gribskov, M. and Devereux, J., eds.), Stockton Press, New York (1991). Preferred methods of determining identity are designed to give the best match between the sequences tested. Methods of determining identity are codified in publicly available computer programs.Sequence alignment and percent identity calculations are performed using sequence analysis software such as the Megalign program of the LASERGENE bioinformatics computing software package (DNASTAR, Inc., Madison, WI), the GCG program package (Wisconsin Package Version 9.0, Genetics Computer Group (GCG), Madison, WI); BLASTP, BLASTN, BLASTX (Altschul et al., J. Mol. Biol., 215:403 (1990)); and DNASTAR (DNASTAR, Inc. 1228 S. Park St. Madison, WI 53715 USA). In the context of this application, when sequence analysis software is used for analysis, it will be understood that the results of the analysis will be based on the "default values" of the program referenced, unless otherwise specified. As used herein, "default values" refers to any set of values ​​or parameters that were originally loaded by the software when it was first initialized. For purposes of determining the percent identity between a query sequence (e.g., a nucleic acid sequence) and a reference sequence, only those nucleotides in the query sequence that match nucleotides in the reference sequence are used in calculating the percent identity. Thus, in determining the percent identity between a query sequence, or a specified portion thereof (e.g., nucleotides 1-522), and a reference sequence, the percent identity will be calculated by dividing the number of matched nucleotides by the total number of nucleotides in the complete query sequence.

[0086] As used herein, nucleotides corresponding to nucleotides in a particular sequence of the present disclosure are identified by alignment of the sequences of the present disclosure to maximize identity with respect to the reference sequence. The numbers used to identify equivalent amino acids in the reference sequence are based on the numbers used to identify the corresponding amino acids in the sequences of the present disclosure.

[0087] As used herein, treating, treatment, or treating refers to, for example, reducing the severity of a disease or condition; shortening the duration of the course of a disease; ameliorating one or more symptoms associated with a disease or condition; imparting a beneficial effect to a subject with a disease or condition, without necessarily curing the disease or condition; or preventing one or more symptoms associated with a disease or condition.

[0088] As used herein, "administering" refers to administering a pharma- ceutically acceptable nucleic acid molecule, a polypeptide expressed therefrom, or a vector comprising a nucleic acid molecule of the present disclosure to a subject via a pharma- ceutically acceptable route. The route of administration can be intravenous, e.g., intravenous injection and intravenous infusion. Additional routes of administration include, e.g., subcutaneous, intramuscular, oral, nasal, and pulmonary administration. The nucleic acid molecule, polypeptide, and vector are administered as part of a pharmaceutical composition that includes at least one excipient.

[0089] As used herein, "pharmaceutical acceptable" refers to molecular entities and compositions that are physiologically acceptable and typically do not cause toxic or allergic or similar undesirable reactions, such as heartburn, dizziness, etc., when administered to humans.In some cases, the term "pharmaceutical acceptable" as used herein means approved by a regulatory agency of the United States Federal or State Government for use in animals, more particularly for use in humans, or listed in the United States Pharmacopeia or other generally recognized pharmacopoeias.

[0090] As used herein, the phrase "subject in need thereof" includes a subject, such as a mammalian subject, who will benefit from administration of a nucleic acid molecule, polypeptide, or vector of the present disclosure. In some embodiments, the subject is a human subject. In some embodiments, the subject is an individual with hemophilia. The subject may be an adult or a juvenile (e.g., under the age of 12).

[0091] As used herein, the term "therapeutic protein" refers to any polypeptide known in the art that is administered to a subject. In some embodiments, the therapeutic protein comprises a protein selected from a clotting factor, a growth factor, an antibody, a functional fragment thereof, or a combination thereof. As used herein, the term "clotting factor" refers to a naturally occurring or recombinantly produced molecule or analog thereof that prevents or reduces the persistence of bleeding episodes in a subject. In other words, the term "clotting factor" refers to a molecule that has procoagulant activity, i.e., a molecule that contributes to the conversion of fibrinogen into a mesh of insoluble fibrin that causes blood to coagulate (coagulate or clot). As used herein, "clotting factor" includes activated clotting factors, their zymogens, or activatable clotting factors. An "activatable clotting factor" is a clotting factor in an inactive form (e.g., in its zymogen form) that can be converted to an active form. The term "clotting factor" refers to factor I (FI), factor II (FII), factor III (FIII), factor IV (FIV), factor V (FV), factor FVI (FVII), factor FVIII (FVIII), factor FIX (FIX), factor X (FX), factor XI (FXI), factor XII (FXII), factor XIII (FXIII), von Willebrand factor (VWF), prekallikrein, high molecular weight kininogen, fibronectin, amyloidosis, fibronectin, erythrocyte serine / erythrocyte sarcoma, fibronectin, ... Examples of antibodies include, but are not limited to, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2 antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor 1 (PAI-1), plasminogen activator inhibitor 2 (PAI-2), zymogens thereof, activated forms thereof, or any combination thereof.

[0092] As used herein, "clotting activity" means the ability to participate in the cascade of biochemical reactions that lead to the formation of a fibrin clot and / or to reduce the severity, duration, or frequency of bleeding or bleeding episodes.

[0093] As used herein, the term "growth factor" includes any growth factor known in the art, including cytokines and hormones.

[0094] In some embodiments, the therapeutic protein is encoded by a gene selected from the X-linked gene dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha glucosidase), or any combination thereof.

[0095] As used herein, the term "heterologous" or "exogenous" refers to a molecule that is not normally found in a given context, e.g., within a cell or polypeptide. For example, an exogenous or heterologous molecule is introduced into a cell and is present only after manipulation of the cell, e.g., by transfection or other form of genetic engineering, whereas a heterologous amino acid sequence may be present within a protein where it is not found in nature.

[0096] As used herein, the term "optimized" in relation to a nucleotide sequence refers to a polynucleotide sequence that codes for a polypeptide, where the polynucleotide sequence is mutated to enhance the properties of the polynucleotide sequence. In some embodiments, the optimization is performed to increase transcription levels, increase translation levels, increase steady-state mRNA levels, increase or decrease binding to regulatory proteins such as general transcription factors, increase or decrease splicing, or increase the yield of the polypeptide produced by the polynucleotide sequence. Examples of changes that can be made to a polynucleotide sequence to optimize the nucleotide sequence include codon optimization, G / C content optimization, removal of repeat sequences, removal of AT-rich elements, removal of cryptic splice sites, removal of cis-activating elements that suppress transcription or translation, addition or removal of poly-T or poly-A sequences, addition of sequences near the transcription start site that enhance transcription, such as Kozak consensus sequences, removal of sequences that can form stem-loop structures, removal of destabilizing sequences, and combinations of two or more of these.

[0097] II. Nucleic acid molecules Certain aspects of the present disclosure aim to overcome the deficiencies of AAV vectors for gene therapy. In particular, certain aspects of the present disclosure are directed to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette, for example, encoding a therapeutic protein and / or a therapeutic miRNA. In some embodiments, the first ITR and the second ITR flank a gene cassette comprising a heterologous polynucleotide sequence. In some embodiments, the nucleic acid molecule does not comprise a gene encoding a capsid protein, a replication protein, and / or an assembly protein. In some embodiments, the gene cassette encodes a therapeutic protein. In some embodiments, the therapeutic protein comprises a clotting factor. In some embodiments, the gene cassette encodes a miRNA. In certain embodiments, the gene cassette is disposed between the first ITR and the second ITR. In some embodiments, the nucleic acid molecule further comprises one or more non-coding regions. In certain embodiments, the one or more non-coding regions comprise a promoter sequence, an intron, a post-transcriptional regulatory element, a 3'UTR poly(A) sequence, or any combination thereof.

[0098] In one embodiment, the gene cassette is a single-stranded nucleic acid. In another embodiment, the gene cassette is a double-stranded nucleic acid. In another embodiment, the gene cassette is a closed-end double-stranded ceDNA.

[0099] In some embodiments, the nucleic acid molecule comprises: (a) a first ITR, which is an ITR from a member of the Parvoviridae family other than AAV (e.g., B19 or GPV ITR); (b) a tissue-specific promoter sequence, e.g., the TTP promoter or the TTR promoter; (c) an intron, e.g., a synthetic intron; (d) a nucleotide encoding a miRNA or a therapeutic protein, e.g., a clotting factor; (e) a post-transcriptional regulatory element, e.g., a WPRE; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) a second ITR, which is an ITR from a member of the Parvoviridae family other than AAV (e.g., B19 or GPV ITR). In some embodiments, the nucleic acid molecule comprises: (a) a first ITR, which is an ITR from a member of the Parvoviridae family other than AAV (e.g., B19 or GPV ITR); (b) a tissue-specific promoter sequence, e.g., the mTTR promoter; (c) an intron, e.g., a synthetic intron; (d) a nucleotide encoding a miRNA or a therapeutic protein, e.g., a clotting factor; (e) a post-transcriptional regulatory element, e.g., a WPRE; (f) a 3'UTR poly(A) tail sequence, e.g., bGHpA; and (g) a second ITR, which is an ITR from a member of the Parvoviridae family other than AAV (e.g., B19 or GPV ITR).

[0100] In some embodiments, the nucleic acid molecule comprises a first ITR, a second ITR, and a gene cassette encoding a target sequence, where the target sequence encodes a therapeutic protein, and the therapeutic protein comprises a factor VIII (FVIII) polypeptide.

[0101] In some embodiments, the gene cassette comprises a nucleotide sequence encoding a codon-optimized FVIII driven by an mTTR promoter. In some embodiments, the mTTR promoter comprises the nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the gene cassette further comprises an A1MB2 enhancer element. In some embodiments, the A1MB2 enhancer element comprises the nucleic acid sequence of SEQ ID NO: 30. In some embodiments, the gene cassette further comprises a chimeric intron or a synthetic intron. In some embodiments, the chimeric intron consists of a chicken beta-actin intron / rabbit beta-globin intron, modified to eliminate five existing ATG sequences to reduce false translation initiation. In some embodiments, the intron sequence is located 5' to the nucleic acid sequence encoding the FVIII polypeptide. In some embodiments, the chimeric intron is located 5' to a promoter sequence, such as the mTTR promoter. In some embodiments, the chimeric intron comprises the nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene cassette further comprises a woodchuck post-transcriptional regulatory element (WPRE). In some embodiments, the WPRE comprises the nucleic acid sequence of SEQ ID NO: 33. In some embodiments, the gene cassette further comprises a bovine growth hormone polyadenylation (bGHpA) signal. In some embodiments, the bGHpA signal comprises the nucleic acid sequence of SEQ ID NO: 34. In some embodiments, the gene cassette comprises a nucleotide sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity to SEQ ID NO: 27. In some embodiments, the gene cassette comprises the nucleotide sequence of SEQ ID NO: 27.

[0102] In some embodiments herein, an isolated nucleic acid molecule is disclosed that encodes a FVIII protein. In some embodiments herein, an isolated nucleic acid molecule is disclosed that encodes a FVIII protein and comprises a nucleotide sequence that is at least about 75% identical to SEQ ID NO: 28. In some embodiments herein, an isolated nucleic acid molecule is disclosed that encodes a FVIII protein and comprises a nucleotide sequence shown in SEQ ID NO: 28.

[0103] In some embodiments herein, an isolated nucleic acid molecule is disclosed that encodes a FVIII protein. In some embodiments herein, an isolated nucleic acid molecule is disclosed that encodes a FVIII protein and comprises a nucleotide sequence that is at least about 75% identical to SEQ ID NO: 29. In some embodiments herein, an isolated nucleic acid molecule is disclosed that encodes a FVIII protein and comprises a nucleotide sequence shown in SEQ ID NO: 29.

[0104] A. Inverted terminal repeat (ITR) Certain embodiments of the present disclosure are directed to nucleic acid molecules that include a first ITR, e.g., a 5' ITR, and a second ITR, e.g., a 3' ITR. Typically, ITRs are involved in the replication and rescue or excision of parvovirus (e.g., AAV) DNA from prokaryotic plasmids (Samulski et al., 1983, 1987; Senapathy et al., 1984; GottliebandMuzyczka, 1988). In addition, ITRs are also considered to be the minimal sequences required for the integration of AAV provirus and packaging of AAV DNA into virions (McLaughlin et al., 1988; Samulski et al., 1989). These elements are essential for efficient replication of parvovirus genomes. It is hypothesized that the minimal canonical elements essential for ITR function are Rep binding sites and terminal separation sites plus a variable palindrome that allows hairpin formation. Palindromic nucleotide regions usually function together in cis as origins of DNA replication and packaging signals for viruses. Complementary sequences in ITRs fold into hairpin structures during DNA replication. In some embodiments, ITRs fold into T-shaped hairpin structures. In other embodiments, ITRs fold into hairpin structures other than T-shaped, such as U-shaped hairpin structures. Data suggest that the T-shaped hairpin structure of AAV ITRs can inhibit the expression of transgenes flanked by ITRs. See, for example, Zhou et al., Scientific Reports, 7:5432 (July 14, 2017). By utilizing ITRs that do not form T-shaped hairpin structures, this form of inhibition is avoided. Thus, in certain aspects, polynucleotides that include non-AAV ITRs have improved transgene expression compared to polynucleotides that include AAV ITRs that form T-shaped hairpins.

[0105] In some embodiments, the ITR comprises a naturally occurring ITR, for example, the ITR comprises all or a portion of an ITR from a member of the Parvoviridae family. In some embodiments, the ITR comprises a synthetic sequence. In one embodiment, the first ITR or the second ITR comprises a synthetic sequence. In another embodiment, each of the first ITR and the second ITR comprises a synthetic sequence. In some embodiments, the first ITR or the second ITR comprises a naturally occurring sequence. In another embodiment, each of the first ITR and the second ITR comprises a naturally occurring sequence.

[0106] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, e.g., a truncated ITR. In some embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, where the fragment is at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, at least about 50 nucleotides, at least about 55 nucleotides, at least about 60 nucleotides, at least about 65 nucleotides, at least about 70 nucleotides, at least about 75 nucleotides, at least about 80 nucleotides, at least about 85 nucleotides, at least about 90 nucleotides, at least about 95 nucleotides, at least about 100 nucleotides, At least about 125 nucleotides, at least about 150 nucleotides, at least about 175 nucleotides, at least about 200 nucleotides, at least about 225 nucleotides, at least about 250 nucleotides, at least about 275 nucleotides, at least about 300 nucleotides, at least about 325 nucleotides, at least about 350 nucleotides, at least about 375 nucleotides, at least about 400 nucleotides, at least about 425 nucleotides, at least about 450 nucleotides, at least about 475 nucleotides, at least about 500 nucleotides, at least about 525 nucleotides, at least about 550 nucleotides, at least about 575 nucleotides, or at least about 600 nucleotides; the ITR retains the functional properties of the naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, where the fragment comprises at least about 129 nucleotides; the ITR retains the functional properties of the naturally occurring ITR. In certain embodiments, the ITR comprises or consists of a fragment of a naturally occurring ITR, where the fragment comprises at least about 102 nucleotides; the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR retains the Rep binding element (RBE) of the wild-type ITR from which it is derived.In some embodiments, an ITR retains at least one of the RBEs of the wild-type ITR from which it is derived. In some embodiments, an ITR retains at least one of the RBEs or a functional portion thereof of the wild-type ITR from which it is derived. Preservation of the RBE can be important for ITR stability and manufacturing purposes.

[0107] In some embodiments, the ITR comprises or consists of a portion of a naturally occurring ITR, where the fragment comprises at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the length of the naturally occurring ITR; and the fragment retains the functional properties of the naturally occurring ITR.

[0108] In certain embodiments, the ITRs, when properly aligned, have a sequence similar to that of the homologous naturally occurring ITR portion, at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, at least 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 100%, at least 101%, at least 102%, at least 103%, at least 104%, at least 105%, at least 106%, at least 107%, at least 108%, at least 109%, at least 110%, at least 111%, at least 112%, at least 113%, at least 114%, at least 115%, at least 116%, at least 117%, at least 118%, at least 119%, at least 120%, at least 121%, at least 122%, at least 123%, at least 124%, at least 125%, In another embodiment, the ITR comprises or consists of a sequence having at least 5%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to the naturally occurring ITR; in which case the ITR retains the functional properties of the naturally occurring ITR. In another embodiment, the ITR comprises or consists of a sequence having at least 90% sequence identity to the portion of the homologous naturally occurring ITR when properly aligned; in which case the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 80% sequence identity to a portion of a homologous naturally occurring ITR; in this case, the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 70% sequence identity to a portion of a homologous naturally occurring ITR; in this case, the ITR retains the functional properties of the naturally occurring ITR.In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 60% sequence identity to a portion of a homologous naturally occurring ITR; in this case, the ITR retains the functional properties of the naturally occurring ITR. In some embodiments, the ITR comprises or consists of a sequence that, when properly aligned, has at least 50% sequence identity to a portion of a homologous naturally occurring ITR; in this case, the ITR retains the functional properties of the naturally occurring ITR.

[0109] In some embodiments, ITR comprises ITR from AAV genome.In some embodiments, ITR is ITR of AAV genome selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10 and any combination thereof.In some embodiments, ITR is ITR of any AAV genome known to those skilled in the art, including natural isolate, for example natural human isolate.In certain embodiments, ITR is ITR of AAV2 genome.In another embodiment, ITR is a synthetic sequence that is engineered to comprise ITR from one or more of AAV genomes at its 5' end and 3' end.

[0110] In some embodiments, the ITRs are not derived from the AAV genome (i.e., the ITRs are derived from a virus that is not AAV). In some embodiments, the ITRs are non-AAV ITRs. In some embodiments, the ITRs are non-AAV ITRs from the Parvoviridae family, including but not limited to the following: Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Abeparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, Penstildensovirus, and any combination thereof. In certain embodiments, the ITRs are derived from Erythrovirus B19 (a human virus), a parvovirus. In another embodiment, the ITRs are derived from the Muscovy Duck Parvovirus (MDPV) strain. In certain embodiments, the MDPV strain is an attenuated MDPV strain, such as MDPV FZ91-30 strain. In other embodiments, the MDPV strain is a pathogenic MDPV strain, such as MDPV YY strain. In some embodiments, the ITRs are derived from porcine parvovirus, such as porcine parvovirus U44978 strain. In some embodiments, the ITRs are derived from minute virus of mice, such as minute virus of mice U34256 strain. In some embodiments, the ITRs are derived from canine parvovirus, such as canine parvovirus M19296 strain. In some embodiments, the ITRs are derived from mink enteritis virus, such as mink enteritis virus D00765 strain. In some embodiments, the ITRs are derived from the genus Dependoparvovirus. In one embodiment, the genus Dependoparvovirus is a Dependovirus goose parvovirus (GPV) strain. In a specific embodiment, the GPV strain is an attenuated GPV strain, such as GPV 82-0321V strain. In another specific embodiment, the GPV strain is a pathogenic GPV strain, such as a GPV B strain.

[0111] The first and second ITRs of a nucleic acid molecule may be derived from the same genome, e.g., the genome of the same virus, or may be derived from different genomes, e.g., the genomes of two or more different viral genomes (also known as "hybrid" ITRs). In certain embodiments, the first and second ITRs are derived from the same AAV genome. In specific embodiments, the two ITRs present in the nucleic acid molecule of the present invention may be the same, in particular the AAV2 ITR. In other embodiments, the first ITR is derived from the AAV genome and the second ITR is not derived from the AAV genome (e.g., derived from a genome other than AAV). In other embodiments, the first ITR is not derived from the AAV genome (e.g., derived from a genome other than AAV) and the second ITR is derived from the AAV genome. In yet other embodiments, neither the first ITR nor the second ITR is derived from the AAV genome (e.g., derived from a genome other than AAV). In a particular embodiment, the first and second ITRs are identical.

[0112] In some embodiments, the first ITR is derived from a genome other than AAV and the second ITR is derived from a genome other than AAV, where the first ITR and the second ITR are derived from the same genome. Non-limiting examples of viral genomes other than AAV are from the genera Bocavirus, Dependovirus, Erythrovirus, Amdovirus, Parvovirus, Densovirus, Iteravirus, Contravirus, Abeparvovirus, Copiparvovirus, Protoparvovirus, Tetraparvovirus, Ambidensovirus, Brevidensovirus, Hepandensovirus, and Penstildensovirus. In some embodiments, the first ITR is derived from a genome other than AAV and the second ITR is derived from a genome other than AAV, where the first ITR and the second ITR are derived from different viral genomes (also known as "hybrid" ITRs). In some embodiments, the first ITR is derived from the B19 genome and the second ITR is derived from GPV. In some embodiments, the first ITR is derived from the GPV genome and the second ITR is derived from B19.

[0113] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from a parvovirus, erythrovirus B19 (a human virus). In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from a parvovirus, erythrovirus B19 (a human virus).

[0114] In some embodiments, the first ITR comprises or consists of the whole or part of the ITR from AAV genome or a genome other than AAV, and the second ITR comprises or consists of the whole or part of the ITR from AAV genome or a genome other than AAV.In some embodiments, the part of the ITR from AAV genome or a genome other than AAV is a truncated form of the ITR from naturally occurring AAV genome or a genome other than AAV.In some embodiments, the part of the ITR from AAV genome or a genome other than AAV comprises the part of the ITR from naturally occurring AAV genome or a genome other than AAV.For example, the part of the ITR from AAV genome or a genome other than AAV comprises the part of the ITR from naturally occurring AAV genome or a genome other than AAV, and at least one RBE or its functional part is conservative.

[0115] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a part of the ITR from B19. In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a part of the ITR from B19. In some embodiments, the second ITR is the reverse complement of the first ITR. In some embodiments, the first ITR is the reverse complement of the second ITR. In some embodiments, the first ITR and / or the second ITR from B19 can form a hairpin structure. In certain embodiments, the hairpin structure does not include a T-shaped hairpin.

[0116] In some embodiments, the first ITR and / or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NOs: 1-8, 17, or 18, wherein the first ITR and / or second ITR retains the functional properties of the B19 ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 1-8, 17, or 18, where the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0117] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NO: 1-8, 17, or 18. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 1. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 2. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 3. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 4. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 5. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 6. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 7. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 8. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 17. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 18.

[0118] In some embodiments, the first ITR is derived from the AAV genome and the second ITR is derived from goose parvovirus (GPV). In other embodiments, the second ITR is derived from the AAV genome and the first ITR is derived from goose parvovirus (GPV).

[0119] In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a part of an ITR derived from a GPV. In certain embodiments, the first ITR and / or the second ITR comprises or consists of all or a part of an ITR derived from a GPV. In some embodiments, the second ITR is the reverse complement of the first ITR. In some embodiments, the first ITR is the reverse complement of the second ITR. In some embodiments, the first ITR and / or the second ITR derived from a GPV can form a hairpin structure. In certain embodiments, the hairpin structure does not include a T-shaped hairpin.

[0120] In some embodiments, the first and / or second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to the nucleotide sequence set forth in SEQ ID NOs:9-16 or 19-22, wherein the first and / or second ITR retains the functional properties of the GPV ITR from which it is derived. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence that is at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to a nucleotide sequence selected from SEQ ID NOs: 9-16 or 19-22, where the first ITR and / or the second ITR is capable of forming a hairpin structure. In certain embodiments, the hairpin structure does not comprise a T-shaped hairpin.

[0121] In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 9-16 or 19-22. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 9. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 10. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 11. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 12. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 13. In some embodiments, the first ITR and / or the second ITR comprises or consists of a nucleotide sequence set forth in SEQ ID NO: 14. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 15. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 16. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 19. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 20. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 21. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 22. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 23. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO:24.In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 25. In some embodiments, the first ITR and / or the second ITR comprises or consists of the nucleotide sequence set forth in SEQ ID NO: 26.

[0122] In some embodiments, the first ITR and / or second ITR comprises a polynucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to nucleotides 1-49, 50-58, and 59-125 of SEQ ID NO:1, or nucleotides 1-27 and 50-114 of SEQ ID NO:15, SEQ ID NO:23, or SEQ ID NO:25. In some embodiments, the first ITR and / or the second ITR comprises a polynucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to nucleotides 1-67, 68-76, and 77-125 of SEQ ID NO:2, or nucleotides 1-65 and 88-114 of SEQ ID NO:16, SEQ ID NO:23, or SEQ ID NO:26.

[0123] Those skilled in the art will appreciate that any of the first ITR sequences described herein can be matched with any of the second ITR sequences described herein. In some embodiments, the first ITR sequence described herein is the 5' ITR sequence. In some embodiments, the second ITR sequence described herein is the 3' ITR sequence. In some embodiments, the second ITR sequence described herein is the 5' ITR sequence. In some embodiments, the first ITR sequence described herein is the 3' ITR sequence. Those skilled in the art will be able to determine the appropriate orientation of the first ITR and second ITR described herein for the construction of the gene cassette.

[0124] In another specific embodiment, the ITR is a synthetic sequence engineered to contain an ITR at its 5'-end and 3'-end that is not derived from the AAV genome. In another specific embodiment, the ITR is a synthetic sequence engineered to contain an ITR at its 5'-end and 3'-end that is derived from one or more non-AAV genomes. The two ITRs present in the nucleic acid molecule of the present invention can be from the same non-AAV genome or from different non-AAV genomes. In particular, the ITRs can be from the same non-AAV genome. In a specific embodiment, the two ITRs present in the nucleic acid molecule of the present invention can be the same, in particular, AAV2 ITR.

[0125] In some embodiments, the ITR sequence comprises one or more palindromic sequences. The palindromic sequences of the ITRs disclosed herein include, but are not limited to, naturally occurring palindromic sequences (i.e., sequences found in nature), synthetic sequences such as pseudopalindromic sequences (i.e., sequences not found in nature), and combinations or modifications thereof. A "pseudopalindromic sequence" is a palindromic DNA sequence, including imperfect palindromic sequences, that share less than 80% or no nucleic acid sequence identity, including less than 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5%, with sequences in natural AAV or non-AAV palindromic sequences that form secondary structures. Natural palindromic sequences are obtained or derived from any genome disclosed herein. Synthetic palindromic sequences can be based on any genome disclosed herein.

[0126] The palindrome may be a continuous sequence or an interrupted sequence. In some embodiments, the interrupted sequence comprises an insertion of a second sequence. In some embodiments, the second sequence comprises a promoter, an enhancer, an integration site for an integrase (e.g., a site for Cre recombinase or Flp recombinase), an open reading frame for a gene product, or a combination thereof.

[0127] In some embodiments, the ITR forms a hairpin loop structure. In one embodiment, the first ITR forms a hairpin structure. In another embodiment, the second ITR forms a hairpin structure. In yet another embodiment, both the first ITR and the second ITR form a hairpin structure. In some embodiments, the first ITR and / or the second ITR do not form a T-shaped hairpin structure. In certain embodiments, the first ITR and / or the second ITR form a hairpin structure other than a T-shaped structure. In some embodiments, the hairpin structure other than a T-shaped structure comprises a U-shaped hairpin structure.

[0128] In some embodiments, the ITRs in the nucleic acid molecules described herein may be transcriptionally activating ITRs. The transcriptionally activating ITRs may comprise all or part of the wild-type ITRs that have been transcriptionally activated by the incorporation of at least one transcriptionally active element. Various types of transcriptionally active elements are suitable for use in this context. In some embodiments, the transcriptionally active element is a constitutive transcriptionally active element. Constitutive transcriptionally active elements provide sustained levels of gene transcription and are preferred when it is desired that the transgene be expressed on a sustained basis. In other embodiments, the transcriptionally active element is an inducible transcriptionally active element. Inducible transcriptionally active elements generally exhibit low activity in the absence of an inducer (or an inducing condition) and are upregulated in the presence of an inducer (or a switch to an inducing condition). Inducible transcriptionally active elements may be preferred when expression is desired only at a certain time or at a certain location, or when it is desired to titrate the expression level using an inducer. Transcriptionally active elements can also be tissue specific; that is, active only in certain tissues or cell types.

[0129] The transcriptionally active element is incorporated into the ITR in various ways. In some embodiments, the transcriptionally active element is incorporated 5' to any part of the ITR or 3' to any part of the ITR. In other embodiments, the transcriptionally active element of the transcriptionally activating ITR is between two ITR sequences. If the transcriptionally active element contains two or more elements that must be separated, these elements alternate with parts of the ITR. In some embodiments, the hairpin structure of the ITR is deleted and replaced by an inverted repeat of the transcription element. This latter arrangement would create a hairpin that mimics the deleted part in the structure. There may be multiple tandem transcriptionally active elements in the transcriptionally activating ITR, which may be adjacent or separated. In addition, protein binding sites (e.g., Rep binding sites) are also introduced into the transcriptionally active element of the transcriptionally activating ITR. The transcriptionally active element may include any sequence that allows the control of transcription of DNA by RNA polymerase to form RNA, and may include, for example, the transcriptionally active elements defined below.

[0130] Transcriptionally activating ITRs provide both transcriptional activation and ITR functions to a nucleic acid molecule in a relatively limited nucleotide sequence length, which effectively maximizes the length of the transgene that is carried and expressed from the nucleic acid molecule. The incorporation of transcriptionally activating elements into ITRs can be accomplished in a variety of ways. Comparison of ITR sequences and sequence requirements of transcriptionally activating elements can provide insight into the manner of encoding elements within the ITR. For example, transcriptional activity is added to an ITR through the introduction of specific changes in the ITR sequence that duplicate the functional elements of the transcriptionally activating element. There are numerous techniques in the art that efficiently add, delete, and / or change specific nucleotide sequences at specific sites (see, for example, Deng and Nickoloff (1992), Anal. Biochem., 200:81-88). Another way of creating transcriptionally activating ITRs involves the introduction of restriction sites at desired positions within the ITR. In addition, multiple transcriptionally activating elements are incorporated into transcriptionally activating ITRs using methods known in the art.

[0131] By way of example, transcriptionally activating ITRs are created by the incorporation of one or more transcriptionally active elements, such as a TATAbox, a GCbox, a CCAATbox, an Sp1 site, an Inr region, a CRE (cAMP regulatory element) site, an ATF-1 / CRE site, an APBβbox, an APBαbox, a CArGbox, a CCACbox, or any other element involved in transcription known in the art.

[0132] An embodiment of the present disclosure provides a method for cloning the nucleic acid molecules described herein, comprising inserting the nucleic acid molecule capable of complex secondary structures into a suitable vector and introducing the resulting vector into a suitable bacterial host strain. As is known in the art, complex secondary structures of nucleic acids (e.g., long palindromic regions) can be unstable and difficult to clone in bacterial host strains. For example, nucleic acid molecules comprising the first and second ITRs of the present disclosure (e.g., parvovirus ITRs other than AAV, e.g., B19 ITR or GPV ITR) can be difficult to clone using conventional methods. Long DNA palindromic sequences inhibit DNA replication and are unstable in the genomes of E. coli, Bacillus, Streptococcus, Streptomyces, Saccharomyces cerevisiae, mice, and humans. These effects result from the formation of hairpin or cruciform structures by intrastrand base pairing. In E. coli, inhibition of DNA replication can be significantly overcome in SbcC or SbcD mutants. SbcD is the nuclease subunit and SbcC is the ATPase subunit of the SbcCD complex. The E. coli SbcCD complex is an exonuclease complex that contributes to blocking the replication of long palindromic sequences. The SbcCD complex is a core with ATP-dependent double-stranded DNA exonuclease activity and ATP-independent single-stranded DNA endonuclease activity. SbcCD can collapse replication forks by recognizing DNA palindromic sequences and attacking the resulting hairpin structures.

[0133] In certain embodiments, suitable bacterial host strains are unable to degrade cruciform DNA structures. In certain embodiments, suitable bacterial host strains comprise disruption in the SbcCD complex. In some embodiments, disruption in the SbcCD complex comprises gene disruption in the SbcC gene and / or in the SbcD gene. In certain embodiments, disruption in the SbcCD complex comprises gene disruption in the SbcC gene. In the art, various bacterial host strains are known that comprise gene disruption in the SbcC gene. For example, without limitation, bacterial host strain PMC103 comprises the genotypes sbcC, recD, mcrA, ΔmcrBCF; bacterial host strain PMC107 comprises the genotypes recBC, recJ, sbcBC, mcrA, ΔmcrBCF; bacterial host strain SURE comprises the genotypes recB, recJ, sbcC, mcrA, ΔmcrBCF, umuC, uvrC. Thus, in some embodiments, the method of cloning a nucleic acid molecule described herein comprises inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector and introducing the resulting vector into the host strain PMC103, PMC107, or SURE. In certain embodiments, the method of cloning a nucleic acid molecule described herein comprises inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector and introducing the resulting vector into the host strain PMC103.

[0134] Suitable vectors are known in the art and are described elsewhere herein.In certain embodiments, the vector suitable for use in the cloning method of the present disclosure is a low copy vector.In certain embodiments, the vector suitable for use in the cloning method of the present disclosure is pBR322.

[0135] Thus, the disclosure provides a method of cloning a nucleic acid molecule, comprising inserting a nucleic acid molecule capable of complex secondary structures into a suitable vector and introducing the resulting vector into a bacterial host strain that contains a disruption in the SbcCD complex, wherein the nucleic acid molecule comprises a first inverted terminal repeat (ITR) and a second ITR, and wherein the first ITR and / or the second ITR comprises a nucleotide sequence that is at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NOs: 1-22 or a functional derivative thereof.

[0136] B. Therapeutic Proteins Certain aspects of the present disclosure are directed to nucleic acid molecules comprising a gene cassette encoding a first ITR, a second ITR, and a target sequence, the target sequence encoding a therapeutic protein. In some embodiments, the gene cassette encodes one therapeutic protein. In some embodiments, the gene cassette encodes more than one therapeutic protein. In some embodiments, the gene cassette encodes two or more copies of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more variants of the same therapeutic protein. In some embodiments, the gene cassette encodes two or more different therapeutic proteins.

[0137] Certain embodiments of the present disclosure are directed to a nucleic acid molecule comprising a first ITR, a second ITR, and a gene cassette encoding a therapeutic protein, the therapeutic protein comprising a coagulation factor. In some embodiments, the coagulation factor is selected from the group consisting of FI, FII, FIII, FIV, FV, FVI, FVII, FVIII, FIX, FX, FXI, FXII, FXIII, VWF, prekallikrein, high molecular weight kininogen, fibronectin, antithrombin III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitor (ZPI), plasminogen, alpha 2 antiplasmin, tissue plasminogen activator (tPA), urokinase, plasminogen activator inhibitor 1 (PAI-1), plasminogen activator inhibitor 2 (PAI2), any zymogen thereof, any active form thereof, and any combination thereof. In one embodiment, the coagulation factor comprises FVIII or a variant or fragment thereof. In another embodiment, the clotting factor comprises FIX or a variant or fragment thereof. In another embodiment, the clotting factor comprises FVII or a variant or fragment thereof. In another embodiment, the clotting factor comprises VWF or a variant or fragment thereof.

[0138] In some embodiments, the nucleic acid molecule comprises a gene cassette encoding a first ITR, a second ITR, and a target sequence, where the target sequence encodes a therapeutic protein, and the therapeutic protein comprises a factor VIII polypeptide. As used herein, "factor VIII", abbreviated as "FVIII" throughout this application, means a functional FVIII polypeptide under its normal role in coagulation, unless otherwise specified. Thus, the term "FVIII" includes mutant polypeptides that are functional. "FVIII protein" is used interchangeably with FVIII polypeptide (or protein) or FVIII. Examples of FVIII functions are the ability to activate coagulation, act as a cofactor for factor IX, or bind Ca2 +and the ability to form a tenase complex with factor IX in the presence of phospholipids, which then converts factor X to its activated form, Xa.

[0139] As used herein, a FVIII moiety in a therapeutic protein has FVIII activity, which is measured by any method known in the art. Numerous tests are available to assess the function of the coagulation system: activated partial thromboplastin time (aPTT) test, chromogenic assays, ROTEM assay, prothrombin time (PT) test (also used to determine INR), fibrinogen test (often via the Clauss method), platelet count, platelet function test (often via PFA-100), TCT, bleeding time, mixing test (whether abnormalities are corrected when the patient's plasma is mixed with normal plasma), clotting factor assays, antiphospholipid antibodies, D-dimers, genetic tests (e.g., prothrombin mutation G20210A, which is the factor V Leiden mutation), dilute Russell's snake venom time (dRVVT), other platelet function tests, thromboelastography (TEG or Sonoclot), thromboelastometry (TEM®, e.g., ROTEM®), or euglobulin lysis time (ELT).

[0140] The aPTT test is a performance indicator that measures the efficacy of both the "intrinsic" pathway (also called the contact activation pathway) and the common coagulation pathway. This test is commonly used to measure the clotting activity of commercially available recombinant clotting factors, such as FVIII. The aPTT test is used in conjunction with the prothrombin time (PT), which measures the extrinsic pathway.

[0141] ROTEM analysis provides information about the overall dynamics of hemostasis: clotting time, clot formation, clot stability, and lysis. In thromboelastometry, the different parameters depend on many factors that affect the activity of the plasma coagulation system, platelet function, fibrinolysis, or their interactions. This assay can provide a complete picture about secondary hemostasis.

[0142] The chromogenic assay mechanism is based on the principle of the blood coagulation cascade, where activated FVIII accelerates the conversion of factor X to factor Xa in the presence of activated factor IX, phospholipids, and calcium ions. Factor Xa activity is assessed by hydrolysis of the p-nitroanilide (pNA) substrate, specific for factor Xa. The initial release rate of p-nitroaniline, measured at 405 nM, is directly proportional to the factor Xa activity, and therefore also to the FVIII activity in the sample. The chromogenic assay is recommended by the FVIII and Factor IX Subcommittee of the Scientific and Standardization Committee (SSC) of the International Society on Thrombosis and Hemostatsis (ISTH). Since 1994, the chromogenic assay is also the reference method by the European Pharmacopoeia for the potency assignment of FVIII concentrates.

[0143] In some embodiments, the gene cassette comprises a nucleotide sequence encoding a FVIII polypeptide, where the nucleotide sequence is codon-optimized. In some embodiments, the gene cassette comprises a nucleotide sequence encoding a codon-optimized FVIII driven by an mTTR promoter and a synthetic intron. In some embodiments, the gene cassette comprises a nucleotide sequence disclosed in International Application No. PCT / US2017 / 015879, which is incorporated by reference in its entirety. In some embodiments, the gene cassette is "hFVIIIco6XTEN", a gene cassette described in PCT / US2017 / 015879.

[0144] In some embodiments, the nucleic acid molecule comprises a gene cassette encoding a first ITR, a second ITR, and a target sequence, the target sequence encoding a therapeutic protein, and the therapeutic protein comprises a growth factor. The growth factor is selected from any growth factor known in the art. In some embodiments, the growth factor is a hormone. In other embodiments, the growth factor is a cytokine. In some embodiments, the growth factor is a chemokine.

[0145] In some embodiments, the growth factor is adrenomedullin (AM). In some embodiments, the growth factor is angiopoietin (Ang). In some embodiments, the growth factor is an autocrine motility factor. In some embodiments, the growth factor is a bone morphogenetic protein (BMP). In some embodiments, the BMP is selected from BMP2, BMP4, BMP5, and BMP7. In some embodiments, the growth factor is a ciliary neurotrophic factor family member. In some embodiments, the ciliary neurotrophic factor family member is selected from ciliary neurotrophic factor (CNTF), leukemia inhibitory factor (LIF), and interleukin 6 (IL-6). In some embodiments, the growth factor is a colony stimulating factor. In some embodiments, the colony stimulating factor is selected from macrophage colony stimulating factor (m-CSF), granulocyte colony stimulating factor (G-CSF), and granulocyte macrophage colony stimulating factor (GM-CSF). In some embodiments, the growth factor is epidermal growth factor (EGF). In some embodiments, the growth factor is an ephrin. In some embodiments, the ephrin is selected from ephrinA1, ephrinA2, ephrinA3, ephrinA4, ephrinA5, ephrinB1, ephrinB2, and ephrinB3. In some embodiments, the growth factor is erythropoietin (EPO). In some embodiments, the growth factor is a fibroblast growth factor (FGF). In some embodiments, the FGF is selected from FGF1, FGF2, FGF3, FGF4, FGF5, FGF6, FGF7, FGF8, FGF9, FGF10, FGF11, FGF12, FGF13, FGF14, FGF15, FGF16, FGF17, FGF18, FGF19, FGF20, FGF21, FGF22, and FGF23. In some embodiments, the growth factor is fetal bovine somatotrophin (FBS). In some embodiments, the growth factor is a GDNF family member. In some embodiments, the GDNF family member is selected from glial cell line-derived neurotrophic factor (GDNF), neurturin, persephin, and artemin. In some embodiments, the growth factor is growth differentiation factor 9 (GDF9). In some embodiments, the growth factor is hepatocyte growth factor (HGF).In some embodiments, the growth factor is hepatoma-derived growth factor (HDGF). In some embodiments, the growth factor is insulin. In some embodiments, the growth factor is an insulin-like growth factor. In some embodiments, the insulin-like growth factor is insulin-like growth factor 1 (IGF-1) or IGF-2. In some embodiments, the growth factor is an interleukin (IL). In some embodiments, the IL is selected from IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, and IL-7. In some embodiments, the growth factor is keratinocyte growth factor (KGF). In some embodiments, the growth factor is migration stimulating factor (MSF). In some embodiments, the growth factor is macrophage stimulating protein (MSP) or hepatocyte growth factor-like protein (HGFLP). In some embodiments, the growth factor is myostatin (GDF-8). In some embodiments, the growth factor is neuregulin. In some embodiments, the neuregulin is selected from neuregulin 1 (NRG1), NRG2, NRG3, and NRG4. In some embodiments, the growth factor is a neurotrophic factor. In some embodiments, the growth factor is brain-derived neurotrophic factor (BDNF). In some embodiments, the growth factor is nerve growth factor (NGF). In some embodiments, the NGF is neurotrophic factor 3 (NT-3) or NT-4. In some embodiments, the growth factor is placental growth factor (PGF). In some embodiments, the growth factor is platelet-derived growth factor (PDGF). In some embodiments, the growth factor is renalase (RNLS). In some embodiments, the growth factor is T cell growth factor (TCGF). In some embodiments, the growth factor is thrombopoietin (TPO). In some embodiments, the growth factor is transforming growth factor. In some embodiments, the transforming growth factor is transforming growth factor alpha (TGF-α) or TGF-β. In some embodiments, the growth factor is tumor necrosis factor alpha (TNF-α). In some embodiments, the growth factor is vascular endothelial growth factor (VEGF).

[0146] C. Expression Control Sequences In some embodiments, the nucleic acid molecule of the present disclosure further comprises at least one expression control sequence.An expression control sequence as used herein is any regulatory nucleotide sequence, such as a promoter sequence or a promoter-enhancer combination, that facilitates efficient transcription and translation of an operably linked coding nucleic acid.For example, the nucleic acid molecule of the present disclosure is operably linked to at least one transcription control sequence.The gene expression control sequence can be, for example, a mammalian promoter or a viral promoter, such as a constitutive promoter or an inducible promoter.

[0147] Constitutive mammalian promoters include, but are not limited to, promoters for the following genes: hypoxanthine phosphoribosyltransferase (HPRT), adenosine deaminase, pyruvate kinase, beta-actin promoter, and other constitutive promoters. Exemplary viral promoters that function constitutively in eukaryotic cells include, for example, promoters derived from cytomegalovirus (CMV), simian viruses (e.g., SV40), papillomavirus, adenovirus, human immunodeficiency virus (HIV), Rous sarcoma virus, cytomegalovirus, Moloney leukemia virus long terminal repeat (LTR), and other retroviruses, as well as the thymidine kinase promoter of herpes simplex virus. Other constitutive promoters are known to those skilled in the art. Promoters useful for the gene expression sequences of the present disclosure also include inducible promoters. Inducible promoters are expressed in the presence of an inducer. For example, the metallothionein promoter is induced to promote transcription and translation in the presence of certain metal ions. Other inducible promoters are known to those skilled in the art.

[0148] In one embodiment, the disclosure includes expression of a transgene under the control of a tissue-specific promoter and / or enhancer. In another embodiment, the promoter or other expression control sequence selectively enhances expression of the transgene in hepatocytes. In certain embodiments, the promoter or other expression control sequence selectively enhances expression of the transgene in hepatocytes, sinusoidal cells, and / or endothelial cells. In a particular embodiment, the promoter or other expression control sequence selectively enhances expression of the transgene in endothelial cells. In certain embodiments, the promoter or other expression control sequence selectively enhances expression of the transgene in muscle cells, the central nervous system, the eye, the liver, the heart, or any combination thereof. Examples of liver-specific promoters include, but are not limited to, the mouse thyretin promoter (mTTR), the endogenous human factor VIII promoter, the human alpha 1 antitrypsin promoter (hAAT), the human albumin minimal promoter, and the mouse albumin promoter. In certain embodiments, the promoter comprises the mTTR promoter. The mTTR promoter is described in RH Costa et al., 1986, Mol. Cell. Biol., 6:4697. The FVIII promoter is described in Figueiredo and Brownlee, 1995, J. Biol. Chem., 270:11828-11838. In some embodiments, the promoter is selected from a liver-specific promoter (e.g., alpha 1 antitrypsin (AAT) promoter), a muscle-specific promoter (e.g., muscle creatine kinase (MCK) promoter, myosin heavy chain alpha (αMHC) promoter, myoglobin (MB) promoter, and desmin (DES) promoter), a synthetic promoter (e.g., SPc5-12 promoter, 2R5Sc5-12 promoter, dMCK promoter, and tMCK promoter), and any combination thereof.

[0149] In one embodiment, the promoter is selected from the group consisting of mouse transthyretin promoter (mTTR), native human factor VIII promoter, human alpha 1 antitrypsin promoter (hAAT), human albumin minimal promoter, mouse albumin promoter, TTPp, CASI promoter, CAG promoter, cytomegalovirus (CMV) promoter, alpha 1 antitrypsin (AAT) promoter, muscle creatine kinase (MCK) promoter, myosin heavy chain alpha (αMHC) promoter, myoglobin (MB) promoter, desmin (DES) promoter, SPc5-12 promoter, 2R5Sc5-12 promoter, dMCK promoter, and tMCK promoter, phosphoglycerate kinase (PGK) promoter, and any combination thereof. In some embodiments, the promoter is a TTP promoter. In some embodiments, the promoter is a mouse transthyretin promoter (mTTR) promoter.

[0150] To achieve therapeutic efficacy, the expression level is further enhanced using one or more enhancer elements. One or more enhancers may be administered alone or in conjunction with one or more promoter elements. Typically, the expression control sequence includes multiple enhancer elements and tissue-specific promoters. In one embodiment, the enhancer includes one or more copies of the alpha-1-microglobin / bikunin enhancer (Rouet et al., 1992, J. Biol. Chem., 267:20765-20773; Rouet et al., 1995, Nucleic Acids Res., 23:395-404; Rouet et al., 1998, Biochem. J. 334:577-584; Ill et al., 1997, Blood Coagulation Fibrinolysis, 8:S23-S30). In another embodiment, the enhancer is derived from liver-specific transcription factor binding sites such as EBP, DBP, HNF1, HNF3, HNF4, HNF6, including HNF1, (sense)-HNF3, (sense)-HNF4, (antisense)-HNF1, (antisense)-HNF6, (sense)-EBP, (antisense)-HNF4 (antisense), along with Enh1.

[0151] In one embodiment, the nucleic acid molecule of the present disclosure further comprises an intron sequence. In some embodiments, the intron sequence is located 5' to the nucleic acid sequence encoding the FVIII polypeptide. In some embodiments, the intron sequence is a naturally occurring intron sequence. In some embodiments, the intron sequence is a synthetic sequence. In some embodiments, the intron sequence is derived from a naturally occurring intron sequence. In certain embodiments, the intron sequence comprises the SV40 small T intron.

[0152] In some embodiments, the nucleic acid molecule further comprises a post-transcriptional regulatory element. In certain embodiments, the regulatory element comprises a woodchuck hepatitis virus regulatory element (WPRE). In some embodiments, the WPRE is mutated.

[0153] In some embodiments, the nucleic acid molecule comprises a microRNA (miRNA) binding site. In one embodiment, the miRNA binding site is a miRNA binding site for miR-142-3p. In other embodiments, the miRNA binding site is selected from the miRNA binding sites disclosed by Rennie et al., RNA Biol., 13(6):554-560 (2016) and STarMirDB, available at http: / / sfold.wadsworth.org / starmirDB.php, which are incorporated herein by reference in their entireties.

[0154] In some embodiments, the nucleic acid molecule comprises one or more DNA nuclear targeting sequences (DTS). The DTS facilitates the translocation of DNA molecules containing such sequences to the nucleus. In certain embodiments, the DTS comprises an SV40 enhancer sequence. In certain embodiments, the DTS comprises a c-Myc enhancer sequence. In some embodiments, the DTS is between the first ITR and the second ITR. In some embodiments, the DTS is 3' to the first ITR and 5' to the therapeutic protein. In other embodiments, the DTS is 3' to the therapeutic protein and 5' to the second ITR.

[0155] In some embodiments, the nucleic acid molecule further comprises a 3'UTR poly(A) tail sequence. In one embodiment, the 3'UTR poly(A) tail sequence comprises bGH poly(A). In one embodiment, the 3'UTR poly(A) tail comprises an actin poly(A) site. In one embodiment, the 3'UTR poly(A) tail comprises a hemoglobin poly(A) site.

[0156] In some embodiments, transgene expression is targeted to the liver. In certain embodiments, transgene expression is targeted to hepatocytes. In other embodiments, transgene expression is targeted to endothelial cells. In a particular embodiment, transgene expression is targeted to any tissue that naturally expresses endogenous FVIII.

[0157] In some embodiments, transgene expression is targeted to the central nervous system. In certain embodiments, transgene expression is targeted to neurons. In some embodiments, transgene expression is targeted to afferent neurons. In some embodiments, transgene expression is targeted to efferent neurons. In some embodiments, transgene expression is targeted to interneurons. In some embodiments, transgene expression is targeted to glial cells. In some embodiments, transgene expression is targeted to astrocytes. In some embodiments, transgene expression is targeted to oligodendrocytes. In some embodiments, transgene expression is targeted to microglia. In some embodiments, transgene expression is targeted to ependymal cells. In some embodiments, transgene expression is targeted to Schwann cells. In some embodiments, transgene expression is targeted to satellite cells.

[0158] In some embodiments, transgene expression is targeted to muscle tissue. In some embodiments, transgene expression is targeted to smooth muscle. In some embodiments, transgene expression is targeted to cardiac muscle. In some embodiments, transgene expression is targeted to skeletal muscle.

[0159] In some embodiments, expression of the transgene is targeted to the eye. In some embodiments, expression of the transgene is targeted to photoreceptor cells. In some embodiments, expression of the transgene is targeted to retinal ganglion cells.

[0160] III.Host cells The present disclosure also provides a host cell comprising the nucleic acid molecule or vector of the present disclosure. As used herein, the term "transformation" is used broadly to refer to the introduction of DNA into a recipient host cell, resulting in a change in the genotype and, as a result, in the change of the recipient cell.

[0161] "Host cell" refers to a cell that is transformed with a vector constructed using recombinant DNA methods and encoding at least one heterologous gene. The host cell of the present disclosure is preferably of mammalian origin; most preferably of human or murine origin. Those skilled in the art are considered to be capable of preferentially determining the particular host cell line that is best suited for their purpose. Exemplary host cell lines include, but are not limited to, CHO, DG44, and DUXB11 (Chinese hamster ovary cell line, DHFR deleted), HELA (human cervical carcinoma), CVI (monkey kidney cell line), COS (a derivative of CVI cells with SV40 T antigen), R1610 (Chinese hamster fibroblast) BALBC / 3T3 (mouse fibroblast), HAK (hamster kidney cell line), SP2 / O (mouse myeloma), P3x63-Ag8.653 (mouse myeloma), BFA-1c1BPT (bovine endothelial cells), RAJI (human lymphocytes), PER.C6®, NS0, CAP, BHK21, and HEK293 (human kidney). In a particular embodiment, the host cell is selected from the group consisting of CHO cells, HEK293 cells, BHK21 cells, PER.C6® cells, NS0 cells, CAP cells, and any combination thereof. In some embodiments, the host cell of the present disclosure is derived from an insect. In a particular embodiment, the host cell is an SF9 cell. Host cell lines are typically available from commercial services, the American Tissue Culture Collection, or published literature.

[0162] Introduction of the nucleic acid molecules or vectors of the present disclosure into host cells can be accomplished by a variety of techniques well known to those of skill in the art. These include, but are not limited to, transfection (including electrophoresis and electroporation), protoplast fusion, calcium phosphate precipitation, cell fusion with enveloped DNA, microinjection, and infection with intact viruses. See Ridgway, AAG, "Mammalian Expression Vectors," Chapter 24.2, pages 470-472, "Vectors," edited by Rodriguez and Denhardt (Butterworths, Boston, Mass. 1988). Most preferably, plasmid introduction into the host is via electroporation. The transformed cells are grown under conditions appropriate for the production of light and heavy chains, and assayed for the synthesis of heavy and / or light chain proteins. Exemplary assay methods include enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), or fluorescence-activated cell sorting analysis (FACS), immunohistochemistry, and the like.

[0163] The host cells containing the isolated nucleic acid molecule or vector of the present disclosure are grown in an appropriate growth medium. As used herein, the term "appropriate growth medium" refers to a medium containing nutrients required for cell growth. Nutrients required for cell growth may include a carbon source, a nitrogen source, essential amino acids, vitamins, minerals, and growth factors. Optionally, the medium may contain one or more selection factors. Optionally, the medium may contain calf serum or fetal calf serum (FCS). In one embodiment, the medium is substantially free of IgG. The growth medium will generally select for cells containing the DNA construct, for example, by drug selection or deficiency of essential nutrients, complemented by a selectable marker on the DNA construct or co-transfected with the DNA construct. Cultured mammalian cells are generally grown in commercially available serum-containing or serum-free media (e.g., MEM, DMEM, DMEM / F12). In one embodiment, the medium is CDoptiCHO (Invitrogen, Carlsbad, Calif.). In another embodiment, the medium is CD17 (Invitrogen, Carlsbad, Calif.) Selection of an appropriate medium for the particular cell line used is within the level of one of ordinary skill in the art.

[0164] IV. Pharmaceutical Compositions A composition containing a nucleic acid molecule of the present disclosure, a polypeptide encoded by the nucleic acid molecule, a vector, or a host cell may contain a suitable pharma- ceutically acceptable carrier. For example, the composition may contain excipients and / or adjuvants that facilitate processing of the active compound into a preparation designed for delivery to a site of action.

[0165] In one embodiment, the present disclosure is directed to a pharmaceutical composition comprising (a) a nucleic acid molecule, vector, polypeptide, or host cell disclosed herein; and (b) a pharma- ceutically acceptable excipient.

[0166] In some embodiments, the pharmaceutical composition further comprises a delivery agent. In certain embodiments, the delivery agent comprises a lipid nanoparticle (LNP). In other embodiments, the pharmaceutical composition further comprises a liposome, other polymer molecules, and exosomes.

[0167] As used herein, "lipid nanoparticle" refers to a nanoparticle that comprises a plurality of lipid molecules that are physically associated with each other by intermolecular forces.Lipid nanoparticles can be, for example, microspheres (including single-phase vesicles and multi-lamellar vesicles, e.g., liposomes), the dispersed phase in an emulsion, micelles, or the internal phase in a suspension.

[0168] In some embodiments, the present disclosure provides an encapsulated nucleic acid molecule composition, which may comprise lipid nanoparticles that encapsulate the nucleic acid molecules of the present invention. Lipid nanoparticles may comprise one or more lipids (e.g., cationic lipids, non-cationic lipids, and PEG-modified lipids). In certain embodiments, lipid nanoparticles of the present disclosure are formulated to deliver one or more nucleic acid molecules of the present invention to one or more target cells. Examples of suitable lipids include, but are not limited to, phosphatidyl compounds (e.g., phosphatidylethanolamine, sphingolipids, phosphatidylcholine, phosphatidylserine, phosphatidylglycerol, gangliosides, and cerebrosides). "Cationic lipid" refers to any lipid molecular species that carries a net positive charge at a certain pH (e.g., physiological pH).

[0169] In certain embodiments, the lipid nanoparticles of the present disclosure have a certain N / P ratio. As used herein, "N / P ratio" or "NP ratio" refers to the ratio of positively charged polymeric amine groups to negatively charged nucleic acid phosphate groups. The N / P characteristics of lipid nanoparticle / nucleic acid molecule complexes can affect properties such as net surface charge, stability, and size. The NP ratio of the lipid nanoparticles described herein can be about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, and any ratio therebetween. For example, the NP ratio of the lipid nanoparticles described herein can be about 18, about 36, or about 72.

[0170] Thus, in certain embodiments, a pharmaceutical composition comprises a nucleic acid molecule of the present disclosure encapsulated within a lipid nanoparticle and a pharma- ceutically acceptable excipient.

[0171] The pharmaceutical composition is formulated for parenteral administration (i.e., intravenous, subcutaneous, or intramuscular) by bolus injection. The formulation for injection is presented in unit dosage form, for example, in ampoules or multi-dose containers with added preservatives. The composition may take the form of a suspension, solution, or emulsion in an oily or aqueous medium, and may contain formulating agents, such as suspending, stabilizing, and / or dispersing agents. Alternatively, the active ingredient may be in powder form for constitution with a suitable medium, for example, pyrogen-free water.

[0172] Preparations suitable for parenteral administration also include aqueous solutions of the active compound in water-soluble form, for example, in water-soluble salt form. In addition, suspensions of the active compound as appropriate oily injection suspensions are also administered. Suitable lipophilic solvents or vehicles include fatty oils, for example, sesame oil, or synthetic fatty acid esters, for example, ethyl oleate or triglycerides. Aqueous injection suspensions may contain substances that increase the viscosity of the suspension, including, for example, sodium carboxymethylcellulose, sorbitol, and dextran. Optionally, the suspension may also contain a stabilizer. Liposomes are also used to encapsulate the molecules of the present disclosure for delivery to cells or interstitial spaces. Exemplary pharmaceutically acceptable carriers are physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, water, saline, phosphate buffered saline, dextrose, glycerol, ethanol, and the like. In some embodiments, the composition includes an isotonic agent, for example, a sugar, a polyalcohol such as mannitol, sorbitol, or sodium chloride. In other embodiments, the composition includes a pharma- ceutically acceptable substance, such as a humectant, or minor amounts of auxiliary substances, such as humectants or emulsifiers, preservatives or buffers, which enhance the shelf life or efficacy of the active ingredient.

[0173] The compositions of the present disclosure may be in a variety of forms, including, for example, liquid (e.g., injectable and infusible solutions), dispersions, suspensions, semi-solids, and solids. The preferred form depends on the mode of administration and therapeutic application.

[0174] The composition is formulated as a solution, microemulsion, dispersion, liposome, or other ordered structure suitable for high drug concentration. Sterile injectable solutions are prepared by incorporating the active ingredient in the required amount in a suitable solvent with one or a combination of the above-listed ingredients as required, followed by filtration sterilization. In general, dispersions are prepared by incorporating the active ingredient into a sterile medium containing a basic dispersion medium and other required ingredients from the above-listed ingredients. In the case of sterile powders for preparing sterile injectable solutions, the preferred preparation method is vacuum drying and freeze-drying, which produces a powder of the active ingredient plus any additional desired ingredients from a previously sterile-filtered solution. The proper fluidity of the solution can be maintained by using a coating such as lecithin, or by maintaining the required particle size in the case of dispersions, or by using surfactants. Prolonged absorption of injectable compositions can be achieved by incorporating an agent that delays absorption, such as monostearate salts and gelatin, into the composition.

[0175] The active ingredient is formulated with controlled release formulation or device.The examples of such formulation and device include implant, transdermal patch and microencapsulated delivery system.Biodegradable polymer, biocompatible polymer, such as ethylene vinyl acetate, polyanhydride, polyglycolic acid, collagen, polyorthoester and polylactic acid are used.The method for producing such formulation and device is known in the art.See, for example, "Sustained and Controlled Release Drug Delivery Systems", edited by JR Robinson, Marcel Dekker, Inc., New York, 1978.

[0176] Injectable depot preparations are made by forming microencapsulated matrices of drugs in biodegradable polymers such as polylactide-polyglycolide.Depending on the drug-to-polymer ratio and the nature of the polymer used, the drug release rate is controlled.Other exemplary biodegradable polymers are polyorthoesters and polyanhydrides.Injectable depot preparations are also made by encapsulating drugs in liposomes or microemulsions.

[0177] The composition may also incorporate an auxiliary active compound. In one embodiment, the nucleic acid molecule of the present disclosure is formulated with a clotting factor, or a variant, fragment, analog, or derivative thereof. For example, the clotting factor includes, but is not limited to, factor V, factor VII, factor VIII, factor IX, factor X, factor XI, factor XII, factor XIII, prothrombin, fibrinogen, von Willebrand factor, or recombinant soluble tissue factor (rsTF), or an activated form of any of the above. The clotting factor of the hemostatic agent may also include an antifibrinolytic agent, such as epsilon-aminocaproic acid, tranexamic acid.

[0178] Dosage regimen is adjusted to obtain the desired optimal response. For example, a single bolus may be administered, or a number of divided doses may be administered over time, and the dose may be reduced or increased accordingly as indicated by the exigencies of the therapeutic situation. For ease of administration and uniformity of dosage, it is advantageous to formulate parenteral compositions in unit dosage form. For example, see "Remington's Pharmaceutical Sciences" (Mack Pub.Co., Easton, Pa., 1980).

[0179] In addition to the active compound, liquid dosage forms may contain inactive ingredients such as water, ethyl alcohol, ethyl carbonate, ethyl acetate, benzyl alcohol, benzyl benzoate, propylene glycol, 1,3-butylene glycol, dimethylformamide, oils, glycerol, tetrahydrofururyl alcohol, polyethylene glycol, and fatty acid esters of sorbitan.

[0180] Non-limiting examples of suitable pharmaceutical carriers are also described in "Remington's Pharmaceutical Sciences" by EW Martin. Some examples of excipients include starch, glucose, lactose, sucrose, gelatin, malt, rice, flour, chalk, silica gel, sodium stearate, glycerol monostearate, talc, sodium chloride, nonfat dry milk, glycerol, propylene glycol, water, ethanol, etc. The composition may also contain a pH buffering agent and a humectant or emulsifier.

[0181] For oral administration, the pharmaceutical composition may take the form of a tablet or capsule, which is prepared by conventional means. The composition may also be prepared as a liquid, for example, a syrup or suspension. The liquid may contain a suspending agent (e.g., sorbitol syrup, cellulose derivatives, or hydrogenated edible fats), an emulsifying agent (lecithin or gum acacia), a non-aqueous medium (e.g., almond oil, oily esters, ethyl alcohol, or fractionated vegetable oils), and a preservative (e.g., methyl-p-hydroxybenzoate or propyl-p-hydroxybenzoate, or sorbic acid). The preparation may also contain flavorings, colorings, and sweetening agents. Alternatively, the composition may be provided as a dry product for constitution with water or another suitable vehicle.

[0182] For buccal administration, the composition may take the form of tablets or lozenges following conventional protocols.

[0183] For inhalation administration, the compound for use according to the present disclosure is conveniently delivered in the form of a nebulized aerosol, with or without excipients, or in the form of an aerosol spray from a pressurized pack or nebulizer, optionally with a propellant, such as dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoromethane, carbon dioxide, or other suitable gas.In the case of a pressurized aerosol, the dosage unit is determined by providing a valve that delivers a metered amount.Capsules and cartridges of, for example, gelatin, for use in an inhaler or insufflator are formulated containing a powder mix of the compound and a suitable powder base, such as lactose or starch.

[0184] Pharmaceutical compositions can also be formulated for rectal administration as suppositories or retention enemas, e.g., containing conventional suppository bases such as cocoa butter or other glycerides.

[0185] In some embodiments, the composition is administered by a route selected from the group consisting of topical administration, intraocular administration, parenteral administration, intrathecal administration, subdural administration, and oral administration. Parenteral administration can be intravenous or subcutaneous administration.

[0186] V. Treatment In some aspects, the present disclosure is directed to a method of treating a disease or condition in a subject in need thereof, the method comprising administering a nucleic acid molecule, vector, polypeptide, or pharmaceutical composition disclosed herein.

[0187] In some embodiments, the present disclosure is directed to a method of treating a bleeding disorder. In some embodiments, the present disclosure is directed to a method of treating hemophilia A.

[0188] The isolated nucleic acid molecule, vector, or polypeptide is administered intravenously, subcutaneously, intramuscularly, or via any mucosal surface, for example, via oral, sublingual, buccal, sublingual, intranasal, rectal, vaginal, or pulmonary routes. The coagulation factor protein may be implanted within or linked to a biopolymeric solid support that allows for sustained release of the chimeric protein to the desired site.

[0189] For oral administration, the pharmaceutical composition may take the form of a tablet or capsule, which is prepared by conventional means. The composition may also be prepared as a liquid, for example, a syrup or suspension. The liquid may contain a suspending agent (e.g., sorbitol syrup, cellulose derivatives, or hydrogenated edible fats), an emulsifying agent (lecithin or gum acacia), a non-aqueous medium (e.g., almond oil, oily esters, ethyl alcohol, or fractionated vegetable oils), and a preservative (e.g., methyl-p-hydroxybenzoate or propyl-p-hydroxybenzoate, or sorbic acid). The preparation may also contain flavorings, colorings, and sweetening agents. Alternatively, the composition may be provided as a dry product for constitution with water or another suitable vehicle.

[0190] For buccal and sublingual administration, the compositions may take the form of tablets, lozenges, or fast-dissolving films following conventional protocols.

[0191] For inhalation administration, the polypeptides having clotting factor activity for use according to the present disclosure are conveniently delivered in the form of an aerosol spray (e.g., in PBS) from pressurized packs or nebulizers, for example with dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoromethane, carbon dioxide, or other suitable gas. In the case of a pressurized aerosol, the dosage unit is determined by providing a valve to deliver a metered amount. Capsules and cartridges of, for example, gelatin, for use in an inhaler or insufflator are formulated containing a powder mix of the compound and a suitable powder base, such as lactose or starch.

[0192] In one embodiment, the route of administration of the isolated nucleic acid molecule, vector, or polypeptide is a parenteral route. As used herein, the term "parenteral" includes intravenous, intraarterial, intraperitoneal, intramuscular, subcutaneous, rectal, or vaginal administration. The intravenous form of parenteral administration is preferred. All of these forms of administration are expressly contemplated to be within the scope of this disclosure, but the form for administration will be an injectable solution, particularly for intravenous or intraarterial injection or instillation. Typically, pharmaceutical compositions suitable for injection may include buffers (e.g., acetate, phosphate, or citrate buffers), surfactants (e.g., polysorbates), and optionally stabilizers (e.g., human albumin). However, in other methods compatible with the teachings of this specification, the isolated nucleic acid molecule, vector, or polypeptide is delivered directly to the site of the harmful cell population, thereby increasing the exposure of the affected tissue to the therapeutic agent.

[0193] Preparations for parenteral administration include sterile aqueous or non-aqueous solutions, sterile aqueous or non-aqueous suspensions, and emulsions. Examples of non-aqueous solvents are propylene glycol, polyethylene glycol, vegetable oils such as olive oil, and injectable organic esters such as ethyl oleate. Aqueous carriers include water, alcohol / aqueous solutions, emulsions, or suspensions including saline and buffered media. In the present disclosure, pharma- ceutically acceptable carriers include, but are not limited to, phosphate buffer, 0.01-0.1M, preferably 0.05M, or 0.8% saline. Other common parenteral vehicles include sodium phosphate, dextrose Ringer's, dextrose, and sodium chloride, lactated Ringer's, or fixed oils. Intravenous vehicles include fluid and nutrient replenishers, electrolyte replenishers, such as intravenous vehicles based on dextrose Ringer's, and the like. Preservatives and other additives, such as, for example, antimicrobial agents, antioxidants, chelating agents, and inert gases, may also be present.

[0194] More specifically, pharmaceutical compositions suitable for injectable use include sterile aqueous solutions (if water soluble) or dispersions, and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. In such cases, the composition must be a sterile composition, and should be fluid to the extent that easy syringability exists. The composition should be stable under the conditions of manufacture and storage, and preferably be preservative against the contaminating action of microorganisms, such as bacteria and fungi. The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, and liquid polyethylene glycol, and the like), and suitable mixtures thereof. Proper fluidity can be maintained, for example, by the use of a coating such as lecithin, or, in the case of dispersions, by maintaining the required particle size, or by the use of surfactants.

[0195] Pharmaceutical compositions can also be formulated for rectal administration as suppositories or retention enemas, e.g., containing conventional suppository bases such as cocoa butter or other glycerides.

[0196] The effective dose of the composition of the present disclosure for treating a condition varies depending on many different factors, including the means of administration, the target site, the physiological condition of the patient, whether the patient is a human or an animal, other medicines administered, and whether the treatment is a preventive or therapeutic treatment.Usually, the patient is a human, but non-human mammals, including transgenic mammals, may also be treated.Treatment dosages are titrated using routine methods known to those skilled in the art that optimize safety and efficacy.

[0197] The nucleic acid molecules, vectors, or polypeptides of the disclosure are optionally administered in combination with other agents that are effective in treating the disorder or condition in need of treatment (e.g., prophylactic or therapeutic treatment).

[0198] As used herein, administration of the disclosed isolated nucleic acid molecule, vector, or polypeptide with or in combination with adjunctive therapy refers to sequential administration or application, simultaneous administration or application, co-administration or application, co-administration or application, or parallel administration or application of the therapy and the disclosed polypeptide. Those skilled in the art will recognize that the administration or application of various components of the combined therapy regimen is timed to enhance the efficacy of treatment. Those skilled in the art (e.g., physicians) will be able to easily identify an effective combined therapy regimen based on the selected adjunctive therapy and the teachings of this specification without undue experimentation.

[0199] It will be further appreciated that the isolated nucleic acid molecules, vectors, or polypeptides of the present disclosure may be used with or in combination with one or more drugs (e.g., to provide a combination therapeutic regimen). Exemplary drugs to be combined with the polypeptides or polynucleotides of the present disclosure include drugs that represent the current standard of care for the particular disorder being treated. Such drugs may be chemicals or biopharmaceuticals in nature. The term "biopharmaceutical" or "biopharmaceutical agent" refers to any pharmacologic active agent made from living organisms and / or their products that is intended for use as a therapeutic agent.

[0200] The amount of drugs used in combination with the polynucleotides or polypeptides of the present disclosure may vary from subject to subject and may be administered according to what is known in the art. See, for example, Bruce A Chabner et al., "Antineoplastic Agents," GOODMAN and GILMAN, "PHARMACOLOGICAL BASIS OF THERAPEUTICS," pp. 1233-1287 (Joel G. Hardman et al., eds., 9th ed., 1996). In another embodiment, amounts of such drugs are administered that are consistent with standard of care.

[0201] In one embodiment, also disclosed herein is a kit comprising the nucleic acid molecule disclosed herein and instructions for administering the nucleic acid molecule to a subject in need thereof. In another embodiment, disclosed herein is a baculovirus system for producing the nucleic acid molecule provided herein. The nucleic acid molecule is produced in insect cells. In another embodiment, provided is a nanoparticle delivery system for an expression construct. The expression construct comprises the nucleic acid molecule disclosed herein.

[0202] VI. Gene Therapy Certain aspects of the present disclosure provide methods of expressing a gene expression construct in a subject, comprising administering to a subject in need thereof an isolated nucleic acid molecule of the present disclosure. In some aspects, the present disclosure provides methods of increasing expression of a polypeptide in a subject, comprising administering to a subject in need thereof an isolated nucleic acid molecule of the present disclosure.

[0203] Gene therapy has been explored as a possible treatment for various conditions, including but not limited to hemophilia A. Gene therapy is a particularly attractive treatment for hemophilia due to its potential to cure the disease through sustained endogenous production of clotting factors, such as FVIII, after a single administration of a vector. Hemophilia A is well suited to gene replacement approaches, since its clinical symptoms are entirely attributable to the lack of a single gene product (e.g., FVIII), which circulates in minute amounts (200 ng / ml) in plasma.

[0204] All of the various aspects, embodiments, and options described herein may be combined in any and all variations.

[0205] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.

[0206] Working Example Having provided the foregoing disclosure, a further understanding can be obtained by reference to the examples provided herein, which are intended for purposes of illustration only and are not intended to be limiting. EXAMPLES

[0207] Design and construction of modified GPV ITRs and modified B19 ITRs The ITR sequences of AAV2 (Gene Bank Accession No.: NC_001401.2), Dependovirus GPV (Gene Bank Accession No.: U25749.1), and Erythrovirus B19 (Gene Bank Accession No.: KY940273.1) were analyzed. Based on this analysis, modified derivatives of wild-type GPV ITR and wild-type B19 ITR were designed to explore which nucleic acid sequences of GPV ITR and B19 ITR are required for sustained transduction of eukaryotic cells by genetic constructs carrying modified ITR. Nucleic acid sequences of exemplary modified ITRs are presented in Table 2. The predicted DNA structure for each of the modified ITRs is shown in Figures 4-9.

[0208] The modified ITR sequences of SEQ ID NOs: 1-8 are truncated derivatives of the wild-type B19 ITR. Graphical representations of the predicted structures of the truncated B19 ITR derivatives are shown in Figures 1A and 2A. The modified ITR sequences of SEQ ID NOs: 9-16 are truncated derivatives of the wild-type GPV ITR. Graphical representations of the predicted structures of the truncated GBV ITR derivatives are shown in Figures 1B and 2B. These truncated ITRs were designed to maintain the hairpin structure of the wild-type ITR and preserve one or more of the Rep binding elements (RBEs) in the sequence. The bases removed in these sequences are nucleotides located between the RBE and the dyad with their corresponding binding partner on the other side of the hairpin. The modified ITR sequences of SEQ ID NOs: 1-16 have hairpin lengths and thermodynamic stabilities in the ranges predicted to contribute to improved in vivo stability and potency, as well as manufacturability and product stability.

[0209] The modified ITR sequence of SEQ ID NO: 1 and 2 (B19_min) contains an additional truncation of the B19 ITR sequence. A graphical representation of the predicted structure of the derivative B19_min ITR is shown in Figure 2A.

[0210] The modified ITR sequence (GPV_min) of SEQ ID NO: 15 and 16 contains an additional truncation of the GPV ITR sequence that removes nucleotides that are between, but do not contribute to, the RBE. A graphical representation of the predicted structure of the derivative GPV_min ITR is shown in Figure 2B.

[0211] The modified ITR sequences of SEQ ID NOs: 17-22 contain truncated ITR sequences with additional nucleotides inserted in the dyad to improve de novo synthesis and subsequent cloning attempts.

[0212] It was hypothesized that due to the presence of retained RBEs within the ITRs, gene expression constructs carrying these modified derivative ITRs would remain functional and efficiently transduce eukaryotic cells. EXAMPLES

[0213] Generation of FVIII expression constructs with modified ITRs The gene construct was created using a plasmid encoding a FVIII gene cassette containing a codon-optimized FVIII with an mTTR promoter and a synthetic intron.

[0214] Novel constructs were generated that contain a codon-optimized FVIII gene cassette flanked by modified ITRs set forth in SEQ ID NOs: 1-22 at either or both of the 5' or 3' ITR positions. Exemplary constructs and modified ITR pairs are shown in Table 1.

[0215] These new constructs were propagated and maintained in E. coli strain PMC103, which contains a deleted sbcC gene that encodes an exonuclease that recognizes and eliminates cruciform DNA structures. PMC103 E. coli was able to support the growth of the codon-optimized FVIII constructs with modified ITRs after optimizing the temperature and growth conditions of the bacterial culture. All new constructs were selected on ampicillin resistance plates and screened by restriction enzyme mapping to confirm the correct gene structure. The modified ITRs were also subjected to Sanger sequencing after restriction digestion and gel extraction.

[0216] [Table 1] EXAMPLES

[0217] Preparation of a single-stranded DNA fragment containing a FVIII expression cassette flanked by modified ITRs It was hypothesized that the formation of hairpin structures within the ITR regions flanking the codon-optimized FVIII expression cassette would drive sustained transduction of target cells. For each plasmid construct, single-stranded (ss) DNA fragments with hairpin ITR structures were generated by denaturing the double-stranded DNA fragment products (FVIII expression cassette and plasmid backbone) of PvuII or LguI digestion at 95°C, followed by cooling at 4°C to allow the palindromic ITR sequences to form hairpins. EXAMPLES

[0218] In vivo assessment of FVIII expression from ssDNA and dsDNA constructs of the B19 minimal ITR. To verify the ability of expression constructs carrying modified ITR regions to mediate sustained transgene expression in vivo, 34.7ug of DNA of pFVIII.B19_min construct containing a codon-optimized FVIII transgene flanked by B19 minimal ITRs (5'ITR: SEQ ID NO:1 and 3'ITR: SEQ ID NO:2) was delivered systemically to 8-12 week old male hemophilic (HemA) mice via hydrodynamic injection (HDI). DNA was administered in double stranded (dsDNA) format without preformed hairpins or as single stranded DNA (ssDNA) with preformed ITR hairpins as prepared in Example 3. Plasma samples were collected from experimental animals 3 and 7 days after injection. Plasma FVIII activity in blood was analyzed by FVIII activity chromogenic assay.

[0219] As shown in Figure 3, analysis of plasma FVIII levels 3 and 7 days after injection revealed that ssDNA resulted in higher levels of FVIII than dsDNA. On day 7, FVIII levels from ssDNA remained stable, while FVIII levels from dsDNA were reduced by about 30%. On day 14, FVIII levels were near zero for both ssDNA and dsDNA. ELISA on samples on day 14 revealed the presence of anti-FVIII antibodies (inhibitors) in all mice. Without being bound by theory, the near-zero levels on day 14 may be due to the presence of anti-FVIII antibodies, likely formed due to the expression levels of FVIII above physiological levels. This example supports that ssDNA and dsDNA constructs containing modified ITRs, such as the B19 ITR, can express FVIII. EXAMPLES

[0220] Assessment of gene expression in vivo from ssDNA constructs To verify the functionality of the ssDNA constructs containing the codon-optimized FVIII expression cassette flanked by modified ITRs, each ssDNA construct was subjected to a 5'-nucleotide sequence analysis of hFVIIIR593C, as well as the ssFVIII.B19_min construct. + / + / HemA mice (see Example 4). Exemplary ITR sequences are the set shown as SEQ ID NOs: 1-22 in Table 2. GPV and B19 combinations of wild-type and modified ITRs were created and examined.

[0221] Single stranded DNA was generated from the codon-optimized FVIII expression construct with flanking modified ITRs as in Example 3, and designated hFVIIIR593C + / + Liver-directed FVIII expression driven by the mTTR promoter within a codon-optimized FVIII expression cassette was investigated in / HemA mice (5-12 weeks of age). + / + / HemA mice were injected with 10, 20, 37, 50 μg or other predefined quantities of ssDNA containing the expression cassettes described above via hydrodynamic injection (HDI), and FVIII was measured from mouse plasma collected at 1, 3, 7, 14, 21, and 28 days after injection, or at other predefined time intervals. FVIII expression and lifespan in mice administered these expression cassettes with modified ITRs were directly compared to FVIII expression and lifespan in mice administered ssDNA constructs with B19Δ135 ITR, GPVΔ162 ITR, and / or the corresponding wild-type ITR expression cassettes. FVIII activity in blood was analyzed by a FVIII activity chromogenic assay. EXAMPLES

[0222] Use of the baculovirus expression system to generate ceDNA expression constructs carrying modified ITRs Systemic delivery of closed-end DNA (ceDNA) expression cassettes has been shown to establish persistent transduction of hepatocytes and drive long-term, stable transgene expression in the liver.To investigate transduction efficiency and gene expression, the baculovirus expression system described in Example 1D of WO2020 / 033863 is used to generate FVIII expression cassettes for gene constructs carrying modified ITRs in the form of closed-end DNA (ceDNA) molecules in insect cells.This baculovirus expression system is based on the system described in Li et al., PLoS ONE, 8(8):e69879 (2013). EXAMPLES

[0223] Assessment of in vivo gene expression from ceDNA constructs carrying modified ITRs in HemA mice The ceDNA construct containing the optimized FVIII expression cassette flanked by the modified 5' and 3' ITRs was designated hFVIIIR593C, as was the pFVIII.B19_min construct. + / + The ITR sequences were examined in / HemA mice (see Example 4). Exemplary ITR sequences are the set shown in Table 2 as SEQ ID NOs: 1-22. ceDNA constructs were generated as described in Example 6.

[0224] The ceDNA derived from the codon-optimized FVIII expression construct with flanking modified ITRs was designated hFVIIIR593C + / + Liver-directed FVIII expression was investigated in / HemA mice (5-12 weeks old). + / + / HemA mice were injected with 10, 20, 37, 50 μg or other predefined quantities of ceDNA containing the expression cassettes described above via hydrodynamic injection (HDI), and FVIII was measured from mouse plasma collected at 1, 3, 7, 14, 21, and 28 days after injection, or at other predefined time intervals. FVIII expression and lifespan in mice administered these expression cassettes with modified ITRs was directly compared to FVIII expression and lifespan in mice administered ceDNA constructs with B19Δ135 ITR, GPVΔ162 ITR, and / or the corresponding wild-type ITR expression cassettes. FVIII activity in blood was analyzed by a FVIII activity chromogenic assay. EXAMPLES

[0225] Generation of reporter gene constructs carrying ITRs from B19 or GPV To demonstrate the utility of the modified ITR-based gene expression system as a platform for general use in gene therapy applications, expression cassettes were generated containing reporter constructs with green fluorescent protein (GFP) or luciferase (luc) flanked by modified 5' and 3' ITRs. The same procedure as in Example 2 was used, but the FVIII open reading frame (ORF) in the codon-optimized FVIII expression cassette was replaced with the GFP or luc ORF by conventional molecular cloning methods.

[0226] We also created a modified ITR-flanked expression cassette containing mouse phenylalanine hydroxylase (PAH) transgene, which is used to assess the expression of PAH and the reduction of blood phenylalanine concentration in a mouse model of phenylketonuria.Using this model, PKU mice were administered 200μg of modified ITR-flanked ssDNA via hydrodynamic injection (HDI) for expression in the liver.After 3, 7, 14, 28, 42, 56, 70, and 81 days, blood samples were collected and plasma was isolated for the determination of phenylalanine concentration.To confirm the presence of mouse PAH protein in the liver, Western blot was performed on liver lysates taken from treated mice on the 81st day after injection. EXAMPLES

[0227] Preparation of ssDNA reporter gene constructs carrying modified ITRs The ssDNA reporter or PAH constructs were prepared as described in Example 3. Briefly, the plasmids were digested with the restriction enzymes LguI, MscI, and Eco53kI. The ssDNA fragments with hairpin ITR structures were generated by denaturing the double-stranded DNA fragment products of the restriction enzyme digestion (reporter expression cassette and plasmid backbone) at 95°C and then cooling at 4°C to allow the palindromic ITR sequences to form hairpins. The resulting ssDNA constructs were examined for their ability to establish persistent transduction of liver, muscle tissue, photoreceptors in the eye, central nervous system (CNS), or other tissues in mice by detection of the reporter gene or PAH. EXAMPLES

[0228] In vivo assessment of ssDNA-mediated reporter expression To verify the ability of the ssDNA reporter constructs described in Example 13 to mediate sustained transgene expression in vivo, 5-12 week old mice (at least 4 animals per group) were injected systemically and / or locally into target tissues with 5, 10, or 20 μg of reporter ssDNA per mouse. Blood samples were collected at predetermined time points and reporter transgene levels were detected and / or measured using routine techniques. EXAMPLES

[0229] Modified V2.0 FVIIIXTEN expression cassette with engineered parvovirus ITRs We hypothesized that the expression level of transgenes would be increased by codon-optimizing cDNA for the target host. Previous studies, described in US Publication No. 20190185543, have supported physiological levels of FVIII expression from the V1.0 FVIIIco6XTEN expression cassette. However, to further improve the expression level of transgenes and reduce immunogenicity, FVIIIXTEN cDNA was codon-optimized by deleting CpG repeats to avoid the natural immune response elicited against the DNA vector that codes for FVIIIXTEN expression cassette together with parvovirus ITR. A modified V2.0 FVIIIXTEN expression cassette was created, which comprises B-domain deleted (BDD) codon-optimized human factor VIII (BDDcoFVIII) fused with XTEN 144 peptide (FVIIIXTEN), under the control of liver-specific modified mouse transthyretin (mTTR) promoter (mTTR482) with enhancer element (A1MB2), hybrid synthetic intron (chimeric intron), woodchuck posttranscriptional regulatory element (WPRE) and bovine growth hormone polyadenylation (bGHpA) signal. The V2.0 FVIIIXTEN expression cassette comprises the nucleotide sequence of SEQ ID NO:27. A graphical depiction of an exemplary V2.0 FVIIIXTEN expression cassette is shown in FIG. 10.

[0230] Initial in vivo efficacy studies showed significant improvement in FVIII activity compared to the V1.0 FVIIIXTEN expression cassette (data not shown). Thus, engineered parvovirus ITRs were cloned, including AAV2 WT (FIG. 10A), HBoV1 WT (SEQ ID NO:25, SEQ ID NO:26) (FIG. 10B), B19 WT (SEQ ID NO:23, SEQ ID NO:24), B19 Minimal (SEQ ID NO:1, SEQ ID NO:2) (FIG. 10C), and GPVΔ186 ITR (SEQ ID NO:9, SEQ ID NO:10), GPVΔ120 ITR (SEQ ID NO:13, SEQ ID NO:14), or GPV Minimal ITR (SEQ ID NO:15, SEQ ID NO:16) (FIG. 10D), such that the ITRs flank the V2.0 FVIIIXTEN expression cassette. The in vivo functionality of the modified V2.0 FVIIIXTEN expression cassettes with different parvoviral ITRs in the form of single-stranded (ss) DNA or closed-end (ce) DNA was also demonstrated by systemic administration of hFVIIIR593C by hydrophilic tail vein injection as shown below. + / + This is also supported in / HemA mice. EXAMPLES

[0231] Assessment of single-stranded FVIIIXTEN (ssFVIIIXTEN) DNA in vivo It was hypothesized that the hairpin formed within the ITR region would drive long-term, sustained transgene expression at high levels. To validate the functionality of the modified FVIIIXTEN expression cassette in vivo, single-stranded DNA (ssDNA) containing the V2.0 FVIIIXTEN gene cassette with engineered parvovirus ITRs was transformed into hFVIIIR593C + / +hFVIIIR593C / HemA mice were examined. These mice contain a human FVIII-R593C transgene designed with a mouse albumin (Alb) promoter driving the expression of a modified human coagulation factor VIII (FVIII) cDNA carrying a mutation frequently observed in patients with mild hemophilia A. These mice also carry a knockout of the FVIII gene and are deficient for endogenous FVIII protein. These double mutant mice tolerate injections of human FVIII and have no FVIII activity. These double mutant mice produce only trace amounts of inhibitory antibodies after treatment with human FVIII and lack FVIII-responsive T or B cells. + / + The / HemA mice are further described in Bril et al. (2006), Thromb. Haemost., 95(2):341-7.

[0232] ssFVIIIXTEN with different parvovirus ITRs were generated by digesting the plasmid DNA construct with a restriction enzyme that recognizes ITR-related sequences and results in blunt-ended DNA. The double-stranded DNA products of the digestion (FVIII expression cassette and plasmid backbone) were heat denatured at 95° C. (denaturation) followed by cooling at 4° C. (renaturation) to allow the palindromic ITR sequences to form hairpins (FIG. 11). The resulting ssFVIIIXTEN (ssDNA) was then transferred to hFVIIIR593C via hydrodynamic tail vein injection. + / + FVIII activity was measured by Chromogenix Coatest® SP Factor VIII chromogenic assay according to the manufacturer's instructions.

[0233] Plasma FVIII activity normalized to percent of normal for animals injected with ssFVIIIXTEN is shown in FIG. 12. The results showed sustained FVIIIXTEN expression over time in all tested parvovirus ITRs, albeit at variable levels. All tested variants of GPV ITR or hybrid GPV ITR showed sustained reduction in FVIIIXTEN expression levels compared to other parvovirus ITRs. In contrast, HBoV1 ITR and B19 ITR showed an initial reduction in FVIIIXTEN by day 56, then stabilized through day 168, suggesting ITR-dependent persistence of the FVIIIXTEN transgene in vivo. Unlike GPV ITR, both B19 ITR and HBoV1 ITR showed significantly higher levels of FVIII expression, regardless of the variant tested. Among the different tested parvovirus ITRs, HBoV1 ITR showed significantly higher levels of FVIII expression than hFVIIIR593C. + / + In / HemA mice, the FVIII activity was significantly elevated (>1000%) relative to normal FVIII activity (Figure 10B). These results validate the functionality of the modified FVIIIXTEN expression cassettes with different parvoviral ITRs and support the ITR-dependent stability as well as sustained transgene expression in vivo. EXAMPLES

[0234] Assessment of closed-end FVIIIXTEN (ceFVIIIXTEN) DNA in vivo Although ssFVIIIXTEN (ssDNA) was effective in expressing modified FVIIIXTEN expression cassettes in vivo, there are several limitations associated with ssDNA used as a non-viral gene therapy vector. One of them is the level of endotoxin contamination due to the prokaryotic host (E. coli) used to generate plasmid DNA, which also contains foreign sequences such as antibiotic resistance genes and prokaryotic origin of replication required for selection and amplification in E. coli. To address these challenges and limitations, a eukaryotic cell-based system was developed to generate DNA therapeutic drug substance in the form of closed-end DNA (ceDNA) composed of FVIIIXTEN expression cassettes with parvovirus ITRs. Gene organization by ceDNA is similar to recombinant AAV vector DNA, but the conformation is different.

[0235] To generate this DNA vector, we utilized the baculovirus insect cell line, which is the only platform for influenza vaccine production that is widely used and FDA approved for biopharmaceutical production. As described in US Patent Application No. 63 / 069,073, three different approaches to ceDNA production were employed in the baculovirus system. Exemplary agarose gel images of purified ceDNA encoding modified V2.0 FVIIIXTEN flanked by AAV2 WT ITR or HBoV1 WT ITR compared to starting material (SM) are shown in Figure 13A.

[0236] To verify the functionality of the modified FVIIIXTEN expressed from ceDNA, purified ceFVIIIXTEN was administered to hFVIIIR593C via hydrodynamic tail vein injection. + / + In / HemA mice, 0.3 μg, 1.0 μg, or 2.0 μg per mouse were injected systemically, which is equivalent to 12 μg, 40 μg, and 80 μg / kg, respectively. Plasma samples from injected mice were collected at the indicated intervals, and FVIII activity was measured by the chromogenic assay described above.

[0237] Plasma FVIII activity normalized to percent normal for animals injected with ceFVIIIXTEN is shown in Figure 13B. Results showed a dose-dependent response in HemA mice, with FVIII expression observed at supraphysiological levels (>500% of normal levels) at the highest dose of AAV2 ceDNA or HBoV1 ceDNA tested. However, in the ceFVIIIXTEN AAV2 ITR injected cohort, a gradual decline in FVIII expression was observed until 140 days after injection, after which the levels stabilized. Interestingly, the FVIII expression levels observed with ceFVIIIXTEN HBoV1 ITR showed a similar trend to that observed in the ceFVIIIXTEN AAV2 ITR injected cohort.

[0238] These in vivo efficacy studies validate the functionality of ceFVIIIXTEN DNA and provide proof of concept that parvoviral ITRs can be used to generate functional non-viral gene therapy vectors encoding a transgene of interest in a baculovirus-insect cell line.

[0239] array

[0240] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7]

Table 2-8

Table 2-9

Table 2-10

Table 2-11

Table 2-12

Table 2-13

Claims

1. A nucleic acid molecule comprising a first inverted terminal repeat (ITR) and a second ITR flanking a gene cassette comprising a heterologous polynucleotide sequence, The first ITR is nucleotides 1 to 49, 50 to 58, and 59 to 125 of SEQ ID NO:1; Nucleotides 1-27 and 50-114 of SEQ ID NO: 15, or SEQ ID NO: 25 comprising a polynucleotide sequence that is at least about 75% identical to The second ITR is nucleotides 1 to 67, 68 to 76, and 77 to 125 of SEQ ID NO:2; Nucleotides 1-65 and 88-114 of SEQ ID NO: 16, or SEQ ID NO: 26 A nucleic acid molecule comprising a polynucleotide sequence that is at least about 75% identical to

2. The first ITR is nucleotides 1 to 49, 50 to 58, and 59 to 125 of SEQ ID NO:1; Nucleotides 1-27 and 50-114 of SEQ ID NO: 15, or SEQ ID NO: 25 2. The nucleic acid molecule of claim 1, comprising a polynucleotide sequence that is at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to

3. 3. The nucleic acid molecule of claim 1, wherein the first ITR comprises nucleotides 1 to 49, 50 to 58, and 59 to 125 of SEQ ID NO: 1, nucleotides 1 to 27 and 50 to 114 of SEQ ID NO: 15, or SEQ ID NO:

25.

4. The nucleic acid molecule of claim 3, wherein the first ITR comprises the polynucleotide sequence set forth in SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 5, SEQ ID NO: 9, SEQ ID NO: 13, SEQ ID NO: 15, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 21, or SEQ ID NO:

25.

5. The second ITR is nucleotides 1 to 67, 68 to 76, and 77 to 125 of SEQ ID NO:2; Nucleotides 1-65 and 88-114 of SEQ ID NO: 16, or SEQ ID NO: 26 3. The nucleic acid molecule of claim 1 or 2, comprising a polynucleotide sequence that is at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to

6. 6. The nucleic acid molecule of claim 1, wherein the second ITR comprises nucleotides 1 to 67, 68 to 76, and 77 to 125 of SEQ ID NO:2, nucleotides 1 to 65 and 88 to 114 of SEQ ID NO:16, or SEQ ID NO:

26.

7. The nucleic acid molecule of claim 6, wherein the second ITR comprises the polynucleotide sequence set forth in SEQ ID NO:2, SEQ ID NO:4, SEQ ID NO:6, SEQ ID NO:10, SEQ ID NO:14, SEQ ID NO:16, SEQ ID NO:18, SEQ ID NO:20, SEQ ID NO:22, SEQ ID NO:

26.

8. The nucleic acid molecule of claim 1 , wherein the gene cassette is a single-stranded nucleic acid or a double-stranded nucleic acid.

9. The nucleic acid molecule of claim 1 , wherein the heterologous polynucleotide sequence encodes a therapeutic protein.

10. 10. The nucleic acid molecule of claim 1, wherein the heterologous polynucleotide sequence encodes a clotting factor, a growth factor, a hormone, a cytokine, an antibody, a fragment thereof, or any combination thereof.

11. 2. The nucleic acid molecule of claim 1, wherein the heterologous polynucleotide sequence encodes X-linked dystrophin, MTM1 (myotubularin), tyrosine hydroxylase, AADC, cyclohydrolase, SMN1, FXN (frataxin), GUCY2D, RS1, CFH, HTRA, ARMS, CFB / CC2, CNGA / CNGB, Prf65, ARSA, PSAP, IDUA (MPS I), IDS (MPS II), PAH, GAA (acid alpha glucosidase), GALT, OTC, CMD1A, LAMA2, or any combination thereof.

12. 2. The nucleic acid molecule of claim 1, wherein the heterologous polynucleotide sequence encodes a microRNA (miRNA), and the miRNA downregulates expression of a target gene comprising SOD1, HTT, RHO, CD38, or any combination thereof.

13. The nucleic acid molecule of claim 1 , wherein the heterologous polynucleotide sequence is codon-optimized for expression in humans.

14. The nucleic acid molecule of claim 1 formulated with a delivery agent.

15. The nucleic acid molecule of claim 14 , wherein the delivery agent comprises a lipid nanoparticle (LNP).

16. 15. The nucleic acid molecule of claim 14, wherein the delivery agent comprises a liposome, a non-lipid polymer molecule, an endosome, or any combination thereof.

17. 10. The nucleic acid molecule of claim 1, formulated for intravenous administration, transdermal administration, intradermal administration, intraneuronal administration, intraocular administration, intrathecal administration, subcutaneous administration, intrapulmonary administration, oral administration, or any combination thereof.

18. 10. The nucleic acid molecule of claim 1, formulated for administration by in situ injection or by inhalation.