Stuffer sequences

EP4731774A2Pending Publication Date: 2026-04-29ASKBIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
ASKBIO INC
Filing Date
2024-06-21
Publication Date
2026-04-29

AI Technical Summary

Technical Problem

Current stuffer sequences in DNA molecules used for AAV vector production can elicit an immune response due to CpG motifs and may lead to the reconstitution of replicative virus, posing regulatory and safety concerns.

Method used

Development of short stuffer sequences with at least 85% identity to specific nucleotide sequences (SEQ ID NO: 1-6 or their complements), which can be incorporated into nucleic acids as linear or circular DNA, positioned relative to protelomerase binding sites and AAV ITR sequences to minimize immune stimulation and viral reactivation.

Benefits of technology

The short stuffer sequences reduce immune stimulation and prevent the reconstitution of replicative virus, providing a safer raw material for viral production while maintaining effective vector production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024034993_26122024_PF_FP_ABST
    Figure US2024034993_26122024_PF_FP_ABST
Patent Text Reader

Abstract

The technology described herein relates generally to recombinant AAV preparation and production.
Need to check novelty before this filing date? Find Prior Art

Description

STUFFER SEQUENCES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No.63 / 522,536 filed June 22, 2023, the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0002] The technology described herein relates generally to recombinant AAV preparation and production. BACKGROUND

[0003] The presence of a universal stuffer sequence in DNA molecules e.g., plasmid DNA, closed ended DNA molecules (clDNA) used for AAV vector production (using rep-cap, Ad helper, and ITR-transgene) allows for the use of a unique qPCR or ddPCR assay to quantify total residual DNA in the purified AAV preparation. Some stuffer sequences elicit an unintentional immune response. Recently, these stuffer sequences have raised regulatory and safety concerns because potential immune-stimulatory CpG motifs have been detected in the sequence. In addition, some stuffers can result in reconstitution of potentially replicative virus during production and / or in previously infected hosts.

[0004] Thus, there remains a need in the art for new synthetic stuffers that can provide a safe raw material for viral production. The disclosure addresses this need. SUMMARY

[0005] In one aspect provided herein is a nucleic acid comprising a short stuffer sequence. In some embodiments, the short stuffer sequence comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-6. It is noted that the “short stuffer sequence” is also referred to as “short stuffer” herein.

[0006] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequence is linear DNA. For example, the nucleic acid comprising the short stuffer sequence can be a close ended linear duplexed DNA (clDNA). The nucleic acid comprising the short stuffer sequence can be a circular DNA e.g., plasmid DNA. For example, the plasmid DNA can be a precursor plasmid DNA template to make clDNA.

[0007] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequence further comprises at least one protelomerase binding site. The short stuffer can be located upstream (5’-end) or downstream (3’-end) of the protelomerase binding site. Accordingly, in some embodiments of any one of the aspects described herein, theshort stuffer is located downstream of the protelomerase binding site. In some embodiments of any one of the aspects described herein, the short stuffer is located upstream of the protelomerase binding site.

[0008] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer comprises two protelomerase binding sites. The short stuffer can be located between or outside the two protelomerase binding sites. In some embodiments of any one of the aspects described herein, the short stuffer is located between the two protelomerase binding sites. In some embodiments, the distance between the protelomerase binding site and the stuffer sequence is between 2 to 15 nucleotides. In some embodiments, the distance between the protelomerase binding site and the stuffer sequence is between 5 to 10 nucleotides. In some embodiments, the distance between the protelomerase binding site and the stuffer sequence is between 6 to 9 nucleotides.

[0009] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequence further comprises a heterologous transgene operably linked to one or more regulatory elements. It is noted the short stuffer can be located upstream (5’-end) or downstream (3’-end) of the heterologous transgene. Accordingly, in some embodiments of any one of the aspects described herein, the short stuffer is located upstream (i.e., 5’) of the heterologous transgene. In some other embodiments of any one of the aspects described herein, the short stuffer is located downstream (i.e., 3’) of the heterologous transgene.

[0010] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequence further comprises at least one adeno-associated virus (AAV) inverted terminal repeat (ITR) sequence. The short stuffer can be located upstream (5’-end) or downstream (3’-end) of the ITR. Accordingly, in some embodiments of any one of the aspects described herein, the short stuffer is located upstream (e.g., 5’-end) of the ITR sequence. In some embodiments of any one of the aspects described herein, the short stuffer is located downstream (e.g., 3’-end) of the ITR sequence. In some embodiments, the distance between the 5’ ITR and the stuffer sequence is from about 8 nucleotides to about 25 nucleotides. For example, the distance between the 5’ ITR and the stuffer sequence is from about 10 nucleotides to about 20 nucleotides the distance. In some embodiments, the distance between the 5’ ITR and the stuffer sequence is from about 16 nucleotides. In some embodiments, the distance between the 3’ ITR and the stuffer sequence is from about 8 nucleotides to about 25 nucleotides. For example, the distance between the 3’ ITR and the stuffer sequence is from about 10 nucleotides to about 20 nucleotides the distance. In some other embodiments, the distance between the 3’ ITR and the stuffer sequence is 16bp.

[0011] In some embodiments of any one of the aspects described herein, the nucleic acidcomprising the short stuffer further comprises at least one ITR sequence and a heterologous transgene. The short stuffer, the at least one ITR sequence and the heterologous transgene can be positioned in any combination in the nucleic acid. For example, the at least one ITR sequence can be located between the short stuffer and the heterologous transgene. In another example, the heterologous transgene can be located between the short stuffer and the at least one ITR.

[0012] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises at least one protelomerase binding site and at least one ITR sequence. It is noted that the short stuffer, the at least one protelomerase binding site and the at least one ITR sequence can be positioned in any combination in the nucleic acid. For example, the short stuffer can be located between the at least one protelomerase binding site and the at least one ITR.

[0013] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer comprises at least two ITRs. The short stuffer can be located between the two ITRs or outside the two ITRs. In some embodiments, the short stuffer located outside the two ITRs. For example, the short stuffer can be located upstream (e.g., 5’) of the two ITRs. In another example, the short stuffer can be located downstream (e.g., 3’) of the two ITRs.

[0014] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer comprises two ITRs and a heterologous transgene. Generally, the heterologous transgene is located between the two ITRs. For example, the heterologous transgene is located between the two ITRs, and one of the ITRs is located between the short stuffer and the heterologous transgene. It is noted that the short stuffer can be located upstream (e.g., 5’) or down stream of one or both of the two ITRs. For example, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, the heterologous transgene is located between the first and second ITR sequences, and the short stuffer is upstream of the first ITR sequence. In another example, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, the heterologous transgene is located between the first and second ITR sequences, and the short stuffer is down stream of the second ITR sequence.

[0015] In some embodiments of any one of the aspects described herein, one of the ITRs is a mutant ITR such that the Rep nicking site (trs) is deleted and thereby promotes generation of self- complimentary (sc) genome of AAV as described in Granted US Patent US 7,790,154, which is incorporated herein by reference in its entirety. In a non-limiting example, a vector construct has a mutation in one TR, such that the Rep nicking site (trs) is deleted, while the other TR is wild type. This results in rolling hairpin replication initiating from the wild type (wt) end of the genome, proceeding through the mutant end without terminal resolution, and then continuing back across the genome again to create the dimer. The end product is a self-complementary genome with themutant TR in the middle and wt TRs now at each end. Replication and packaging of this molecule then proceeds as normal from the wt TRs, except that the dimeric structure is maintained in each round. In some embodiments, the short stuffer sequence of the invention can be located upstream (e.g, 5’) or downstream (e.g, 3’) of the mutant ITR. In one preferred embodiment, the short stuffer sequences are located upstream (e.g., 5’) of the mutant ITR and / or, downstream (e.g, 3’) of the wt ITR.

[0016] In some embodiments of one aspect described herein, a template for preferentially producing duplexed vector is generated with a resolvable AAV TR at one end and a modified AAV TR is produced by inserting a sequence into the TR as described in Granted US Patent US 7,790,154, which is incorporated herein by reference in its entirety. In some embodiments, the short stuffer sequence of the invention can be located upstream (e.g, 5’) or downstream (e.g, 3’) of the modified ITR comprising the insertion. Insertion of linkers displaces the wt AAV nicking site inward away from the native position, resulting in an inability to be resolved by the AAV Rep protein after replication. These substrates accumulate a dimeric intermediate until gene conversion takes place .These molecules that are produced are dimeric in form (covalently linked through the modified TR), more specifically, because they are self-complementary they provide a unique source of parvovirus vectors carrying double-stranded substrates In a non-limiting example, the wt AAV plasmid psub201, is used to produce this template as described in Samulski et al., (1987) J. Virology 61:3096.

[0017] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequences, comprise a single ITR and a 56-bp recognition sequence of protelomerase (TelN) to covalently join the top and bottom strands, allowing the nucleic acid encoding vector to be generated with just a single ITR as described as close-end double-stranded AAV (cceAAV) in Zhang et al., Molecular Therapy, Methods and Clinical development, Volume 32, Issue 1, 101206, March 14, 2024, which is incorporated herein by reference in its entirety. In some embodiments, the short stuffer sequences are located upstream (e.g. 5’) or downstream (e.g 3’) of the ITR sequence. In a preferred embodiment, the short stuffer sequences of the invention are located upstream (e.g., 5’) of the single ITR. In certain embodiments, the short stuffer sequences are located upstream (e.g. 5’) or downstream (e.g 3’) of the 56bp recognition sequence of protelomerase (TelN) sequence. In one preferred embodiment, the short stuffer sequences are located downstream (e.g 3’) of the 56bp recognition sequence of protelomerase (TelN) sequence.In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequences, comprise one ITR.

[0018] In the nucleic acid described herein, each ITR sequence can be selected independently from an ITR sequence of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10,AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4A and / or any chimeras thereof. It is noted when a nucleic acid described herein comprises two or more ITR sequences, the ITR sequences can be from the same AAV serotype or from differ AAV serotypes. Thus, in some embodiments of any one of the aspects described herein, the ITR sequences are from the same AAV serotype. In some other embodiments of any one of the aspects described herein, the ITR sequences are from the different AAV serotypes.

[0019] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer sequence further comprises at least one AAV ITR sequence comprising a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82-89.

[0020] In some embodiments of any one of the aspects described herein, the short stuffer is located upstream (e.g., 5’-end) of the ITR sequence, and the ITR comprises a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82-89. In some embodiments of any one of the aspects described herein, the short stuffer is located downstream (e.g., 3’-end) of the ITR sequence, and the ITR comprises a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82-89.

[0021] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises at least one ITR sequence and a heterologous transgene, and the ITR comprises a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82- 89. In these embodiments, the short stuffer, the at least one ITR sequence and the heterologous transgene can be positioned in any combination in the nucleic acid. For example, the at least one ITR sequence can be located between the short stuffer and the heterologous transgene. In another example, the heterologous transgene can be located between the short stuffer and the at least one ITR.

[0022] In some embodiments of any one of the aspects described herein, the nucleic acidcomprising the short stuffer further comprises at least one protelomerase binding site and at least one ITR sequence, where the ITR comprises a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82-89.

[0023] In these embodiments, the short stuffer, the at least one protelomerase binding site and the at least one ITR sequence can be positioned in any combination in the nucleic acid. For example, the short stuffer can be located between the at least one protelomerase binding site and the at least one ITR.

[0024] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer comprises at least two ITRs, where each ITR independently comprises a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82-89. In these embodiments, the short stuffer can be located between the two ITRs or outside the two ITRs. In some embodiments, the short stuffer is located outside the two ITRs. For example, the short stuffer can be located upstream (e.g., 5’) of the two ITRs. In another example, the short stuffer can be located downstream (e.g., 3’) of the two ITRs.

[0025] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer comprises two ITRs and a heterologous transgene, where each ITR independently comprises a nucleotide sequence having at least 85% (e.g., at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100%, i.e., fully identical) identity to a nucleotide sequence of any one of SEQ ID NO: 71 or 82-89. In these embodiments, the heterologous transgene can be located between the two ITRs. For example, the heterologous transgene is located between the two ITRs, and one of the ITRs is located between the short stuffer and the heterologous transgene. It is noted that the short stuffer can be located upstream (e.g., 5’) or downstream of one or both of the two ITRs. For example, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, the heterologous transgene is located between the first and second ITR sequences, and the short stuffer is upstream of the first ITR sequence. In another example, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, the heterologous transgene is located between the first and second ITR sequences, and the short stuffer is down stream of the second ITR sequence.

[0026] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises a nucleic acid sequence encoding one or more helperproteins that assist in AAV replication and thereby, rAAV production.

[0027] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises a nucleic acid sequence encoding one or more helper proteins that assist rAAV production and at least one protelomerase binding site. It is noted that the short stuffer, the at least one protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins can be positioned in any combination in the nucleic acid. For example, the short stuffer can be located between the protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins. Further, the short stuffer can be upstream (5’) or downstream (3’) of the nucleic acid sequence encoding one or more helper proteins. In some embodiments, the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding one or more helper proteins. In some other embodiments, the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding one or more helper proteins.

[0028] Helper proteins that assist rAAV production can comprise one or more of an E2A region, an E4 region, and a virus associated (VA) RNA region. In some embodiments, helper proteins that assist rAAV production can optionally comprise one or more of an E1 region, an E3 region and / or a Major Late Promoter (MLP) region. In several embodiments of any aspect described herein, the helper proteins are Adenoviral helper proteins. In some embodiments of any one of the aspects described herein, the nucleic acid sequence encoding one or more helper proteins comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 67-70.

[0029] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises a nucleic acid sequence encoding a AAV rep protein and a AAV cap protein.

[0030] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises at least one protelomerase binding site and a nucleic acid sequence encoding a AAV rep protein and a AAV cap protein.

[0031] It is noted that the short stuffer, the at least one protelomerase binding site and the nucleic acid sequence encoding the AAV rep and cap proteins can be positioned in any combination in the nucleic acid. For example, the short stuffer can be located between the at least one protelomerase binding site and the nucleic acid sequence encoding the AAV rep and cap proteins. Further, the short stuffer can be upstream (5’) or downstream (3’) of the nucleic acid sequence encoding the AAV rep and cap proteins. In some embodiments, the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding the AAV rep and cap proteins. For example, the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequenceencoding the AAV rep and AAV cap proteins. In some other embodiments, the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding the AAV rep and cap proteins. For example, the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

[0032] The AAV rep protein and the AAV cap protein can be independently from any AAV serotype. For example, AAV rep and cap proteins are independently from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof. It is noted that the AAV rep and cap proteins can be from the same AAV serotype or from different serotypes. Accordingly, in some embodiments, the AAV Rep protein and the AAV Cap protein are from the same AAV serotype. In some other embodiments, the AAV Rep protein and the AAV Cap protein are from different AAV serotype.

[0033] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises a stop codon (e.g., TAA, TAG or TGA). When the nucleic acid comprises a stop codon, the short stuffer can be located downstream (3’-end) or upstream (5’-end) of the stop codon. Accordingly, in some embodiments, the nucleic acid comprises a stop codon upstream (5’-end) of the short stuffer. In some other embodiments, the nucleic acid comprises a stop codon upstream (3’-end) of the short stuffer.

[0034] In some embodiments of any one of the aspects described herein, the nucleic acid comprises at least one ITR sequence and a protelomerase binding site upstream of the ITR sequence, and the short stuffer comprises the sequence of SEQ ID NO: 1 or 4 or a nucleotide sequence complementary to any one of SEQ ID NO: 1 or 4, and the short stuffer is located between the ITR and the protelomerase binding site.

[0035] In some embodiments of any one of the aspects described herein, the nucleic acid comprises at least one ITR sequence and a protelomerase binding site downstream of the ITR sequence, the short stuffer comprises the sequence of SEQ ID NO: 1 or 4 or a nucleotide sequence complementary to any one of SEQ ID NO: 1 or 4, and the short stuffer is located between the ITR and the protelomerase binding site.

[0036] In another aspect described herein is a nucleic acid comprising a large stuffer. Generally, the large stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotidesequence of any one of SEQ ID NO: 7-9 or 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81. It is noted that the “large stuffer” is also referred to as “large stuffer sequence” herein.

[0037] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the large stuffer further comprises a nucleic acid sequence encoding an AAV Rep protein and AAV cap protein operably linked to one or, more promoters. It is noted the large stuffer can be located upstream (e.g., 5’-end) or downstream (e.g., 3’-end) of the nucleic acid sequence encoding an AAV Rep protein. For example, the large stuffer is located upstream (e.g., 5’-end) of the nucleic acid sequence encoding the AAV Rep protein. In another example, the large stuffer is located downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep protein.

[0038] In some embodiments, the large stuffer is located within the nucleic acid sequence encoding an AAV Rep protein. For example, the large stuffer is located in an intron in the nucleic acid sequence encoding an AAV Rep protein.

[0039] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the larger stuffer further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter and the large stuffer is located upstream of the promoter. In some other embodiments of any one of the aspects described herein, the nucleic acid comprising the larger stuffer further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter and the large stuffer is located downstream of the promoter.

[0040] Exemplary AAV rep promotors include, but are not limited to, p5 promoter and p19 promoter. Thus, in some embodiments, the nucleic acid comprising the larger stuffer further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a p19 promoter and the large stuffer is located upstream of the promoter. In some other embodiments, the nucleic acid comprising the larger stuffer further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a p19 promoter and the large stuffer is located downstream of the promoter.

[0041] It is noted the AAV rep protein can be a large AAV rep protein or a small AAV rep protein. Thus, in some embodiments, the AAV rep protein is a large AAV rep protein. For example, the AAV rep protein is Rep68 or Rep78. In some embodiments of any one of the aspects described herein, the AAV Rep is large Rep (Rep68).

[0042] It is noted the AAV Rep protein can be a modified AAV Rep protein.

[0043] In some embodiments of the various aspects, the nucleic acid comprising the larger stuffer further comprises a nucleic acid sequence encoding an AAV Rep protein and an AAV cap protein operably linked to a promoter. It is noted that the sequence encoding the AAV Rep protein can be upstream or downstream of the sequence encoding the AAV cap protein. For example, thesequence encoding the AAV Rep protein can be upstream (e.g., 5’-end) of the sequence encoding the AAV Cap protein. In another example, the sequence encoding the AAV Rep protein can be downstream (e.g., 3’-end) of the sequence encoding the AAV Cap protein.

[0044] Generally, the large stuffer is not of a mammalian origin. In other words, the large stuffer does not comprise a nucleotide sequence of mammalian origin.

[0045] In some embodiments of any one of the aspects described herein, the large stuffer comprises a nucleotide sequence of non-mammalian origin.

[0046] In some embodiments of any one of the aspects described herein, the large stuffer is synthetic.

[0047] In some embodiments of any one of the aspects described herein, the large stuffer does not comprise more than one (e.g., 1, 2, 3, 4, 5, 6, 7 or 8) of the following: a transcription factor binding site; a regulatory element; an AAV Rep binding site; donor or acceptor splicing site; an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; a repetitive or palindrome sequence longer than 5 nucleotides; a strong secondary structure; or a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0048] In some embodiments of any one of the aspects described herein, the large stuffer comprises a GC content of less than about 50%, e.g., less than about 45%, or less than about 40%.

[0049] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the large stuffer is larger than 5.5kb.

[0050] In some embodiments of any one of the aspects described herein, the nucleic acid comprises the sequence MAG, where M is A or C, upstream of the large stuffer.

[0051] In some embodiments of any one of the aspects described herein, the large stuffer is of a size at least 2kb.

[0052] In yet another aspect described herein is a nucleic acid comprising a nucleotide sequence having at least 85% identity to SEQ ID NO: 10. In some embodiments, the nucleic acid comprising the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a large stuffer of a size at least 2 kb.

[0053] In some embodiments of any one of the aspects described herein, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 comprises one or more nucleotides between positions 45 and 46 of SEQ ID NO: 10. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 comprises from about 10 to about 10,000 nucleotides between positions 45 and 46 of SEQ ID NO: 10. In some embodiments, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 comprises from about 2,000 to about 10,000 nucleotides between positions 45 and 46 of SEQ ID NO: 10.

[0054] In some embodiments of any one of the aspects described herein, the nucleic acid comprising a nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter. It is noted the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 can be located upstream (e.g., 5’-end) or downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep protein. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located upstream (e.g., 5’-end) of the nucleic acid sequence encoding the AAV Rep protein. In another example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep protein.

[0055] In some embodiments of any one of the aspects described herein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located in the nucleic acid sequence encoding the AAV Rep protein. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located in an intron in the nucleic acid sequence encoding an AAV Rep protein.

[0056] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter and the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is upstream of the promoter. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is upstream of the p19 promoter.

[0057] In some other embodiments of any one of the aspects described herein, the nucleic acid comprising the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter and the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is downstream of the promoter. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is downstream of the p19 promoter.

[0058] In some embodiments of any one of the aspects described herein, the nucleic acid comprising a nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a nucleic acid sequence encoding an AAV Rep protein and a AAV Cap protein operably linked to a promoter. It is noted the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 can be located upstream (e.g., 5’-end) or downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep and Cap proteins. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located upstream (e.g., 5’-end) of the nucleic acid sequence encoding the AAV Rep and AAV Cap proteins. In another example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep and Cap proteins.

[0059] In some embodiments of any one of the aspects described herein, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located in the nucleic acid sequence encoding the AAV Rep and Cap proteins. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located in an intron in the nucleic acid sequence encoding the AAV Rep and Cap proteins.

[0060] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a nucleic acid sequence encoding an AAV Rep protein and a AAV Cap protein operably linked to a promoter and the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is upstream of the promoter. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is upstream of the p19 promoter.

[0061] In some other embodiments of any one of the aspects described herein, the nucleic acid comprising the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 further comprises a nucleic acid sequence encoding an AAV Rep protein and a Cap protein operably linked to a promoter and the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is downstream of the promoter. For example, the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is downstream of the p19 promoter.

[0062] In some embodiments of any one of the aspects described herein, the large stuffer of a size at least 2 kb is located between positions 44 and 45 of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10.

[0063] In some embodiments of any one of the aspects described herein, the large stuffer of size at least 2 kb comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81.

[0064] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is larger than 5.5 kb.

[0065] In some embodiments of any one of the aspects described herein, the large stuffer of a size at least 2 kb does not comprise a nucleotide sequence of mammalian origin.

[0066] In some embodiments of any one of the aspects described herein, the large stuffer of a size at least 2 kb comprises a nucleotide sequence of non-mammalian origin.

[0067] In some embodiments of any one of the aspects described herein, the large stuffer of a size at least 2 kb is synthetic.

[0068] In some embodiments of any one of the aspects described herein, the large stuffer of a size at least 2 kb does not comprise more than one (e.g., 1, 2, 3, 4, 5, 6, 7 or 8) of the following: atranscription factor binding site; a regulatory element; an AAV Rep binding site; a donor or acceptor splicing site; an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; a repetitive or palindrome sequence longer than 5 nucleotides; and / or a strong secondary structure; or a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0069] In some embodiments of any one of the aspects described herein, the large stuffer of a size at least 2 kb comprises a nucleotide sequence having at least 85% to any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81.

[0070] In some embodiments of any one of the aspects described herein, the 5’-end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to the sequence MAG, where M is A or C. In some embodiments, the 3’-end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to A or G. For example, the 5’-end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to the sequence MAG, where M is A or C, and the 3’-end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to A or G

[0071] In some embodiments of any one of the aspects described herein, the large stuffer of size at least 2 kb is located downstream of the sequence MAG, where M is A or C.

[0072] In some embodiments of any one of the aspects described herein, the nucleic acid comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 11-13 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-13. For example, the nucleic acid comprises a nucleotide sequence having at least 85% identity to SEQ ID NO: 13 or a nucleotide sequence complementary to SEQ ID NO: 13.

[0073] In some embodiments of any one of the aspects described herein, the nucleic acid described herein, e.g., a nucleic acid comprising a large stuffer prevents production of replication competent rAAV vector.

[0074] In several embodiments of any of the aspects described herein, the nucleic acids described herein can be single-stranded or double-stranded. In some embodiments of any one of the aspects described herein, the nucleic acid is a linear DNA. For example, the nucleic acid is closed ended linear duplex DNA (clDNA). Alternatively, clDNA is termed as no-end DNA or, neDNA. Exemplary closed linear duplexed DNA molecules, or no-end DNA molecules include, but are not limited to, doggybone DNA (dbDNA), and dumbbell shaped DNA.

[0075] In some embodiments of any one of the aspects described herein, the nucleic acid described herein is a vector, such as a plasmid, a bacmid, or a cosmid.

[0076] In another aspect described herein is a host cell comprising the nucleic acid described herein. In some embodiments, the host cell can be a prokaryotic cell, such as a bacterial cell.Examples of bacterial cells that can be host cells include, but are not limited to, E. coli. In some embodiments, the host cell can be eukaryotic cell. For example, the host cell can be an insect cell or a mammalian cell. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a HEK293 or a HeLa cell.

[0077] In some embodiments of any one of the aspects described herein, the host cell comprises at least one nucleic acid encoding one or more helper proteins that assist rAAV production, at least one nucleic acid encoding AAV rep and AAV cap proteins, and at least one nucleic acid encoding a transgene of interest. For example, the host cell comprises at least one nucleic acid encoding one or more helper proteins that assist in AAV replication and thereby rAAV production, at least one nucleic acid encoding AAV rep and AAV cap proteins, and at least one nucleic acid encoding a transgene of interest, wherein the nucleic acid encoding AAV rep and AAV cap proteins comprises a large stuffer described herein.

[0078] In another example, the host cell comprises at least one nucleic acid encoding one or more helper proteins that assist rAAV production, at least one nucleic acid encoding AAV rep and AAV cap proteins, and at least one nucleic acid encoding a transgene of interest, wherein at least one (e.g., only one or both) of the nucleic acid encoding one or more helper proteins and the nucleic acid encoding the transgene comprises a short stuffer described herein.

[0079] In yet another example, the host cell comprises at least one nucleic acid encoding one or more helper proteins that assist rAAV production, at least one nucleic acid encoding AAV rep and AAV cap proteins, and at least one nucleic acid encoding a transgene of interest, wherein at least one (e.g., only one or both) of the nucleic acid encoding one or more helper proteins and the nucleic acid encoding the transgene comprises a short stuffer described herein, and the nucleic acid encoding AAV rep and AAV cap proteins comprises a large stuffer described herein.

[0080] In another aspect described herein is a use of a nucleic acid described herein or a host cell described herein in a method of producing a plurality of viral particles. For example, use of a nucleic acid comprising a short stuffer or a cell comprising the same in a method of producing a plurality of viral particles, e.g., recombinant AAV (rAAV) particles. Without wishing to be bound by a theory, nucleic acid comprising a large stuffer and / or a sequence having at last 85% identity to SEQ ID NO: 10 described herein can prevent producing replication competent rAAV.

[0081] In another aspect described herein is a method for producing a plurality of viral particles. The method comprises culturing a host cell comprising a nucleic acid described herein, e.g., a nucleic acid comprising a short stuffer in a culture medium under conditions in which viral particle, e.g., rAAV particles are produced.

[0082] Without wishing to be bound by a theory, the stuffer sequences described herein don’t interfere with vector production.BRIEF DESCRIPTION OF THE DRAWINGS

[0083] FIGS.1A and 1B are schematics showing fragments containing short stuffer for functional validation. The selected stuffers were cloned upstream of the luciferase reporter gene, to detect any potential transcriptional activity. A negative control (no stuffer sequence) and positive control (CMV promoter) sequence were also designed for comparison purposes. The fragments with a stop codon include stuffer sequences #3, #6, and #7. The fragments without a stop codon include stuffer sequences #2, #8, and #9.

[0084] FIGS.2A-2C show schematics of the stuffer sequences being subcloned into the precursor backbone. FIG. 2A, controls (Negative, Positive, original stuffer), FIG. 2B, tests (Stuffer #3, #6, #7_with stop codon), and FIG.2C, tests (Stuffer #2, #8, #9_w / o stop codon).

[0085] FIGS.3A and 3B examine the neDNA production of the stuffer sequences which includes amplification of the precursor plasmid and manufacturing of neDNA containing the new stuffer sequences at a 20mL scale. FIG.3A, identity by size confirmation, and FIG.3B, identity by Sanger sequencing.

[0086] FIG. 4 shows a schematic of the stuffer sequence-luciferase neDNA transfection into HEK293 cells.

[0087] FIG.5 shows luciferase enzymatic activity of the different neDNA at 48h post-transfection in HEK293 cells relative to total proteins. Results are expressed as relative light units per micrograms of total proteins. Mock: no neDNA. Statistical analysis: One-way ANOVA multiple comparisons. ****: P<0,0001, ns: non-significant

[0088] FIG. 6 shows luciferase mRNA expressionat 48h post-transfection relative to β-actin mRNA.

[0089] FIG. 7 shows a schematic of the experimental design to examine whether the stuffer sequences —located immediately downstream of the TelRL sequence (ATCAGCACACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGTGCTGATA, SEQ ID NO: 14) — have promoter activity in vivo, as determined by luciferase expression (mRNA and enzymatic activity).

[0090] FIG. 8 shows all the stuffer sequences drive a similar level of luciferase expression (mRNA) compared with the construct without stuffer. While luciferase expression was observed at the mRNA level, no protein activity was detected. This was consistent with the results from the in vitro cell culture experiment.

[0091] FIG. 9 shows there are no clear differences between stuffers in terms of expression, even normalized per plasmid copy number. The expression of luciferase mRNA in the animals that received DNA injection is substantially above background (mice that received vehicle showed no expression). TelRL sequence seems to drive a low level of expression in liver.

[0092] FIG.10 examines 4 different AAV products to validate in an in vivo study.

[0093] FIG.11 identifies the AAV92L batches.

[0094] FIG.12 shows the genomic titers by ddPCR.

[0095] FIG.13 examines genome titer by ITR-ddPCR versus particle titer by SEC-HPLC.

[0096] FIG.14 examines ITR ddPCR and SEC-HPLC values.

[0097] FIGS. 15A and 15B show the residual neDNA and pDNA in copies / mL (FIG. 15A) and % relative to ITR ddPCR (FIG.15A) titers in purified AAV vectors. Residual neDNA are detected by ddPCR using primers and probes targeting the stuffer 2, stuffer 7, 5’TelRS or 3’telRS sequence, depending on the neDNA used for AAV production. Residual pDNA is detected by ddPCR using primers and probe targeting the kanamycin-resistance coding sequence.

[0098] FIG.16 shows the vectors genome integrity by AGE.

[0099] FIG.17 depicts the design of the synthetic stuffers with a “neutral” sequence.

[0100] FIG.18 shows the design of the final rep-cap cassettes with the synthetic stuffers.

[0101] FIG. 19 shows the experimental design for selection of the rep-cap cassettes versus the 2 / 8 rep-cap.

[0102] FIG. 20 shows the comparison of ddPCR titers and residual plasmid for 2 / 8, pRC2- 8_SynINT2.2, pRC2-8_SynINT3.1, and pRC2-8_SynINT3.2.

[0103] FIG. 21 examines the residual plasmid ratios for 2 / 8, pRC2-8_SynINT2.2, pRC2- 8_SynINT3.1, and pRC2-8_SynINT3.2.

[0104] FIG.22 examines the SEC-HPLC data for 2 / 8, pRC2-8_SynINT2.2, pRC2-8_SynINT3.1, and pRC2-8_SynINT3.2.

[0105] FIG.23 shows a schematic of the transcription map of the AAV genome.

[0106] FIG. 24 examines how a short optimized intronic sequence was slightly modified by introduction of unique restriction sites, and removal of other restriction sites for further cloning purposes and reached a final size of 167 bp. Sequences shown are SEQ ID NO: 90 (top) and SEQ ID NO: 91 (bottom).

[0107] FIG. 25 shows the intronic sequence that was used to select the insertion sites to reconstitute consensus splice donor and acceptor sites. Sequence shown is MAGGTRAG(N)xYNYYRAY(N)y(Y)10NYAGR (SEQ ID NO: 92), wherein N is A, G, C or T; M is C or A; R is A or G; Y is C or T; x is 10 to 10,000; and y is 0 to 20.

[0108] FIG.26 depicts a representation of the initial rep-cap cassette used for production of AAV serotype 8, and the new rep-cap cassettes obtained following the synthetic intron insertion downstream and upstream the p19 promoter.

[0109] FIG.27 depicts a plasmid map of the pUC19_beta-actin (SEQ ID NO: 72).

[0110] FIG.28 depicts a plasmid map of the _Positive Control DNA (SEQ ID NO: 73).

[0111] FIGS.29A and 29B show schematics of elements comprised in XX85 further comprising the protelomerase sites. Exemplary elements include stuffer sequences 2 and / or 7 (SEQ ID NOs: 1 and 4) can be included upstream of the 5’ end of the E4 region (i.e. a 5’ stuffer) and / or downstream of the 3’ end of the E2A region (i.e. a 3’ stuffer). Further, a 5’ stuffer can be located at the AscI site 5’ of the E4 region and / or a 3’ stuffer can be located at the NotI site 3’ of the E2A region.

[0112] FIGS.30A-30C present data showing Rep proteins expression in Pro10TMcells analyzed with the WES capillary and immuno-blotting system (Protein Simple) at harvest for AAV3B- lux2A-GFP 2L scale productions. Rep proteins were detected with monoclonal antibody 303.9 (Progen) and expression was normalized with beta-actin. R-511 is pUC_RC3b; R-512 is pRC3b_SynINT 3.1 (downstream p19); and R-513 is pRC3b_SynINT 3.2 (upstream p19).

[0113] FIG. 31 present data showing migration, purity and VP proteins ratio for purified AAV3B-lux2A-GFP 2L scale productions analyzed by CE-SDS with the Maurice capillary electrophoresis system (Protein Simple). R-511 is pUC_RC3b; R-512 is pRC3b_SynINT 3.1 (downstream p19); and R-513 is pRC3b_SynINT 3.2 (upstream p19).

[0114] FIG. 32 present bar graphs showing vector genome titers of purified AAV3B 2L scale productions with two different transgenes obtained by ITR ddPCR.. R-511 is pUC_RC3b; R-512 is pRC3b_SynINT 3.1 (downstream p19); and R-513 is pRC3b_SynINT 3.2 (upstream p19).

[0115] FIG.33 present bar graphs showing capsid titer and A260 / A280 ratio by SEC-HPLC for purified AAV3B-lux2A-GFP 2L scale productions with two different transgens. R-511 is pUC_RC3b; R-512 is pRC3b_SynINT 3.1 (downstream p19); and R-513 is pRC3b_SynINT 3.2 (upstream p19).

[0116] FIG. 34A is a schematic of cell plating and passages required for amplification of replication competent AAV (rcAAV) and detection by qPCR targeting the Rep coding sequence..

[0117] FIG.34B shows that rcAAV can be detected in a specific AAV3B vector when produced with a standard rep-cap plasmid, while rcAAV wereundetected when the same vector was produced with the oversized rep-cap constructs comprising large stuffer sequences.

[0118] FIG. 35 depicts a bar graph showing final product productivity (VG / L) by ITR-ddPCR obtained with the 3 batches. UC Pool (ultracentrifugation pool) is the pool of the iodixanol density gradient fractions collected after ultracentrifugation (containing full AAV particles). BV pool (bulk virus pool) is the pool of the fractions that are eluted from the affinity capture column (containing purified AAV particles) after pH adjustment. The BV pool is the input for the density gradient

[0119] FIG.36 depicts a bar graph showing final product productivity (VP / L) by ELISA obtained with the three batches.

[0120] FIG.37 depicts a bar graph showing final product productivity (VG / L) by ITR-qPCR obtained with the three batches. UC Pool (ultracentrifugation pool) is the pool of the iodixanol density gradient fractions collected after ultracentrifugation (containing full AAV particles). BV pool (bulk virus pool) is the pool of the fractions that are eluted from the affinity capture column (containing purified AAV particles) after pH adjustment. The BV pool is the input for the density gradient

[0121] FIG. 38 depicts SDS-PAGE and Silver staining analysis of the final products of R-511, R-512, R-513.

[0122] FIG.39 depicts western blot analysis of the final products of R-511, R-512 and R-513.

[0123] FIG. 40 depicts a bar graph showing HCDNA-qPCR data obtained with the 123bp amplicon or the 254bp amplicon, with (+) and without (-) DNase pre-treatment, normalized using the titer 1E+09 VG obtained by ddPCR.

[0124] Figure 41 depicts a bar graph showing data obtained with the KanR-ddPCR assay, with (+) and without (-) DNase pre-treatment, presented as the percentage (right) between KanR sequence copy number and the VG titer obtained by ITR-qPCR.

[0125] FIG.42 depicts the design of an exemplary AAV-GFP vector according to the disclosure.

[0126] FIG.43 depicts a plasmid map of the ssAAV9-GFP vector (SEQ ID NO: 93).

[0127] FIGS. 44A and 44B depicts results of ITR-ddPCR and SEC-HPLC (FIG. 44A) and sedimentation velocity analytical ultracentrifugation (FIG. 44B) analysis of rAAV production using the ssAAV-GFP vector depicted in FIG.43. DETAILED DESCRIPTION

[0128] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. All documents, or portions of documents, cited in this application, including, but not limited to, patents, patent applications, articles, books, and treatises, are hereby expressly incorporated by reference in their entirety for any purpose. Short stuffer sequences

[0129] Embodiments of the various aspects described herein include a stuffer sequences, e.g., a short stuffer or a large stuffer. As used herein, “stuffer” sequence refers to a non-coding sequence of non-viral origin. A stuffer sequence preferably has no or minimal regulatory effect on coding sequences in the same nucleic acid molecule.

[0130] In one aspect provided herein is a nucleic acid comprising a short stuffer sequence. Generally, the short stuffer sequence is less than 500 nucleotides in length. For example, the short stuffer is from about 50 nucleotides to about 400 nucleotides in length. In some embodiments, the short stuffer is from about 100 nucleotides to about 300 nucleotides in length. For example, the short stuffer is from about 150 nucleotides to about 250 nucleotides in length. In some embodiments, the short stuffer is less than about 200 nucleotides in length.

[0131] In some embodiments, the short stuffer does not comprise one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8 or all 9) of a TATA box, CpG dinucleotides, an endonuclease cleavage site, a transcription factor binding site, an open reading frame (ORF) longer than 10 amino acids, a donor or acceptor splicing site, a regulatory element, a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides, and / or an AAV Rep binding site.

[0132] In some embodiments, the short stuffer has a GC content of less than about 50%. For example, the short stuffer has a GC content of less than about 45%. In some embodiments, the short stuffer has a GC content of less than about 40%, e.g., less than about 35%.

[0133] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-6. For example, the short stuffer comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-6. In some embodiments, the short stuffer comprises a nucleotide sequence having 100% identity with any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-6. For example, the short stuffer consists of a nucleotide sequence having 100% identity with any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-6.

[0134] In some embodiments, the short stuffer comprises a nucleotide sequence having at least 85% identity with SEQ ID NO: 1 or 4 or a nucleotide sequence complementary to SEQ ID NO: 1 or 4. For example, the short stuffer comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with SEQ ID NO: 1 or 4 or a nucleotide sequence complementary to SEQ ID NO: 1 or 4. In some embodiments, the short stuffer comprises a nucleotide sequence having 100% identity with SEQ ID NO: 1 or 4 or a nucleotide sequence complementary to SEQ ID NO: 1 or 4. For example, the short stuffer consists of a nucleotide sequence having 100% identity with SEQ ID NO: 1 or 4 or a nucleotide sequence complementary to SEQ ID NO: 1 or 4.

[0135] In some embodiments, the short stuffer comprises a nucleotide sequence comprising atleast 75 contiguous nucleotides of any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to at least 75 contiguous nucleotides of any one of SEQ ID NO: 1-6. For example, the short stuffer comprises a nucleotide sequence comprising at least 100 contiguous nucleotides of any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to at least 100 contiguous nucleotides of any one of SEQ ID NO: 1-6. In some embodiments, the short stuffer comprises a nucleotide sequence comprising at least 125 contiguous nucleotides of any one of SEQ ID NO: 1- 6 or a nucleotide sequence complementary to at least 125 contiguous nucleotides of any one of SEQ ID NO: 1-6. For example, the short stuffer comprises a nucleotide sequence comprising at least 150 contiguous nucleotides of any one of SEQ ID NO: 1-6 or a nucleotide sequence complementary to at least 150 contiguous nucleotides of any one of SEQ ID NO: 1-6.

[0136] Short stuffer sequences as described herein are preferably neutral sequences e.g, having no or, minimum effect on transcriptional activity of the nucleic acid molecule they are present in. Larger Stuffer

[0137] Embodiments of the various aspects described herein include a large stuffer. As used herein, a large stuffer or large stuffer sequence is a non-coding sequence that can increase the size of an expression cassette over the DNA packaging limit of the AAV capsid, which is around 5 kb. Preferably, the large stuffer has no or minimal regulatory effect e.g., transcriptional modulation effect on coding sequences in the same nucleic acid molecule. Large stuffer sequences are preferably neutral sequences.

[0138] Generally, the large stuffer has length greater than 2 kb. For example, the large stuffer has a length greater than 2.1kb, greater than 2.2kb, greater than 2.3kb, greater than 2.4kb, greater than 2.5kb, greater than 2.6kb, greater than 2.7kb, greater than 2.8kb, greater than 2.9kb, or greater than 3.0kb. In some embodiments, the large stuffer has a length greater than 3.1kb, greater than 3.2kb, greater than 3.3kb, greater than 3.4kb, greater than 3.5kb, greater than 3.6kb, greater than 3.7kb, greater than 3.8kb, greater than 3.9kb, or greater than 4.0kb. For example, the large stuffer has a length greater than 4.1kb, greater than 4.2kb, greater than 4.3kb, greater than 4.4kb, greater than 4.5kb, greater than 4.6kb, greater than 4.7kb, greater than 4.8kb, greater than 4.9kb, or greater than 5.0kb. In some embodiments, the large stuffer has a length greater than 5.1kb, greater than 5.2kb, greater than 5.3kb, greater than 5.4kb, 5.5kb, such as greater than 5.6kb, greater than 5.7kb, greater than 5.8kb, greater than 5.9kb, or greater than 6.0kb, In some embodiments, the large stuffer has a length greater than 6.5kb, greater than 6.6kb, greater than 6.7kb, greater than 6.8kb, greater than 6.9kb, greater than 7.0kb, greater than 7.1kb, greater than 7.2kb, greater than 7.3kb, greater than 7.4kb, greater than 7.5kb, greater than 7.6kb, greater than 7.7kb, greater than 7.8kb, greater than 7.9kb, or greater than 8.0kb.

[0139] Generally, the large stuffer is not of a mammalian origin. In other words, the large stuffer does not comprise a nucleotide sequence of mammalian origin. In some embodiments of any one of the aspects described herein, the large stuffer comprises a nucleotide sequence of non- mammalian origin. In some embodiments of any one of the aspects described herein, the large stuffer is synthetic.

[0140] Typically, the large stuffer does not comprise more than one (e.g., 1, 2, 3, 4, 5, 6, 7 or 8) of the following: a transcription factor binding site; a regulatory element; an AAV Rep binding site; a donor or acceptor splicing site; an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; a repetitive or palindrome sequence longer than 5 nucleotides; and / or a strong secondary structure; or a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0141] In some embodiments, the large stuffer has a GC content of less than about 50%. For example, the large stuffer has a GC content of less than about 45%. In some embodiments, the large stuffer has a GC content of less than about 40%, e.g., less than about 35%.

[0142] In some embodiments of any one of the aspects described herein, the large stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81. For example, the large stuffer comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81. In some embodiments, the large stuffer comprises a nucleotide sequence having 100% identity with any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81. For example, the large stuffer consists of a nucleotide sequence having 100% identity with any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to any one of SEQ ID NO: 7-9 or 81.

[0143] In some embodiments, the large stuffer comprises a nucleotide sequence having at least 85% identity with SEQ ID NO: 9 or a nucleotide sequence complementary to SEQ ID NO: 9. For example, the large stuffer comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with SEQ ID NO: 9 or a nucleotide sequence complementary to SEQ ID NO: 9. In some embodiments, the large stuffer comprises a nucleotide sequence having 100% identity with SEQ ID NO: 9 or a nucleotide sequence complementary to SEQ ID NO: 9. For example, the large stuffer consists of a nucleotide sequence having 100% identity with SEQ ID NO: 9 or a nucleotide sequence complementary to SEQ ID NO: 9.Short Stuffer: Fragments of larger stuffer

[0144] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1250 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1250 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1000 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1000 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 950 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 950 nuceleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 900 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 900 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 850 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 850 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 800 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 800 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 750 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 750 nucleotides of any one of SEQ ID NO: 7-9 or 81.

[0145] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 250 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 250 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 200 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 200 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 150 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 150 nucleotides of any one of SEQ ID NO: 7-9 or 81. Forexample, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 100 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 100 nucleotides of any one of SEQ ID NO: 7-9 or 81.

[0146] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 25 nucleotides to about 100 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 25 nucleotides to about 100 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 100 nucleotides to about 900 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 100 nucleotides to about 900 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 150 nucleotides to about 800 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 150 nucleotides to about 800 nucleotides any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 250 nucleotides to about 750 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 250 nucleotides to about 750 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 275 nucleotides to about 700 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 275 nucleotides to about 700 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 300 nucleotides to about 675 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 300 nucleotides to about 675 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 325 nucleotides to about 650 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 325 nucleotides to about 650 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 350 nucleotides to about 600 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 325 nucleotides to about 600 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 375 nucleotides to about 550 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 375 nucleotides to about 550nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 400 nucleotides to about 500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 400 nucleotides to about 500 nucleotides of any one of SEQ ID NO: 7-9 or 81.

[0147] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 45 nucleotide, about 50 nucleotides, about 55 nucleotides, about 60 nucleotides, about 65 nucleotide, about 70 nucleotides, about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotide, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotide, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, about 200 nucleotides, about 205 nucleotides, about 210 nucleotides, about 215 nucleotides, about 220 nucleotides, about 225 nucleotides, about 230 nucleotides, about 235 nucleotides, about 240 nucleotides, about 245 nucleotide, about 250 nucleotides, about 255 nucleotides, about 260 nucleotides, about 265 nucleotide, about 270 nucleotides, about 275 nucleotides, about 280 nucleotides, about 285 nucleotides, about 290 nucleotides, about 295 nucleotides, about 300 nucleotides, about 305 nucleotides, about 310 nucleotides, about 315 nucleotides, about 320 nucleotides, about 325 nucleotides, about 330 nucleotides, about 335 nucleotides, about 340 nucleotides, about 345 nucleotide, about 350 nucleotides, about 355 nucleotides, about 360 nucleotides, about 365 nucleotide, about 370 nucleotides, about 375 nucleotides, about 380 nucleotides, about 385 nucleotides, about 390 nucleotides, about 395 nucleotides, about 400 nucleotides, about 405 nucleotides, about 410 nucleotides, about 415 nucleotides, about 420 nucleotides, about 425 nucleotides, about 430 nucleotides, about 435 nucleotides, about 440 nucleotides, about 445 nucleotide, about 450 nucleotides, about 455 nucleotides, about 460 nucleotides, about 465 nucleotide, about 470 nucleotides, about 475 nucleotides, about 480 nucleotides, about 485 nucleotides, about 490 nucleotides, about 495 nucleotides, about 500 nucleotides, about 505 nucleotides, about 510 nucleotides, about 515 nucleotides, about 520 nucleotides, about 525 nucleotides, about 530 nucleotides, about 535 nucleotides, about 540 nucleotides, about 545 nucleotide, about 550 nucleotides, about 555 nucleotides, about 560 nucleotides, about 565 nucleotide, about 570 nucleotides, about 575 nucleotides, about 580 nucleotides, about 585 nucleotides, about 590 nucleotides, about 595 nucleotides, about 600 n , about 605 nucleotides, about 610nucleotides, about 615 nucleotides, about 620 nucleotides, about 625 nucleotides, about 630 nucleotides, about 635 nucleotides, about 640 nucleotides, about 645 nucleotide, about 650 nucleotides, about 655 nucleotides, about 660 nucleotides, about 665 nucleotide, about 670 nucleotides, about 675 nucleotides, about 680 nucleotides, about 685 nucleotides, about 690 nucleotides, about 695 nucleotides, about 700 nucleotides, about 705 nucleotides, about 710 nucleotides, about 715 nucleotides, about 720 nucleotides, about 725 nucleotides, about 730 nucleotides, about 735 nucleotides, about 740 nucleotides, about 745 nucleotide, about 750 nucleotides, about 755 nucleotides, about 760 nucleotides, about 765 nucleotide, about 770 nucleotides, about 775 nucleotides, about 780 nucleotides, about 785 nucleotides, about 790 nucleotides, about 795 nucleotides, about 800 nucleotides, about 805 nucleotides, about 810 nucleotides, about 815 nucleotides, about 820 nucleotides, about 825 nucleotides, about 830 nucleotides, about 835 nucleotides, about 840 nucleotides, about 845 nucleotide, about 850 nucleotides, about 855 nucleotides, about 860 nucleotides, about 865 nucleotide, about 870 nucleotides, about 875 nucleotides, about 880 nucleotides, about 885 nucleotides, about 890 nucleotides, about 895 nucleotides, about 900 nucleotides, , about 905 nucleotides, about 910 nucleotides, about 915 nucleotides, about 920 nucleotides, about 925 nucleotides, about 930 nucleotides, about 935 nucleotides, about 940 nucleotides, about 945 nucleotide, about 950 nucleotides, about 955 nucleotides, about 960 nucleotides, about 965 nucleotide, about 970 nucleotides, about 975 nucleotides, about 980 nucleotides, about 985 nucleotides, about 990 nucleotides, about 995 nucleotides, about 1000 nucleotides, , about 1005 nucleotides, about 1010 nucleotides, about 1015 nucleotides, about 1020 nucleotides, about 1025 nucleotides, about 1030 nucleotides, about 1035 nucleotides, about 1040 nucleotides, about 1045 nucleotide, about 1050 nucleotides, about 1055 nucleotides, about 1060 nucleotides, about 1065 nucleotide, about 1070 nucleotides, about 1075 nucleotides, about 1080 nucleotides, about 1085 nucleotides, about 1090 nucleotides, about 1095 nucleotides, about 1100 nucleotides, about 1105 nucleotides, about 1110 nucleotides, about 1115 nucleotides, about 1120 nucleotides, about 1125 nucleotides, about 1130 nucleotides, about 1135 nucleotides, about 1140 nucleotides, about 1145 nucleotide, about 1150 nucleotides, about 1155 nucleotides, about 1160 nucleotides, about 1165 nucleotide, about 1170 nucleotides, about 1175 nucleotides, about 1180 nucleotides, about 1185 nucleotides, about 1190 nucleotides, about 1195 nucleotides, about 1200 nucleotides, about 1205 nucleotides, about 1210 nucleotides, about 1215 nucleotides, about 1220 nucleotides, about 1225 nucleotides, about 1230 nucleotides, about 1235 nucleotides, about 1240 nucleotides, about 1245 nucleotide, about 1250 nucleotides, about 1255 nucleotides, about 1260 nucleotides, about 1265 nucleotide, about 1270 nucleotides, about 1275 nucleotides, about 1280 nucleotides, about 1285 nucleotides, about 1290 nucleotides, about 1295 nucleotides, about 1300 nucleotides, about 1305 nucleotides, about 1310nucleotides, about 1315 nucleotides, about 1320 nucleotides, about 1325 nucleotides, about 1330 nucleotides, about 1335 nucleotides, about 1340 nucleotides, about 1345 nucleotide, about 1350 nucleotides, about 1355 nucleotides, about 1360 nucleotides, about 1365 nucleotide, about 1370 nucleotides, about 1375 nucleotides, about 1380 nucleotides, about 1385 nucleotides, about 1390 nucleotides, about 1395 nucleotides, about 1400 nucleotides, about 1405 nucleotides, about 1410 nucleotides, about 1415 nucleotides, about 1420 nucleotides, about 1425 nucleotides, about 1430 nucleotides, about 1435 nucleotides, about 1440 nucleotides, about 1445 nucleotide, about 1450 nucleotides, about 1455 nucleotides, about 1460 nucleotides, about 1465 nucleotide, about 1470 nucleotides, about 1475 nucleotides, about 1480 nucleotides, about 1485 nucleotides, about 1490 nucleotides, about 1495 nucleotides, or about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of about 25 nucleotides, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 45 nucleotide, about 50 nucleotides, about 55 nucleotides, about 60 nucleotides, about 65 nucleotide, about 70 nucleotides, about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotide, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotide, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, about 200 nucleotides, about 205 nucleotides, about 210 nucleotides, about 215 nucleotides, about 220 nucleotides, about 225 nucleotides, about 230 nucleotides, about 235 nucleotides, about 240 nucleotides, about 245 nucleotide, about 250 nucleotides, about 255 nucleotides, about 260 nucleotides, about 265 nucleotide, about 270 nucleotides, about 275 nucleotides, about 280 nucleotides, about 285 nucleotides, about 290 nucleotides, about 295 nucleotides, about 300 nucleotides, about 305 nucleotides, about 310 nucleotides, about 315 nucleotides, about 320 nucleotides, about 325 nucleotides, about 330 nucleotides, about 335 nucleotides, about 340 nucleotides, about 345 nucleotide, about 350 nucleotides, about 355 nucleotides, about 360 nucleotides, about 365 nucleotide, about 370 nucleotides, about 375 nucleotides, about 380 nucleotides, about 385 nucleotides, about 390 nucleotides, about 395 nucleotides, about 400 nucleotides, about 405 nucleotides, about 410 nucleotides, about 415 nucleotides, about 420 nucleotides, about 425 nucleotides, about 430 nucleotides, about 435 nucleotides, about 440 nucleotides, about 445 nucleotide, about 450 nucleotides, about 455 nucleotides, about 460 nucleotides, about 465 nucleotide, about 470 nucleotides, about 475 nucleotides, about 480 nucleotides, about 485 nucleotides, about 490 nucleotides, about 495 nucleotides, about 500 nucleotides, about 505 nucleotides, about 510 nucleotides, about 515nucleotides, about 520 nucleotides, about 525 nucleotides, about 530 nucleotides, about 535 nucleotides, about 540 nucleotides, about 545 nucleotide, about 550 nucleotides, about 555 nucleotides, about 560 nucleotides, about 565 nucleotide, about 570 nucleotides, about 575 nucleotides, about 580 nucleotides, about 585 nucleotides, about 590 nucleotides, about 595 nucleotides, about 600 nucleotides, , about 605 nucleotides, about 610 nucleotides, about 615 nucleotides, about 620 nucleotides, about 625 nucleotides, about 630 nucleotides, about 635 nucleotides, about 640 nucleotides, about 645 nucleotide, about 650 nucleotides, about 655 nucleotides, about 660 nucleotides, about 665 nucleotide, about 670 nucleotides, about 675 nucleotides, about 680 nucleotides, about 685 nucleotides, about 690 nucleotides, about 695 nucleotides, about 700 nucleotides, about 705 nucleotides, about 710 nucleotides, about 715 nucleotides, about 720 nucleotides, about 725 nucleotides, about 730 nucleotides, about 735 nucleotides, about 740 nucleotides, about 745 nucleotide, about 750 nucleotides, about 755 nucleotides, about 760 nucleotides, about 765 nucleotide, about 770 nucleotides, about 775 nucleotides, about 780 nucleotides, about 785 nucleotides, about 790 nucleotides, about 795 nucleotides, about 800 nucleotides, about 805 nucleotides, about 810 nucleotides, about 815 nucleotides, about 820 nucleotides, about 825 nucleotides, about 830 nucleotides, about 835 nucleotides, about 840 nucleotides, about 845 nucleotide, about 850 nucleotides, about 855 nucleotides, about 860 nucleotides, about 865 nucleotide, about 870 nucleotides, about 875 nucleotides, about 880 nucleotides, about 885 nucleotides, about 890 nucleotides, about 895 nucleotides, about 900 nucleotides, , about 905 nucleotides, about 910 nucleotides, about 915 nucleotides, about 920 nucleotides, about 925 nucleotides, about 930 nucleotides, about 935 nucleotides, about 940 nucleotides, about 945 nucleotide, about 950 nucleotides, about 955 nucleotides, about 960 nucleotides, about 965 nucleotide, about 970 nucleotides, about 975 nucleotides, about 980 nucleotides, about 985 nucleotides, about 990 nucleotides, about 995 nucleotides, about 1000 nucleotides, , about 1005 nucleotides, about 1010 nucleotides, about 1015 nucleotides, about 1020 nucleotides, about 1025 nucleotides, about 1030 nucleotides, about 1035 nucleotides, about 1040 nucleotides, about 1045 nucleotide, about 1050 nucleotides, about 1055 nucleotides, about 1060 nucleotides, about 1065 nucleotide, about 1070 nucleotides, about 1075 nucleotides, about 1080 nucleotides, about 1085 nucleotides, about 1090 nucleotides, about 1095 nucleotides, about 1100 nucleotides, about 1105 nucleotides, about 1110 nucleotides, about 1115 nucleotides, about 1120 nucleotides, about 1125 nucleotides, about 1130 nucleotides, about 1135 nucleotides, about 1140 nucleotides, about 1145 nucleotide, about 1150 nucleotides, about 1155 nucleotides, about 1160 nucleotides, about 1165 nucleotide, about 1170 nucleotides, about 1175 nucleotides, about 1180 nucleotides, about 1185 nucleotides, about 1190 nucleotides, about 1195 nucleotides, about 1200 nucleotides, about 1205 nucleotides, about 1210 nucleotides, about 1215nucleotides, about 1220 nucleotides, about 1225 nucleotides, about 1230 nucleotides, about 1235 nucleotides, about 1240 nucleotides, about 1245 nucleotide, about 1250 nucleotides, about 1255 nucleotides, about 1260 nucleotides, about 1265 nucleotide, about 1270 nucleotides, about 1275 nucleotides, about 1280 nucleotides, about 1285 nucleotides, about 1290 nucleotides, about 1295 nucleotides, about 1300 nucleotides, about 1305 nucleotides, about 1310 nucleotides, about 1315 nucleotides, about 1320 nucleotides, about 1325 nucleotides, about 1330 nucleotides, about 1335 nucleotides, about 1340 nucleotides, about 1345 nucleotide, about 1350 nucleotides, about 1355 nucleotides, about 1360 nucleotides, about 1365 nucleotide, about 1370 nucleotides, about 1375 nucleotides, about 1380 nucleotides, about 1385 nucleotides, about 1390 nucleotides, about 1395 nucleotides, about 1400 nucleotides, about 1405 nucleotides, about 1410 nucleotides, about 1415 nucleotides, about 1420 nucleotides, about 1425 nucleotides, about 1430 nucleotides, about 1435 nucleotides, about 1440 nucleotides, about 1445 nucleotide, about 1450 nucleotides, about 1455 nucleotides, about 1460 nucleotides, about 1465 nucleotide, about 1470 nucleotides, about 1475 nucleotides, about 1480 nucleotides, about 1485 nucleotides, about 1490 nucleotides, about 1495 nucleotides, or about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81.

[0148] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 7. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 7.

[0149] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 8. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 8.

[0150] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 9. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 9.

[0151] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary a contiguous fragment of about 25 nucleotides, about 50 nucleotides, about 75 nucleotides, about 100 nucleotides, about 125 nucleotides, about 150 nucleotides, about 175 nucleotides, about 200 nucleotides, about 225 nucleotides, about 250 nucleotides, about 275 nucleotides, about 300 nucleotides, about 325 nucleotides, about 350 nucleotides, about 375 nucleotides, about 400 nucleotides, about 425 nucleotides, about 450 nucleotides, about 475 nucleotides, about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, about 1000 nucleotides, about 1025 nucleotides, about 1050 nucleotides, about 1075 nucleotides, about 1100 nucleotides, about 1125 nucleotides, about 1150 nucleotides, about 11075 nucleotides, about 1200 nucleotides, about 1325 nucleotides, about 1350 nucleotides, about 1375 nucleotides, about 1400 nucleotides, about 1425 nucleotides, about 1450 nucleotides, about 1475 nucleotides, or about 1500 nucleotides of to SEQ ID NO: 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of about 500 nucleotides, about 525 nucleotides, about 550 nucleotides, about 575 nucleotides, about 600 nucleotides, about 625 nucleotides, about 650 nucleotides, about 675 nucleotides, about 725 nucleotides, about 750 nucleotides, about 775 nucleotides, about 800 nucleotides, about 825 nucleotides, about 850 nucleotides, about 875 nucleotides, about 900 nucleotides, about 925 nucleotides, about 950 nucleotides, about 975 nucleotides, or about 1000 nucleotides of SEQ ID NO: 81.

[0152] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a continuous fragment of from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of any one of SEQ ID NO: 7-9 or 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of any one of SEQ ID NO: 7-9 or 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotidesequence complementary to a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of any one of SEQ ID NO: 7-9 or 81.

[0153] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 9. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 9. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 9. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 9.

[0154] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 9. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 9. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200nucleotides of SEQ ID NO: 9. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 9 or a nucleotide sequence complementary to a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 9.

[0155] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 7. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 7. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 7. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 7.

[0156] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 7. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 7. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 7. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 7 or a nucleotide sequence complementary to a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 7.

[0157] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 8. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 8. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 8. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 8.

[0158] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 8. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 8. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of from about 75nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 8. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 8 or a nucleotide sequence complementary to a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 8.

[0159] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of less than about 250 nucleotides of SEQ ID NO: 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of less than about 200 nucleotides of SEQ ID NO: 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of less than about 150 nucleotides of SEQ ID NO: 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to  a contiguous fragment of less than about 100 nucleotides of SEQ ID NO: 81.

[0160] In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of SEQ ID NO: 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of from about 100 nucleotides to about 200 nucleotides of SEQ ID NO: 81. In some embodiments, the short stuffer comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of from about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 115 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides, about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides of SEQ ID NO: 81. For example, the short stuffer comprises a nucleotide sequence of a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 81 or a nucleotide sequence complementary to a contiguous fragment of about 100 nucleotides, about 150 nucleotides, about 175 nucleotides, or about 200 nucleotides of SEQ ID NO: 81.

[0161] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least one protelomerase binding site.

[0162] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least two protelomerase binding sites.

[0163] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises a heterologous transgene operably linked to one or more regulatory elements.

[0164] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least one adeno-associated virus (AAV) inverted terminal repeat (ITR) sequence.

[0165] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least one ITR sequence and a heterologous transgene.

[0166] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least one protelomerase binding site and at least one ITR sequence.

[0167] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least two ITRs.

[0168] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least one two ITRs and a heterologous transgene

[0169] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises at least two protelomerase binding sites and at least two ITRs.

[0170] In some embodiments of any one of the aspects described herein, the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81, and the nucleic acid comprising the short stuffer further comprises a stop codon (e.g., TAA, TAG or TGA).

[0171] In some embodiments of any one of the aspects described herein, the nucleic acid comprises a short stuffer, at least one ITR sequence and a protelomerase binding site upstream of the ITR sequence, and where the short stuffer is located between the ITR and the protelomerase binding site and comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to lessthan about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81)

[0172] In some embodiments of any one of the aspects described herein, the nucleic acid comprises a short stuffer, at least one ITR sequence and a protelomerase binding site downstream of the ITR sequence, where the short stuffer is located between the ITR and the protelomerase binding site and comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81).

[0173] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer comprises a nucleic acid sequence encoding one or more helper proteins that assist in AAV replication and thereby, rAAV production, and wherein the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81).

[0174] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises a nucleic acid sequence encoding one or more helper proteins that assist rAAV production and at least one protelomerase binding site, and wherein the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81).

[0175] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the short stuffer further comprises a nucleic acid sequence encoding a AAV rep protein and / or a AAV cap protein, and wherein the short stuffer comprises a nucleotide sequence of a contiguous fragment of less than about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to less than about 1500 nucleotides of any one of SEQ ID NO: 7-9 or 81).

[0176] Without wishing to be bound by a theory, stuffer sequences of variable size can be used to increase the size of a short transgene sequence. For example, a stuffer sequence can be added upstream or downstream of a short transgene expression cassette to increase the sequence size between AAV ITRs so that the whole vector genome has a size about the same as the size of a wild- type AAV genome, e.g., a size of about 4.5 kb to about 4.8 kb, such as a size of about 4.7 kb.

[0177] Generally, the short stuffer has a unique sequence, region or domain permitting its identification. In other words, the short stuffer comprises a unique sequence, region or domain that can be detected. Each of the short stuffer sequences as described herein has unique sequence, and therefore the entire short stuffer sequence can be used as the unique identifier. Without wishing tobe bound by a theory, short stuffer sequences can be used in nucleic acid sequences for viral production in order to detect residual DNA in that viral preparation. In non-limiting examples, short stuffer sequences are used in DNA sequences for AAV, e.g, recombinant AAV (rAAV) production, and, lentiviral production in order to detect residual DNA in rAAV preparation, and lentiviral preparation. Intron

[0178] In some embodiments, the nucleic acid comprises a nucleotide sequence having at least 85% identity to SEQ ID NO: 10. The nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is also referred to as an intron herein.

[0179] As used herein, an intron is any nucleotide sequence that resides within a gene but does not remain in the final mature mRNA molecule following transcription of that gene and does not code for amino acids that make up the protein encoded by that gene. There are four main types of intron tRNA introns, self-splicing group I introns, self-splicing group II introns, self-splicing group III introns, and splicesome introns. The frequency of introns within different genomes is observed to vary widely across the spectrum of biological organisms.

[0180] In some embodiments, the nucleic acid comprising the intron comprises one or more nucleotides between positions 45 and 56 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10. For example, the intron comprises from about 10 to about 10,000 nucleotides between positions 45 and 46 of the nucleotide sequence having at least 85% identity, at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10.

[0181] In some embodiments, the nucleic acid comprises from about 2,000 to about 5,000 nucleotides, from about 2,500 to about 5,000 nucleotides, from about 3,000 to about 5,000 nucleotides, from about 3,500 to about 5,000 nucleotides, from about 4,000 to about 5,000 nucleotides, from about 4,500 to about 5,000 nucleotides, from about 2,500 to about 4,500 nucleotides, from about 2,500 to about 4,000 nucleotides, from about 2,500 to about 3,500 nucleotides, from about 2,500 to about 3,000 nucleotides between positions 45 and 46 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10.

[0182] In some embodiments, a larger stuffer described herein is located between positions 45 and 46 of the nucleotide sequence having a having at least 85% identity at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, atleast 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10.

[0183] In some embodiments, the nucleic acid further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter.

[0184] In some embodiments, the nucleotide sequence having a having at least 85% identity at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10 is located in the nucleic acid sequence encoding the AAV Rep protein.

[0185] In some embodiments, the nucleic acid sequence encoding an AAV Rep protein comprises an intron comprising a nucleotide sequence having at least 85% identity, at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10.

[0186] In some embodiments, the nucleotide sequence having at least 85% identity, at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10 is located upstream of a promoter operably linked to the nucleic acid encoding the Rep protein.

[0187] In some embodiments, the nucleotide sequence having at least 85% identity, at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10 is located downstream of the promoter, for example, downstream of a p19 promoter.

[0188] In some embodiments, the promoter is a p19 promoter.

[0189] In some embodiments, the AAV Rep is large Rep (Rep68).

[0190] In some embodiments, the large stuffer does not comprise a nucleotide sequence of mammalian origin. In some embodiments, the large stuffer comprises a nucleotide sequence of non-mammalian origin. In some embodiments, the large stuffer is synthetic.

[0191] In some embodiments, the large stuffer does not comprise more than one of the following: a transcription factor binding site; a regulatory element; an AAV Rep binding site; a donor oracceptor splicing site; an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; a repetitive or palindrome sequence longer than 5 nucleotides; and / or a strong secondary structure; or a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0192] In some embodiments, the 5’-end of the nucleotide sequence having at least 85% identity at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10 is linked to the sequence MAG, where M is A or C, and wherein the 5’-end of the nucleotide sequence having at least 85% identity at least 86% identity, at least 87% identity, at least 88% identity, at least 89% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 100% identity to SEQ ID NO: 10 is linked to A or G. Intron+large spacer (SynINT)

[0193] In some embodiments, the nucleic acid comprising the large stuffer comprises a synthetic intron, wherein the synthetic intron comprises a nucleotide sequence having at least 85% identity to nucleotides 1-45 of SEQ ID NO: 10 and linked at its 3’-end to a synthetic intron of size 2kb, and wherein the 3’-end of the synthetic intron of size at least 2kb is linked to 5’-end of a nucleotide sequence having at least 85% identity to nucleotides 46-167 of SEQ ID NO: 10. The synthetic intron is also referred to as SynINT herein.

[0194] As used herein, an intron is any nucleotide sequence that resides within a gene but does not remain in the final mature mRNA molecule following transcription of that gene and does not code for amino acids that make up the protein encoded by that gene. There are four main types of intron tRNA introns, self-splicing group I introns, self-splicing group II introns, self-splicing group III introns, and splicesome introns. The frequency of introns within different genomes is observed to vary widely across the spectrum of biological organisms.

[0195] In some embodiments, the intron of a size at least 2kb does not comprise a nucleotide sequence of mammalian origin. In other embodiments, the intron of a size at least 2kb does not comprise a nucleotide sequence of non-mammalian origin. In further embodiments, the intron of a size at least 2kb is synthetic.

[0196] In some embodiments, the intron of a size at least 2kb does not comprise more than one of the following: a transcription factor binding site; a regulatory element; an AAV Rep binding site; a donor or acceptor splicing site; an endonuclease cleavage site, optionally where theendonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; a repetitive or palindrome sequence longer than 5 nucleotides; and / or a strong secondary structure; or a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0197] In some embodiments, the intron comprises a GC content of less than about 50%, e.g., less than about 45%, or less than about 40%.

[0198] In some embodiments of any one of the aspects described herein, the synthetic intron comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 11-13. For example, the synthetic intron comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any one of SEQ ID NO: 11-13. In some embodiments, the synthetic intron comprises a nucleotide sequence having 100% identity with any one of SEQ ID NO: 11-13. For example, the synthetic intron consists of a nucleotide sequence having 100% identity with any one of SEQ ID NO: 11-13.

[0199] In some embodiments, the synthetic intron comprises a nucleotide sequence having at least 85% identity with SEQ ID NO: 13. For example, the synthetic intron comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with SEQ ID NO: 13. In some embodiments, the synthetic intron comprises a nucleotide sequence having 100% identity with SEQ ID NO: 13. For example, the synthetic intron consists of a nucleotide sequence having 100% identity with SEQ ID NO: 13.

[0200] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the synthetic intron further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter. It is noted the synthetic intron can be located upstream (e.g., 5’-end) or downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep protein. For example, the synthetic intron can is located upstream (e.g., 5’-end) of the nucleic acid sequence encoding the AAV Rep protein. In another example, the synthetic intron is located downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep protein.

[0201] In some embodiments of any one of the aspects described herein, the synthetic intron is located in the nucleic acid sequence encoding the AAV Rep protein. For example, the synthetic intron is located in an intron in the nucleic acid sequence encoding an AAV Rep protein.

[0202] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the synthetic intron further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter and the synthetic intron is upstream of the promoter. For example, the synthetic intron is upstream of the p19 promoter.

[0203] In some other embodiments of any one of the aspects described herein, the nucleic acidcomprising the synthetic intron further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter and the synthetic intron is downstream of the promoter. For example, the synthetic intron is downstream of the p19 promoter.

[0204] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the synthetic intron further comprises a nucleic acid sequence encoding an AAV Rep protein and a AAV Cap protein operably linked to a promoter. It is noted the synthetic intron can be located upstream (e.g., 5’-end) or downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep and Cap proteins. For example, the synthetic intron is located upstream (e.g., 5’-end) of the nucleic acid sequence encoding the AAV Rep and AAV Cap proteins. In another example, the synthetic intron is located downstream (e.g., 3’-end) of the nucleic acid sequence encoding the AAV Rep and Cap proteins.

[0205] In one embodiment of any of the aspects describes herein SEQ ID NO: 65 comprising any of the SEQ ID NOs: 11-13 sequences, wherein the AAV Cap sequence is replaced with other AAV Cap sequences known in the art. One skilled in the art can use appropriate restriction enzyme present in SEQ ID NO: 11-13 and replace AAV Cap sequence e.g, AAV8 with other AAV cap sequences known in the art. Similarly, SEQ ID NO: 65 can be inserted in the Rep sequence from any AAV serotype known in the art.

[0206] In one embodiment of any of the aspects describes herein SEQ ID NO: 66 comprising any of the SEQ ID NOs: 11-13 sequences, wherein the AAV Cap sequence is replaced with other AAV Cap sequences known in the art. One skilled in the art can use appropriate restriction enzyme present in SEQ ID NO: 11-13 and replace AAV Cap sequence e.g, AAV8 with other AAV cap sequences known in the art. Similarly, SEQ ID NO: 66 can be inserted in the Rep sequence from any AAV serotype known in the art.

[0207] In some embodiments of any one of the aspects described herein, the synthetic intron is located in the nucleic acid sequence encoding the AAV Rep and Cap proteins. For example, the synthetic intron is located in an intron in the nucleic acid sequence encoding the AAV Rep and Cap proteins.

[0208] In some embodiments of any one of the aspects described herein, the nucleic acid comprising the synthetic intron further comprises a nucleic acid sequence encoding an AAV Rep protein and an AAV Cap protein operably linked to a promoter and the synthetic intron is upstream of the promoter. For example, the synthetic intron is upstream of the p19 promoter.

[0209] In some other embodiments of any one of the aspects described herein, the nucleic acid comprising the synthetic intron further comprises a nucleic acid sequence encoding an AAV Rep protein and an Cap protein operably linked to a promoter and the synthetic intron is downstream of the promoter. For example, the synthetic intron is downstream of the p19 promoter.

[0210] Generally, the nucleic acid comprising a synthetic intron has a length greater than 5.5kb, such as greater than 5.6kb, greater than 5.7kb, greater than 5.8kb, greater than 5.9kb, or greater than 6.0kb, In some embodiments, the nucleic acid comprising a synthetic intron has a length greater than 6.5kb, greater than 6.6kb, greater than 6.7kb, greater than 6.8kb, greater than 6.9kb, greater than 7.0kb, greater than 7.1kb, greater than 7.2kb, greater than 7.3kb, greater than 7.4kb, greater than 7.5kb, greater than 7.6kb, greater than 7.7kb, greater than 7.8kb, greater than 7.9kb, or greater than 8.0kb. For example, the nucleic acid comprising a synthetic intron has a length greater than 8.1kb, greater than 8.2kb, greater than 8.3kb, greater than 8.4kb, greater than 8.5kb, greater than 8.6kb, greater than 8.7kb, greater than 8.8kb, greater than 8.9kb, greater than 9.0kb, greater than 9.1kb, greater than 9.2kb, greater than 9.3kb, greater than 9.4kb, greater than 9.5kb, greater than 9.6kb, greater than 9.7kb, greater than 9.8kb, greater than 9.9kb, or greater than 10.0kb. In some embodiments, the nucleic acid comprising a synthetic intron has a length greater than 10.1kb, greater than 10.1kb, greater than 10.2kb, greater than 10.3kb, greater than 10.4kb, greater than 10.5kb, greater than 10.6kb, greater than 10.7kb, greater than 10.8kb, greater than 10.9kb, greater than 11.0kb, greater than 11.1kb, greater than 11.2kb, greater than 11.3kb, greater than 11.4kb, greater than 11.5kb, greater than 11.6kb, greater than 11.7kb, greater than 11.8kb, greater than 11.9kb, or greater than 12.0kb. For example, the nucleic acid comprising a synthetic intron has a length greater than 12.1kb, greater than 12.2kb, greater than 12.3kb, greater than 12.4kb, greater than 12.5kb, greater than 12.6kb, greater than 12.7kb, greater than 12.8kb, greater than 12.9kb, greater than 13.0kb, greater than 13.1kb, greater than 13.2kb, greater than 13.3kb, greater than 13.4kb, or greater than 13.5kb. In some embodiments, the nucleic acid comprising a synthetic intron has a length greater than 13.6kb, greater than 13.7kb, greater than 13.8kb, greater than 13.9kb, greater than 14.0kb, greater than 14.1kb, greater than 14.2kb, greater than 14.3kb, greater than 14.4kb, greater than 14.5kb, greater than 14.6kb, greater than 14.7kb, greater than 14.8kb, greater than 14.9kb, or greater than 15.0kb or more.

[0211] In some embodiments of any one of the aspects described herein, the synthetic intron does not comprise a nucleotide sequence of mammalian origin. In some embodiments of any one of the aspects described herein, the synthetic intron comprises a nucleotide sequence of non-mammalian origin. In some embodiments, the synthetic intron is synthetic.

[0212] Generally, synthetic intron intron is located downstream of the sequence MAG, where M is A or C. Thus, in some embodiments of any one of the aspects described herein, the 5’-end of the synthetic intron is linked to the sequence MAG, where M is A or C. In some embodiments, the 3’-end of the synthetic intron is linked to A or G. For example, the 5’-end of the synthetic intron is linked to the sequence MAG, where M is A or C, and the 3’-end of the synthetic intron is linked to A or G.

[0213] In some embodiments of any one of the aspects described herein, the synthetic intron does not comprise more than one (e.g., 1, 2, 3, 4, 5, 6, 7 or 8) of the following: a transcription factor binding site; a regulatory element; an AAV Rep binding site; a donor or acceptor splicing site; an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; a repetitive or palindrome sequence longer than 5 nucleotides; and / or a strong secondary structure; or a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides. Endonuclease cleavage sites

[0214] The short spacer can be flanked by one or more endonuclease cleavage sites. For example, the short spacer can comprise an endonuclease cleavage site at its 5’-end. In another example, the short spacer can comprise an endonuclease cleavage site at its 3’-end. In yet another example, the short spacer comprises an endonuclease cleavage site at its 5’-end and an endonuclease cleavage site at its 3’-end. Exemplary endonucleases include, but are not limited to, BglII, SbfI, PacI, SwaI, ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, and AflII.

[0215] In some embodiments, the short spacer comprises at its 5’-end an endonuclease cleavage site for an endonuclease selected from the group consisting of BglII, SbfI, PacI, SwaI, ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, and AflII. For example, the short spacer comprises at its 5’-end an endonuclease cleavage site for BglII.

[0216] In some embodiments, the short spacer comprises at its 3’-end an endonuclease cleavage site for an endonuclease selected from the group consisting of BglII, SbfI, PacI, SwaI, ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, and AflII. For example, the short spacer comprises at its 3’-end an endonuclease cleavage site for SbfI, PacI, or SwaI.

[0217] In some embodiments, the short spacer comprises an endonuclease cleavage site at its 5’- end and an endonuclease cleavage site at its 3’-end, where each endonuclease cleavage site is for an endonuclease selected independently from the group consisting of BglII, SbfI, PacI, SwaI, ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, and AflII. For example, the short spacer comprises at its 5’-end an endonuclease cleavage site for BglII and at its 3’-end an endonuclease cleavage site for SbfI, PacI, or SwaI. No end DNA (neDNA): DNA with closed linear shaped structure is described herein as No end DNA or neDNA. neDNA is alternatively termed as close ended linear duplexed DNA (clDNA, or, celDNA). Closed ended linear duplexed DNA molecules typically comprise covalently closed ends also described as hairpin loops, where base-pairing between complementary DNA strands is not present. The hairpin loops join the ends of complementary DNA strands. Structures of this type typically form at the telomericends of chromosomes in order to protect against loss or damage of chromosomal DNA by sequestering the terminal nucleotides in a closed structure. In examples of closed linear DNA molecules described herein, hairpin loops flank complementary base-paired DNA strands, forming a closed linear (cl) DNA shaped structure. In some examples, neDNA further comprises at least one, e.g., two protelomerase binding sites. Non limiting examples of closed linear duplexed DNA, or, no-end DNA (neDNA) include doggybone DNA (dbDNA), and / or dumbbell shaped DNA.

[0165]

[0205] Alternate methods of generating covalently closed ended linear duplexed DNA that lack bacterial sequences are known in the art e.g., by formation of mini-circle DNA from plasmids (e.g., as described in U.S. Patent 8,828,726, and U.S. Patent 7,897,380, the contents of each of which are incorporated by reference in their entirety). Protelomerase binding sites

[0218] One method of generating a closed ended linear duplex nucleic acids is by incorporation of protelomerase binding sites (also referred to as protelomerase target site) in a precursor molecule such that the protelomerase binding sites flank the nucleic acid of interest. The nucleic acid of interest can be exposed to protelomerase to thereby cleave and ligate the DNA at the site. Non- limiting examples of protelomerase binding sites are e.g., described in US 9,109,250; US 6,451,563; Nucleic Acids Res. 2015 Oct 15; 43(18): e120; US 9499847; 15 / 508,766; PCT / GB2017 / 052413; and Antisense & nucleic acid drug development 11:149–153 (2001); herein incorporated by reference in their entirety. Generally, closed linear DNA comprises half of protelomerase binding site.

[0219] In some embodiments, a protelomerase binding site from a prokaryotic system can be used. In lysogenic bacteria, the bacteriophage N15 exists as a linear extrachromosomal DNA with covalently closed ends (see Rybchin VN, Svarchevsky AN (1999) The plasmid prophage N15: a linear DNA with covalently closed ends. Mol Microbiol 33:895–903). This DNA arises by a cleaving-joining reaction, which is exerted by a single enzyme, a protelomerase, for example, TelN (prokaryotic telomerase) [Deneke J, Ziegelin G, Lurz R, Lanka E (2000) The protelomerase of temperate Escherichia coli phage N15 has cleaving-joining activity. Proc Natl Acad Sci U S A 97:7721–7726]. A protelomerase such as TelN recognizes a target sequence in double-stranded DNA. The target site is an imperfect palindromic structure termed telRL, which is formed by the two halves telR and telL, corresponding to the covalently closed ends of the linear prophage. The enzyme cleaves both DNA strands and joins the resulting ends to form covalently closed hairpin structures. The resulting DNA molecule has two hairpin loops. TelN is able to linearize a recombinant plasmid harboring the telRL site [Deneke J, et al., (2000). Proc Natl Acad Sci U S A 97:7721–7726]. Therefore, one can employ this enzyme on a plasmid DNA for expression in higher organisms.

[0220] The TelN / telRL system can be used to produce the closed linear DNA fragments either by linearizing a parental plasmid containing one telRL site or by excising the DNA fragment, or non-viral vector fragment, comprising a promoter, the gene of interest, a polyadenylation signal from the parental plasmid with two flanking ITRs, further having two telRL sites flanking the respective segment. The resulting linear covalently closed DNA molecules are functional in vivo.

[0221] The host cell is designed to encode at least one recombinase. The host cell may also be designed to encode two or multiple recombinases. The term “recombinase” refers to an enzyme that catalyzes DNA exchange at a specific target site, for example, a palindromic sequence, by excision / insertion, inversion, translocation and exchange. Examples of suitable recombinases for use in the present system include, but are not limited to, TelN, Tel, Tel (gp26 K02 phage) Cre, Flp, phiC31, Int and other lambdoid phage integrases, e.g., phi 80, HK022 and HP1 recombinases. The target sequences for each of these recombinases are, respectively: the telRL site: TATCAGCACACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGTGCTGATA (SEQ ID NO: 14); the pal site: ACCTATTTCAGCATACTACGCGCGTAGTATGCTGAAATAGGT (SEQ ID NO: 15); the φK02 telRL site: CCATTATACGCGCGTATAATGG (SEQ ID NO: 16); the loxP site: TAACTTCGTATAGCATACATTATACGAAGTTAT (SEQ ID NO: 17); the FRT site: GAAGTTCCTATTCTCTAGAAAGTATAGGAACTTC (SEQ ID NO: 18); the phiC31 attP site: CCCAGGTCAGAAGCGGTTTTCGGGAGTAGTGCCCCAACTGGGGTAACCTTTGA GTTCTCTCAGTT GGGGGCGTAGGGTCGCCGACAYGACACAAGGGGTT (SEQ ID NO: 19); and the λ attP site: TGATAGTGACCTGTTCGTTGCAACACATTGATGAGCAATGCTTTTTTATAATGC CAACTTTGTACAA AAAAGCTGAACGAGAAACGTAAAATGATATAAA (SEQ ID NO: 20).

[0222] In some embodiments, the 56 bp protelomerase (Tel N) binding site as described in Zhang et al., Molecular Therapy, Methods and Clinical development, Volume 32, Issue 1, 101206, March 14, 2024 is TATCAGCACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGCTGATA (SEQ ID NO: 95)

[0223] Expression of the recombinase is under the control of any regulated or inducible promoter, i.e., a promoter which is activated under a particular physical or chemical condition or stimulus. Examples of suitable promoters include thermally-regulated promoters such as the λpL promoter, the IPTG regulated lac promoter, the glucose regulated ara promoter, the T7 polymerase regulated promoter, cold-shock inducible cspA promoter, pH inducible promoters, or combinations thereof, such as tac (T7 and lac) dual regulated promoter.

[0224] Alternate methods of generating covalently closed ended linear duplexed DNA that lack bacterial sequences are known in the art e.g., by formation of mini-circle DNA from plasmids (e.g., as described in U.S. Patent 8,828,726, and U.S. Patent 7,897,380, the contents of each of which are incorporated by reference in their entirety). For example, one method of cell-free synthesis combines the use of two enzymes – Phi29 DNA polymerase and a protelomerase, and generates high fidelity, covalently closed, linear DNA constructs. The constructs contain no antibiotic resistance markers, and therefore eliminate the packaging of these sequences. The process can amplify AAV genome DNA in a 2-week process at commercial scale and maintain the ITR sequences required for virus production.

[0225] Phi29 DNA polymerase is used to amplify double-stranded DNA by rolling circle amplification, and a protelomerase to generate covalently closed ended linear duplexed DNA, which coupled with a streamlined purification process, results in a pure DNA product containing just the sequence of interest. Phi29 DNA polymerase has high fidelity (1×10⁶–1×10⁷) and high processivity (approximately 70 kbp). These features make this polymerase particularly suitable for the large-scale production of GMP DNA. Protelomerases (also known as telomere resolvases) catalyze the formation of covalently closed hairpin ends on linear DNA and have been identified in some phages, bacterial plasmids and bacterial chromosomes. A pair of protelomerases recognizes inverted palindromic DNA recognition sequences and catalyzes strand breakage, strand exchange and DNA ligation to generate closed linear hairpin ends. The formation of these closed ended structures makes the DNA resistant to exonuclease activity, allowing for simple purification and can improve stability and duration of expression.

[0226] In one embodiment, the DNA construct comprises a protelomerase binding site and the covalently closed ends are formed by protelomerase enzyme activity (e.g., in vitro). Protelomerase binding sites and corresponding protelomerases for use in the invention are provided in U.S. Patent No. 9,499,847, the contents of which are incorporated herein by reference in their entirety. A protelomerase target sequence as used in the invention preferably comprises a double stranded palindromic (perfect inverted repeat) sequence of at least 14 base pairs in length. Preferred perfect inverted repeat sequences include the sequences of SEQ ID NOs: 22 to 26 and variants thereof. SEQ ID NO: 21 (NCATNNTANNCGNNTANNATGN) is a 22 base consensus sequence for amesophilic bacteriophage perfect inverted repeat. Base pairs of the perfect inverted repeat are conserved at certain positions between different bacteriophages, while flexibility in sequence is possible at other positions. Thus, SEQ ID NO: 21 is a minimum consensus sequence for a perfect inverted repeat sequence for use with a bacteriophage protelomerase in the process of the present invention.

[0227] Within the consensus defined by SEQ ID NO: 21, SEQ ID NO: 22 (CCATTATACGCGCGTATAATGG) is a perfect inverted repeat sequence for use with E. coli phage N15, and Klebsiella phage Phi KO2 protelomerases. Also, within the consensus defined by SEQ ID NO: 21 and / or SEQ ID NOs: 23 to 26 (SEQ ID NO: 23 (GCATACTACGCGCGTAGTATGC), SEQ ID NO: 24 (CCATACTATACGTATAGTATGG), SEQ ID NO: 25 (GCATACTATACGTATAGTATGC)), are particularly preferred perfect inverted repeat sequences for use respectively with protelomerases from Yersinia phage PY54, Halomonas phage phiHAP-1, and Vibrio phage VP882. SEQ ID NO: 26 (ATTATATATATAAT) is a particularly preferred perfect inverted repeat sequence for use with a Borrelia burgdorferi protelomerase. This perfect inverted repeat sequence is from a linear covalently closed plasmid, lpB31.16 comprised in Borrelia burgdorferi. This 14 base sequence is shorter than the 22 bp consensus perfect inverted repeat for bacteriophages (SEQ ID NO: 21), indicating that bacterial protelomerases may differ in specific target sequence requirements to bacteriophage protelomerases. However, all protelomerase target sequences share the common structural motif of a perfect inverted repeat.

[0228] The perfect inverted repeat sequence may be greater than 22 bp in length depending on the requirements of the specific protelomerase used in the process as described herein. Thus, in some embodiments, the perfect inverted repeat may be at least 30, at least 40, at least 60, at least 80 or at least 100 base pairs in length. Examples of such perfect inverted repeat sequences include SEQ ID NOs: 27 to 29 and variants thereof. SEQ ID NO: 27 (GGCATAC TATACGTATAGTATGCC); SEQ ID NO: 28 (ACCTATTTCAGCATACTACGCGCG- TAGTATGCTGAAATAGGT); SEQ ID NO: 29 (CCTATATTGGGCCACCTATGTATG- CACAGTTCGCCCATACTATACGTATAGTATGGGCGAACTGTGCATACATAGGTGGCC CAATATAGG). SEQ ID NOs: 27 to 29 and variants thereof are particularly preferred for use respectively with protelomerases from Vibrio phage VP882, Yersinia phage PY54 and Halomonas phage phi HAP-1.

[0229] The perfect inverted repeat may be flanked by additional inverted repeat sequences. The flanking inverted repeats may be perfect or imperfect repeats i.e. may be completely symmetrical or partially symmetrical. The flanking inverted repeats may be contiguous with or non-contiguous with the central palindrome. The protelomerase target sequence may comprise an imperfectinverted repeat sequence which comprises a perfect inverted repeat sequence of at least 14 base pairs in length. An example is SEQ ID NO: 34. The imperfect inverted repeat sequence may comprise a perfect inverted repeat sequence of at least 22 base pairs in length. An example is SEQ ID NO: 30.

[0230] In certain embodiments, the protelomerase target sequences comprise the sequences of SEQ ID NOs: 30 to 34 or variants thereof. SEQ ID NO: 30: (TATCAGCACACAATTGCCCATTATACG-CGCGTATAATGGACTATTG TGTGCTGATA); SEQ ID NO: 31 (ATGCGCGCATCCATTATACGCGCGTATAATGGCGATAATACA); SEQ ID NO: 32 (TAGTCACCTATTTCAGCATACTACGCGCGTAGTATGCTGAAATAGG TTACTG); SEQ ID NO: 33: (GGGATCCCGTTCCATACATACATGTATCCATGTGGCATACTATACG TATAGTATGCCGATGTTACATATGGTATCATTCGGGATCCCGTT); SEQ ID NO: 34 (TACTAAATAAATATTATATATATAATTTTTTATTAGTA).

[0231] The sequences of SEQ ID NOs: 30 to 34 comprise perfect inverted repeat sequences as described above, and additionally comprise flanking sequences from the relevant organisms. A protelomerase target sequence comprising the sequence of SEQ ID NO: 30 or a variant thereof is preferred for use in combination with E. coli N15 TelN protelomerase and variants thereof. A protelomerase target sequence comprising the sequence of SEQ ID NO: 31 or a variant thereof is preferred for use in combination with Klebsiella phage Phi K02 protelomerase and variants thereof. A protelomerase target sequence comprising the sequence of SEQ ID NO: 32 or a variant thereof is preferred for use in combination with Yersinia phage PY54 protelomerase and variants thereof. A protelomerase target sequence comprising the sequence of SEQ ID NO: 33 or a variant thereof is preferred for use in combination with Vibrio phage VP882 protelomerase and variants thereof. A protelomerase target sequence comprising the sequence of SEQ ID NO: 34 or a variant thereof is preferred for use in combination with a Borrelia burgdorferi protelomerase.

[0232] Variants of any of the palindrome or protelomerase target sequences described above include homologues or mutants thereof. Mutants include truncations, substitutions or deletions with respect to the native sequence. A variant sequence is any sequence whose presence in the DNA template allows for its conversion into a closed ended linear duplexed DNA by the enzymatic activity of protelomerase. This can readily be determined by use of an appropriate assay for the formation of closed linear DNA. Any suitable assay described in the art may be used. An example of a suitable assay is described in Deneke et al., PNAS (2000) 97, 7721-7726. In certain embodiments, the variant allows for protelomerase binding and activity that is comparable to that observed with the native sequence. Examples of preferred variants of palindrome sequences described herein include truncated palindrome sequences that preserve the perfect repeat structureand remain capable of allowing for formation of closed linear DNA. However, variant protelomerase target sequences may be modified such that they no longer preserve a perfect palindrome, provided that they are able to act as substrates for protelomerase activity.

[0233] It should be understood that the skilled person would readily be able to identify suitable protelomerase target sequences for use in the invention on the basis of the structural principles outlined above. Candidate protelomerase target sequences can be screened for their ability to promote formation of closed linear DNA using the assays described above.

[0234] The covalently closed vectors described herein may be generated in vitro or in vivo. The vectors are covalently closed linear double stranded vectors capable of expressing transgene in a target cell. One example of an in vitro process for the production of a closed linear expression cassette DNA, e.g., containing the ITRs described herein, comprises a) contacting a DNA template comprising at least one expression cassette flanked on either side by a protelomerase target sequence with at least one DNA polymerase in the presence of one or more primers under conditions promoting amplification of said template; and b) contacting amplified DNA produced in a) with at least one, protelomerase under conditions promoting formation of a closed linear expression cassette DNA. The closed linear expression cassette DNA product may comprise, consist or consist essentially of a eukaryotic promoter operably linked to a coding sequence of interest, and optionally a eukaryotic transcription termination sequence. The closed linear expression cassette DNA product may additionally lack one or more bacterial or vector sequences, typically selected from the group consisting of: (i) bacterial origins of replication; (ii) bacterial selection markers (typically antibiotic resistance genes) and (iii) unmethylated CpG motifs.

[0235] As outlined above, any DNA template comprising at least one protelomerase target sequence may be amplified according to the process as described herein. Thus, although production of therapeutic DNA molecules, e.g., for DNA vaccines or other therapeutic proteins and nucleic acid is preferred, the process as described herein may be used to produce any type of closed linear DNA. The DNA template may be a double stranded (ds) or a single stranded (ss) DNA. A double stranded DNA template may be an open circular double stranded DNA, a closed circular double stranded DNA, an open linear double stranded DNA or a closed linear double stranded DNA. Preferably, the template is a closed circular double stranded DNA. Closed circular dsDNA templates are particularly preferred for use with RCA (rolling circle amplification) DNA polymerases. A circular dsDNA template may be in the form of a plasmid or other vector typically used to house a gene for bacterial propagation. Thus, the process as described herein may be used to amplify any commercially available plasmid or other vector, such as a commercially available DNA medicine, and then convert the amplified vector DNA into closed linear DNA.

[0236] An open circular dsDNA may be used as a template where the DNA polymerase is a strand displacement polymerase which can initiate amplification from at a nicked DNA strand. In this embodiment, the template may be previously incubated with one or more enzymes which nick a DNA strand in the template at one or more sites. A closed linear dsDNA may also be used as a template. The closed linear dsDNA template (starting material) may be identical to the closed linear DNA product. Where a closed linear DNA is used as a template, it may be incubated under denaturing conditions to form a single stranded circular DNA before or during conditions promoting amplification of the template DNA. In one embodiment, the closed ended linear duplex DNA is produced in eukaryotic cells for example insect cells as described in PCT publications WO 2019032102 and WO 2019169233, which are incorporated herein by reference in their entireties. In one embodiment, the DNA is not produced in eukaryotic cells and DNA lacks eukaryotic sequences. In one embodiment, the closed ended liner duplex DNA vectors are produced as described in PCT publication WO 2019143885, which is incorporated herein by reference in its entirety.

[0237] In some embodiments of any one of the aspects described herein, the nucleic acid described herein, e.g., the nucleic acid comprising the short stuffer further comprises a protelomerase binding site. In some embodiments, the protelomerase binding site is linked, e.g., operably liked to the 5’-end of a stuffer sequence, e.g., 5’-end of theshort stuffer. In some other embodiments, the protelomerase binding site is linked, e.g., operably liked to the 3’-end of a stuffer sequence, e.g., 3’-end of the short stuffer.

[0238] In some embodiments, the nucleic acid comprises a first protelomerase binding site linked, e.g., operably liked to the 5’-end of the short stuffer, and a second protelomerase binding site linked, e.g., operably liked to the 3’-end of the short stuffer. For example, the nucleic acid comprises a first partial protelomerase binding site, which only comprises either telR or telL, linked, e.g., operably linked to the 5’-end of the short stuffer, and a second partial protelomerase binding site linked, e.g., operably linked to the 3’-end of the short stuffer. The first partial protelomerase binding site and the second partial protelomerase binding site together form a functional protelomerase binding site.

[0239] In some embodiments, the short stuffer is located downstream of the protelomerase binding site. In other embodiments, the short stuffer is located upstream of the protelomerase binding site. In a preferred embodiment, the nucleic acid comprises two protelomerase binding sites and the short stuffer is located between the two protelomerase binding sites.

[0240] In some embodiments, the short stuffer is located between the first protelomerase binding site and an ITR (e.g., ITR1).

[0241] In some embodiments, the nucleic acid comprises a second protelomerase binding site(telLR), and wherein the protelomerase binding site is located 3’ of the short stuffer.

[0242] In some embodiments, the short stuffer is located between the second protelomerase binding site and an ITR (e.g., ITR2). Transgene

[0243] Embodiments of the various aspects described herein include a transgene, e.g., a heterologous transgene operably linked to one or more regulatory elements. For example, the nucleic acid comprising the short stuffer sequence further comprises a heterologous transgene operably linked to one or more regulatory elements.

[0244] A “transgene” is used herein to refer to a polynucleotide or a nucleic acid that is intended or has been introduced into a cell or organism. Transgenes include any nucleic acid, such as a gene that encodes a polypeptide or protein. Suitable transgenes, for example, for use in gene therapy are well known to those of skill in the art. Exemplary transgenes include, but are not limited to, those described in U.S. Pat. Nos. 6,547,099; 6,506,559; and 4,766,072; Published U.S. Application No. 20020006664; 20030153519; 20030139363; and published PCT applications of WO 01 / 68836 and WO 03 / 010180, and e.g., miRNAs and other transgenes of WO2017 / 152149; each of which are hereby incorporated herein by reference in their entirety.

[0245] The composition of the transgene sequence will depend upon the use to which the resulting vector will be put. Suitable transgenes may be readily selected by one of skill in the art. The selection of the transgene is not considered to be a limitation of this invention. In some embodiments, the transgene is a nucleic acid sequence encoding a product which is useful in biology and medicine, such as proteins, peptides, RNA, enzymes, dominant negative mutants, or catalytic RNAs. Desirable RNA molecules include tRNA, dsRNA, ribosomal RNA, catalytic RNAs, siRNA, small hairpin RNA, trans-splicing RNA, and antisense RNAs. One example of a useful RNA sequence is a sequence which inhibits or extinguishes expression of a targeted nucleic acid sequence in the treated animal. Typically, suitable target sequences include oncologic targets and viral diseases. See, for examples of such targets the oncologic targets and viruses identified below in the section relating to immunogens.

[0246] The transgene can be used to correct or ameliorate gene deficiencies, which may include deficiencies in which normal genes are expressed at less than normal levels or deficiencies in which the functional gene product is not expressed. A preferred type of transgene sequence encodes a therapeutic protein or polypeptide which is expressed in a host cell. In some embodiments, the transgene is a heterologous protein, and this heterologous protein is a therapeutic protein. Exemplary therapeutic proteins include, but are not limited to, hemophilia related clotting proteins, such as Factor VIII, Factor IX, Factor X; glial cell derived neurotrophic factor (GDNF); Acid alphaglucosidase (GAA); cytochrome P450 family 46 subfamily A member 1 (CYP46A1); Protein Phosphatase 1 Inhibitor 1-constitutively active form (I1c); fukutin-related protein (FKRP); Lysosome-associated membrane protein 2B (LAMP2B); blood factors, such as β-globin, hemoglobin, tissue plasminogen activator, and coagulation factors; colony stimulating factors (CSF); interleukins, such as IL-l, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-9, etc.; growth factors, such as keratinocyte growth factor (KGF), stem cell factor (SCF), fibroblast growth factor (FGF, such as basic FGF and acidic FGF), hepatocyte growth factor (HGF), insulin-like growth factors (IGFs), bone morphogenetic protein (BMP), epidermal growth factor (EGF), growth differentiation factor-9 (GDF-9), hepatoma derived growth factor (HDGF), myostatin (GDF-8), nerve growth factor (NGF), neurotrophins, platelet-derived growth factor (PDGF), thrombopoietin (TPO), transforming growth factor alpha (TGF-α), transforming growth factor beta (TGF-β), and the like; soluble receptors, such as soluble TNF-α receptors, soluble VEGF receptors, soluble interleukin receptors (e.g., soluble IL-l receptors and soluble type II IL-l receptors), soluble g / d T cell receptors, ligand-binding fragments of a soluble receptor, and the like; enzymes, such as a- glucosidase, imiglucarase, b-glucocerebrosidase, and alglucerase; enzyme activators, such as tissue plasminogen activator; chemokines, such as 1P-10, monokine induced by interferon-gamma (Mig), Groa / IL-8, RANTES, MIP-la, MIR-1b., MCP-l, PF-4, and the like; angiogenic agents, such as vascular endothelial growth factors (VEGFs, e.g., VEGF121, VEGF165, VEGF-C, VEGF-2), glioma-derived growth factor, angiogenin, angiogenin-2; and the like; anti-angiogenic agents, such as a soluble VEGF receptor; protein vaccine; neuroactive peptides, such as nerve growth factor (NGF), bradykinin, cholecystokinin, gastin, secretin, oxytocin, gonadotropin-releasing hormone, beta-endorphin, enkephalin, substance P, somatostatin, prolactin, galanin, growth hormone- releasing hormone, bombesin, dynorphin, warfarin, neurotensin, motilin, thyrotropin, neuropeptide Y, luteinizing hormone, calcitonin, insulin, glucagons, vasopressin, angiotensin II, thyrotropin- releasing hormone, vasoactive intestinal peptide, a sleep peptide, and the like; thrombolytic agents; atrial natriuretic peptide; relaxin; glial fibrillary acidic protein; follicle stimulating hormone (FSH); human alpha- 1 antitrypsin; leukemia inhibitory factor (LIF); tissue factors, luteinizing hormone; macrophage activating factors; tumor necrosis factor (TNF); neutrophil chemotactic factor (NCF); tissue inhibitors of metalloproteinases; vasoactive intestinal peptide; angiogenin; angiotropin; fibrin; hirudin; IF-l receptor antagonists; and the like. Some other non-limiting examples of protein of interest include ciliary neurotrophic factor (CNTF); brain-derived neurotrophic factor (BDNF); amyloid precursor proteins including sAPPα; neurotrophins 3 and 4 / 5 (NT-3 and 4 / 5); aromatic amino acid decarboxylase (AADC); dystrophin family genes e.g, minidystrophin, microdystrophin, nanodystrophin or variant thereof, ; lysosomal acid lipase; phenylalanine hydroxylase (PAH); glycogen storage disease-related enzymes, such as glucose-6- phosphatase, acid maltase, glycogendebranching enzyme, muscle glycogen phosphorylase, liver glycogen phosphorylase, muscle phosphofructokinase, phosphorylase kinase (e.g., PHKA2), glucose transporter (e.g., GFUT2), aldolase A, b-enolase, and glycogen synthase; lysosomal enzymes (e.g., beta-N- acetylhexosaminidase A); and any variants thereof; nucleases including Cas, e.g, spcas9, sacas9, or variant therof;inhibitory RNAs such as miRNAs, shRNAs, siRNAs.

[0247] In some embodiments, the transgene sequence includes a reporter sequence, which upon expression produces a detectable signal. Such reporter sequences include, without limitation, DNA sequences encoding b-lactamase, b-galactosidase (LacZ), alkaline phosphatase, thymidine kinase, green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), luciferase, membrane bound proteins including, for example, CD2, CD4, CD8, the influenza hemagglutinin protein, and others well known in the art, to which high affinity antibodies directed thereto exist or can be produced by conventional means, and fusion proteins comprising a membrane bound protein appropriately fused to an antigen tag domain from, among others, hemagglutinin or Myc.

[0248] These coding sequences, when associated with regulatory elements which drive their expression, provide signals detectable by conventional means, including enzymatic, radiographic, colorimetric, fluorescence or other spectrographic assays, fluorescent activating cell sorting assays and immunological assays, including enzyme linked immunosorbent assay (ELISA), radioimmunoassay (RIA) and immunohistochemistry. For example, where the marker sequence is the LacZ gene, the presence of the vector carrying the signal is detected by assays for beta- galactosidase activity. Where the transgene is green fluorescent protein or luciferase, the vector carrying the signal may be measured visually by color or light production in a luminometer.

[0249] In some embodiments, the transgene is flanked by ITRs. For example, the transgene is operably linked to an ITR on its 5’-end. In another example, the transgene is operably linked to an ITR on its 3’-end. In yet another example, the transgene is operably linked to an ITR on its 5’-end and is operably linked to an ITR on its 3’-end.

[0250] Regulatory elements are regions of non-coding DNA which regulate the transcription of genes. There are two types of regulatory elements: cis-regulatory elements and trans-regulatory elements. Cis-regulatory elements are regions of non-coding DNA which regulate the transcription of neighboring genes. They are found close to the genes that they regulate and typically regulate gene transcription by binding to transcription factors. Cis-regulatory elements are usually between 100 and 1000 base pairs in length. Trans-regulatory elements are DNA sequences encoding upstream regulators which may modify or regulate the expression of distant genes. Unlike the cis- regulatory elements that work through an intramolecular interaction between different parts of the same molecule: (1) a gene; and (2) an adjacent regulatory element for that gene in the same DNAmolecule, trans-regulatory elements work through (1) a transcribed and translated transcription factor protein derived from the trans-regulatory element; and a (2) DNA regulatory element that is adjacent to the regulated gene.

[0251] In some embodiments, the short stuffer is located upstream of the heterologous transgene. In other embodiments, the short stuffer is located downstream of the heterologous transgene.

[0252] In some embodiments, at least one ITR is located between the short stuffer and the heterologous transgene. For example, the short stuffer is located upstream of the heterologous transgene and at least one ITR is located between the short stuffer and the heterologous transgene. In another example, the short stuffer is located downstream of the heterologous transgene and at least one ITR is located between the short stuffer and the heterologous transgene. In some embodiments, the heterologous transgene is flanked on each end by an ITR and one of the ITR is located between the short stuffer and the heterologous transgene. ITRs

[0253] In some embodiments, a nucleic acid described herein comprises at least one adeno- associated virus (AAV) inverted terminal repeat (ITR) sequence. For example, the nucleic acid comprising the short stuffer further comprises at least one ITR sequence.

[0254] A “rAAV vector” or “rAAV genome” is an AAV genome (i.e., vDNA) that comprises one or more heterologous nucleotide sequences. rAAV vectors generally require only the 145 base terminal repeat(s) (TR(s)) in cis to generate virus. All other viral sequences are dispensable and may be supplied in trans (Muzyczka, (1992) Curr. Topics Microbiol. Immunol.158:97). Typically, the rAAV vector genome will only retain the minimal TR sequence(s) so as to maximize the size of the transgene that can be efficiently packaged by the vector. The structural and non- structural protein coding sequences may be provided in trans (e.g., from a vector, such as a plasmid, or by stably integrating the sequences into a packaging cell). The rAAV vector genome comprises at least one TR sequence (e.g., AAV TR sequence, synthetic, or other parvovirus TR sequence), optionally two TRs (e.g., two AAV TRs), which typically will be at the 5' and 3' ends of the heterologous nucleotide sequence(s), but need not be contiguous thereto. The TRs can be the same or different from each other.

[0255] An inverted terminal repeat (ITR) is referred to as a sequence of nucleotide found at each end of some linear replicons, such as transposons, viruses, and plasmids. These sequences have the ability to form a hairpin structure, which contributes to self-priming that allows primase- independent synthesis of the second DNA strand. The ITRs are also important for both integration of the AAV DNA into the host cell genome and rescue from it, as well as for efficient encapsidation of the AAV DNA combined with generation of a fully-assembled, deoxyribonuclease-resistantAAV particles.The inverted terminal repeat (ITR) sequences may or may not be of equal length. ITR sequence of equal length are more efficient in multiplication of the AAV genome.

[0256] As used herein, the term “inverted terminal repeat” or “ITR” includes any wild type viral terminal repeat and synthetic sequences that form hairpin structures and can function as an inverted terminal repeat (ITR). One non-limiting example of synthetic ITR is 165 bp “double-D sequence” as described in U.S. Pat. No.5,478,745 Samulski et al. The capsid structures of autonomous parvoviruses and AAV are described in more detail in BERNARD N. FIELDS et al., VIROLOGY, volume 2, chapters 69 & 70 (4th ed., Lippincott-Raven Publishers). See also, description of the crystal structure of AAV2 (Xie et al., (2002) Proc. Nat. Acad. Sci.99: 10405- 10), AAV4 (Padron et al., (2005) I. Virol.79: 5047-58), AAVS (Walters et al., (2004) I. Virol.78: 3361-71) and CPV (Xie et al., (1996) I. Mol. Biol.6:497-520 and Tsao et al., (1991) Science 251: 1456-64). An “AAV inverted terminal repeat” or “AAV ITR” may be from any AAV, including but not limited to serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, AAV4A, or any other AAV now known or later discovered. The AAV ITRS need not have a wild-type terminal repeat sequence (e.g., a wild-type sequence may be altered by insertion, deletion, truncation or missense mutations), as long as at least one of the ITRs mediates the desired functions, a functional ITR, e.g., replication, virus packaging, integration, and / or provirus rescue, and the like. One of skill in the art understands to choose a Rep protein that is functional for replication of the functional ITR. In some embodiments, the ITR is synthetic or, mutant ITR or, restrictive ITR for example as described in WO2014143932, US 9447433; WO2011088081, US 9169494; WO 2019143950, all of which are incorporated in entirety by reference herein. ITRs used in the nucleic acid comprising stuffer sequence of the invention can be of varying length, e.g, 145 bp or, less. In certain embodiments, one or, more ITRs is 145 bp long. In other embodiments, at least one ITR is 130 bp long. In another embodiment, at least one ITR is 119 bp long.

[0257] In some embodiment, a nucleic acid described herein comprises at least two ITRs. For example, the nucleic acid comprises a heterologous transgene flanked by an ITR at each end. When two or more ITRs are present, they can be same or different. In some preferred embodiments, the ITR sequence is the 130bp ITRs both in flop orientation (SEQ ID NO: 71), as described in Samulski RJ, Chang LS, Shenk T. A recombinant plasmid from which an infectious adeno- associated virus genome can be excised in vitro and its use to study viral replication. J Virol.1987Oct;61(10):3096-101. doi: 10.1128 / JVI.61.10.3096-3101.1987. PMID: 3041032; PMCID: PMC255885, which is incorporated by reference herein in its entirety.

[0258] In some embodiments, the short stuffer is upstream of the at least one ITR sequence. In other embodiments, the short stuffer is downstream of the at least one ITR sequence.

[0259] In some embodiments, the at least one ITR sequence is located between the short stuffer and the heterologous transgene. In some embodiments, the nucleic acid further comprises at least one protelomerase binding site and the short stuffer is located between the at least one protelomerase binding site and the at least ITR.

[0260] In some embodiments, the nucleic acid comprises at least two ITRs and the short stuffer located outside the two ITRs. In some embodiments, the heterologous transgene is located between the two ITRs.

[0261] In some embodiments, the heterologous transgene is located between the two ITRs, and one of the ITRs is located between the short stuffer and the heterologous transgene.

[0262] In some embodiments, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the heterologous polynucleotide sequence is located between the first and second ITR sequences, and wherein the short stuffer is upstream of the first ITR sequence.

[0263] In some embodiments, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the heterologous polynucleotide sequence is located between first and second ITR sequences, wherein the short stuffer is upstream of the first ITR sequenc, and wherein the nucleic acid does not comprise a short stuffer downstream of the second ITR sequence.

[0264] In some embodiments, the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the short stuffer is located upstream of the 5’-end of the first and second ITRs. Helper proteins

[0265] In some embodiments of any of the aspects, the nucleic acid comprises a nucleic acid sequence encoding one or more helper proteins that assist rAAV production. Generally, helper proteins that assist rAAV production are from a helper virus.

[0266] As used herein, the term “helper virus” refers to a virus used when producing copies of a helper virus-dependent viral vector, such as adeno-associated virus, which does not have the ability to replicate on its own. The helper virus is used to co-infect cells alongside the viral vector and provides the necessary proteins for replication of the genome of the viral vector. Helper virusescommonly used to produce rAAV particles include adenovirus, herpes simplex virus, cytomegalovirus, Epstein-Barr virus, and vaccinia virus.

[0267] Helper viruses include Adenovirus (AV), and herpes simplex virus (HSV), and systems exist for producing AAV in insect cells using baculovirus and mammalian cells. It has also been proposed that papilloma viruses may also provide a helper function for AAV (See, e.g., Hermonat et al., Molecular Therapy 9, 5289-S290(2004)). Helper viruses include any virus capable of creating an allowing AAV replication. AV is a nonenveloped nuclear DNA virus with a double- stranded DNA genome of approximately 36 kb. AV is capable of rescuing latent AAV provirus in a cell by providing Ela, Elb55K, E2a, E4orf6, and VA genes, allowing AAV replication and encapsidation. HSV is a family of viruses that have a relatively large double-stranded linear DNA genome encapsidated in an icosahedral capsid, which is wrapped in a lipid bilayer envelope. HSV are infectious and highly transmissible. The following HSV-1 replication UL8, and UL52) and the DNA binding protein ICP8 encoded by the UL29 gene, with other proteins enhancing the helper function.

[0268] In some embodiments, the nucleic acid sequence encoding the one or more helper proteins one or more of an E2A region, an E4 region, and a virus associated (VA) RNA region, and optionally, an E1 region, an E3 region and / or a Major Late Promoter (MLP) region.

[0269] The E1 regions helps to make the AAV produced to be replication competent. The E3 region contains genes encoding proteins that modulate the immune response following wild-type adenovirus infection. The E3 functions are only activated when the E1 region is functional. The Major Late Promoter is important for late phase-specific stimulation of transcription.

[0270] In some embodiments, each helper protein is selected independently from the other helper proteins.

[0271] In some embodiments, the nucleic acid sequence encoding one or more helper proteins comprises the nucleotide sequence having at least 85% identity to any one of SEQ ID NOs: 67-68 (XX-680) or 69-70 (XX85). For example, the nucleic acid sequence encoding one or more helper protein comprises a nucleotide sequence having at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with any one of SEQ ID NO: 68-70. In some embodiments, the nucleic acid sequence encoding one or more helper proteins comprises a nucleotide sequence having 100% identity with any one of SEQ ID NO: 68-70. For example, the nucleic acid sequence encoding one or more helper proteins consists of a nucleotide sequence having 100% identity with any one of SEQ ID NO: 68-70.

[0272] Exemplary sequences of helper proteins that assist in rAAV production are described for example, in U.S. Provisional Application No.63 / 354,304 filed June 22, 2022 (Adenovirus-BasedNucleic Acids and Methods Thereof), the content of which is incorporated herein by reference in its entirety.

[0273] It is noted that the helper proteins can be from a wildtype adenovirus serotypes or variants thereof.

[0274] In some embodiments, the nucleic acid comprises at least one protelomerase binding site and wherein the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins.

[0275] In some embodiments, the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding one or more helper proteins and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins. In other embodiments, the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding one or more helper proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins. In some preferred embodiments, the Ad helper sequence is any one of SEQ ID NOs: 67-70. AAV Rep

[0276] Embodiments of the various aspects described herein nucleic acid sequence encoding the AAV rep protein. As used herein, a “nucleic acid sequence encoding an AAV rep protein,” also referred to as “Rep encoding sequence,” indicate the nucleic acid sequences that encode the parvoviral or AAV non-structural proteins that mediate viral replication and the production of new virus particles. The parvovirus and AAV replication genes and proteins have been described in, e.g., Fields et al., VIROLOGY, volume 2, chapters 69 & 70 (4th ed., Lippincott-Raven Publishers), content of which is incorporated herein by reference in its entirety.

[0277] The nucleic acid sequence encoding the AAV rep protein need not encode all of the parvoviral or AAV Rep proteins. For example, with respect to AAV, the Rep encoding sequences do not need to encode all four AAV Rep proteins (Rep78, Rep 68, Rep52 and Rep40), in fact, it is believed that AAV5 only expresses the spliced Rep68 and Rep40 proteins. In some embodiments, the Rep encoding sequences encode at least those replication proteins that are necessary for viral genome replication and packaging into new virions. The AAV Rep protein encoding sequences will generally encode at least one large Rep protein (i.e., Rep78 / 68) and one small Rep protein (i.e., Rep52 / 40).

[0278] In some embodiments, the Rep encoding sequences encodes the Rep68 protein. For example, the Rep encoding sequence encode the Rep68 and the Rep52 and / or Rep40 proteins. In some embodiments, the Rep encoding sequence encode the Rep68 and Rep52 proteins. In some other embodiments, the Rep encoding sequence Rep68 and Rep40 proteins.

[0279] In some embodiments, the Rep encoding sequences encodes Rep78. For example, the Rep encoding sequence encode the Rep78 and the Rep52 and / or Rep40 proteins. In some embodiments, the Rep encoding sequence encode the Rep78 and Rep52 proteins. In some other embodiments, the Rep encoding sequence Rep78 and Rep40 proteins.

[0280] As used herein, the term “large Rep protein” refers to Rep68 and / or Rep78.

[0281] It is noted the Rep protein, e.g., large Rep protein can be either wild-type or synthetic. A wild-type Rep protein, e.g., large Rep protein can be from any parvovirus or AAV, including but not limited to serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, AAV4, or any other AAV now known or later discovered. A synthetic Rep protein, e.g., large Rep protein may be altered by insertion, deletion, truncation and / or missense mutations. In some preferred embodiments, the AAV rep-cap sequences are p2 / 8, p2 / 9, or pUC_RCX. AAV Cap

[0282] Embodiments of the various aspects described herein nucleic acid sequence encoding a AAV Cap protein. As used herein, a “nucleic acid sequence encoding an AAV cap protein,” also referred to as “Cap encoding sequence,” indicate the nucleic acid sequences that encode the structural proteins that form a functional parvovirus or AAV capsid (i.e., can package DNA and infect target cells). Typically, the cap encoding sequences will encode all of the parvovirus or AAV capsid subunits, but less than all of the capsid subunits may be encoded as long as a functional capsid is produced. Viral capsid proteins (VP; VP1 / VP2 / VP3) form the outer capsid shell that protects the viral genome, as well as being actively involved in cell binding and internalization (Samulski RJ, Muzyczka N. AAV-mediated gene therapy for research and therapeutic purposes. Annu Rev Virol.2014;1(1):427–451. doi: 10.1146 / annurev-virology-031413-085355). The capsid structure of autonomous parvoviruses and AAV are described in more detail in BERNARD N. FIELDS et al., VIROLOGY, volume 2, chapters 69 & 70 (4th ed., Lippincott-Raven Publishers).

[0283] In some embodiments, the AAV cap protein encoding sequence encodes VP1. In some embodiments, the AAV cap protein encoding sequence encodes VP2. In some embodiments, the AAV cap protein encoding sequence encodes VP3. In some embodiments, the AAV cap protein encoding sequence encodes two of VP1, VP2 and VP3. For example, the AAV cap protein encoding sequence encodes VP1 and VP2. In another example, the AAV cap protein encodingsequence encodes VP1 and VP3. In some embodiments, the AAV cap protein encoding sequence encodes all three of VP1, VP2, and VP3.

[0284] It is noted the Cap protein can be either wild-type or synthetic. A wild-type Cap protein can be from any parvovirus or AAV, including but not limited to serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, AAV4, or any other AAV now known or later discovered. A synthetic Cap protein may be altered by insertion, deletion, truncation and / or missense mutations.

[0285] Exemplary AAV cap protein (capsid) sequences are listed in Table 1. Stuffer sequences of the invention is used in one or, more helper nucleic acids e.g, Any of the short or, large stuffers described in the invention are used in one or, more of i) nucleic acid encoding transgene flanked by AAV ITRs, ii)nucleic acid encoding Adenoviral helper proteins, iii) nucleic acid encoding AAV Rep-Cap proteins to produce recombinant AAV (rAAV) vector that has at least one of the viral structural proteins VP1, VP2, or, VP3 selected from AAV serotypes listed in Table 1. TABLE 1 AAV Serotypes and exemplary published corresponding capsid sequence Each of the non-patent literature and patent literature references, that are recited in this Table are herein incorporated by reference in their entirety. *Representative AAV VP1sequences are provided, which further contain the respective VP2 and VP3 sequences as known in the art.US20150315612) AAVhu.29 (See SEQ ID NO: 225 iAAVhu.29R (See SEQ ID NO: 42 with G396Eh i h

[0286] In some embodiments, the AAV cap sequence can be from any AAV. In some preferred embodiments, the AAV cap sequences are from AAV3B, AAV6, or AAV8.

[0287] In some embodiments of the various aspects described herein, the nucleic acid comprises a nucleic acid sequence encoding both an AAV Rep protein and an AAV Cap protein. In such embodiments, the Rep and Cap proteins can be independently selected from any parvovirus or AAV, including but not limited to serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i,AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, AAV4, or any other AAV now known or later discovered.

[0288] In some embodiments, AAV Rep protein and the AAV Cap protein are from the same AAV serotype. In other embodiments, AAV Rep protein and the AAV Cap protein are from different AAV serotypes. It is noted that one or both of the AAV rep and AAV cap proteins can be synthetic.

[0289] In some embodiments, the nucleic acid comprises at least one protelomerase binding site and wherein the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

[0290] In some embodiments, the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins. In other embodiments, the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins. Length

[0291] In some embodiments, the nucleic acid comprising a short stuffer has a length sufficient for packaging of the nucleic acid sequence into a replication competent AAV particle. For example, the nucleic acid comprising a short stuffer has a length less than 6.5kb. In some embodiments, the nucleic acid comprising a short stuffer has a length less than 6.4kb, less than 6.3kb, less than 6.2kb, less than 6.1kb, less than 6.0kb, less than 5.9kb, less than 5.8kb, less than 5.7kb, less than 5.6kb, less than 5.5kb, less than 5.4kb, less than 5.3kb, less than 5.2kb, less than 5.1kb, or less than 5kb. For example, the nucleic acid comprising a short stuffer has a length less than 4.9kb, less than 4.8kb, less than 4.7kb, less than 4.6kb, less than 4.5kb, less than 4.4kb, less than 4.3kb, less than 4.2kb, less than 4.1kb, less than 4kb, less than 3.9kb, less than 3.8kb, less than 3.7kb, less than 3.6kb, less than 3.5kb, less than 3.4kb, less than 3.3kb, less than 3.2kb, less than 3.1kb, or less than 3.0kb. In some embodiments, the nucleic acid comprising ashort stuffer has a length less than 2.9kb, less than 2.8kb, less than 2.7kb, less than 2.6kb, less than 2.5kb, less than 2.4kb, less than 2.3kb, less than 2.2kb, less than 2.1kb, less than 2.0kb, less than 1.9kb, less than 1.8kb, less than 1.7kb, less than 1.6kb, less than 1.5kb, less than 1.4kb, less than 1.3kb, less than 1.2kb, less than 1.1kb, or less than 1.0kb. For example, the nucleic acid comprising ashort stuffer has a length less than 0.9kb, less than 0.8kb, less than 0.7kb, less than 0.6kb, less than 0.5kb, less than 0.4kb, less than 0.3kb, or less than 0.2kb.

[0292] In some embodiments, the nucleic acid comprising a large stuffer has a length sufficient to prevent production of replication competent rAAV vector. For example, the nucleic acid comprising a large stuffer has sufficient length to prevent packaging of the nucleic acid sequence into a replication competent AAV particle. For example, the nucleic acid comprising a large stuffer has a length greater than 5.5kb, such as greater than 5.6kb, greater than 5.7kb, greater than 5.8kb, greater than 5.9kb, greater than 6.0kb, greater than 6.1kb, greater than 6.2kb, or greater than 6.4kb. In some embodiments, the nucleic acid comprising a large stuffer has a length greater than 6.5kb, greater than 6.6kb, greater than 6.7kb, greater than 6.8kb, greater than 6.9kb, greater than 7.0kb, greater than 7.1kb, greater than 7.2kb, greater than 7.3kb, greater than 7.4kb, greater than 7.5kb, greater than 7.6kb, greater than 7.7kb, greater than 7.8kb, greater than 7.9kb, or greater than 8.0kb. For example, the nucleic acid comprising a large stuffer has a length greater than 8.1kb, greater than 8.2kb, greater than 8.3kb, greater than 8.4kb, greater than 8.5kb, greater than 8.6kb, greater than 8.7kb, greater than 8.8kb, greater than 8.9kb, greater than 9.0kb, greater than 9.1kb, greater than 9.2kb, greater than 9.3kb, greater than 9.4kb, greater than 9.5kb, greater than 9.6kb, greater than 9.7kb, greater than 9.8kb, greater than 9.9kb, or greater than 10.0kb. In some embodiments, the nucleic acid comprising a large stuffer has a length greater than 10.1kb, greater than 10.1kb, greater than 10.2kb, greater than 10.3kb, greater than 10.4kb, greater than 10.5kb, greater than 10.6kb, greater than 10.7kb, greater than 10.8kb, greater than 10.9kb, greater than 11.0kb, greater than 11.1kb, greater than 11.2kb, greater than 11.3kb, greater than 11.4kb, greater than 11.5kb, greater than 11.6kb, greater than 11.7kb, greater than 11.8kb, greater than 11.9kb, or greater than 12.0kb. For example, the nucleic acid comprising a large stuffer has a length greater than 12.1kb, greater than 12.2kb, greater than 12.3kb, greater than 12.4kb, greater than 12.5kb, greater than 12.6kb, greater than 12.7kb, greater than 12.8kb, greater than 12.9kb, greater than 13.0kb, greater than 13.1kb, greater than 13.2kb, greater than 13.3kb, greater than 13.4kb, or greater than 13.5kb. In some embodiments, the nucleic acid comprising a large stuffer has a length greater than 13.6kb, greater than 13.7kb, greater than 13.8kb, greater than 13.9kb, greater than 14.0kb, greater than 14.1kb, greater than 14.2kb, greater than 14.3kb, greater than 14.4kb, greater than 14.5kb, greater than 14.6kb, greater than 14.7kb, greater than 14.8kb, greater than 14.9kb, or greater than 15.0kb or more.

[0293] In some embodiments, the nucleic acid further comprises a stop codon (e.g., TAA, TAG or TGA) upstream of the short stuffer. In other embodiments, the nucleic acid further comprises a stop codon (e.g., TAA, TAG or TGA) downstream of the short stuffer.

[0294] In some embodiments, the nucleic acid has sufficient length to prevent packaging of the nucleic acid sequence into a replication competent AAV particle.Vectors

[0295] In some embodiment of any of the aspects, the nucleic acid described herein is a vector. As used herein, a “vector” refers to a compound used as a vehicle to carry foreign genetic material into another cell, where it can be replicated and / or expressed. A cloning vector containing foreign nucleic acid is termed a recombinant vector. Exemplary vectors include, but are not limited to, plasmids, phagemids, bacmids, cosmids, viral vectors, and artificial chromosomes (e.g., bacterial or yeast artificial chromosome). In some embodiments of any one of the aspects described herein, the nucleic acid described herein is a plasmid. In some other embodiments, the nucleic acid described herein is a bacmid. In yet other embodiments, the nucleic acid described herein is a cosmid. Recombinant vectors typically contain an origin of replication, a multicloning site, and a selectable marker. The nucleic acid sequence typically consists of an insert (recombinant nucleic acid or transgene) and a larger sequence that serves as the “backbone” of the vector. The purpose of a vector which transfers genetic information to another cell is typically to isolate, multiply, or express the insert in the target cell. Expression vectors (expression constructs) are for the expression of the exogenous gene in the target cell, and generally have a promoter sequence that drives expression of the exogenous gene / ORF. Insertion of a vector into the target cell is referred to transformation or transfection for bacterial and eukaryotic cells, although insertion of a viral vector is often called transduction. The term “vector” may also be used in general to describe items to that serve to carry foreign genetic material into another cell, such as, but not limited to, a transformed cell or a nanoparticle. Closed-ended linear duplex DNA (clDNA)

[0296] In some embodiments of any one of the aspects described herein, the nucleic acid described herein is a closed ended linear duplex DNA (clDNA). The term “clDNA” or “close ended linear duplexed DNA” as used herein, refers to closed-linear nucleic acid constructs that eliminates the need for bacterial cells and thus, eliminates bacterial sequences (e.g., an antibiotic resistance gene) that are needed for large scale growth in bacteria.

[0297] Closed ended linear duplexed DNA molecules typically comprise covalently closed ends also described as hairpin loops, where base-pairing between complementary DNA strands is not present. The hairpin loops join the ends of complementary DNA strands. Structures of this type typically form at the telomeric ends of chromosomes in order to protect against loss or damage of chromosomal DNA by sequestering the terminal nucleotides in a closed structure. In examples of closed linear DNA molecules described herein, hairpin loops flank complementary base-paired DNA strands, forming a closed linear (cl) DNA shaped structure. DNA with closed linear shaped structure is described herein as close ended linear duplexed DNA (clDNA, or, celDNA).Alternatively, clDNA is termed herein as no-end DNA (neDNA). In some examples, clDNA or, neDNA further comprises at least one, e.g., two protelomerase binding sites.Non limiting examples of closed linear duplexed DNA, or, no-end DNA (neDNA) include doggybone DNA (dbDNA), and / or dumbbell shaped DNA.

[0298] In some embodiments, one or more nucleic acids may be present on close ended linear duplex nucleic acids. Such nucleic acids can be generated by a variety of known methods, including in vitro cell-free synthesis and in vivo methods.

[0299] In certain embodiments, one or more nucleic acid sequence is an amplified linear open ended DNA, with blunt ends or with overhangs, and a synthesized hairpin molecule is ligated to one or both ends to form the closed ended linear duplex DNA comprising one or more of the nucleic acids as described herein. Unligated hairpins are purified away using means well known to those of skill in the art. The DNA may be amplified by PCR and ligated to double stranded form.

[0300] One method of generating the covalently closed ended linear duplex nucleic acid is by incorporation of protelomerase binding sites in a precursor molecule such that the protelomerase binding sites flank the nucleic acid of interest. The nucleic acid of interest can be exposed to protelomerase to thereby cleave and ligate the DNA at the site. Non-limiting examples of cell free in vitro synthesis of dumbbell shaped DNA and dbDNA are e.g., as described in US 9,109,250; US 6,451,563; Efficient production of superior dumbbell-shaped DNA minimal vectors for small hairpin RNA expression Nucleic Acids Res. 2015 Oct 15; 43(18): e120, High-Purity Preparation of a Large DNA Dumbbell-Antisense & nucleic acid drug development 11:149–153 (2001);; US 9,109,250; U.S. Patent No. 9,499,847; U.S. Patent No. 10,501,782; and International Publication No. WO 2018033730 A1; all of which are herein incorporated by reference in their entireties.

[0301] In some embodiments, the nucleic acid is a covalently closed-ended linear duplex DNA. Cells

[0302] The disclosure also provides a host cell comprising a nucleic acid described herein described herein. As used herein, the term “cell” refers to a single cell as well as to a population of (i.e., more than one) cells. A host cell can be a prokaryotic or eukaryotic host cell. Exemplary host cells include, but are not limited to, bacterial cells, yeast cells, plant cell, animal (including insect) or human cells. Exemplary host cells include, but are not limited to, HEK293, CHO, A549 , BHK 21 (clone 13), CV-1 , HeLa, LLCMK2, McCoy, MDCK , MRC-5, NCI-H292 , Vero, Vero76, Wi 38, A549 , Sf9, HepG2, MCF-7, MEF, NS0, HUVEC, Jurkat, Cos-7, 3T3 , HL60, ML-1, KG-1, U- 937, THP-1, K-562, Molt-4, TF-1, Sf9, Sf21, and Hi-5.

[0303] In some embodiments, the host cell is a eukaryotic cell. For example, the host cell is an insect or mammalian cell. In some embodiments, the host cells is a HEK293 cell. In some other embodiments, the host cell is a HeLa cell.

[0304] In some embodiments, the host cell is a microbial cell, for example, bacterial cells such as a E. coli cell, and yeast cell such as S. cerevisiae cell.

[0305] In some embodiments of any one of the aspects described herein, the host cell comprises a nucleic acid comprising a short stuffer, at least one ITR and a nucleic acid sequence encoding a transgene.

[0306] In some embodiments of any one of the aspects described herein, the host cell comprises a nucleic acid comprising a short stuffer and a nucleic acid sequence encoding one or more helper proteins e.g, Adenoviral based helper proteins, AAV helper Rep-Cap proteins) that assist in rAAV production.

[0307] In some embodiments of any one of the aspects described herein, the host cell comprises a nucleic acid comprising a short stuffer and a nucleic acid sequence encoding AAV Rep and / or AAV Cap proteins.

[0308] In some embodiments of any one of the aspects described herein, the host cell comprises a nucleic acid comprising a large stuffer and a nucleic acid sequence encoding AAV Rep and / or AAV Cap proteins.

[0309] In some embodiments of any one of the aspects described herein, the host cell comprises: (i) a nucleic acid comprising a short stuffer, at least one AAV ITR and a nucleic acid sequence encoding a transgene; (ii) a nucleic acid comprising a short stuffer and a nucleic acid sequence encoding one or more helper proteins (e.g., Adenoviral helper proteins) that assist in rAAV production; and (iii) a nucleic acid comprising a large stuffer and a nucleic acid sequence encoding AAV Rep and / or AAV Cap proteins.

[0310] The host cells can be employed in a method of producing viral particles, e.g., rAAV particles. Generally, the method comprises: culturing a host cell comprising a nucleic acid described herein under conditions such that viral particles are produced; and optionally recovering the viral particles from the culture medium. The viral particles can be concentrated and purified by a variety of biochemical and chromatographic methods, including methods utilizing differences in size, charge, hydrophobicity, solubility, specific affinity, etc. between the viral particles and other substances in the cell culture medium.

[0311] The cells can be cultured in suspension and the cells can be cultured in animal component-free conditions. The animal component-free medium can be any animal component- free medium (e.g., serum-free medium) compatible with a given cell line, for example, HEK293 cells. Examples include, without limitation, SFM4Transfx-293 (HYCLONE), Ex-Cell 293 (JRHBIOSCIENCES), LC-SFM (INVITROGEN), and Pro293-S (LONZA) Pro-10 cells (as described in US Patent Application 9,441,206, which is incorporated by reference in its entirety). Uses

[0312] The nucleic acids and host cell described herein can be used in manufacturing or rAAV. For example, nucleic acids comprising a large stuffer can be used in a method of preventing manufacturing of replication competent rAAV.

[0313] As used herein, “replication competent” refers to the nucleic acid containing the viral genome including, but not limited to the ITRs, transgene, and promoter, packaging components, including but not limited to Rep / Cap, and the helper components, including but not limited to E1, E2A, E4, and VA RNA. The E2A region in the ad helper produces the L4-100K protein. The L4- 100K protein is involved in hexon assembly and transport of the hexon structure to the nucleus as well as other proteins that interact with hexon in the final formation of the capsid. This region can also produce adenovirus L4-22K and adenovirus L4-33K. L4-22K is a multifunctional protein involved in packaging of the viral genome into an empty capsid as well as the temporal switch from the early to late phase of infection by regulating both early and late gene expression. L4-33K functions as an alternative splicing factor involved in genome packaging. The E4 region contributes to the expression of early and late genes in virion packaging. Early genes support viral replication inside host cells; late genes support host cell lysis, viral assembly, and virion release. The viral associated (VA) RNA region is a type of non-coding RNA found in adenoviruses. It has a role in regulating translation for both early and late-stage genes. In some embodiments, there is at least one copy of VA RNA present in a replication competent nucleic acid.

[0314] Also, provided herein is a method for producing a plurality of viral particles. Generally, the method comprises culturing a host cell comprising a nucleic acid described herein in a culture medium under conditions in which viral particles are produced.

[0315] In some embodiments, the nucleic acid as described herein produces recombinant AAV (rAAV) by a method comprising: transfecting cells with i) the nucleic acid encoding rAAV genome, ii) a Adenoviral helper nucleic acid and iii) helper nucleic acid encoding AAV capsid and non-structural replication genes, allowing cells sufficient time to produce rAAV particles, and producing clarified lysate comprising rAAV capsid particles, wherein the rAAV capsid particles in the clarified lysate comprises at least about 15% full capsid particles. In certain embodiments, the rAAV in the clarified lysate comprises at least about 15% full capsid particles, at least about 18% full capsid particles, at least about 20% full capsid particles, at least about 22% full capsid particles, at least about 25% full capsid particles, at least about 30% full capsid particles, at least about 35% full capsid particles, or a higher percentage of full capsid particles. In some embodiments, thenucleic acids used in all steps i), ii), and, iii) comprise any of the short stuffer sequences described in the invention. In some embodiments, at least one of nucleic acids used in steps, i), ii), or, iii) comprise any of the short stuffer sequences described herein.

[0316] In certain aspects of the embodiment, the copy number of the nucleic acid as described herein, that is used to produce rAAV, is at least about 2000 copies per cell to at least about 20,000 copies per cell. In some embodiments, the copy number of the nucleic acid as described herein is at least about 1000 copies per cell, at least about 1500 copies per cell, at least about 2000 copies per cell, at least about 2500 copies per cell, at least about 3000 copies per cell, at least about 3500 copies per cell, at least about 4000 copies per cell, at least about 4500 copies per cell, at least about 5000 copies per cell, at least about 5500 copies per cell, at least about 6000 copies per cell, at least about 6500 copies per cell, at least about 7000 copies per cell, at least about 7500 copies per cell, at least about 8000 copies per cell, at least about 8500 copies per cell, at least about 9000 copies per cell, at least about 9500 copies per cell, at least about 10000 copies per cell, at least about 12000 copies per cell, at least about 14000 copies per cell, at least about 16000 copies per cell, at least about 18000 copies per cell, at least about 20000 copies per cell or higher.

[0317] In some embodiments, the copy number of the nucleic acid as described herein is at least about 5000 copies per cell to at least about 12000 copies per cell. In some embodiments, the nucleic acid as described herein is used to produce rAAV particles comprising at least about 20% to at least about 35% full capsid particles. In an exemplary method of producing recombinant AAV, the method comprises A) first transfecting cells with (i) the nucleic acid as described herein present on a recombinant AAV genome comprising an AAV endogenous genome flanked by left inverted terminal repeat (L-ITR) or a recombinant AAV genome comprising nucleic acid encoding any transgene flanked by left and right ITRs, (ii) a helper nucleic acid, and (iii) AAV capsid and non- structural replication (AAVRep-Cap) nucleic acid; B) producing clarified lysate out of a bioreactor, wherein the clarified lysate comprises rAAV particles, C) enriching (or purifying) the rAAV in the clarified lysate (e.g., by chromatography purification methods).

[0318] In some embodiments, the enriching step increases the percentage of full viral particles (e.g., by removing at least some of the partially full or empty viral particles). Without wishing to be bound by theory, the enriched solution comprising full viral particles (e.g., as measured by % full AAV particles, % full rAAV particles) may still comprise partially full viral particles and / or empty viral particles; however, the percentage of partially full viral particles and / or empty viral particles is substantially decreased compared to a clarified lysate that is not enriched.

[0319] As used herein, “transfection” refers to the insertion of a nucleic acid into a target cell. In some embodiments, the target cell is a mammalian cell. In some embodiments, the target cell is a suspension HEK293 cell. There are two different types of transfection: stable transfection andtransient transfection. Stable transfection incorporates exogenous nucleic acids into the transfected cell’s genome whereas in transient transfection, the exogenous nucleic acids are present only for a limited time in the cell and do not integrate with the transfected cell’s genome. In some embodiments, the transfection method used is transient transfection. In some embodiments, the transfection method used is stable transfection. Transfection can be performed with a variety of methods including, but not limited to, calcium phosphate, electroporation, and / or cationic lipid- mediated methods (e.g., LIPOFECTAMINE, polyethylenimine (PEI)). In some embodiments, the transfection method uses polyethylenimine. Transfection can require an optimal cell density based on the cell type, application, and / or transfection technology. Additional description of transfection can be found in Shin et al. Recombinant Adeno-Associated Viral Vector Production and Purification. Methods Mol Biol. 2012; 798: 267–284; Grieger et al. Production of Recombinant Adeno-associated Virus Vectors Using Suspension HEK293 Cells and Continuous Harvest of Vector From the Culture Media for GMP FIX and FLT1 Clinical Vector. Mol Ther. 2016 Feb; 24(2): 287–297; and Meier et al. The Interplay between Adeno-Associated Virus and Its Helper Viruses. Viruses. 2020 Jun; 12(6): 662, each of which is incorporated by reference herein in their entireties.

[0320] As used herein, “sufficient cell mass” refers to an optimal cell density for transfection. In some embodiments, suspension HEK293 cells are expanded to produce sufficient cell mass to seed a bioreactor from at least a 25L scale. In order to achieve maximum production of the rAAV virion of interest, cells can require at least 10 hours, at least 11 hours, at least 12 hours, at least 13 hours, at least 14 hours, at least 15 hours, at least 16 hours, at least 17 hours, at least 18 hours, at least 19 hours, at least 20 hours, at least 21 hours, at least 22 hours, at least 23 hours, at least 24 hours, at least 25 hours, at least 26 hours, at least 27 hours, at least 28 hours, at least 29 hours, at least 30 hours, at least 31 hours, at least 32 hours, at least 33 hours, at least 34 hours, at least 35 hours, at least 36 hours, at least 37 hours, at least 38 hours, at least 39 hours, at least 40 hours, at least 41 hours, at least 42 hours, at least 43 hours, at least 44 hours, at least 45 hours, at least 46 hours, at least 47 hours, at least 48 hours, at least 49 hours, at least 50 hours, at least 51 hours, at least 52 hours, at least 53 hours, at least 54 hours, at least 55 hours, at least 56 hours, at least 57 hours, at least 58 hours, at least 59 hours, at least 60 hours, at least 61 hours, at least 62 hours, at least 63 hours, at least 64 hours, at least 65 hours, at least 66 hours, at least 67 hours, at least 68 hours, at least 69 hours, at least 70 hours, at least 71 hours, at least 72 hours, at least 73 hours, at least 74 hours, at least 75 hours, at least 76 hours, at least 77 hours, at least 78 hours, at least 79 hours, at least 80 hours, at least 81 hours, at least 82 hours, at least 83 hours, at least 84 hours, at least 85 hours, at least 86 hours, at least 87 hours, at least 88 hours, at least 89 hours, at least 90 hours, at least 91 hours, at least 92 hours, at least 93 hours, at least 94 hours, at least 95 hours, at least 96hours, at least 97 hours, at least 98 hours, at least 99 hours, at least 100 hours or more post- transfection before harvesting virions from the transfected cells.

[0321] Transfection of multiple nucleic acids into the same cells can occur simultaneously or it can occur within 5 minutes, within 10 minutes, within 15 minutes, within 20 minutes, within 25 minutes, within 30 minutes, within 35 minutes, within 40 minutes, within 45 minutes, within 50 minutes, within 60 minutes, within 65 minutes, within 70 minutes, within 75 minutes, within 80 minutes, within 85 minutes, within 90 minutes, within 95 minutes, within 100 minutes, within 110 minutes, within 120 minutes, within 130 minutes, within 140 minutes, within 150 minutes, within 160 minutes, within 170 minutes, within 180 minutes, within 190 minutes, within 200 minutes, within 210 minutes, within 220 minutes, within 230 minutes, within 240 minutes, within 250 minutes, within 260 minutes, within 270 minutes, within 280 minutes, within 290 minutes, within 300 minutes, within 310 minutes, within 320 minutes, within 330 minutes, within 340 minutes, within 350 minutes, within 360 minutes or more between transfection of the first nucleic acid and transfection of subsequent nucleic acids.

[0322] In some embodiments, the host cell comprises at least one nucleic acid encoding one or more helper proteins sufficient for rAAVproduction, at least one nucleic acid encoding AAV Rep and AAV Cap proteins, and at least one nucleic acid encoding a transgene of interest.

[0323] In some embodiments, the plurality of viral particles comprises rAAV particles.

[0324] A “filled particle” or “full particle” (also interchangeably referred to as “full AAV particle,” “full AAV capsid particle”, or “full rAAV capsid particle”) refers to a viral particle that comprises an intact viral particle (e.g., complete capsid) comprising a genome (e.g., the viral genome or the recombinant genome, which can comprise a heterologous polynucleotide such as a transgene, i.e., a polynucleotide other than a wild-type virus genome). A “filled” or “full” particle can also be interchangeably referred to as a “packaged particle,” “packaged virus,” “packaged AAV,” or “recombinantly expressed AAV”. It is noted that the terms “particle” and “capsid” can be used interchangeably and / or redundantly herein.

[0325] An “empty particle,” which is also interchangeably referred to as “empty AAV particle,” refers to a viral particle that comprises at least one viral protein but lacks all of the genome, e.g., virus genome or recombinant genome. Empty particles do not include, e.g., an intact viral particle comprising a heterologous polynucleotide.

[0326] A “partially full particle,” which is also interchangeably referred to as “partially full AAV particle” or “partially filled AAV particle,” refers to a viral particle that comprises at least one viral protein but lacks at least part of the genome, e.g., virus genome or recombinant genome. As used herein, “partially full particle” also include particles containing DNA from the host cell or pDNA used in transfection.

[0327] The percentage of full AAV particles (“% AAV full” or “% full”) in the clarified lysate produced using the nucleic acids as described herein can be expressed as the number of “full” AAV particles over the total number of AAV particles (including “full,” partially full,” and “empty” AAV particles). Manufacturing

[0328] In several embodiments, the nucleic acid comprising the stuffers of the invention is used to produce recombinant AAV (rAAV). rAAV was manufactured using the method as described in PCT / US2022 / 013279, published as WO2022159679, and / or, as described in PCT / US2021 / 013689, published as WO / 2021 / 146591 which are incorporated herein by reference in its entirety. In some embodiments, the nucleic acid as described herein produces recombinant AAV (rAAV) by a method comprising: transfecting cells with i) the nucleic acid encoding rAAV genome, ii) a Adenoviral helper nucleic acid and iii) helper nucleic acid encoding AAV capsid and non-structural replication genes, allowing cells sufficient time to produce rAAV particles. In some embodiments, the nucleic acids used in all steps i), ii), and, iii) comprise any of the short stuffer sequences described in the invention. In some embodiments, at least one of nucleic acids used in steps, i), ii), or, iii) comprise any of the short stuffer sequences described herein. In several embodiments of any one aspect described herein the nucleic acid used in i, ii, and iii are plasmid DNA. In several embodiments of any one aspect described herein the nucleic acid used in i, ii, and iii are clDNA or, neDNA. In some embodiments, clDNA or, neDNA used in i), ii), or, iii) further comprises protelomerase binding site e.g., TelRL.

[0329] In some embodiments, the nucleic acid as described herein is used to manufacture haploid, rational haploid, or rational polyploid AAV, e.g., as described in US 10,550,405, International Patent applications PCT / US2018 / 022725, PCT / US2018 / 044632, all of which are incorporated herein by reference in their entirety. In some embodiments, the nucleic acid (which is used to manufacture AAV and or recombinant AAV) is plasmid DNA or closed ended linear duplexed DNA (clDNA). The clDNA described herein is alternatively termed as no-end DNA (neDNA). In some exemplary aspects, the clDNA described herein is no-end DNA (neDNA). In some embodiments, the nucleic acid as described herein is used to manufacture recombinant AAV that comprises one or both of the ITRs that is 145 nucleotides long, or less than 145 nucleotides long. In some embodiments, the nucleic acid as described herein is used to manufacture recombinant AAV that comprises one or both of the ITRs that is 140 nucleotides long, 135 nucleotides long, 130 nucleotides long, 125 nucleotides long or less than 125 nucleotides long.

[0330] In some embodiments, the nucleic acid as described herein produces recombinant AAV (rAAV) by a method comprising: transfecting cells with i) Ad helper nucleic acid of invention, ii)rAAV genome and iii) AAV capsid and non-structural replication genes, allowing cells sufficient time to produce rAAV particles, and producing clarified lysate comprising rAAV capsid particles, wherein the rAAV capsid particles in the clarified lysate comprises at least about 15% full capsid particles. In certain embodiments, the rAAV in the clarified lysate comprises at least about 15% full capsid particles, at least about 18% full capsid particles, at least about 20% full capsid particles, at least about 22% full capsid particles, at least about 25% full capsid particles, at least about 30% full capsid particles, at least about 35% full capsid particles, or a higher percentage of full capsid particles.

[0331] In certain aspects of the embodiment, the copy number of the nucleic acid as described herein, that is used to produce rAAV, is at least about 2000 copies per cell to at least about 20,000 copies per cell. In some embodiments, the nucleic acid as described herein is at least about 1000 copies per cell, at least about 1500 copies per cell, at least about 2000 copies per cell, at least about 2500 copies per cell, at least about 3000 copies per cell, at least about 3500 copies per cell, at least about 4000 copies per cell, at least about 4500 copies per cell, at least about 5000 copies per cell, at least about 5500 copies per cell, at least about 6000 copies per cell, at least about 6500 copies per cell, at least about 7000 copies per cell, at least about 7500 copies per cell, at least about 8000 copies per cell, at least about 8500 copies per cell, at least about 9000 copies per cell, at least about 9500 copies per cell, at least about 10000 copies per cell, at least about 12000 copies per cell, at least about 14000 copies per cell, at least about 16000 copies per cell, at least about 18000 copies per cell, at least about 20000 copies per cell or higher.

[0332] In some embodiments, the nucleic acid as described herein is at least about 5000 copies per cell to at least about 12000 copies per cell. In some embodiments, the nucleic as described herein is used to produce rAAV particles comprising at least about 20% to at least about 35% full capsid particles. In an exemplary method of producing recombinant AAV, the method comprises A) first transfecting cells with (i) the helper nucleic acid as described herein, (ii) a recombinant AAV genome comprising an AAV endogenous genome flanked by left inverted terminal repeat (L-ITR) or a recombinant AAV genome comprising nucleic acid encoding any transgene flanked by left and right ITRs, and (iii) AAV capsid and non-structural replication (AAVRep-Cap) nucleic acid; B) producing clarified lysate out of a bioreactor, wherein the clarified lysate comprises rAAV particles, C) enriching (or purifying) the rAAV in the clarified lysate (e.g., by chromatography purification methods).

[0333] In some embodiments, the enriching step increases the percentage of full viral particles (e.g., by removing at least some of the partially full or empty viral particles). Without wishing to be bound by theory, the enriched solution comprising full viral particles (e.g., as measured by % full AAV particles, % full rAAV particles) may still comprise partially full viral particles and / orempty viral particles; however, the percentage of partially full viral particles and / or empty viral particles is substantially decreased compared to a clarified lysate that is not enriched.

[0334] In some embodiments, the nucleic acid as described herein produces purified recombinant AAV (rAAV) particles by a method comprising: A) transfecting cells with i) Ad helper nucleic acids, ii) rAAV genome and iii) AAV capsid and non-structural replication genes, B) allowing cells sufficient time to produce rAAV particles, C) producing clarified lysate, and D) purifying (or enriching) the clarified lysate (e.g., using chromatography purification methods), thereby producing enriched or purified rAAV particles. In some embodiments, the purified or enriched rAAV particles comprise at least about 65% full capsid particles. In certain embodiments, the purified or enriched rAAV particles comprise at least about 70% full capsid particles, at least about 75% full capsid particles, at least about 80% full capsid particles, at least about 85% full capsid particles, at least about 90% full capsid particles, at least about 95% full capsid particles, at least about 98% full capsid particles, at least about 99% full capsid particles or at least about 99.5% full capsid particles or higher. In certain embodiments, the purified or enriched rAAV particles comprise 100% full capsid particles. In certain embodiments, the purified or enriched rAAV particles comprise less than about 10% empty capsid particles, less than about 8% empty capsid particles, less than about 6% empty capsid particles, less than about 5% empty capsid particles, less than about 5% empty capsid particles, less than about 3% empty capsid particles, less than about 2% empty capsid particles, less than about 1% empty capsid particles, less than about 0.8% empty capsid particles, less than about 0.6% empty capsid particles, less than about 0.5% empty capsid particles, less than about 0.4% empty capsid particles, less than about 0.3% empty capsid particles, less than about 0.2% empty capsid particles, less than about 0.1% empty capsid particles, less than about 0.08% empty capsid particles, less than about 0.06% empty capsid particles, less than about 0.05% empty capsid particles, less than about 0.03% empty capsid particles, less than about 0.02% empty capsid particles, or less than about 0.01% empty capsid particles, or fewer % empty capsid particles. In some embodiments, the purified or enriched rAAV is substantially devoid of empty capsid particle. Residual DNA

[0335] Residual DNA as used herein is non-viral genome detected in a viral population, or, in a plurality of viral particles. Residual DNA can be a part of the backbone of plasmid DNA, precursor plasmid DNA, or, clDNA that are used to produce the viral vector. As non-limiting examples, residual DNA are DNA encoding Rep genes, needed for viral replication, or any fragment thereof; part of promoters operably linked to Rep genes; DNA encoding helper virus proteins, needed for viral replication, or any fragment thereof; antibiotic resistance genes, e.g., bacterial sequences. Theshort stuffer sequences as described in the invention can be used to detect residual DNA in a population of any viral vector known in the art. In one embodiment, SEQ ID NO: 1-6, or, sequences with at least 85% identity thereto, or, a contiguous fragment from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NO: 7-9 are used to detect residual DNA in a population of viral vector. Non limiting examples of viral vector are, Lentiviral vectors, Retroviral vectors, Adenoviral vectors, Adeno associated viral vectors (AAV), Herpes, Simplex viral vectors (HSV), or, any chimeric or hybrid viral vectors known in the art. The residual DNA can be expressed as copies / ml of the viral vector population.

[0336] The invention as described herein further describes a method of detecting residual DNA by using oligonucleotide primer or, oligonucleotide probes that can anneal to and thereby detects any of the short stuffer sequences in a viral vector preparation or, viral vector population. In one embodiment, oligonucleotide primers or, probes anneal to and thereby detect any of the sequences selected from SEQ ID NO: 1-6, or, sequences with at least 85% identity thereto, a fragment therof, or, a contiguous fragment from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NO: 7-9, in a viral population.

[0337] The oligonucleotide primer or probe sequences are unique as described in the invention that anneals to the unique short stuffer sequences described in the invention. The oligonucleotide primer and / or, probe as described herein is about 15 nucleotides to about 35 nucleotides long. The method of detecting residual DNA in a viral preparation includes providing two or, more of oligonucleotide primers or, probes.

[0338] Further aspects of the invention include providing a kit to detect residual DNA or, in other words to detect purity of a viral preparation where the kit comprises at least one oligonucleotide primer and / or, probe that anneals to and thereby detect any sequence from SEQ ID NO: 1-6, a fragment therof or, sequences with at least 85% identity thereto, , or, a contiguous fragment from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NO: 7-9 in a viral population. The kit further comprises at least one of helper nucleic acids comprising the short stuffer and / or, large stuffer sequences as described herein to produce viral vector.

[0339] Some exemplary aspects of the disclosure are described by one or more of following numbered Embodiments:

[0340] Embodiment 1: A nucleic acid comprising a short stuffer sequence, wherein the short stuffer sequence comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 1-6.

[0341] Embodiment 2: The nucleic acid of Embodiment 1, wherein the nucleic acid is linear DNA.

[0342] Embodiment 3: The nucleic acid of Embodiment 1 or 2, wherein the nucleic acid is a closed ended linear duplexed DNA (clDNA).

[0343] Embodiment 4: The nucleic acid of any one of Embodiments 1-3, wherein the nucleic acid further comprises at least one protelomerase binding site.

[0344] Embodiment 5: The nucleic acid of Embodiment 4, wherein the short stuffer is located downstream of the protelomerase binding site.

[0345] Embodiment 6: The nucleic acid of Embodiment 4, wherein the short stuffer is located upstream of the protelomerase binding site.

[0346] Embodiment 7: The nucleic acid of Embodiment 4, wherein the nucleic acid comprises two protelomerase binding sites and the short stuffer is located between the two protelomerase binding sites.

[0347] Embodiment 8: The nucleic acid of any one of Embodiments 1-7, wherein the nucleic acid further comprises a heterologous transgene operably linked to one or more regulatory elements.

[0348] Embodiment 9: The nucleic acid of Embodiment 8, wherein the short stuffer is located upstream of the heterologous transgene.

[0349] Embodiment 10: The nucleic acid of Embodiment 9, wherein the short stuffer is located downstream of the heterologous transgene.

[0350] Embodiment 11: The nucleic acid of any one of Embodiments 8-10 wherein the nucleic acid further comprises at least one adeno-associated virus (AAV) inverted terminal repeat (ITR) sequence.

[0351] Embodiment 12: The nucleic acid of Embodiment 11, wherein the short stuffer is upstream of the at least one ITR sequence.

[0352] Embodiment 13: The nucleic acid of Embodiment 11, wherein the short stuffer is downstream of the at least one ITR sequence.

[0353] Embodiment 14: The nucleic acid of any one of Embodiments 11-13, wherein the at least one ITR sequence is located between the short stuffer and the heterologous transgene.

[0354] Embodiment 15: The nucleic acid of any one of Embodiments 10-13, wherein the nucleic acid further comprises at least one protelomerase binding site and the short stuffer is located between the at least one protelomerase binding site and the at least ITR.

[0355] Embodiment 16: The nucleic acid of any one of Embodiments 11-15, wherein the nucleic acid comprises at least two ITRs and the short stuffer located outside the two ITRs.

[0356] Embodiment 17: The nucleic acid of Embodiment 16, wherein the heterologous transgene is located between the two ITRs.

[0357] Embodiment 18: The nucleic acid of Embodiments 16, wherein the heterologous transgene is located between the two ITRs, and one of the ITRs is located between the short stuffer and the heterologous transgene.

[0358] Embodiment 19: The nucleic acid of Embodiment 16, wherein the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the heterologous polynucleotide sequence is located between the first and second ITR sequences, and wherein the short stuffer is upstream of the first ITR sequence.

[0359] Embodiment 20: The nucleic acid of Embodiment 16, wherein the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the heterologous polynucleotide sequence is located between first and second ITR sequences, wherein the short stuffer is upstream of the first ITR sequenc, and wherein the nucleic acid does not comprise a short stuffer down stream of the second ITR sequence.

[0360] Embodiment 21: The nucleic acid of Embodiment 16, wherein the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the short stuffer is located upstream of the 5’-end of the first and second ITRs.

[0361] Embodiment 22: The nucleic acid of any one of Embodiments 7-21, wherein each ITR sequence is selected independently from an ITR sequence of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

[0362] Embodiment 23: The nucleic acid of Embodiment 22, wherein the ITR sequences are from the same AAV serotype.

[0363] Embodiment 24: The nucleic acid of Embodiment 22, wherein the ITR sequences are from the different AAV serotype.

[0364] Embodiment 25: The nucleic acid of any one of Embodiments 1-7, wherein the nucleic acid comprises a nucleic acid sequence encoding one or more helper proteins that assist rAAV replication.

[0365] Embodiment 26: The nucleic acid of Embodiment 25, wherein the nucleic acid comprises at least one protelomerase binding site and wherein the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins.

[0366] Embodiment 27: The nucleic acid of Embodiment 26, wherein the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding one or more helper proteins.

[0367] Embodiment 28: The nucleic acid of Embodiment 26, wherein the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding one or more helper proteins.

[0368] Embodiment 29: The nucleic acid of any one of Embodiments 25-28, wherein the helper protein sufficient for rAAV replication comprises one or more of an E2A region, an E4 region, and a virus associated (VA) RNA region, and optionally, an E1 region, an E3 region and / or a Major Late Promoter (MLP) region.

[0369] Embodiment 30: The nucleic acid of any one of Embodiments 25-29, wherein the nucleic acid sequence encoding one or more helper proteins comprises the nucleotide sequence of any one of SEQ ID NOs: 67-70.

[0370] Embodiment 31: The nucleic acid of any one of Embodiments 1-7, wherein the nucleic acid further comprises nucleic acid sequence encoding a AAV rep protein and a AAV cap protein.

[0371] Embodiment 32: The nucleic acid of Embodiment 31, wherein the nucleic acid comprises at least one protelomerase binding site and wherein the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

[0372] Embodiment 33: The nucleic acid of Embodiment 31, wherein the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

[0373] Embodiment 34: The nucleic acid of Embodiment 31, wherein the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

[0374] Embodiment 35: The nucleic acid of any one of Embodiments 31-34, wherein AAV Rep protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

[0375] Embodiment 36: The nucleic acid of any one of Embodiments 31-35, wherein AAV Cap protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

[0376] Embodiment 37: The nucleic acid of any one of Embodiments 31-36, wherein AAV Rep protein and the AAV Cap protein are from the same AAV serotype.

[0377] Embodiment 38: The nucleic acid of any one of Embodiments 31-36, wherein AAV Rep protein and the AAV Cap protein are from different AAV serotypes.

[0378] Embodiment 39: The nucleic acid of any one of Embodiments 1-6, wherein the nucleic acid comprises an ITR sequence and a protelomerase binding upstream of the ITR sequence, wherein the nucleic acid comprises the sequence of SEQ ID NOs: 1 or 4 between the ITR and the protelomerase binding site.

[0379] Embodiment 40: The nucleic acid of any one of Embodiments 1-6, wherein the nucleic acid comprises an ITR sequence and a protelomerase binding downstream of the ITR sequence, wherein the nucleic acid comprises the sequence of SEQ ID NOs: 1 or 4 between the ITR and theprotelomerase binding site.

[0380] Embodiment 41: The nucleic acid of any one of Embodiments 1-40, wherein the nucleic acid further comprises a stop codon (e.g., TAA, TAG or TGA) upstream of the short stuffer.

[0381] Embodiment 42: The nucleic acid of any one of Embodiments 1-40, wherein the nucleic acid further comprises a stop codon (e.g., TAA, TAG or TGA) downstream of the short stuffer.

[0382] Embodiment 43: The nucleic acid of any one of Embodiments 1-42, wherein the nucleic acid is a vector.

[0383] Embodiment 44: The nucleic acid of any one of Embodiments 1-43, wherein the nucleic acid is a plasmid.

[0384] Embodiment 45: A host cell comprising the nucleic acid of any one of Embodiments 1- 44.

[0385] Embodiment 46: The host cell of Embodiment 45, wherein the host cell is an insect cell or a mammalian cell.

[0386] Embodiment 47: The host cell of Embodiment 46, wherein the mammalian cell is a HEK293 cell or a HeLa cell.

[0387] Embodiment 48: Use of a nucleic acid of any one of Embodiments 1-44 in a method of producing a plurality of viral particles.

[0388] Embodiment 49: Use of Embodiment 48, wherein the viral particles are recombinant AAV (rAAV) particles.

[0389] Embodiment 50: A method for producing a plurality of viral particles, the method comprising culturing a host cell of Embodiment 46 or 47 in a culture medium under conditions in which viral particles are produced.

[0390] Embodiment 51: A method for producing a plurality of viral particles, the method comprises culturing a host cell comprising a nucleic acid of any one of Embodiments 1-44 in a culture medium under conditions in which viral particles are produced.

[0391] Embodiment 52: The method of Embodiment 50 or 51, wherein the host cell comprises at least one nucleic acid encoding one or more helper proteins sufficient for rAAV replication, at least one nucleic acid encoding AAV rep and AAV cap proteins, and at least one nucleic acid encoding a transgene of interest.

[0392] Embodiment 53: The method of any one of Embodiments 50-52, wherein the host cell comprises at least one nucleic acid of any one of Embodiments 7-24.

[0393] Embodiment 54: The method of any one of Embodiments 50-53, wherein the host cell comprises at least one nucleic acid of any one of Embodiments 25-30.

[0394] Embodiment 55: The method of any one of Embodiment 50-54, wherein the host cell comprises at least one nucleic acid of any one of Embodiments 31-38.

[0395] Embodiment 56: The method of any one of Embodiments 50-55, wherein the host cell comprises at least one nucleic acid of any one of Embodiments 7-24, at least one nucleic acid of any one of Embodiments 25-30, and at least one nucleic acid of any one of Embodiments 31-38.

[0396] Embodiment 57: The method of any one of Embodiments 50-56, wherein the plurality of viral particles comprises rAAV particles.

[0397] Embodiment 58: A nucleic acid comprising a large stuffer, wherein the large stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 7-9.

[0398] Embodiment 59: The nucleic acid of Embodiment 58, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter.

[0399] Embodiment 60: The nucleic acid of Embodiment 59, wherein the large stuffer is located in the nucleic acid sequence encoding the AAV Rep protein.

[0400] Embodiment 61: The nucleic acid of Embodiment 60, wherein the large stuffer is located in an intron in the encoding the AAV Rep protein.

[0401] Embodiment 62: The nucleic acid of any one of Embodiments 59-61, wherein the large stuffer is upstream of the promoter.

[0402] Embodiment 63: The nucleic acid of any one of Embodiments 59-61, wherein the large stuffer is downstream of the promoter.

[0403] Embodiment 64: The nucleic acid of any one of Embodiments 59-63, wherein the promoter is a p19 promoter.

[0404] Embodiment 65: The nucleic acid of any one of Embodiments 59-64, wherein the AAV Rep is large Rep.

[0405] Embodiment 66: The nucleic acid of any one of Embodiments 59-65, wherein the AAV Rep is Rep68 or Rep78.

[0406] Embodiment 67: The nucleic acid of any one of Embodiments 59-66, wherein the AAV Rep is Rep68.

[0407] Embodiment 68: The nucleic acid of any one of Embodiments 59-67, wherein the AAV Rep protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

[0408] Embodiment 69: The nucleic acid of any one of Embodiments 59-68, wherein the AAVRep protein is a modified Rep protein.

[0409] Embodiment 70: The nucleic acid of any one of Embodiments 59-69, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Cap protein.

[0410] Embodiment 71: The nucleic acid of Embodiment 70, wherein the nucleic acid sequence encoding the AAV Cap protein is downstream of the nucleic acid sequence encoding the AAV Rep protein.

[0411] Embodiment 72: The nucleic acid of Embodiment 70, wherein the nucleic acid sequence encoding the AAV Cap protein is upstream of the nucleic acid sequence encoding the AAV Rep protein.

[0412] Embodiment 73: The nucleic acid of any one of Embodiments 70-72, wherein the nucleic acid sequence encoding the AAV Cap protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4A and / or any chimeras thereof.

[0413] Embodiment 74: The nucleic acid of any one of Embodiments 70-73, wherein AAV Rep protein and the AAV Cap protein are from the same AAV serotype.

[0414] Embodiment 75: The nucleic acid of any one of Embodiments 70-73, wherein AAV Rep protein and the AAV Cap protein are from different AAV serotypes.

[0415] Embodiment 76: The nucleic acid of any one of Embodiments 70-75, wherein the nucleic acid sequence encoding the AAV Rep protein is upstream of the nucleic acid sequence encoding the AAV Cap protein.

[0416] Embodiment 77: The nucleic acid of any one of Embodiments 70-75, wherein the nucleic acid sequence encoding the AAV Rep protein is downstream of the nucleic acid sequence encoding the AAV Cap protein.

[0417] Embodiment 78: The nucleic acid of any one of 58-77, wherein the large stuffer does not comprise a nucleotide sequence of mammalian origin.

[0418] Embodiment 79: The nucleic acid of any one of 58-78, wherein the large stuffer comprises a nucleotide sequence of non-mammalian origin.

[0419] Embodiment 80: The nucleic acid of any one of 58-79, wherein the large stuffer is synthetic.

[0420] Embodiment 81: The nucleic acid of any one of 58-80, wherein the large stuffer does not comprise more than one of the following: (a) a transcription factor binding site; (b) a regulatoryelement; (c) an AAV Rep binding site; (d) a donor or acceptor splicing site; (e) an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; (f) a repetitive or palindrome sequence longer than 5 nucleotides; (g) a strong secondary structure; and / or (h) a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0421] Embodiment 82: The nucleic acid of any one of Embodiments 58-81, wherein the large stuffer comprises a GC content of less than about 50%, e.g., less than about 45%, or less than about 40%.

[0422] Embodiment 83: The nucleic acid of any one of Embodiment 58-82, wherein the nucleic acid is larger than 5.5kb.

[0423] Embodiment 84: The nucleic acid of any one of Embodiments 58-83, wherein the nucleic acid comprises the sequence MAG, where M is A or C, upstream of the large stuffer.

[0424] Embodiment 85: The nucleic acid of any one of Embodiments 58-84, wherein the nucleic acid is a vector.

[0425] Embodiment 86: The nucleic acid of any one of Embodiments 58-85, wherein the nucleic acid is a plasmid.

[0426] Embodiment 87: The nucleic acid of any one of Embodiments 58-86, wherein the nucleic acid is linear DNA.

[0427] Embodiment 88: The nucleic acid of any one of Embodiments 58-87, wherein the nucleic acid is closed ended linear duplexed DNA (clDNA).

[0428] Embodiment 89: A host cell comprising a nucleic acid of any one of Embodiments 58-88.

[0429] Embodiment 90: The host cell of Embodiment 89, wherein the host cell is an insect cell or a mammalian cell, optionally the mammalian cell is a HEK293 cell or a HeLa cell.

[0430] Embodiment 91: Use of a nucleic acid of any one of Embodiments 59-88 or a host cell of Embodiment 89 or 90 in a method of preventing manufacturing of a replication competent rAAV.

[0431] Embodiment 92: A nucleic acid comprising a nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10 and a large stuffer of a size at least 2 kb.

[0432] Embodiment 93: The nucleic acid of Embodiment 92, wherein the nucleic acid comprises one or more nucleotides between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10.

[0433] Embodiment 94: The nucleic acid of Embodiment 93, wherein the nucleic acid comprises from about 10 to about 10,000 nucleotides between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10.

[0434] Embodiment 95: The nucleic acid of Embodiment 94, wherein the nucleic acid comprises from about 2,000 to about 5,000 nucleotides between positions 44 and 45 of the nucleotidesequence having a having at least 85% identity to SEQ ID NO: 10.

[0435] Embodiment 96: The nucleic acid of any one of Embodiments 92-95, wherein the large stuffer is located between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10.

[0436] Embodiment 97: The nucleic acid of any one of Embodiments 92-96, wherein the large stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 7-9.

[0437] Embodiment 98: The nucleic acid of any one of Embodiments 92-97, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter.

[0438] Embodiment 99: The nucleic acid of Embodiment 98, wherein the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10 is located in the nucleic acid sequence encoding the AAV Rep protein.

[0439] Embodiment 100: The nucleic acid of Embodiment 98 or 99, wherein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located in an intron in the nucleic acid sequence encoding an AAV Rep protein.

[0440] Embodiment 101: The nucleic acid of any one of Embodiments 98-100, wherein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is upstream of the promoter.

[0441] Embodiment 102: The nucleic acid of any one of Embodiments 98-100, wherein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is downstream of the promoter.

[0442] Embodiment 103: The nucleic acid of any one of Embodiments 98-102, wherein the promoter is a p19 promoter.

[0443] Embodiment 104: The nucleic acid of any one of Embodiments 98-103, wherein the AAV Rep is large Rep (Rep68).

[0444] Embodiment 105: The nucleic acid of any one of Embodiments 98-104, wherein the AAV Rep protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

[0445] Embodiment 106: The nucleic acid of any one of Embodiments 98-105, wherein the AAV Rep protein is a modified Rep protein.

[0446] Embodiment 107: The nucleic acid of any one of Embodiments 98-106, wherein thenucleic acid further comprises a nucleic acid sequence encoding an AAV Cap protein.

[0447] Embodiment 108: The nucleic acid of Embodiment 107, wherein the nucleic acid sequence encoding the AAV Cap protein is downstream of the nucleic acid sequence encoding the AAV Rep protein.

[0448] Embodiment 109: The nucleic acid of Embodiment 107, wherein the nucleic acid sequence encoding the AAV Cap protein is upstream of the nucleic acid sequence encoding the AAV Rep protein.

[0449] Embodiment 110: The nucleic acid of any one of Embodiments 107-109, wherein the nucleic acid sequence encoding the AAV Cap protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV2 3xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

[0450] Embodiment 111: The nucleic acid of any one of Embodiments 107-110, wherein AAV Rep protein and the AAV Cap protein are from the same AAV serotype.

[0451] Embodiment 112: The nucleic acid of any one of Embodiments 107-110, wherein AAV Rep protein and the AAV Cap protein are from different AAV serotypes.

[0452] Embodiment 113: The nucleic acid of any one of Embodiments 92-112, which is larger than 5.5 kb.

[0453] Embodiment 114: The nucleic acid of any one of 92-113, wherein the large stuffer does not comprise a nucleotide sequence of mammalian origin.

[0454] Embodiment 115: The nucleic acid of any one of 92-114 wherein the large stuffer comprises a nucleotide sequence of non-mammalian origin.

[0455] Embodiment 166: The nucleic acid of any one of 92-115, wherein the large stuffer is synthetic.

[0456] Embodiment 117: The nucleic acid of any one of 92-116, wherein the large stuffer does not comprise more than one of the following: (a) a transcription factor binding site; (b) a regulatory element; (c) an AAV Rep binding site; (d) a donor or acceptor splicing site; (e) an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; (f) a repetitive or palindrome sequence longer than 5 nucleotides; (g) a strong secondary structure; and / or (h) a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

[0457] Embodiment 118: The nucleic acid of any one of Embodiments 92-117, wherein the largestuffer comprises a GC content of less than about 50%, e.g., less than about 45%, or less than about 40%.

[0458] Embodiment 119: The nucleic acid of any one of Embodiment 92-118, wherein the nucleic acid is larger than 5.5kb.

[0459] Embodiment 120: The nucleic acid of any one of Embodiments 92-119, wherein the 5’- end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to the sequence MAG, where M is A or C.

[0460] Embodiment 121: The nucleic acid of any one of Embodiments 92-120, wherein the 3’- end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to A or G.

[0461] Embodiment 122: The nucleic acid of any one of Embodiments 92-121, wherein the large stuffer is located downstream of the sequence MAG, where M is A or C.

[0462] Embodiment 123: The nucleic acid of any one of Embodiments 92-122, wherein the nucleic acid comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 11-13.

[0463] Embodiment 124: The nucleic acid of Embodiment 123, wherein 5’-end of the nucleotide sequence having at least 85% identity to any of SEQ ID NOs: 11-13 is linked to the sequence MAG, where M is A or C.

[0464] Embodiment 125: The nucleic acid of Embodiment 123 or 124, wherein the 3’-end of the nucleotide sequence having at least 85% identity to any of SEQ ID NOs: 11-13 is linked to A or G.

[0465] Embodiment 126: The nucleic acid of any one of Embodiments 92-125, wherein the nucleic acid comprises a nucleotide sequence having at least 85% identity to SEQ ID NO: 11.

[0466] Embodiment 127: The nucleic acid of any one of Embodiments 92-126, wherein the nucleic acid prevents production of replication competent rAAV vector.

[0467] Embodiment 128: The nucleic acid of any one of Embodiments 92-127, wherein the nucleic acid is a vector.

[0468] Embodiment 129: The nucleic acid of any one of Embodiments 92-128, wherein the nucleic acid is a plasmid.

[0469] Embodiment 130: The nucleic acid of any one of Embodiments 92-129, wherein the nucleic acid is linear DNA.

[0470] Embodiment 131: The nucleic acid of any one of Embodiments 92-130, wherein the nucleic acid is closed ended linear duplexed DNA (clDNA).

[0471] Embodiment 132: A host cell comprising a nucleic acid of any one of Embodiments 92- 131.

[0472] Embodiment 133: The host cell of Embodiment 132, wherein the host cell is an insect cell or a mammalian cell, optionally the mammalian cell s a HEK293 cell or a HeLa cell.

[0473] Embodiment 134: Use of a nucleic acid of any one of Embodiments 92-131 or a host cell of any one of Embodiments 132-133 in a method of preventing manufacturing of a replication competent rAAV.

[0474] Embodiment 135: A method of detecting residual DNA in a population of viral particles , the method comprising using at least one oligonucleotide sequence that anneals to any one of the nucleic acid sequences selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and SEQ ID NO:6.

[0475] Embodiment 136: The method of Embodiment 135, the method comprises using at least two oligonucleotide sequences that anneal to any one of the nucleic acid sequences selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and SEQ ID NO:6.

[0476] Embodiment 137: The method of Embodiment 136, wherein the oligonucleotide is about 20 nucleotides long.

[0477] Embodiment 138: The method of Embodiment 136, wherein the oligonucleotide is about 25 nucleotides long. Selected definitions

[0478] Unless otherwise defined herein, scientific and technical terms used in connection with the present application shall have the meanings that are commonly understood by those of ordinary skill in the art. The terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is defined solely by the claims. Definitions of common terms in cell biology, immunology, and molecular biology can be found in The Merck Manual of Diagnosis and Therapy, 20th Edition, published by Merck Sharp & Dohme Corp., 2018 (ISBN 0911910190, 978-0911910421); Robert S. Porter et al. (eds.), The Encyclopedia of Molecular Cell Biology and Molecular Medicine, published by Blackwell Science Ltd., 1999-2012 (ISBN 9783527600908); and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 1-56081-569-8); Immunology by Werner Luttmann, published by Elsevier, 2006; Janeway's Immunobiology, Kenneth Murphy, Allan Mowat, Casey Weaver (eds.), W. W. Norton & Company, 2016 (ISBN 0815345054, 978-0815345053); Lewin's Genes XI, published by Jones & Bartlett Publishers, 2014 (ISBN-1449659055); Michael Richard Green and Joseph Sambrook, Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., USA (2012) (ISBN 1936113414); Davis et al., Basic Methods in MolecularBiology, Elsevier Science Publishing, Inc., New York, USA (2012) (ISBN 044460149X); Laboratory Methods in Enzymology: DNA, Jon Lorsch (ed.) Elsevier, 2013 (ISBN 0124199542); Current Protocols in Molecular Biology (CPMB), Frederick M. Ausubel (ed.), John Wiley and Sons, 2014 (ISBN 047150338X, 9780471503385), Current Protocols in Protein Science (CPPS), John E. Coligan (ed.), John Wiley and Sons, Inc., 2005; and Current Protocols in Immunology (CPI) (John E. Coligan, ADA M Kruisbeek, David H Margulies, Ethan M Shevach, Warren Strobe, (eds.) John Wiley and Sons, Inc., 2003 (ISBN 0471142735, 9780471142737), the contents of which are all incorporated by reference herein in their entireties.

[0479] Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular.

[0480] In some embodiments of any of the aspects, the disclosure described herein does not concern a process for cloning human beings, processes for modifying the germ line genetic identity of human beings, uses of human embryos for industrial or commercial purposes or processes for modifying the genetic identity of animals which are likely to cause them suffering without any substantial medical benefit to man or animal, and also animals resulting from such processes.

[0481] Groupings of alternative elements or embodiments disclosed herein are not to be construed as limitations. Each group member can be referred to and claimed individually or in any combination with other members of the group or other elements found herein. One or more members of a group can be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is herein deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.

[0482] The abbreviation, "e.g.," is derived from the Latin exempli gratia, and is used herein to indicate a non-limiting example. Thus, the abbreviation "e.g.," is synonymous with the term "for example."

[0483] Other than in the operating examples, or where otherwise indicated, all numbers expressing quantities of ingredients or reaction conditions used herein should be understood as modified in all instances by the term “about.” The term “about” when used to describe the present invention, in connection with percentages means ^1%.

[0484] The term “variant,” when used in the context of a polynucleotide sequence, may encompass a polynucleotide sequence related to a wild type gene. This definition may also include, for example, “allelic,” “splice,” “species,” or “polymorphic” variants. A splice variant may have significant identity to a reference molecule, but will generally have a greater or lesser number of polynucleotides due to alternate splicing of exons during mRNA processing. The corresponding polypeptide may possess additional functional domains or an absence of domains. Species variantsare polynucleotide sequences that vary from one species to another. Of particular utility in the technology are variants of wild type gene products. Variants may result from at least one mutation in the nucleic acid sequence and may result in altered mRNAs or in polypeptides whose structure or function may or may not be altered. Any given natural or recombinant gene may have none, one, or many allelic forms. Common mutational changes that give rise to variants are generally ascribed to natural deletions, additions, or substitutions of nucleotides. Each of these types of changes may occur alone, or in combination with the others, one or more times in a given sequence.

[0485] The term "nucleic acid" as used herein typically refers to an oligomer or polymer (preferably a linear polymer) of any length composed essentially of nucleotides. A nucleotide unit commonly includes a heterocyclic base, a sugar group, and at least one, e.g., one, two, or three, phosphate groups, including modified or substituted phosphate groups. Heterocyclic bases may include inter alia purine and pyrimidine bases such as adenine (A), guanine (G), cytosine (C), thymine (T) and uracil (U) which are widespread in naturally-occurring nucleic acids, other naturally-occurring bases (e.g., xanthine, inosine, hypoxanthine) as well as chemically or biochemically modified (e.g., methylated), non-natural or derivatised bases. Sugar groups may include inter alia pentose (pentofuranose) groups such as preferably ribose and / or 2-deoxyribose common in naturally-occurring nucleic acids, or arabinose, 2-deoxyarabinose, threose or hexose sugar groups, as well as modified or substituted sugar groups. Nucleic acids as intended herein may include naturally occurring nucleotides, modified nucleotides or mixtures thereof. A modified nucleotide may include a modified heterocyclic base, a modified sugar moiety, a modified phosphate group or a combination thereof. Modifications of phosphate groups or sugars may be introduced to improve stability, resistance to enzymatic degradation, or some other useful property. The term "nucleic acid" further preferably encompasses DNA, RNA and DNA RNA hybrid molecules. A "nucleic acid" can be double-stranded, partly double stranded, or single-stranded. Where single-stranded, the nucleic acid can be the sense strand or the antisense strand. In addition, nucleic acid can be circular or linear.

[0486] A variant amino acid or DNA sequence can be at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more, identical to a native or reference sequence. The degree of homology (percent identity) between a native and a mutant sequence can be determined, for example, by comparing the two sequences using freely available computer programs commonly employed for this purpose on the world wide web (e.g., BLASTp or BLASTn with default settings).

[0487] The terms "identity" and "identical" and the like refer to the sequence similarity between two polymeric molecules, e.g., between two nucleic acid molecules, such as between two DNA molecules. Sequence alignments and determination of sequence identity can be done, e.g., usingthe Basic Local Alignment Search Tool (BLAST) originally described by Altschul et al. 1990 (J Mol Biol 215: 403-10), such as the "Blast 2 sequences" algorithm described by Tatusova and Madden 1999 (FEMS Microbiol Lett 174: 247-250).

[0488] Methods for aligning sequences for comparison are well-known in the art. Various programs and alignment algorithms are described in, for example: Smith and Waterman (1981) Adv. Appl. Math.2:482; Needleman and Wunsch (1970) J. Mol. Biol.48:443; Pearson and Lipman (1988) Proc. Natl. Acad. Sci. U.S.A.85:2444; Higgins and Sharp (1988) Gene 73:237-44; Higgins and Sharp (1989) CABIOS 5:151-3; Corpet et al. (1988) Nucleic Acids Res.16:10881-90; Huang et al. (1992) Comp. Appl. Biosci. 8:155-65; Pearson et al. (1994) Methods Mol. Biol. 24:307-31; Tatiana et al. (1999) FEMS Microbiol. Lett. 174:247-50. A detailed consideration of sequence alignment methods and homology calculations can be found in, e.g., Altschul et al. (1990) J. Mol. Biol.215:403-10.

[0489] The National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST™; Altschul et al. (1990)) is available from several sources, including the National Center for Biotechnology Information (Bethesda, MD), and on the internet, for use in connection with several sequence analysis programs. A description of how to determine sequence identity using this program is available on the internet under the "help" section for BLAST™. For comparisons of nucleic acid sequences, the "Blast 2 sequences" function of the BLAST™ (Blastn) program may be employed using the default parameters. Nucleic acid sequences with even greater similarity to the reference sequences will show increasing percentage identity when assessed by this method. Typically, the percentage sequence identity is calculated over the entire length of the sequence.

[0490] For example, a global optimal alignment is suitably found by the Needleman-Wunsch algorithm with the following scoring parameters: Match score: +2, Mismatch score: -3; Gap penalties: gap open 5, gap extension 2. The percentage identity of the resulting optimal global alignment is suitably calculated by the ratio of the number of aligned bases to the total length of the alignment, where the alignment length includes both matches and mismatches, multiplied by 100.

[0491] The nucleic acids described herein can be synthetic. “Synthetic” in the present application means a nucleic acid molecule that does not occur in nature. Synthetic nucleic acid are produced artificially, typically by recombinant technologies. Such synthetic nucleic acids may contain naturally occurring sequences (e.g., promoter, enhancer, intron, and other such regulatory sequences), but these are present in a non-naturally occurring context. For example, a synthetic gene (or portion of a gene) typically contains one or more nucleic acid sequences that are not contiguous in nature (chimeric sequences), and / or may encompass substitutions, insertions, anddeletions and combinations thereof. The term “synthetic promoter” as used herein relates to a promoter that does not occur in nature.

[0492] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.” It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include the plural reference unless the context clearly dictates otherwise.

[0493] As used herein, the terms “comprising,” “comprise” or “comprised,” and variations thereof, in reference to defined or described elements of an item, composition, apparatus, method, process, system, etc. are meant to be inclusive or open ended, permitting additional elements, thereby indicating that the defined or described item, composition, apparatus, method, process, system, etc. includes those specified elements--or, as appropriate, equivalents thereof--and that other elements can be included and still fall within the scope / definition of the defined item, composition, apparatus, method, process, system, etc.

[0494] As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the content clearly dictates otherwise.

[0495] Specific elements of any of the foregoing embodiments can be combined or substituted for elements in other embodiments. Furthermore, while advantages associated with certain embodiments of the disclosure have been described in the context of these embodiments, other embodiments may also exhibit such advantages, and not all embodiments need necessarily exhibit such advantages to fall within the scope of the disclosure.

[0496] It should be understood that this invention is not limited to the particular methodology, protocols, and reagents, etc., described herein and as such may vary. The terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is defined solely by the claims. EXAMPLES

[0497] The following non-limiting examples are provided for illustrative purposes only in order to facilitate a more complete understanding of representative embodiments now contemplated. Example 1: Helper adenovirus plasmids

[0498] The stuffer sequences SEQ ID NO: 1 and SEQ ID NO: 4 can be included in a nucleic acid as described herein, e.g., a pxx85 sequence as described herein. When Stuffers 2 andStuffers 7 are included in such a sequence, they will not induce gene expression, e.g., of the E2 or E4 proteins.

[0499] Adenovirus 5 (hAd5) based nucleic acids described herein, e.g., XX85, further comprising Stuffer 2 and / or 7 (SEQ ID NOs: 1 and 4) can be hydrodynamically injected in mice and expression of E2 and / or E4 in the liver can be measured. Plasmids containing either one of the stuffer sequences can be administered via hydrodynamic tail vein injection to 7-week-old C57BL / 6JOlaHsd male mice. Twenty-four hours after the injection, animals can be euthanized, and gene expression quantified (mRNA levels and protein levels) in liver. Plasmid copy number (PCN) can also be determined in liver to normalize the levels of expression. The level of gene and protein expression from plasmids comprising one or both Stuffers will be similar to the levels observed in a control construct with no stuffer.

[0500] SEQ ID NO: 1 and / or SEQ ID NO: 4 can be included in a nucleic acid sequence described herein e.g., XX85 further comprising the protelomerase sites. In particular, stuffer sequences 2 and / or 7 (SEQ ID NOs: 1 and 4) can be included upstream of the 5’ end of the E4 region (i.e. a 5’ stuffer) and / or downstream of the 3’ end of the E2A region (i.e. a 3’ stuffer). Further, a 5’ stuffer can be located at the AscI site 5’ of the E4 region and / or a 3’ stuffer can be located at the NotI site 3’ of the E2A region. These locations are shown schematically in FIG. 29. Nucleic acids comprising one or both stuffers, in a 5’ and / or 3’ position can be tested, e.g., as shown in the following table. Nucleic acids comprising only a 5’ Stuffer can display superior performance, and nucleic acids comprising only a 5’ Stuffer 7 can display particularly superior performance. For example, XX85 Ad helper nucleic acid further comprising a stuffer 7 at the 5’ end will produce rAAV with higher titer when compared to rAAV produced with Stuffer 7 at 5’ end and stuffer 2 at 3’ end.

[0501] Adeno-associated virus is a dependovirus, that naturally needs co-infection with a helper virus, such as adenovirus or herpes virus, to efficiently complete its life cycle. However, only a small portion of the helper virus genes are needed to achieve efficient AAV genome replication and encapsidation (see Meier et al. Viruses 2020 for review).

[0502] Recombinant AAV vectors used for gene transfer application are made of a DNA sequence of interest (the transgene) flanked by AAV inverted terminal repeats (ITR), which are packaged into an AAV capsid made of AAV structural proteins. The widely used method for rAAV production consists in transfection of HEK293 cells with three plasmids, the first one containing the transgene flanked by AAV ITR, the second one encoding the AAV Rep and capsid proteins, and the third one encoding the adenovirus type 5 (Ad5) helper proteins, e.g., VA RNA (virus- associated small non-coding RNA), E2A (single-stranded DNA binding protein), and E4 helper functions, where the host cell e.g., Pro10 cells provide the E1 helper function. In certain methods, rAAV is manufactured using two nucleic acids instead of three nucleic acids, e.g, one single nucleic acid encodes AAV helper Rep-Cap gene, and hAd5 based helper nucleic acid of the invention. In this instance, rAAV is manufactured using i) one single nucleic acid encoding Ad helper function and AAV helper Rep-Cap gene, and ii) rAAV genome (e.g, AAV ITR to ITR encompassing transgene,) e.g as the method with two plasmids described in European Patent EP1412510B1.

[0503] Adenovirus helper plasmids available up to now are the following (not necessarily exhaustive list):

[0504] pXX6-80, which was created through successive deletions of the Ad5 genome, to maintain expression of the small virus-associated RNA (VA RNA I and VA RNA II), the single-stranded DNA binding protein (E2A or DBP) and the E4 proteins, while avoiding expression of the adenovirus capsid structural proteins (Xiao et al. J Virol 1998). Thus, the plasmid keeps the relative orientation of the different sequence elements from the Ad5 viral genome. It contains the following nucleotides positions of Ad5 (RefSeq # AC_000008): 9847-13258, 21444-28119, and 30819- 35919. One drawback of this plasmid is its large size (almost 19 kb), which leads to significant cost for production and quality control. In addition, this plasmid still contains still contains the full- length coding sequence of some structural proteins, in particular the adenovirus fiber which could lead to toxicity and immune response if expressed in treated patients. In a non limiting example of the invention, the Adenoviral helper XX-680 plasmid or, XX-680 neDNA comprises the short stuffer sequences SEQ ID NOs: 1-6, or, 85% identity thereto, or, comprises a nucleotide sequence of a contiguous fragment from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NOs: 7-9.

[0505] pXX85, which was recently developed at Asklepios BioPharmaceutical, Inc has a smaller size (10.6 kb), and it doesn’t contain the fiber sequence. The relative orientation of the Ad5sequence elements has been also changed in this plasmid compared to the wild-type Ad5 viral genome. It contains the following nucleotides positions of Ad5 (RefSeq # AC_000008): 32721- 35915, 10562-11072, and 22316-27173. In a non-limiting example of the invention, the Adenoviral helper XX-85 plasmid (SEQ ID NO: 69) or, XX-85 clDNA (SEQ ID NO: 70) comprises the short stuffer sequences SEQ ID NOs: 1-6, or, 85% identity thereto, or, comprises a nucleotide sequence of a contiguous fragment of from about 75 nucleotides to about 250 nucleotides of any one of SEQ ID NOs: 7-9. Example 2: Short stuffers used in the molecules used for recombinant AAV (rAAV) production and detecting residual DNA in the viral preparation:

[0506] An objective of this study was to prepare new stuffer sequences in order to confirm the inert nature of the stuffer candidates, both in vitro and in vivo. rAAV was produced by transfecting cells with i) the nucleic acid encoding rAAV genome,e.g, neDNA molecule comprising transgene sequence flanked by AAV ITRs ii) a Adenoviral helper nucleic acid e.g, neDNA molecule comprising sequence encoding adenoviral helper proteins, iii) helper nucleic acid e.g, neDNA molecule comprising sequence encoding AAV capsid and non-structural replication genes, allowing cells sufficient time to produce rAAV particles. Residual neDNA in the rAAV preparations was measured by a qPCR or ddPCR assay designed to target the stuffer sequence.

[0507] Method for the detection of residual neDNA stuffer sequences in purified AAV preparations.

[0508] Ten microliters of purified AAV sample are first diluted in 1.25X DNase I Incubation Buffer (Roche) and 0.0625% Pluronic F-68 in a final volume of 100 µL. Fifty microliters of diluted sample are added to 50 µL of 2 mM Tris pH8.0, 2 mM EDTA, 0.2% SDS, 1 mg / mL proteinase K, and the 100 µL reaction is incubated for 1 hour at 55˚°C followed by 10 minutes at 80°C. The proteinase K-digested sample is then submitted to six 10-fold dilutions (fom 1 / 10 to 1 / 106) in ddPCR dilution buffer, which is made with 1X GeneAmp PCR buffer I (Applied Biosystems) containing 0.05% Pluronic F-68 and 2 µg / mL sheard salmon sperm DNA. Five microliters of each dilution are finally analyzed by ddPCR in a 25 µL reaction volume. The ddPCR reaction is performed in 1X ddPCR Supermix for Probes (no dUTP), containing 450 nM of each forward and reverse primer, 250 nM of probe, and 0.04 µL of FastDigest BglII restriction endonuclease (Thermo Scientific). Sequences and parameters of the primers and probes are shown in Table 3 below. The ddPCR reaction cycling parameters are provided in Table 4 below. A positive control is prepared with a neDNA containing both stuffer 2 and stuffer 7 sequences,digested with BglII restriction endonuclease, diluted at a concentration of 1.0E+7 copies / µL in AAV formulation buffer. A negative control is prepared with the AAV formulation buffer alone. Each sample, positive control and negative control is processed in quadruplicate. A no template control is prepared with the ddPCR dilution buffer alone.

[0509] The acceptance criteria for the assay are the following: ‐ Negative control for extraction and no template control: ≤ 5 copies / uL. ‐ Positive control recovery: 50 – 150%

[0510] Method for the design of ddPCR primers and probes. The specifications used for the primers are the following: ^ GC content: 40-60% ^ Sequence length: 18-30 bp, optimally 20 bp ^ Absence of hairpin secondary structure ^ No more than 2 G and / or C at the 3’end ^ No more than 3 successive identical nucleotide ^ Tm difference between both primers of maximum 1°C The specification used for the probes are the following: ^ GC content: 20-80% ^ Sequence length: maximum 30 bp, optimally 15-18 bp ^ Tm 3-10°C higher than the Tm of the primers ^ Probe location close to one primer on the same DNA strand ^ No more than 4 successive G ^ No G at the 5’end ^ Avoid hairpin secondary structure

[0511] For selection of the primers and probes sequences, the software used is PrimerQuest Tool, and for analysis of potential secondary structures, the software used is OligoAnalyzer, both from Integrated DNA Technologies (IDT).

[0512] The primers sequences targeting stuffer sequence #2 and stuffer sequence #7 found in Table 3 were used to detect residual DNA in a rAAV preparation. In FIG.15, the residual neDNA and pDNA are in copies / mL and % relative to ITR ddPCR titers in purified AAV vectors. Residual neDNA was detected by ddPCR using primers and probes targeting the stuffer 2, stuffer 7, 5’TelRS or 3’telRS sequence, depending on the neDNA used for AAV production. Residual pDNA was detected by ddPCR using primers and probe targeting the kanamycin-resistance coding sequence.

[0513] Amplification of neDNA precursor plasmid: Glycerol stocks containing neDNA precursor plasmid of the new stuffer candidates and controls were used to amplify DNA. Precursor plasmid was produced at Mega Scale (Qiagen) and following plR002 manufacturing instructions.

[0514] Amplification data of precursor plasmid containing the new stuffer sequences and controls are summarized in Table 5.Table 5. neDNA precursor plasmids amplified in this work.

[0515] In order to confirm precursor plasmid identity, samples were sent for sequencing to Macrogen Inc. (Spain). All products showed 100% sequencing coverage and sequences matched with reference molecule map. In conclusion, all the precursor plasmids were sequenced successfully matching the reference sequence (data not shown).

[0516] Manufacture of neDNA containing new stuffer sequences were performed at a 20 ml scale using the polyethylene glycol (PEG) purification protocol as described in U.S. Patent No. 9,499,847; U.S. Patent No.10,501,782; and International Publication no. WO 2018033730 A1; all of which are herein incorporated by reference in their entireties.

[0517] Briefly, manufacturing was initiated with a small quantity of a neDNAprecursor plasmid that was amplified by an in vitro dual enzyme process: (i) Phi29 polymerase was utilized for rolling cycle amplification (RCA), and (ii) the protelomerase TelN was utilized for covalent closure of both ends. The enzymatic process also includes restriction enzyme digest (ApaLI was used in this process) and exonucleases to remove residual DNA. In this case, PEG purification was used to clear enzymes and DNA residuals from the final product.

[0518] neDNA manufacture data are collected in Table 6. Table 6. neDNA productions done for this work.and the final product was lost.

[0519] Results on the characterization of the neDNA: Characterization of the neDNA molecules was performed by confirming the identity through (i) agarose gel electrophoresis and (ii) Sanger sequence analysis using Macrogen Inc. (Spain). In conclusion, all neDNA molecules containing the stuffer sequences and the controls showed the expected band size by agarose gel electrophoresis, and 100% match with their respective reference molecule map. Table 7 summarizes the characterization results obtained for the neDNA produced for this work.

[0520] In order to examine the luciferase activity from the stuffer-luciferase neDNA molecules, they were transfected into HEK293 cells in 24 well plates (schematic in FIG. 4) with triplicate transfections using LIPOFECTAMINE 3000. Cells were harvested at 48 hours for either luciferase activity measurement (FIG. 5) or RNA quantification (FIG. 6). Experiments were performed blindly with two different operators. Compared to the positive control, there was relatively little to no luciferase enzymatic activity or RNA quantification 48 hours post transfection.

[0521] Examining the use of stuffer sequences in C57BL / 6JOlaHsd mice, FIG. 7 shows the breakdown of how much stuffer sequence was injected into each mouse and the different stuffer sequences and controls that were used. In order to answer the question if the stuffer sequences provide any mRNA or protein activity, all the stuffer sequences drive a similar level of luciferase expression (mRNA) compared with the construct without stuffer. This level is around 5000 times lower vs that driven by the CMV promoter. As shown in FIG. 8, mRNA expression is low but above background (Ct values from 30 to 35). No luciferase protein expression was detected in the groups containing no promoter.

[0522] In order to answer the question if TelRL sequence is driving mRNA expression, the expression of luciferase mRNA in the animals that received DNA injection is substantially above background (mice that received vehicle showed no expression), which is shown in FIG. 9. There are no clear differences between stuffers in terms of expression, even normalized per plasmid copy number. Stuffers do NOT seem to add promoter activity to the cassette, in fact they seem to reduceexpression compared with the neDNA without stuffer. The TelRL sequence seems to drive a low level of expression in liver.

[0523] The ddPCR assay has been optimized in order to detected residual DNA. This has resulted in the selection of two lead candidate stuffer sequences (stuffer sequences #2 and #7 (SEQ ID NOs. 1 and 4)) which have been cloned into a backbone in order to test the manufactuability of neDNA precursors and AAVs. As shown in FIG.10, 4 different AAV products will be tested in an in vivo validation study as AAVs of stuffers #2 and #7 using pJAL130 and pCHATAM as references. We expect to see optimal manufacturability and productivity of the 4 different AAV products. Materials & Methods

[0524] Table 8 is a list of different neDNA candidates. Table 12 is a list of primer and probe set used for qPCR. Table 8: neDNA stuffer candidatesTable 9: Primer and Probes sets used for qPCR

[0525] All the experiments were conducted by two operators and blinded, to avoid sample bias.

[0526] DNA transfections: HEK293 cells were transfected with different amounts of DNA (based on mass or molecular weight) in combination with two amounts of lipofectamine 3000 (0.75 or 1.50 μL) following the manufacturer´s protocol (Invitrogen, Cat# L30000001). Briefly, two different master mixes were prepared. Master Mix 1 (MM1) contained Opti-MEMTM mixed withLipofectamine 3000 Reagent,and Master Mix 2 (MM2) contained the DNA mixed with Opti- MEMTM and P3000 Reagent. MM1 and MM2 were mixed and incubated at room temperature (RT) for 15 minutes to allow DNA-Lipofectamine complex formation. The mixture was added to the HEK293 cells covered with fresh complete DMEM (media with 10% Fetal Bovine Serum and Penicillin / Streptomycin). The transfected cells were incubated for 48 hours at 37°C, and then harvested to proceed with the different analyses.

[0527] RNA extraction: Cells were harvested by up and down pipetting and the total volume of media and cells was transferred to 1.5 mL tubes. Cells were pelleted at 100 x g for 5 min at 4ºC. Supernatants were discarded and cellular RNA were extracted using the Maxwell® RSC simplyRNA Tissue Kit (Promega, Cat.# AS1340). Briefly, cell pellets were homogenized with 200 µL of homogenate solution, then 200 µL of Lysis Buffer were added, and the lysates were loaded into Maxwell® RSC cartridges. The RNA were eluted with 50 µL Nuclease Free Water (NFW), and quantified based on A260 using the software Gen5 in the Synergy HTX Multi-Mode Reader.

[0528] Retro-transcription: (i) DNase Treatment: RNA samples were DNAse-treated using the TURBO DNA-free™ Kit (Invitrogen, Cat.# AM1907). Briefly, 1 μg of RNA was digested with TURBO DNase in a 50 µL reaction volume for 25 minutes at 37ºC . Five microliters of DNase Inactivation Reagent were then added to the samples and incubated for 5 minutes at room temperature. Finally, the samples were centrifuged to remove the DNase Inactivation Reagent and transferred to a 96-well plate; and (ii) Reverse transcription: Reverse transcription was performed with the SuperScript® IV Reverse Transcriptase. 200 ng of RNA samples were mixed with random hexamers and dNTPs, as indicated in Table 10. Table 10: Components of the reaction for the primer-RNA annealing

[0529] The 13 µL reaction volume was incubated for 5 minutes at 65ºC to allow oligonucleotides annealing, then cooled on ice until the next step. The reaction mix containing the SuperScript® IV Reverse Transcriptase was prepared as indicated in Table 11. Table 11: Components of the RT reaction ComponentVolume

[0530] The components of the RT reaction were added to the primer-RNA mixture in a final volume of 20 µL and were then incubated 10 minutes at 23ºC, 10 minutes at 55ºC, and 10 minutes at 80ºC to inactivate the enzyme. The cDNAs were stored at -20ºC.

[0531] Quantitative PCR (qPCR) for human β-actin and Luciferase

[0532] Standard Curve for human β-actin: To generate a standard curve for human β-actin, a 265 bp amplicon was obtained from HEK293 cells cDNA. Briefly, total RNA was extracted with the RNeasy Plus Mini Kit (Qiagen, Cat# 74134), and retrotranscribed with the Superscript One-Step IV system (ThermoFisher, Cat# 12594025). The primers used to amplify the β-actin amplicon are shown in Table 12. Table 12: Primers used to amplify a 265-bp amplicon of b-actin cDNA from HEK293 cells

[0533] The PCR product was purified with a Macherey-Nagel™ Nucleospin™ column, phosphorylated with T4 polynucleotide kinase and cloned at the SmaI site of plasmid pUC19. The selected pUC19_ β-actin clone was confirmed by sequencing.

[0534] For the qPCR standard curve, the pUC19_β-actin plasmid was linearized with EheI, purified using Macherey-Nagel™ Nucleospin™ columns and quantified using the Thermo Fisher Qubit 4. A stock at a concentration of 1E+09 copies / µL was stored at -80ºC. The plasmid map with the qPCR primers / probe and EheI-cutting site is shown in FIG.27.

[0535] Standard Curve preparation for Luciferase: The qPCR standard curve for luciferase was prepared using the pro_Positive Control plasmid, which contains the Luciferase expression cassette under the transcriptional control of the CMV promoter. The plasmid DNA was digested with SdaI and BcuI to remove potential secondary structure. The digested DNA was purified using the Macherey-Nagel™ Nucleospin™ columns and quantified using the Thermo Fisher Qubit 4. Finally, a stock at a concentration of 1E+09 copies / µL was prepared and stored at -80ºC. The plasmid map with the qPCR primers / probe and restriction enzymes sites is shown in FIG.28.

[0536] qPCR Protocol: Five nanograms of the input cDNA were used to perform the qPCR reaction in a total volume of 20 µL. The qPCR mastermix composition and qPCR steps for quantification of human β-actin and luciferase cDNA are described in Tables 13-16.Table 13: Mastermix conditions for Actin-qPCRTable 14: Mastermix conditions for Luc-qPCRTable 15: Human b-actin qPCR conditionsTable 16: Luciferase qPCR conditions

[0537] To calculate the copy number of each target, a standard curve was added in each qPCR plate ranging from 1E+08 to 1E+02 copies / µL.

[0538] Luciferase Activity: Cells were harvested at 48h post-transfection and luciferase gene expression level was measured using the Pierce ® Firefly Luciferase Glow Assay Kit (Thermo Scientific, Cat#16177) following manufacturer´s instructions. Briefly, supernatant was removed from the 24-multiwell plate, cells were washed with 1X PBS, 300 μL of Lysis Buffer were added, and the cells were incubated for 15 min at room temperature with shaking. Then, 15 μL of cell lysate was transferred to a white 96-wells plate, mixed with 50 μL of Working Solution, and incubated for 10 min in darkness. Luciferase activity was measured using the Synergy HTX Multi- Mode Reader with the Gen5software. Results were obtained in relative light units (RLU) and normalized by mg of proteins, as determined by the bicinchoninic acid assay (BCA) assay.

[0539] Protein quantification by BCA: Total proteins concentration was measured using the Pierce BCA Protein Assay Kit (Thermo Scientific, Cat#123225) following manufacturer´s instructions. Briefly, 25 μL of cell lysate was transferred to a flat-bottom clear 96-wells plate, mixed with 200 μL of Working Solution, and incubated for 30 min at 37°C. A BSA standard curve was included in the assay. Protein concentration was measured using the Synergy HTX Multi- Mode Reader with the Gen5 software. Results

[0540] Optimization of transfection conditions: Before testing the neDNA stuffer candidates (Table 8), a preliminary experiment was performed to set up transfection conditions with neDNA. To this end, the neDNA positive control (neDNA-CMV-Luc) and negative control (neDNA-Luc) were transfected based on mass (0.5 μg of each neDNA) or molecular weight (3.85E-07 μmole of each neDNA),using two different volumes of Lipofectamine 3000 (0.75 μl or 1.5 μL), as specified by the manufacturer. Luciferase activity was then measured in transfected cells. The results presented show that luciferase is efficiently expressed from the positive control, and that the negative control results in background luminescence, as well as non-transfected cells. They also show that there is no significant difference between the different transfection conditions.

[0541] In the next experiments, transfections were all performed based on DNA mass (0.5 μg / well) with 0.75 μL of Lipofectamine 3000 to simplify blinded experiments.

[0542] Analysis of Luciferase expression at the protein level: To assess the promoter activity of the different stuffer seq: uences, luciferase gene expression was analyzed by measuring luciferase activity in protein extracts obtained from HEK293 cells transfected with the different neDNA molecules (Table 8). The assay was performed in duplicate by two different operators. As shown, background luciferase activity was measured in cells transfected with all the neDNA molecules containing or not a stuffer sequence, as well as in non-transfected cells, and only the positive control (neDNA-CMV-Luc) shows significant luciferase expression.

[0543] Analysis of Luciferase expression at mRNA level: To further assess potential promoter activity of the different stuffer sequences, luciferase gene expression was quantified at the mRNA level. To this end, luciferase mRNA was measured by RT-qPCR in HEK293 cells transfected with the different neDNA molecules, using human β-actin mRNA as the housekeeping gene for normalized gene expression.

[0544] As shown in FIG. 8, no significant differences were found when comparing luciferase mRNA levels expressed from the different neDNA molecules containing or not a stuffer sequence in transfected HEK293 cells, and only the positive control (neDNA-CMV-Luc) expressed a significantly higher level of luciferase mRNA. Conclusions

[0545] Significant luciferase enzymatic activity was detected only in cells transfected with the positive control neDNA containing a CMV promoter sequence. Only background activity was detected in cells transfected with the neDNA containing the different stuffer sequences, with no significant difference compared to the negative control without stuffer.

[0546] Some levels of luciferase mRNA were detected by RT-qPCR in cells transfected with the neDNA containing the different stuffer sequences, with no significant difference compared to the negative control without stuffer. This might be explained by some low-level transcription of the luciferase sequence in the absence of promoter, or to background promoter activity driven by sequence elements located upstream luciferase (phage N15 telomer sequence, restriction enzyme recognition sites, Kozak translation initiation sequence, …). The levels of luciferase mRNA detected in cells transfected with the positive control neDNA containing a CMV promoter sequence was however much higher (around 2000-3000 fold). See, FIG.5.

[0547] In conclusion, no significant promoter activity was detected in HEK293 cells transfected with neDNA molecules containing the different stuffer sequences. See, FIG.6. neDNA stuffer sequences

[0548] This study was performed to address the regulatory perspective regarding the origin of the DNA stuffers between the telomeric sequence and the ITRs in the neDNAprecursor plasmids and remove any potential off-target effects derived by their presence. This study looked at CpG dinucleotide content, origin of the current stuffer sequences and design criteria and impact of AAV manufacturability due to the presence of potential Rep binding sites.

[0549] The stuffer sequences of the invention were designed that have the following specifications as depicted below: ^ Sequence length ~164 bp^ Motifs / features to avoid: o TATA boxes (TATAWAWR) o CpG dinucleotides o Potential TFBS o Restriction enzymes (BamHI, XbaI, FspI, ApaLI, ClaI, RsrII, DrdI, NdeI, MfeI, SbfI, PacI and SwaI) o No ORFs longer than 10 amino acids o Potential donor or acceptor splicing sites o Repetitive sequences (palindromes,...

Claims

CLAIMS What is claimed is:

1. A nucleic acid comprising a short stuffer sequence, wherein the short stuffer sequence comprises a nucleotide sequence having at least 85% identity to: (i) a nucleotide sequence of any one of SEQ ID NOs: 1-6 or a nucleotide sequence complementary to any one of SEQ ID NO: 1-6; or (ii) a nucleotide sequence of a contiguous fragment of from about 250 nucleotides to about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81 or a nucleotide sequence complementary to a contiguous fragment of from about 250 nucleotides to about 1500 nucleotides of any one of SEQ ID NO: 7-9, 81.

2. The nucleic acid of claim 1, wherein the nucleic acid is linear DNA.

3. The nucleic acid of claim 1 or 2, wherein the nucleic acid is a closed ended linear duplexed DNA (clDNA).

4. The nucleic acid of any one of claims 1-3, wherein the nucleic acid further comprises at least one protelomerase binding site.

5. The nucleic acid of claim 4, wherein the short stuffer is located downstream of the protelomerase binding site.

6. The nucleic acid of claim 4, wherein the short stuffer is located upstream of the protelomerase binding site.

7. The nucleic acid of claim 4, wherein the nucleic acid comprises two protelomerase binding sites and the short stuffer is located between the two protelomerase binding sites.

8. The nucleic acid of any one of claims 1-7, wherein the nucleic acid further comprises a heterologous transgene operably linked to one or more regulatory elements.

9. The nucleic acid of claim 8, wherein the short stuffer is located upstream of the heterologous transgene.

10. The nucleic acid of claim 9, wherein the short stuffer is located downstream of the heterologous transgene.

11. The nucleic acid of any one of claims 8-10 wherein the nucleic acid further comprises at least one adeno-associated virus (AAV) inverted terminal repeat (ITR) sequence.

12. The nucleic acid of claim 11, wherein the short stuffer is upstream of the at least one ITR sequence.

13. The nucleic acid of claim 11, wherein the short stuffer is downstream of the at least one ITR sequence.

14. The nucleic acid of any one of claims 11-13, wherein the at least one ITR sequence is located between the short stuffer and the heterologous transgene.

15. The nucleic acid of any one of claims 10-13, wherein the nucleic acid further comprises at least one protelomerase binding site and the short stuffer is located between the at least one protelomerase binding site and the at least ITR.

16. The nucleic acid of any one of claims 11-15, wherein the nucleic acid comprises at least two ITRs and the short stuffer located outside the two ITRs.

17. The nucleic acid of claim 16, wherein the heterologous transgene is located between the two ITRs.

18. The nucleic acid of claims 16, wherein the heterologous transgene is located between the two ITRs, and one of the ITRs is located between the short stuffer and the heterologous transgene.

19. The nucleic acid of claim 16, wherein the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the heterologous transgene sequence is located between the first and second ITR sequences, and wherein the short stuffer is upstream of the first ITR sequence.

20. The nucleic acid of claim 16, wherein the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the heterologous polynucleotide sequence is located between first and second ITR sequences, wherein the short stuffer is upstream of the first ITR sequence, and wherein the nucleic acid does not comprise a short stuffer downstream of the second ITR sequence.

21. The nucleic acid of claim 16, wherein the nucleic acid comprises a first ITR (e.g., left ITR) sequence and a second ITR (e.g., right ITR) sequence, wherein the short stuffer is located upstream of the 5’-end of the first and second ITRs.

22. The nucleic acid of any one of claims 7-21, wherein each ITR sequence is selected independently from an ITR sequence of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

23. The nucleic acid of claim 22, wherein the ITR sequences are from the same AAV serotype.

24. The nucleic acid of claim 22, wherein the ITR sequences are from the different AAV serotype.

25. The nucleic acid of any one of claims 1-7, wherein the nucleic acid comprises a nucleic acid sequence encoding one or more helper proteins that assist AAV replication in rAAV production.

26. The nucleic acid of claim 25, wherein the nucleic acid comprises at least one protelomerase binding site and wherein the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding one or more helper proteins.

27. The nucleic acid of claim 26, wherein the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding one or more helper proteins.

28. The nucleic acid of claim 26, wherein the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding one or more helper proteins.

29. The nucleic acid of any one of claims 25-28, wherein the helper protein sufficient for AAV replication comprises one or more of an E2A region, an E4 region, and a virus associated (VA) RNA region, and optionally, an E1 region, an E3 region and / or a Major Late Promoter (MLP) region.

30. The nucleic acid of any one of claims 25-29, wherein the nucleic acid sequence encoding one or more helper proteins comprises the nucleotide sequence of any one of SEQ ID NOs: 67-70.

31. The nucleic acid of any one of claims 1-7, wherein the nucleic acid further comprises nucleic acid sequence encoding a AAV rep protein and a AAV cap protein.

32. The nucleic acid of claim 31, wherein the nucleic acid comprises at least one protelomerase binding site and wherein the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

33. The nucleic acid of claim 31, wherein the short stuffer is upstream of the 5’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

34. The nucleic acid of claim 31, wherein the short stuffer is downstream of the 3’-end of the nucleic acid sequence encoding the AAV rep and AAV cap proteins, and the short stuffer is located between the protelomerase binding site and the nucleic acid sequence encoding the AAV rep and AAV cap proteins.

35. The nucleic acid of any one of claims 31-34, wherein AAV Rep protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAVAnc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

36. The nucleic acid of any one of claims 31-35, wherein AAV Cap protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

37. The nucleic acid of any one of claims 31-36, wherein AAV Rep protein and the AAV Cap protein are from the same AAV serotype.

38. The nucleic acid of any one of claims 31-36, wherein AAV Rep protein and the AAV Cap protein are from different AAV serotypes.

39. The nucleic acid of any one of claims 1-6, wherein the nucleic acid comprises an ITR sequence and a protelomerase binding upstream of the ITR sequence, wherein the nucleic acid comprises the sequence of SEQ ID NO: 1 or 4 between the ITR and the protelomerase binding site.

40. The nucleic acid of any one of claims 1-6, wherein the nucleic acid comprises an ITR sequence and a protelomerase binding downstream of the ITR sequence, wherein the nucleic acid comprises the sequence of SEQ ID NO: 1 or 4 between the ITR and the protelomerase binding site.

41. The nucleic acid of any one of claims 1-40, wherein the nucleic acid further comprises a stop codon (e.g., TAA, TAG or TGA) upstream of the short stuffer.

42. The nucleic acid of any one of claims 1-40, wherein the nucleic acid further comprises a stop codon (e.g., TAA, TAG or TGA) downstream of the short stuffer.

43. The nucleic acid of any one of claims 1-42, wherein the nucleic acid is a vector.

44. The nucleic acid of any one of claims 1-43, wherein the nucleic acid is a plasmid.

45. A host cell comprising the nucleic acid of any one of claims 1-44.

46. The host cell of claim 45, wherein the host cell is an insect cell or a mammalian cell.

47. The host cell of claim 46, wherein the mammalian cell is a HEK293 cell or a HeLa cell.

48. Use of a nucleic acid of any one of claims 1-44 in a method of producing a plurality of viral particles.

49. Use of claim 48, wherein the viral particles are recombinant AAV (rAAV) particles.

50. A method for producing a plurality of viral particles, the method comprising culturing a host cell of claim 46 or 47 in a culture medium under conditions in which viral particles are produced.

51. A method for producing a plurality of viral particles, the method comprises culturing a host cell comprising a nucleic acid of any one of claims 1-44 in a culture medium under conditions in which viral particles are produced.

52. The method of claim 50 or 51, wherein the host cell comprises at least one nucleic acid encoding one or more helper proteins sufficient for AAV replication, at least one nucleic acid encoding AAV rep and AAV cap proteins, and at least one nucleic acid encoding a transgene of interest.

53. The method of any one of claims 50-52, wherein the host cell comprises at least one nucleic acid of any one of claims 7-24.

54. The method of any one of claims 50-53, wherein the host cell comprises at least one nucleic acid of any one of claims 25-30.

55. The method of any one of claim 50-54, wherein the host cell comprises at least one nucleic acid of any one of claims 31-38.

56. The method of any one of claims 50-55, wherein the host cell comprises at least one nucleic acid of any one of claims 7-24, at least one nucleic acid of any one of claims 25- 30, and at least one nucleic acid of any one of claims 31-38.

57. The method of any one of claims 50-56, wherein the plurality of viral particles comprises rAAV particles.

58. A nucleic acid comprising a large stuffer, wherein the large stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 7-9 or 81.

59. The nucleic acid of claim 58, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter.

60. The nucleic acid of claim 59, wherein the large stuffer is located in the nucleic acid sequence encoding the AAV Rep protein.

61. The nucleic acid of claim 60, wherein the large stuffer is located in an intron in the nucleic acid encoding the AAV Rep protein.

62. The nucleic acid of any one of claims 59-61, wherein the large stuffer is located upstream of the promoter.

63. The nucleic acid of any one of claims 59-61, wherein the large stuffer is located downstream of the promoter.

64. The nucleic acid of any one of claims 59-63, wherein the promoter is a p19 promoter.

65. The nucleic acid of any one of claims 59-64, wherein the AAV Rep is large Rep.

66. The nucleic acid of any one of claims 59-65, wherein the AAV Rep is Rep68 or Rep78.

67. The nucleic acid of any one of claims 59-66, wherein the AAV Rep is Rep68.

68. The nucleic acid of any one of claims 59-67, wherein the AAV Rep protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

69. The nucleic acid of any one of claims 59-68, wherein the AAV Rep protein is a modified Rep protein.

70. The nucleic acid of any one of claims 59-69, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Cap protein.

71. The nucleic acid of claim 70, wherein the nucleic acid sequence encoding the AAV Cap protein is downstream of the nucleic acid sequence encoding the AAV Rep protein.

72. The nucleic acid of claim 70, wherein the nucleic acid sequence encoding the AAV Cap protein is upstream of the nucleic acid sequence encoding the AAV Rep protein.

73. The nucleic acid of any one of claims 70-72, wherein the nucleic acid sequence encoding the AAV Cap protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9- PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4A and / or any chimeras thereof.

74. The nucleic acid of any one of claims 70-73, wherein AAV Rep protein and the AAV Cap protein are from the same AAV serotype.

75. The nucleic acid of any one of claims 70-73, wherein AAV Rep protein and the AAV Cap protein are from different AAV serotypes.

76. The nucleic acid of any one of claims 70-75, wherein the nucleic acid sequence encoding the AAV Rep protein is upstream of the nucleic acid sequence encoding the AAV Cap protein.

77. The nucleic acid of any one of claims 70-75, wherein the nucleic acid sequence encoding the AAV Rep protein is downstream of the nucleic acid sequence encoding the AAV Cap protein.

78. The nucleic acid of any one of 58-77, wherein the large stuffer does not comprise a nucleotide sequence of mammalian origin.

79. The nucleic acid of any one of 58-78, wherein the large stuffer comprises a nucleotide sequence of non-mammalian origin.

80. The nucleic acid of any one of 58-79, wherein the large stuffer is synthetic.

81. The nucleic acid of any one of 58-80, wherein the large stuffer does not comprise more than one of the following: a. a transcription factor binding site; b. a regulatory element; c. an AAV Rep binding site; d. a donor or acceptor splicing site; e. an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; f. a repetitive or palindrome sequence longer than 5 nucleotides; g. a strong secondary structure; and / or h. a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

82. The nucleic acid of any one of claims 58-81, wherein the large stuffer comprises a GC content of less than about 50%, e.g., less than about 45%, or less than about 40%.

83. The nucleic acid of any one of claim 58-82, wherein the nucleic acid is larger than 5.5kb.

84. The nucleic acid of any one of claims 58-83, wherein the nucleic acid comprises the sequence MAG, where M is A or C, upstream of the large stuffer.

85. The nucleic acid of any one of claims 58-84, wherein the nucleic acid is a vector.

86. The nucleic acid of any one of claims 58-85, wherein the nucleic acid is a plasmid.

87. The nucleic acid of any one of claims 58-86, wherein the nucleic acid is linear DNA.

88. The nucleic acid of any one of claims 58-87, wherein the nucleic acid is closed ended linear duplexed DNA (clDNA).

89. A host cell comprising a nucleic acid of any one of claims 58-88.

90. The host cell of claim 89, wherein the host cell is an insect cell or a mammalian cell, optionally the mammalian cell is a HEK293 cell or a HeLa cell.

91. Use of a nucleic acid of any one of claims 59-88 or a host cell of claim 89 or 90 in a method of preventing manufacturing of a replication competent rAAV.

92. A nucleic acid comprising a nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10 and a large stuffer of a size at least 2 kb.

93. The nucleic acid of claim 92, wherein the nucleic acid comprises one or more nucleotides between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO:

10.

94. The nucleic acid of claim 93, wherein the nucleic acid comprises from about 10 to about 10,000 nucleotides between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO:

10.

95. The nucleic acid of claim 94, wherein the nucleic acid comprises from about 2,000 to about 5,000 nucleotides between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO:

10.

96. The nucleic acid of any one of claims 92-95, wherein the large stuffer is located between positions 44 and 45 of the nucleotide sequence having a having at least 85% identity to SEQ ID NO:

10.

97. The nucleic acid of any one of claims 92-96, wherein the large stuffer comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NO: 7-9 or 81.

98. The nucleic acid of any one of claims 92-97, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Rep protein operably linked to a promoter.

99. The nucleic acid of claim 98, wherein the nucleotide sequence having a having at least 85% identity to SEQ ID NO: 10 is located in the nucleic acid sequence encoding the AAV Rep protein.

100. The nucleic acid of claim 98 or 99, wherein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is located in an intron in the nucleic acid sequence encoding an AAV Rep protein.

101. The nucleic acid of any one of claims 98-100, wherein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is upstream of the promoter.

102. The nucleic acid of any one of claims 98-100, wherein the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is downstream of the promoter.

103. The nucleic acid of any one of claims 98-102, wherein the promoter is a p19 promoter.

104. The nucleic acid of any one of claims 98-103, wherein the AAV Rep is large Rep (Rep68).

105. The nucleic acid of any one of claims 98-104, wherein the AAV Rep protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

106. The nucleic acid of any one of claims 98-105, wherein the AAV Rep protein is a modified Rep protein.

107. The nucleic acid of any one of claims 98-106, wherein the nucleic acid further comprises a nucleic acid sequence encoding an AAV Cap protein.

108. The nucleic acid of claim 107, wherein the nucleic acid sequence encoding the AAV Cap protein is downstream of the nucleic acid sequence encoding the AAV Rep protein.

109. The nucleic acid of claim 107, wherein the nucleic acid sequence encoding the AAV Cap protein is upstream of the nucleic acid sequence encoding the AAV Rep protein.

110. The nucleic acid of any one of claims 107-109, wherein the nucleic acid sequence encoding the AAV Cap protein is from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh74, AAVrh10, po1, AAV9-PHP.B, AAV9-ePHP.B, AAV LK03, AAV Anc80L65, AAVDJ, AAV1A6ii, AAV1P5ii, AAV4A1ii, AAV7P4i, AAV9A1i, AAV9A2i, AAV9A6i, AAV9P1i, AAV9P2i, AAV9P5i, AAVrh10A1i, AAVrh10A2i, AAVrh10P1i, AAV12P2ii, AAVS10P1i, AAV JEA, AAV23xA P2i, AAVDJ P2i, AAV 2i8, AAV2G9, AAV2.5i82g9, AAV2.5, AAVr10pLDB_L2, AAVr10pLDB_P31, AAV4E, and AAV4Aand / or any chimeras thereof.

111. The nucleic acid of any one of claims 107-110, wherein AAV Rep protein and the AAV Cap protein are from the same AAV serotype.

112. The nucleic acid of any one of claims 107-110, wherein AAV Rep protein and the AAV Cap protein are from different AAV serotypes.

113. The nucleic acid of any one of claims 92-112, which is larger than 5.5 kb.

114. The nucleic acid of any one of 92-113, wherein the large stuffer does not comprise a nucleotide sequence of mammalian origin.

115. The nucleic acid of any one of 92-114 wherein the large stuffer comprises a nucleotide sequence of non-mammalian origin.

116. The nucleic acid of any one of 92-115, wherein the large stuffer is synthetic.

117. The nucleic acid of any one of 92-116, wherein the large stuffer does not comprise more than one of the following: a. a transcription factor binding site; b. a regulatory element; c. an AAV Rep binding site; d. a donor or acceptor splicing site; e. an endonuclease cleavage site, optionally where the endonuclease is ApaLI, BamHI, ClaI, DrdI, FspI, RsrII, XbaI, NcoI, SacII, CsiI, AflII, or PacI; f. a repetitive or palindrome sequence longer than 5 nucleotides; g. a strong secondary structure; and / or h. a repetitive or palindrome sequence, optionally a repetitive or palindrome sequence longer than 5 nucleotides.

118. The nucleic acid of any one of claims 92-117, wherein the large stuffer comprises a GC content of less than about 50%, e.g., less than about 45%, or less than about 40%.

119. The nucleic acid of any one of claim 92-118, wherein the nucleic acid is larger than 5.5kb.

120. The nucleic acid of any one of claims 92-119, wherein the 5’-end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to the sequence MAG, where M is A or C.

121. The nucleic acid of any one of claims 92-120, wherein the 3’-end of the nucleotide sequence having at least 85% identity to SEQ ID NO: 10 is linked to A or G.

122. The nucleic acid of any one of claims 92-121, wherein the large stuffer is located downstream of the sequence MAG, where M is A or C.

123. The nucleic acid of any one of claims 92-122, wherein the nucleic acid comprises a nucleotide sequence having at least 85% identity to a nucleotide sequence of any one of SEQ ID NOs: 11-13.

124. The nucleic acid of claim 123, wherein 5’-end of the nucleotide sequence having at least 85% identity to any of SEQ ID NOs: 11-13 is linked to the sequence MAG, where M is A or C.

125. The nucleic acid of claim 123 or 124, wherein the 3’-end of the nucleotide sequence having at least 85% identity to any of SEQ ID NOs: 11-13 is linked to A or G.

126. The nucleic acid of any one of claims 92-125, wherein the nucleic acid comprises a nucleotide sequence having at least 85% identity to SEQ ID NO:

11.

127. The nucleic acid of any one of claims 92-126, wherein the nucleic acid prevents production of replication competent rAAV vector.

128. The nucleic acid of any one of claims 92-127, wherein the nucleic acid is a vector.

129. The nucleic acid of any one of claims 92-128, wherein the nucleic acid is a plasmid.

130. The nucleic acid of any one of claims 92-129, wherein the nucleic acid is linear DNA.

131. The nucleic acid of any one of claims 92-130, wherein the nucleic acid is closed ended linear duplexed DNA (clDNA).

132. A host cell comprising a nucleic acid of any one of claims 92-131.

133. The host cell of claim 132, wherein the host cell is an insect cell or a mammalian cell, ptionally the mammalian cell s a HEK293 cell or a HeLa cell.

134. Use of a nucleic acid of any one of claims 92-131 or a host cell of claim 132or 133 in a method of preventing manufacturing of a replication competent rAAV.

135. A method of detecting residual DNA in a population of viral particles, the method comprising using at least one oligonucleotide sequence that anneals to any one of the nucleic acid sequences selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and SEQ ID NO:

6.

136. The method of claim 135, the method comprises using at least two oligonucleotide sequences that anneal to any one of the nucleic acid sequences selected from the group consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and SEQ ID NO:

6.

137. The method of claim 136, wherein the oligonucleotide is about 20 nucleotides long.

138. The method of claim136, wherein the oligonucleotide is about 25 nucleotides long.