Adeno-associated virus packaging vectors and methods of using same

JP2025512585A5Pending Publication Date: 2026-04-27JANSSEN RESEARCH & DEVELOPMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
JANSSEN RESEARCH & DEVELOPMENT LLC
Filing Date
2023-04-19
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing methods for generating recombinant adeno-associated virus (rAAV) particles face challenges in preventing abnormal packaging of rep and cap genes, which can lead to inefficient or incorrect assembly of viral particles.

Method used

A plasmid is constructed for packaging rAAV5 viruses, incorporating the AAV2 replication gene and the AAV5 capsid gene, with an artificial 2kb intron inserted into the rep gene to prevent abnormal packaging by making the coding regions too large to package correctly.

Benefits of technology

The solution effectively prevents the abnormal packaging of rep and cap genes, ensuring efficient and correct assembly of rAAV5 particles, thereby improving the reliability and efficacy of rAAV production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000014_0000
    Figure 00000014_0000
Patent Text Reader

Abstract

Provided herein are nucleic acid molecules that include artificial introns and modified replicative coding sequences. Also provided are vectors and cells that include the nucleic acid molecules that include artificial introns and modified replicative coding sequences.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 334,455, filed April 25, 2022, and U.S. Provisional Patent Application No. 63 / 334,468, filed April 25, 2022, the disclosures of each of which are incorporated by reference in their entirety herein.

[0002] FIELD OF THEINVENTION Disclosed herein are nucleic acid molecules that include artificial introns and modified replicative coding sequences. Also disclosed herein are vectors and cells that include the nucleic acid molecules that include artificial introns and modified replicative coding sequences.

[0003] (Sequence Listing) This application contains a sequence listing, which was submitted herewith as an xml file (name: "JBI6718-Sequence-Listing-ST26.xml", size: 37,181 bytes, created on March 24, 2023). The sequence listing is incorporated by reference in its entirety. [Background technology]

[0004] Adeno-associated virus (AAV) can be used to generate recombinant AAV particles that contain the DNA sequence of interest for delivery to target cells. Because these AAV particles lack viral genes, viral structural genes and packaging genes are provided in separate packaging plasmids. Typically, the packaging plasmid contains the replication (rep) gene and the capsid (cap) viral gene. Summary of the Invention [Means for solving the problem]

[0005] A plasmid has been constructed for packaging recombinant adeno-associated virus serotype 5 (AAV5) virus. This plasmid contains the AAV replication genes from AAV2 and the capsid gene from AAV5. To prevent aberrant packaging of the rep and cap genes into rAAV capsids, an artificial intron has been inserted into the rep gene. This 2 kb intron makes the rep and cap coding regions of the plasmid too large to be packaged.

[0006] Described herein is a nucleic acid molecule that includes the nucleic acid sequence set forth in SEQ ID NO: 1. Also described herein is a nucleic acid molecule that includes the nucleic acid sequence set forth in SEQ ID NO: 6.

[0007] Further described are vectors comprising a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 6. In certain embodiments, the vector comprises at least one of a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 6, a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 7, and a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 8.

[0008] In a further embodiment, the vector contains a kanamycin resistance gene.

[0009] Also described herein is a vector comprising the nucleic acid sequence set forth in SEQ ID NO:5.

[0010] Also described herein are cells transformed to express a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO:1 or SEQ ID NO:6. In certain embodiments, the cells are transformed to express a vector comprising a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO:1 or SEQ ID NO:6. In certain embodiments, the cells are transformed to express a vector comprising at least one of a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO:8.

[0011] Also described herein is a cell transformed to express a vector comprising a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 5. In certain embodiments, the cell is a virus-producing cell.

[0012] Also described herein is a method for generating an adeno-associated virus packaging vector, comprising introducing a nucleic acid molecule of claim 1 into a replication coding sequence to generate an intron-modified replication gene, and introducing the intron-modified replication gene into a vector backbone. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] The disclosed nucleic acid molecules, vectors, and cells may be more readily understood by reference to the following detailed description, which forms a part of this disclosure: It is to be understood that the disclosed nucleic acid molecules, vectors, and cells are not limited to those specifically described and / or illustrated herein, and that the terminology used herein is for the purpose of describing particular embodiments by way of example only, and is not intended to limit the claimed nucleic acid molecules, vectors, and cells.

[0014] Unless otherwise stated, any description of possible mechanisms or modes of operation or reasons for improvement are intended to be exemplary only, and the disclosed nucleic acid molecules, vectors, and cells are not limited by the merits of any such proposed mechanisms or modes of operation or reasons for improvement.

[0015] When a range of numerical values ​​is recited or established herein, the range includes its endpoints, and all individual integers and fractions within the range, and also includes each of the narrower ranges formed by all the various possible combinations of these endpoints and internal integers and fractions, forming subgroups of the larger group of values ​​within the recited range, as if each of the narrower ranges were explicitly recited. When a range of numerical values ​​is recited herein as being greater than the recited value, the range is nevertheless finite, with its upper limit defined by a value operable within the context of the invention described herein. When a range of numerical values ​​is recited herein as being less than the recited value, the range is nevertheless defined by its lower limit by a non-zero value. It is not intended that the scope of the invention be limited to the specific values ​​recited in defining the range. All ranges are inclusive and combinable.

[0016] When values ​​are expressed as approximations, by use of the antecedent "about," it is to be understood that the particular value forms another embodiment. Reference to a particular numerical value is intended to include at least the particular value unless the context clearly dictates otherwise.

[0017] It is understood that certain features of the invention that are described herein for clarity in the context of separate embodiments may also be provided in combination in a single embodiment. That is, unless expressly incompatible or specifically excluded, each individual embodiment is considered to be combinable with any other embodiment, and such combination is considered to be another embodiment. Conversely, different features of the invention that are described for brevity in the context of a single embodiment may be provided separately or in any subcombination. Finally, although an embodiment may be described as part of a series of steps or as part of a more general structure, each step may be considered to be an independent embodiment in itself that can be combined with the others.

[0018] Various terms relating to the embodiments of the present specification are used throughout the specification and claims. Unless otherwise indicated, such terms are to be given their ordinary meaning in the art. Other specifically defined terms are to be interpreted in a manner consistent with the definitions provided herein.

[0019] Similarly, the term "comprising" is intended to include examples encompassed by the terms "consisting essentially of" and "consisting of." Similarly, the term "consisting essentially of" is intended to include examples encompassed by the term "consisting of."

[0020] When a value is expressed as an approximation by using the descriptor "about", it is understood that the particular value forms another embodiment. In general, the use of the term "about" indicates an approximation that may vary depending on the desired properties to be obtained by the disclosed subject matter and should be interpreted in the specific context in which it is used based on its function. Those skilled in the art can interpret this as a matter of routine. In some cases, the number of significant figures used for a particular value can be one non-limiting way to determine the extent of the term "about". In other cases, the gradations used in a series of values ​​can be used to determine the intended range available for the term "about" for each value.

[0021] Unless otherwise specified, the term "about" refers to a ±10% variation of the associated value. Thus, the term "about" is used to encompass a ±10% or less variation, a ±5% or less variation, a ±1% or less variation, a ±0.5% or less variation, or a ±0.1% or less variation from the stated value.

[0022] When lists are presented, unless otherwise stated, it is to be understood that each individual element of that list and every combination of that list is a separate embodiment. For example, a list of embodiments presented as "A, B, or C" should be interpreted to include the embodiments "A," "B," "C," "A or B," "A or C," "B or C," or "A, B, or C."

[0023] As used herein, the singular forms "a," "an," and "the" are intended to include plurals.

[0024] nucleic acid molecule Disclosed herein are nucleic acid molecules encoding artificial introns that can be inserted into coding sequences, such as AAV replication (rep) coding sequences. In certain embodiments, the nucleic acid molecule encodes an artificial intron comprising SEQ ID NO:1. In certain embodiments, the nucleic acid molecule encodes an artificial intron that is a composite of a 5' intron fragment from human beta actin, such as a nucleic acid sequence comprising SEQ ID NO:2, a synthetic random non-coding sequence, such as a nucleic acid sequence comprising SEQ ID NO:3, and a 3' intron fragment from human beta actin, such as a nucleic acid sequence comprising SEQ ID NO:4. In certain embodiments, the synthetic random non-coding sequence is about 2.1 kb in length. In certain embodiments, the synthetic random non-coding sequence is generated by a random sequence generator.

[0025] Described herein is a nucleic acid molecule comprising (or consisting of, or essentially consisting of) the nucleic acid sequence set forth in SEQ ID NO:1. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 80% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 85% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 90% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Also described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) an artificial nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO: 1. Also described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) an artificial nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO: 1. Also described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) an artificial nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO: 1.

[0026] As used herein, the term "artificial" refers to a nucleic acid molecule that has been modified by man, e.g., does not occur in nature.

[0027] Described herein is a nucleic acid molecule comprising (or consisting of, or essentially consisting of) the nucleic acid sequence set forth in SEQ ID NO:3. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 80% identity to the nucleic acid sequence set forth in SEQ ID NO:3. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 85% identity to the nucleic acid sequence set forth in SEQ ID NO:3. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 90% identity to the nucleic acid sequence set forth in SEQ ID NO:3. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Described herein is also a nucleic acid molecule comprising (or consisting of, or essentially consisting of) an artificial nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:1. Also described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) an artificial nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO: 1. Also described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) an artificial nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO: 1. Also described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) an artificial nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO: 3.

[0028] Further described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) the nucleic acid sequence set forth in SEQ ID NO:2, the nucleic acid set forth in SEQ ID NO:3, and the nucleic acid set forth in SEQ ID NO:4. Also described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:2, a nucleic acid sequence having at least 80%, at least 85%, at least 90%, or at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:3, and a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:4. Also described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:2, a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:3, and a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:4. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:2, a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:3, and a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:4. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:2, a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:3, and a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:4. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:2, a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:3, and a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:4.Also described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO: 2, a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO: 3, and a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO: 4. In certain embodiments, the nucleic acid sequence set forth in SEQ ID NO: 2, the nucleic acid set forth in SEQ ID NO: 3, and the nucleic acid set forth in SEQ ID NO: 4 are operably linked.

[0029] Also disclosed herein is a nucleic acid molecule encoding an intron-modified AAV rep coding sequence that includes an artificial intron. In certain embodiments, the nucleic acid molecule encoding an intron-modified AAV rep coding sequence comprises SEQ ID NO:6. In certain embodiments, the nucleic acid molecule encoding an artificial intron comprises SEQ ID NO:1. In certain embodiments, the artificial intron is inserted into a CAGG sequence in the replicating coding sequence. In certain embodiments, the artificial intron is inserted into a CAGG sequence in the AAV rep coding sequence before the last "G". In certain embodiments, the insertion of the artificial intron does not disrupt the translation of the AAV rep coding sequence.

[0030] Described herein is a nucleic acid molecule comprising (or consisting of, or consisting essentially of) the nucleic acid sequence set forth in SEQ ID NO:6. Described herein is also a nucleic acid molecule comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:6. Described herein is also a nucleic acid molecule comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:6. Described herein is also a nucleic acid molecule comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:6. Described herein is also a nucleic acid molecule comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:6. Described herein is also a nucleic acid molecule comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO:6.

[0031] Also disclosed herein are nucleic acid molecules encoding an intron-modified AAV rep coding sequence, a VP1 coding sequence, and a polyA signal. In certain embodiments, the nucleic acid molecule encoding the intron-modified AAV rep coding sequence comprises SEQ ID NO:6. In some embodiments, the nucleic acid molecule encoding the VP1 coding sequence comprises SEQ ID NO:7. In certain embodiments, the nucleic acid molecule encoding the polyA signal comprises SEQ ID NO:7.

[0032] In certain embodiments, the VP1 coding sequence is an AAV5 (adeno-associated virus type 5) VP1 coding sequence. In certain embodiments, the polyA signal is an AAV2 (adeno-associated virus type 2) polyA signal.

[0033] Further described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) the nucleic acid sequence set forth in SEQ ID NO:6, the nucleic acid set forth in SEQ ID NO:7, and the nucleic acid set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 80% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 80% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 80% identity to the nucleic acid sequence set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or consisting essentially of) a nucleic acid sequence having at least 85% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 85% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 85% identity to the nucleic acid sequence set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 90% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 90% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 90% identity to the nucleic acid sequence set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 95% identity to the nucleic acid sequence set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 96% identity to the nucleic acid sequence set forth in SEQ ID NO:8.Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 97% identity to the nucleic acid sequence set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 98% identity to the nucleic acid sequence set forth in SEQ ID NO:8. Also described herein are nucleic acid molecules comprising (or consisting of, or essentially consisting of) a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO:7, and a nucleic acid sequence having at least 99% identity to the nucleic acid sequence set forth in SEQ ID NO:8. In certain embodiments, the nucleic acid sequence set forth in SEQ ID NO:6, the nucleic acid set forth in SEQ ID NO:7, and the nucleic acid set forth in SEQ ID NO:8 are operably linked.

[0034] The term "identical" or percent "identity," in the context of two or more nucleic acids, refers to two or more sequences or subsequences that have the same or a certain percentage of nucleotides that are the same when compared and aligned for maximum correspondence, as measured using one of the sequence comparison algorithms described below or by visual inspection.

[0035] For sequence comparison, typically one sequence serves as a reference sequence to which test sequence is compared.When using sequence comparison algorithm, test and reference sequences are input into computer, subsequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Then, sequence comparison algorithm calculates the percent sequence identity of test sequence to reference sequence based on designated program parameters.

[0036] Optimal alignment of sequences for comparison can be determined, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by visual inspection (see generally Current Protocols in Molecular Biology, FMAusubel et al., eds., Current Protocols, Greene Publishing Associates, Inc. and John Wiley & Sons, Inc.). This can be done through a joint venture with Ausubel, Sons, Inc. (1995 Supplement) (Ausubel).

[0037] Examples of suitable algorithms for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms described in Altschul et al., (1990) J. Mol. Biol. 215:403-410 and Altschul et al., (1997) Nucleic Acids Res. 25:3389-3402, respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information. This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match when aligned with words of the same length in database sequences or meet some positive threshold score T. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs that contain them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased.

[0038] For nucleotide sequences, the parameters M (reward score for a pair of matching residues, always >0) and N (penalty score for mismatching residues, always <0) are used to calculate the cumulative score. For amino acid sequences, a scoring matrix is ​​used to calculate the cumulative score. Extension of the word hits in each direction is stopped when the cumulative alignment score falls off its maximum achieved value by an amount X, when the accumulation of one or more alignments of negative scoring residues causes the cumulative score to fall below zero, or when the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) of 10, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)).

[0039] In addition to calculating percent sequence identity, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin & Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the minimum sum probability (P(N)), which provides an indication of the probability that a match between two nucleotide sequences or two amino acid sequences will occur by chance. For example, a nucleic acid is considered to be similar to a reference sequence if the minimum sum probability in the comparison of the test nucleic acid to the reference nucleic acid is less than about 0.1, more preferably less than about 0.01, and most preferably less than about 0.001.

[0040] A further indication that two nucleic acid sequences are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross-reactive with the polypeptide encoded by the second nucleic acid, as described below.Thus, a polypeptide is typically substantially identical to a second polypeptide, e.g., the two peptides differ only by conservative substitutions.Another indication that two nucleic acid sequences are substantially identical is that the two molecules hybridize to each other under stringent conditions.

[0041] In some embodiments, the nucleic acid molecule comprises DNA.

[0042] In some embodiments, the nucleic acid molecule comprises RNA.

[0043] In some embodiments, the RNA is mRNA.

[0044] In some embodiments, the nucleic acid molecule comprises a promoter, an enhancer, a polyadenylation site, a Kozak sequence, a stop codon, or any combination thereof.

[0045] Methods for making the nucleic acid molecules of the present disclosure are known in the art and include chemical synthesis, enzymatic synthesis (e.g., in vitro transcription), enzymatic or chemical cleavage of longer precursors, chemical synthesis of smaller fragments of polynucleotides followed by fragment ligation, or known PCR methods.The polynucleotide sequence to be synthesized can be designed with the appropriate codons for the desired amino acid sequence.Generally, preferred codons can be selected for the intended host that the sequence will be used for expression.

[0046] vector The present disclosure also provides vectors comprising, consisting of, or consisting essentially of any of the nucleic acid molecules disclosed herein.The present disclosure also provides vectors comprising a nucleic acid molecule encoding any of the polypeptides disclosed herein.

[0047] The vector disclosed herein can be a packaging vector that contains the AAV coding region (AAV rep and cap genes) without the 145 bp inverted terminal repeat (ITR). The disclosed packaging vector is in a suitable cell line (e.g., human 293 cells) with the AAV ITR chimeric protein, DNA contained in the construct that codes for the AAV ITR chimeric protein. The cells can then be infected with adenovirus. The viral vector, sometimes also referred to as a viral particle, can be purified from cell lysate using methods known in the art (e.g., cesium chloride density gradient ultracentrifugation, etc.) and confirmed to be free of detectable replication-competent AAV or adenovirus (e.g., by cytopathic effect bioassay).

[0048] Described herein are vectors comprising nucleic acid molecules that comprise a nucleic acid sequence set forth in SEQ ID NO: 1. Also described herein are vectors comprising nucleic acid molecules that comprise a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence set forth in SEQ ID NO: 1.

[0049] Further described herein is a vector comprising a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO: 6. Also described herein is a vector comprising a nucleic acid molecule comprising a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence set forth in SEQ ID NO: 6.

[0050] In certain embodiments, vectors include a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO: 6, a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO: 7, and a nucleic acid molecule comprising a nucleic acid sequence set forth in SEQ ID NO: 8. Also described herein are nucleic acid molecules comprising a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence set forth in SEQ ID NO: 6, nucleic acid molecules comprising a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence set forth in SEQ ID NO: 7, and vectors comprising a nucleic acid molecule comprising a nucleic acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the nucleic acid sequence set forth in SEQ ID NO: 8. In certain embodiments, the nucleic acid sequence set forth in SEQ ID NO:6, the nucleic acid molecule comprising SEQ ID NO:7, and the nucleic acid molecule comprising SEQ ID NO:8 are operably linked.

[0051] In still further embodiments, the vector comprises (or consists of, or consists essentially of) a nucleic acid molecule comprising (or consists of, or consists essentially of) the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 80% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 85% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 90% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 95% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 96% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 97% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 98% identity to the nucleic acid sequence set forth in SEQ ID NO:5. In certain embodiments, the vector comprises a nucleic acid sequence having at least about 99% identity to the nucleic acid sequence set forth in SEQ ID NO:5.

[0052] The vector may be a vector intended for expressing the polynucleotide of the present disclosure in any host, such as bacteria, yeast, or mammals. Suitable expression vectors are typically replicable in the host organism, either as episomes or as an integral part of the host chromosomal DNA. Usually, expression vectors contain a selection marker, such as ampicillin resistance, hygromycin resistance, tetracycline resistance, kanamycin resistance, or neomycin resistance, to allow detection of cells transformed or transduced with the desired DNA sequence. In certain embodiments, the vector contains a kanamycin resistance gene.

[0053] Exemplary vectors are plasmids, cosmids, phages, viral vectors, or artificial chromosomes.

[0054] Suitable vectors that may be used include, but are not limited to, bacterial: pUC57, pBs, phagescript, PsiX174, pBluescript SK, pBs KS, pNH8a, pNH16a, pNH18a, pNH46a (Stratagene, La Jolla, Calif., USA), pTrc99A, pKK223-3, pKK233-3, pDR540, and pRIT5 (Pharmacia, Uppsala, Sweden). Eukaryotic: pWLneo, pSV2cat, pOG44, PXR1, pSG (Stratagene), pSVK3, pBPV, pMSG, and pSVL (Pharmacia). In a particular embodiment, the vector is pUC57.

[0055] The present disclosure provides an expression vector comprising a nucleic acid molecule of the present disclosure. The present disclosure also provides an expression vector comprising a nucleic acid molecule encoding a polypeptide of the present disclosure.

[0056] Also described herein is a method for generating an adeno-associated virus (AAV) packaging vector, the method comprising the steps of: (a) introducing a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO:1 into an AAV replication coding sequence to generate an intron-modified AAV replication gene, and (b) introducing the intron-modified AAV replication gene into a vector backbone.

[0057] Engineered Cells Expressing Nucleic Acids of the Disclosure Engineered cells that express the nucleic acid molecules or vectors of this disclosure are within the scope of this disclosure.

[0058] Suitable engineered cells include, for example, cells obtained from animals and humans, hi certain embodiments, the cells are virus-producing cells.

[0059] In certain embodiments, described herein are cells transformed to express a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 1. In certain embodiments, described herein are cells transformed to express a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO: 6.

[0060] In certain embodiments, cells transformed to express a vector comprising a nucleic acid molecule comprising the nucleic acid sequence set forth in SEQ ID NO:6, a nucleic acid molecule comprising SEQ ID NO:7, and a nucleic acid molecule comprising SEQ ID NO:8 are described herein.

[0061] In certain embodiments, a cell transformed to express a vector comprising a nucleic acid molecule comprising SEQ ID NO:5 is described herein. EXAMPLES

[0062] These examples are presented for illustrative purposes only and are not intended to limit the scope of the claims provided herein.

[0063] Example 1. AAV5 RepCap Plasmids Engineering an intron into the AAV2 Rep gene The artificial intron used to disrupt the REP coding sequence (SEQ ID NO:1) was a composite of a 5' intron fragment from human β-actin (SEQ ID NO:2), a 2.1 kb synthetic random non-coding sequence (SEQ ID NO:3), and a 3' intron fragment of human β-actin (SEQ ID NO:4). SEQ ID NO:2 was based on a fragment generated by a random sequence generator (https: / / birc.au.dk / ~palle / php / fabox / random_sequence_generator.php) and subsequently modified to remove potential splicing signals.

[0064] The AAV Rep gene utilizes alternative splicing to produce several different protein isoforms, but there are no pre-existing introns that are always spliced ​​out of all protein-coding transcripts. Therefore, it was not possible to simply increase the size of the pre-existing intron. Since the recognition of splice donors or splice acceptors also depends on the adjacent exon sequences, there are only certain locations in the REP gene where the insertion of an intron results in efficient RNA splicing.

[0065] To identify suitable intron insertion sites, the REP gene was scanned for "CAGG" sequences and composite test sequences were generated in which an intron (SEQ ID NO:1) was inserted before the last G. These composite sequences were subjected to splice site prediction analysis using the NETGene2 algorithm (https: / / services.healthtech.dtu.dk / service.php?NetGene2-2.42). The insertion site used in the final construct is present in a portion of the Rep coding sequence shared by all four REP protein isoforms (Rep78, Rep68, Rep50, Rep40). Another potential spliced ​​donor site was predicted by NetGene2 upstream of the intron insertion. This site was mutated from GT at position 800 of SEQ ID NO:5 to AT to make splicing at the inserted intron more efficient.

[0066] Plasmid P849 (SEQ ID NO:5) was generated by gene synthesis. It contains an intron-modified AAV2 REP gene (SEQ ID NO:6), an AAV5 VP1 coding sequence (SEQ ID NO:7), an AAV2 polyA signal (SEQ ID NO:8) cloned into a plasmid backbone derived from pUC57-KAN containing a kanamycin resistance gene. A variant of this plasmid was generated with the AAV5 VP1 sequence deleted (SEQ ID NO:10). This plasmid can be used to generate plasmids for packaging additional serotypes via cloning in the VP1 coding regions from other AAVs.

[0067] Testing plasmids in AAV packaging assays material Cell culture: virus producing cells (VPCs) (Thermo-Fisher, Waltham, MA, #A35347); DMEM (Thermo-Fisher, #11995-040); fetal bovine serum (FBS) (Cytiva Life Sciences, Marlborough, MA; #SH30079.03IR); Opti-MEM (Thermo-Fisher, #31985-062).

[0068] Chemicals, buffers, and enzymes: 1M Tris, pH7.5 (Millipore-Sigma, St. Louis, MO, #T2319-1L); 1M Magnesium Chloride (Millipore-Sigma, #63069-500ML); Triton-X100 (Millipore-Sigma, T9284-100ML); AAV lysis buffer (10mM Tris-HCl, pH7.5; 20mM MgCl2, 1%Triton-X100 (v / v)); 10% Pluronic F-68 (Thermo-Fisher, #24040-032); 10× GeneAmp PCR buffer I (1.5mM containing MgCl2 (Thermo-Fisher, #4379876); sheared salmon sperm DNA (Thermo-Fisher, #AM9680); virus dilution buffer (VDB) (1× PCR buffer I, 2 μg / ml sheared salmon sperm DNA, and 0.05% Pluronic F-68); benzonase nuclease 250 units / μl (Millipore-Sigma E1014-25K); PEI-MAX transfection reagent (Polysciences, Warrington PA, #24765-2), 1 mg / ml, dissolved in water at pH 7.5.

[0069] Digital droplet PCR: 2x SuperMix for Probes (Bio-Rad, Hercules, California, #186-3026); DG32 AutoDG Cartridge (Bio-Rad, #1864108); Auto Droplet Generator Oil in PBS (Bio-Rad, #1864110); Droplet Reader Oil (Bio-Rad, #1863004); Auto Droplet Generator (Bio-Rad Catalog, #186-4101); QX200 Droplet Reader (Bio-Rad, #186-4003); C1000 Touch Thermal Cycler with Deep Well Reaction Module (Bio-Rad, #185-1197).

[0070] PrimeTime qPCR Assays: 20x stocks of these assays consist of forward and reverse PCR primers (18 μM), as well as a 5' nuclease probe containing the fluorescent quenchers ZEN and Black Hole Quencher 1 (3IABkFQ), and either FAM or HEX fluorescent reporter dyes (5 μM). The assays were synthesized by Integrated DNA Technologies, Inc., Coralville IA.

[0071] Assay VPC cells were plated at 5E6 cells per T25 flask in 5 mL growth medium (DMEM+10% FBS) and incubated at 37° C. in a CO2 incubator for 2 days. 4.2 μg P849 + 6.8 μg adenovirus helper plasmid + 2.6 μg Cis plasmid (mCherry-IRES-fLUC transgene flanked by AAV2 ITRs) were diluted to 80 μl in Opti-MEM and mixed with 80 μl PEI-MAX (diluted to 212.5 μg / ml in OptiMEM) and incubated at room temperature for 30 min. The transfection mixture was added to 12 mL of medium (DMEM+0.5% FBS) and used to replace the growth medium in the T25 flask. Cells were incubated at 37° C. in a CO2 incubator for 72 h. The cells and medium were transferred to a 125 ml Erlenmeyer flask containing 1.2 mL of lysis buffer + 160 units of benzonase, and the flask was shaken at 120 rpm for 1 hour in a CO2 incubator at 37° C. The lysed cells were clarified by centrifugation at 4000 rpm for 15 minutes, and the supernatant was transferred to a new tube.

[0072] AAV titers were measured by digital droplet PCR (ddPCR). A 2 μl aliquot of lysate was diluted in 18 μl virus dilution buffer containing 0.2 units of benzonase and incubated at room temperature for 30 min. Samples were further diluted in triplicate with virus dilution buffer, mixed with ddPCR master mix, and ddPCR assayed for mCherry and droplets were generated with a Bio-Rad auto droplet generator. Droplets were subjected to PCR using the profile (95° C. 10 min; 42× (94° C. 30 sec, 60° C. 1 min, 72° C. 15 sec, all three with cycle times of 2° C. / sec); 98° C. 10 min. Droplets were detected using a QX200 droplet reader (Bio-Rad) according to the manufacturer's instructions.

[0073] result A total of 1.2E+12±8.1E+10 AAV vector genomes (DNase-resistant particles as detected by the mCherry probe) were produced, whereas the negative control sample without lysate produced no signal. These data indicate that P849 encodes the AAV Rep and Cap genes required for packaging the mCherry transgene from the Cis plasmid into DNase-resistant particles, consistent with the production of recombinant AAV. Furthermore, the intron inserted into REP does not appear to disrupt the function of these genes.

[0074] Those skilled in the art will appreciate that changes could be made to the embodiments described above without departing from the broad inventive concept. It is understood therefore that the invention is not limited to the particular embodiments disclosed, but is intended to cover modifications within the spirit and scope of the invention as defined herein.

[0075] array

[0076] [Table 1-1]

[0077] [Table 1-2]

[0078] [Table 1-3]

[0079] [Table 1-4]

[0080] [Table 1-5]

[0081]

Table 1-6

[0082]

Table 1-7

[0083]

Table 1-8

[0084]

Table 1-9

[0085]

Table 1-10

[0086]

Table 1-11

[0087]

Table 1-12

[0088]

Table 1-13

Claims

1. A nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to sequence number 1.

2. A nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to sequence number 6.

3. A vector comprising a nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 1, or a nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to SEQ ID NO:

6.

4. The vector according to claim 3, wherein the vector comprises a nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to sequence number 6, a nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to sequence number 7, and a nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to sequence number 8.

5. The vector according to claim 4, wherein the vector comprises a kanamycin resistance gene.

6. The vector according to claim 3, wherein the vector comprises a nucleic acid molecule having a nucleic acid sequence that is at least 95% identical to sequence number 5.

7. A vector containing a nucleic acid sequence that is at least 95% identical to sequence number 5.

8. Cells transformed to express the nucleic acid molecule described in claim 1 or 2, or the vector described in any one of claims 4 to 7.

9. The cell according to claim 8, wherein the cell is a virus-producing cell.

10. A method for producing an adeno-associated virus packaging vector, (a) the step of introducing the nucleic acid molecule described in claim 1 into a replication coding sequence to generate an intron-modified replication gene, and (b) A method comprising the step of introducing the intron-modified replica gene into a vector skeleton.

11. A nucleic acid molecule containing a nucleic acid sequence that is at least 95% identical to sequence number 10.