Construction of an adeno-associated virus capsid library for insect cells
A nucleic acid construct for insect cells addresses cross-packaging and mosaicism issues in AAV capsid library production, ensuring efficient and complex library generation suitable for insect cell platforms.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- UNIQURE BIOPHARMA BV
- Filing Date
- 2024-04-18
- Publication Date
- 2026-06-02
AI Technical Summary
Existing methods for constructing AAV capsid libraries in mammalian cells face challenges such as cross-packaging and mosaicism, leading to inefficiencies and unsuitability for insect cell-based production platforms.
A nucleic acid construct is developed for insect cells, incorporating a nucleotide sequence encoding an AAV capsid variant, a transcriptional transregulator-dependent enhancer element, and AAV reverse terminal repeats, along with optional reporter genes and barcodes, to minimize cross-packaging and mosaicism, using baculovirus promoters and endonuclease recognition sequences for precise manipulation.
The method enables the production of AAV capsid libraries in insect cells with reduced cross-packaging and mosaicism, maintaining sufficient depth and complexity for efficient capsid variant selection and production.
Smart Images

Figure 2026517698000004 
Figure 2026517698000005 
Figure 2026517698000006
Abstract
Description
[Technical Field]
[0001] This invention relates to the fields of medicine, molecular biology, gene therapy, and insect cell culture. In particular, this invention relates to means and methods for preparing adeno-associated virus capsid libraries for insect cells. [Background technology]
[0002] Adeno-associated viruses (AAVs) are emerging as a favorable platform for gene therapy. However, the success of AAV-based therapeutic strategies is limited by existing immune responses, nonspecific targeting, and inefficient production on a large scale. To address these challenges, efforts have continued to design novel AAV capsids that can evade neutralizing antibodies (NAbs) and / or have improved tissue tropism or modified specificity.
[0003] AAV capsid libraries have been widely and successfully applied, for example in directed evolutionary studies, to select capsids with improved properties. The construction of such AAV libraries utilizes a mammalian production platform in Hek293 cells. Libraries produced on this platform possess sufficient depth and titer to conduct these types of studies. In this approach, the AAV library is generated by transfection of a highly diverse library of plasmids encoding capsid variants expressed in a “natural” AAV genome context, where the two viral genes (Rep and Cap) are adjacent to reverse-terminal repeats. Subsequently, AAV variants with desired properties can be isolated by sequential rounds of phenotypic selection and viral DNA recovery. Therefore, this method relies heavily on the excellent correlation of genome-capsid “identity.” However, the construction of AAV capsid libraries on the Hek293 platform requires the transfection of large amounts of plasmid DNA, which makes them susceptible to cross-packaging and mosaicism of the introduced genes, where the particles are composed of genomes and capsid monomers derived from different library members (see, for example, Schmit et al., 2019, Mol Ther Methods Clin Dev, 17:107-121).
[0004] Furthermore, capsids selected from libraries produced on mammalian platforms may not be suitable for efficient production on insect cell-based platforms.
[0005] Therefore, there is still a need in the art to provide means and methods for generating AAV libraries in insect cells that address the above-mentioned drawbacks and minimize the occurrence of cross-packaging and mosaic phenomena in the resulting libraries. The object of the present invention is to provide such means and methods. [Overview of the Initiative]
[0006] In a first embodiment, a nucleic acid construct is provided comprising: i) a nucleotide sequence encoding an adeno-associated virus (AAV) capsid variant, which encodes mRNA, and whose translation in an insect cell generates an AAV capsid variant and is operably ligated to a promoter for driving expression in the insect cell; ii) an enhancer element operably ligated to a promoter, which is dependent on a transcriptional transregulator, wherein the introduction of the transcriptional transregulator into an insect cell containing the nucleic acid construct induces transcription from the promoter; iii) at least one AAV reverse terminal repeat (ITR) sequence; iv) an optional reporter gene operably ligated to a promoter for driving expression in a mammalian cell; and v) an optional barcode.
[0007] In one embodiment of the nucleic acid construct, the transcriptional transregulator is a baculovirus pre-initial protein (IE1) or its splice variant (IE0), the transcriptional transregulator-dependent enhancer element is a baculovirus homologous region (hr) enhancer element, and preferably the baculovirus is Autographa californica multicapsid polyhedrosis virus. In one embodiment, the hr enhancer element comprises at least one copy of the hr28-mer sequence CTTTACGAGTAGAATTCTACGCGTAAAA and / or at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides identical to the sequence CTTTACGAGTAGAATTCTACGCGTAAAA, and includes at least one copy of a sequence that binds to the baculovirus IE1 protein, wherein the hr enhancer element is operably ligated to an expression cassette containing a reporter gene operably ligated to a polH promoter, then a) under non-inducible conditions, an expression cassette having the hr enhancer element yields fewer results than other identical expression cassettes containing the hr2-0.9 element. a) A cassette that produces a reporter transcript or has an hr enhancer element produces an amount less than 1.1, 1.2, 1.5, 2, 5, or 10 times the amount of reporter transcript produced by another identical expression cassette containing an hr4b element; b) Under induction conditions, an expression cassette having an hr enhancer element produces at least 50, 60, 70, 80, 90, or 100% of the amount of reporter transcript produced by another identical expression cassette containing an hr4b or hr2-0.9 element, where the hr enhancer element is more preferably selected from the group consisting of hr1, hr2-0.9, hr3, hr4b, and hr5, of which hr4b and hr5 are preferred, and of which hr4b is most preferred.
[0008] In one embodiment of the nucleic acid construct, the promoter for driving expression in insect cells is a baculovirus promoter, preferably a late or very late baculovirus promoter, more preferably a promoter selected from the group consisting of polH, p10, p6.9, and pSel120 promoters.
[0009] In one embodiment, the nucleic acid construct further comprises an endonuclease recognition sequence, preferably a site-specific endonuclease recognition sequence. In one embodiment, the recognition sequence is a protelomerase, preferably a phage N15 TelN protelomerase recognition sequence.
[0010] In one embodiment of the nucleic acid construct, the endonuclease recognition sequence is not adjacent to at least one AAV ITR sequence, preferably not between two AAV ITR sequences, and as a result, the recognition sequence is not incorporated into the genome of the rAAV vector produced in insect cells.
[0011] In one embodiment, the nucleic acid construct comprises, in order from 5' to 3', a) a first AAV ITR sequence; b) an enhancer element; c) an expression cassette containing a nucleotide sequence encoding an AAV capsid variant, operably linked to a promoter to drive expression in insect cells; and d) a second AAV ITR, wherein the recognition sequence is located at least one of i) upstream of the first AAV ITR and ii) downstream of the second AAV ITR, and optionally, a reporter gene operably linked to a promoter to drive expression in mammalian cells is located between a) and b), between b) and c), or between c) and d), and optionally, a barcode is located between a) and b), between b) and c), or between c) and d), and preferably, the recognition sequence is a unique recognition sequence.
[0012] In one embodiment of the nucleic acid construct, at least one of the following is: a) a promoter for driving expression in mammalian cells, which is a cell type and / or tissue-specific promoter, preferably a human synapsin promoter (hSynl), a trans tyretin promoter (TTR), a cytokeratin 18, a cytokeratin 19, an unc-45 myosin chaperone B (unc45b) promoter, a cardiac troponin T (cTnT) promoter, a glial fibrillary acidic protein (GFAP) promoter, a myelin basic protein (MBP) promoter, and a methyl-CpG binding protein 2 (Me a) The promoter is selected from the group consisting of cp2) promoters; b) The reporter gene encodes a colorimetric, fluorescent, or bioluminescent reporter protein, preferably selected from the group consisting of secreted alkaline phosphatase (SEAP), β-galactosidase EGFP, mCherry, mCloverS, mRubyS, mApple, iRFP, tdTomato, mVenus, YFP, RFP, firefly luciferase, sea pansy luciferase, and nanoluciferase; and c) The barcode is a nucleotide sequence of 4 to 18 nucleotides in length.
[0013] In a second embodiment, a nucleic acid molecule comprising the nucleic acid construct described herein is provided. In one embodiment, the nucleic acid molecule is a linear nucleic acid molecule.
[0014] In one embodiment, linear nucleic acid molecules can be obtained by digestion with a site-specific endonuclease that cleaves at a recognition sequence in the nucleic acid construct.
[0015] In one embodiment, the linear nucleic acid molecule is a nucleic acid molecule in which at least one end of the linear nucleic acid molecule is protected from exonuclease degradation, preferably both ends of the linear nucleic acid molecule are protected from exonuclease degradation.
[0016] In one embodiment of the linear nucleic acid molecule, the endonuclease is a protelomerase, preferably a phage N15 TelN protelomerase, and at least one end of the linear nucleic acid molecule is closed by a covalent bond, or preferably both ends of the linear nucleic acid molecule are closed by covalent bonds.
[0017] In one embodiment of a linear nucleic acid molecule, the positions of elements i) to v) and the recognition sequence in the construct are such that the expression of the AAV Rep protein in insect cells containing the nucleic acid construct results in the replication of the DNA fragment containing elements i) to v) and its packaging into an AAV capsid.
[0018] In a third aspect, a library is provided comprising a plurality of nucleic acid molecules, including the nucleic acid constructs described herein.
[0019] In one embodiment, the library is such that, in each member of the library, the nucleotide sequence encoding an AAV capsid variant differs by at least one nucleotide from the nucleotide sequence encoding an AAV capsid variant in another member of the library, and preferably, in each member of the library, the nucleotide sequence encoding an AAV capsid variant differs by at least one amino acid from the AAV capsid variant encoded by another member of the library.
[0020] In a fourth aspect, insect cells are provided that include a nucleic acid molecule containing a nucleic acid construct described herein, or a library containing (a plurality of) such nucleic acid molecules containing a nucleic acid construct described herein.
[0021] In one embodiment, the insect cell further comprises a baculovirus vector containing at least one expression cassette for the expression of the AAV Rep protein in the insect cell.
[0022] In a fifth aspect, a library is provided comprising a plurality of AAV capsid variants, wherein the members of the library comprise nucleic acids replicated and packaged from members of a library of nucleic acid molecules comprising nucleic acid constructs, as described herein, for example, in a third aspect.
[0023] A sixth aspect provides a method for preparing a library of AAV capsid variants. In one embodiment, a method for preparing a library of AAV capsid variants includes: a) providing a library comprising a plurality of nucleic acid constructs as defined above, wherein the amino acid sequence of each member AAV capsid variant encoded in the member construct differs by at least one amino acid from the amino acid sequence of the AAV capsid variant encoded in the other member constructs in the library; b) amplifying the library of nucleic acid constructs, preferably obtained by rolling circle amplification; c) digesting the amplified library of nucleic acid constructs obtained in b) by a site-specific endonuclease that cleaves at recognition sequences in the nucleic acid constructs to produce a library comprising unit-length linear amplified nucleic acid constructs; d) transfecting insect cells with the library comprising the linear amplified nucleic acid constructs obtained in c); e) infecting insect cells with a baculovirus vector comprising at least one expression cassette for expressing AAV Rep protein in the insect cells and culturing the cells; and f) recovering the library of AAV capsid variants produced by the insect cells in e).
[0024] In one embodiment, a method for preparing a library of AAV capsid variants is such that the library in a) is provided by: i) providing a nucleic acid construct defined in the first aspect; and ii) replacing a fragment containing at least a part of the nucleotide sequence encoding the AAV capsid variant in the construct with a library of fragments, wherein the amino acid sequence of the AAV capsid variant of each member of the library is encoded by the fragment of that member and differs from the amino acid sequence of the AAV capsid variant encoded by the construct of other members in the library by at least one amino acid.
[0025] In one embodiment, in step ii), the replacement is performed by Gibson assembly, in step b), the rolling circle amplification is performed by using Φ29 DNA polymerase, preferably in combination with random primers, and / or in c), protection against exonuclease degradation is provided at the ends of the linear amplification library of unit-length nucleic acid constructs. In one embodiment, in step c), the endonuclease is phage N15 TelN protelomerase, which provides covalently closed ends to the linear amplification library of unit-length nucleic acid constructs.
[0026] In a seventh aspect, a method for identifying an AAV capsid variant having a desired characteristic is provided. In one embodiment, the method comprises: a) providing a library of AAV capsid variants as defined above or obtained by a method for preparing a library of AAV capsid variants as defined above; b) contacting the library with a cell, organoid, or tissue culture or administering the library to a non-human animal; c) transducing the AAV capsid variants in the library into cells in the cell, organoid, tissue, or animal; and d) identifying at least one AAV capsid variant that has transduced at least one desired cell type in the cell, organoid, tissue, or animal in b) as an AAV capsid variant having the desired characteristic, and optionally recovering the AAV capsid variant having the desired characteristic from the desired cells.
[0027] In one embodiment of the method for identifying an AAV capsid variant having a desired characteristic, step d) comprises detecting transduction of at least the desired cells by detecting expression of a reporter gene in the desired cells. In one embodiment of the method, the variant having the desired characteristic is identified by: i) sequencing a barcode, and ii) at least one region of a DNA construct encoding at least one amino acid difference [Mode for Carrying Out the Invention]
[0028] [Description of the Invention] [Definitions] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. One of ordinary skill in the art will recognize many methods and materials similar or equivalent to those described herein that can be used in the practice of the invention. Indeed, the invention is in no way limited to these methods.
[0029] In this specification and its claims, the verb “includes” and its conjugations are used in their non-restrictive sense to mean that the items following the word are included, but not excluded, except for items not specifically mentioned. Furthermore, references to elements with the indefinite article “a” or “an” do not rule out the possibility of multiple elements being present unless the context explicitly requires that only one or one element exists. Thus, the indefinite article “a” or “an” usually means “at least one.”
[0030] As used herein, the term "and / or" indicates that one or more of the described cases may occur alone or in combination with at least one or up to all of the described cases.
[0031] As used herein, "at least" a particular value means that value or greater than or equal to that particular value. For example, "at least two" is understood to be the same as "two or more," i.e., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, etc.
[0032] When the terms "about" or "approximately" are used in relation to a numerical value (e.g., about 10), it is preferable that the value may be more or less 0.1% of the given value (10).
[0033] As used herein, “effective dose” means the amount of drug required to improve the symptoms of a disease compared to an untreated patient. For example, the effective dose of an active drug used to carry out the present invention for the therapeutic treatment of cancer or neurological disease will vary depending on the mode of administration, the age, weight, and overall health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate dose and administration regimen. Such a dose is referred to as the “effective” dose and may be determined, for example, as genome copy number per kilogram (GC / kg) or as GC per dose. Accordingly, in relation to the administration of a drug that is “effective for ~” in the context of this disclosure, the disease or condition indicates that administration in a clinically appropriate manner will result in a beneficial effect in at least a statistically significant proportion of patients, such as improvement of symptoms, cure, reduction of at least one disease sign or symptom, extension of lifespan, improvement of quality of life, or other effects that are generally recognized as positive by a physician familiar with the treatment of a particular type of disease or condition.
[0034] The use of a substance as a pharmaceutical as described herein may also be interpreted as the use of such substance in the manufacture of a pharmaceutical. Similarly, whenever a substance is used for therapeutic purposes or as a pharmaceutical, it may also be used for the manufacture of a pharmaceutical for therapeutic purposes. Products for use as pharmaceuticals as described herein may be used in therapeutic methods, such therapeutic methods including the administration of products for use.
[0035] The terms "homology" and "sequence identity" are used interchangeably herein. Sequence identity is defined herein as the relationship between two or more amino acid (polypeptide or protein) sequences or two or more nucleic acid (polynucleotide) sequences, determined by comparing sequences. In the art, "identity" and "similarity" also mean the degree of sequence relevance between amino acid or nucleic acid sequences, which may be determined by matching strings of such sequences. "Identity" and "similarity" can be readily calculated by known methods.
[0036] Sequence identity and sequence similarity can be determined by the alignment of two peptide or nucleotide sequences using a global or local alignment algorithm, depending on the lengths of the two sequences. Sequences of similar lengths are preferably aligned using a global alignment algorithm (e.g., Needleman-Wunsch) that optimally aligns the sequences over their entire length, while sequences of substantially different lengths are preferably aligned using a local alignment algorithm (e.g., Smith-Waterman). Sequences may be referred to as "substantially identical" or "essentially similar" when they share at least a certain minimum percentage of sequence identity (as defined below) (e.g., when optimally aligned by the program GAP or BESTFIT using default parameters). GAP uses the Needleman-Wunsch global alignment algorithm to align two sequences over their full length (total length), maximizing the number of matches and minimizing the number of gaps. When two sequences have similar lengths, global alignment is appropriately used to determine sequence identity. Generally, the default GAP parameters are used with a gap creation penalty of 50 (nucleotides) / 8 (proteins) and a gap elongation penalty of 3 (nucleotides) / 2 (proteins). For nucleotides, the default scoring matrix used is nwsgapdna, and for proteins, the default scoring matrix is Blosum62 (Henikoff & Henikoff, 1992, PNAS89, 915-919).Sequence alignment and scoring for sequence identity percentage can be determined using computer programs such as the GCG Wisconsin Package, Version 10.3, available from Accelrys Inc., 9685 Scranton Road, San Diego, CA 92121-3752 USA, or using open-source software such as EmbossWIN version 2.10.0's "needle" program (using the global Needleman-Wunsch algorithm) or "water" program (using the local Smith-Waterman algorithm), using the same parameters as GAP described above, or using default settings (for both "needle" and "water," and for both protein and DNA alignments, the default gap opening penalty is 10.0, the default gap elongation penalty is 0.5, and the default scoring matrix is Blossum62 for protein and DNAFull for DNA). If the sequences have substantially different full lengths, local alignment, such as that using the Smith-Waterman algorithm, is preferred.
[0037] Alternatively, the degree of similarity or identity can be determined by searching public databases using algorithms such as FASTA and BLAST. Therefore, the nucleic acid and protein sequences of the present invention can be further used as "query sequences" for performing searches against public databases, for example, to identify other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) from Altschul, et al. (1990) J.Mol.Biol.215:403-10. A BLAST nucleotide search can be performed using the NBLAST program, score=100, word length=12, to obtain nucleotide sequences homologous to the oxidoreductase nucleic acid molecule of the present invention. A BLAST protein search can be performed using the BLASTx program, score=50, word length=3, to obtain amino acid sequences homologous to the protein molecule of the present invention. To obtain gapped alignments for comparative purposes, gapped BLAST can be used, as described in Altschul et al., (1997) Nucleic Acids Res. 25(17):3389-3402. When using the BLAST and Gapped BLAST programs, the default parameters for each program (e.g., BLASTX and BLASTn) can be used. See the National Center for Biotechnology Information website (http: / / www.ncbi.nlm.nih.gov / ) for more information.
[0038] As used herein, the terms “selectively hybridize,” “selectively hybridize,” and similar terms are intended to describe hybridization and washing conditions in which nucleotide sequences that are at least 66%, at least 70%, at least 75%, at least 80%, more preferably at least 85%, even more preferably at least 90%, more preferably at least 95%, more preferably at least 98%, or more preferably at least 99% homologous to each other typically remain hybridized. That is, such hybridized sequences may share at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, more preferably at least 85%, even more preferably at least 90%, more preferably at least 95%, more preferably at least 98%, or more preferably at least 99% sequence identity.
[0039] A preferred, non-limiting example of such hybridization conditions is hybridization in 6x sodium chloride / sodium citrate (SSC) at about 45°C, followed by one or more washes in 1x SSC, 0.1% SDS at about 50°C, preferably about 55°C, preferably about 60°C, and even more preferably about 65°C.
[0040] Highly stringent conditions include, for example, hybridization in 5x SSC / 5x Denhardt solution / 1.0% SDS at approximately 68°C and washing in 0.2x SSC / 0.1% SDS at room temperature. Alternatively, washing may be carried out at 42°C.
[0041] Those skilled in the art will know the conditions that should be applied to stringent hybridization conditions and highly stringent hybridization conditions. Further guidance on such conditions can be found, for example, in Sambrook et al., 1989, Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press, NY; and Ausubel et al. (eds.), Sambrook and Russell (2001) “Molecular Cloning: A Laboratory Manual (3 rd This is readily available in the field in the book *Current Protocols in Molecular Biology* (John Wiley & Sons, NY), published by Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, New York, 1995.
[0042] Naturally, polynucleotides that hybridize only to poly(A) sequences (such as the 3'-terminal poly(A) region of mRNA) or only to the complementary stretch of T (or U) will not be included in the polynucleotides of the present invention used to specifically hybridize to a portion of the nucleic acids of the present invention. This is because such polynucleotides hybridize to any nucleic acid molecule containing the poly(A) stretch or its complement (e.g., substantially any double-stranded cDNA clone).
[0043] In this specification, “nucleic acid construct” or “nucleic acid vector” is understood to mean an artificial nucleic acid molecule resulting from the use of recombinant DNA technology. Therefore, the term “nucleic acid construct” does not include naturally occurring nucleic acid molecules, although a nucleic acid construct may include (some of) naturally occurring nucleic acid molecules. A “vector” is a nucleic acid construct (typically DNA or RNA) that helps to transfer an exogenous nucleic acid sequence (i.e., DNA or RNA) into a host cell. A vector is preferably maintained in a host by at least one of autonomous replication and integration into the host cell’s genome. The term “expression vector” or “expression construct” refers to a nucleotide sequence that can influence the expression of a gene in a host cell or host organism that is compatible with such a sequence. These expression vectors typically comprise at least one “expression cassette,” which is a functional unit that can influence the expression of a sequence encoding the product to be expressed, the coding sequence being operably linked to at least a suitable transcriptional regulatory sequence and, optionally, a suitable expression regulatory sequence including a 3' transcription termination signal. Further factors necessary or useful for influencing expression, such as expression enhancer elements, may also exist. Expression vectors can be introduced into suitable host cells and influence the expression of coding sequences in in vitro cell cultures of the host cells. Preferred expression vectors are suitable for the expression of viral proteins and / or nucleic acids, particularly recombinant parvovirus proteins and / or nucleic acids, such as baculovirus vectors for the expression of parvovirus proteins and / or nucleic acids in insect cells.
[0044] A “parvovirus vector” is defined as a recombinantly produced parvovirus or parvovirus particle containing a polynucleotide that is delivered to a host cell either in vivo, ex vivo, or in vitro. Adeno-associated virus (AAV) vectors are an example of parvovirus vectors. In this specification, a parvovirus or AAV vector refers to a portion of the parvovirus genome, usually at least one ITR, and a polynucleotide containing a transgene, which is preferably packaged within a parvovirus or AAV capsid.
[0045] As used herein, the terms “promoter” or “transcriptional regulatory sequence” refer to a nucleic acid fragment structurally identified by the presence of a binding site to any other DNA sequence known to those skilled in the art, which functions to control the transcription of one or more coding sequences, is located upstream of the transcription start site of the coding sequence, and includes, but is not limited to, a DNA-dependent RNA polymerase, a transcription start site, and transcription factor binding sites, repressor and activator protein binding sites, and any other nucleotide sequences that act to directly or indirectly regulate the amount of transcription from the promoter. A “component” promoter is a promoter that is active in most tissues under most physiological and developmental conditions. An “inducible” promoter is a promoter that is physiologically or developmentally modulated, for example, by the application of a chemical inducer or biological entity.
[0046] The term "reporter" can be used interchangeably with "marker," but it is primarily used to refer to visible markers such as green fluorescent protein (GFP) or luciferase.
[0047] The terms "protein" and "polypeptide" are used interchangeably to refer to molecules consisting of chains of amino acids, without referring to a specific mode of action, size, three-dimensional structure, or origin.
[0048] The term "gene" refers to a DNA fragment that contains a region (transcription region) that is transcribed into an RNA molecule (e.g., mRNA) within a cell and operably ligated to an appropriate regulatory region (e.g., a promoter). A gene will typically contain several operably ligated fragments, such as a promoter, a 5' leader sequence, a coding region, and a 3' untranslated sequence (3' end) containing polyadenylation sites. "Genetic expression" refers to the process by which the DNA region operably ligated to an appropriate regulatory region, particularly a promoter, is transcribed into RNA that is biologically active, i.e., can be translated into a biologically active protein or peptide.
[0049] When the term "homologous" is used to describe the relationship between a given (recombinant) nucleic acid or polypeptide molecule and a given host organism or host cell, in nature, it is understood to mean that the nucleic acid or polypeptide molecule is produced by a host cell or organism of the same species, preferably the same type or strain. When homologous to a host cell, the nucleic acid sequence encoding the polypeptide is typically (but not necessarily) operably linked to a different (heterogeneous) promoter sequence, and, if applicable, another (heterogeneous) secretory signal sequence and / or terminator sequence, than those in its natural environment. It is understood that regulatory sequences, signal sequences, terminator sequences, etc., may also be homologous to the host cell. In this regard, the use of "homologous" sequence elements alone enables the construction of "self-cloning" genetically modified organisms (GMOs) (self-cloning is defined herein as in Annex II of European Directive 98 / 81 / EC). When used to describe the relationship between two nucleic acid sequences, the term "homologous" means that one single-stranded nucleic acid sequence can hybridize to a complementary single-stranded nucleic acid sequence. The degree of hybridization can depend on many factors, including the amount of identity between sequences and hybridization conditions such as temperature and salt concentration, as will be discussed later.
[0050] The terms “heterogeneous” and “exogenous,” when used in reference to nucleic acids (DNA or RNA) or proteins, refer to nucleic acids or proteins that are not naturally present as part of the organism, cell, genome, or DNA or RNA sequence in which they exist, or that are found in a different cell or location in the genome or DNA or RNA sequence than where they are found naturally. Heterogeneous and exogenous nucleic acids or proteins are not endogenous to the cell into which they are introduced, but are obtained from another cell or produced synthetically or recombinantly. Generally, though not always, such nucleic acids encode proteins that are not normally produced by the cell in which the DNA is transcribed or expressed, i.e., exogenous proteins. Similarly, exogenous RNA encodes proteins that are not normally expressed in the cell in which the exogenous RNA exists. Heterogeneous / exogenous nucleic acids and proteins are sometimes also called foreign nucleic acids or proteins. Any nucleic acid or protein that a person skilled in the art would recognize as foreign to the cell in which it is expressed is included herein in the term heterogeneous or exogenous nucleic acid or protein. The terms heterogeneous and exogenous also apply to unnatural combinations of nucleic acids or amino acid sequences, i.e., combinations in which at least two of the combined sequences are heterogeneous.
[0051] As used herein, the term “not naturally occurring” when used in reference to an organism means that the organism has at least one genetic alteration not typically found in naturally occurring strains of the species mentioned, including wild-type strains of the species mentioned. Genetic alterations include, for example, modifications that introduce expressible nucleic acids encoding proteins or enzymes, other nucleic acid additions, nucleic acid deletions, nucleic acid substitutions, or other functional disruptions of the organism’s genetic material. Such modifications include, for example, coding regions and functional fragments of heterologous or homologous polypeptides of the referenced species. Further modifications include, for example, non-coding regulatory regions in which the modification alters the expression of a gene or operon. Genetic modifications to nucleic acid molecules encoding enzymes or functional fragments thereof can confer biochemical reaction capacity or metabolic pathway capacity to unnaturally occurring organisms that have been altered from their naturally occurring state.
[0052] As used herein, the term “operably linked” refers to the linking of functionally related polynucleotide (or polypeptide) elements. A nucleic acid is “operably linked” if it is functionally related to another nucleic acid sequence. For example, a transcriptional regulatory sequence is operably linked to a coding sequence if it affects the transcription of that coding sequence. Being operably linked means that the linked DNA sequences are typically contiguous and, if necessary, contiguous and within a read frame when two protein-coding regions need to be linked.
[0053] When an expression regulatory sequence controls and regulates the transcription and / or translation of a nucleotide sequence, the expression regulatory sequence is "operably ligated" to the nucleotide sequence. Thus, an expression regulatory sequence may include a promoter, enhancer, internal ribosome entry site (IRES), transcription terminator, start codon prior to a protein-coding gene, splicing signal for an intron, and stop codon.
[0054] The term “expression regulatory sequence” is intended to include, at a minimum, sequences designed so that their presence affects expression, and may also include additional beneficial components. For example, leader sequences and fusion partner sequences are expression regulatory sequences. The term may also include the design of nucleic acid sequences such that undesirable potential start codons inside or outside the frame are removed from the sequence. It may also include the design of nucleic acid sequences such that undesirable potential splice sites are removed. This includes sequences called poly-A tails, i.e., sequences that direct the addition of a series of adenine residues to the 3' end of mRNA, or polyadenylation sequences (pA), or poly-A sequences. They can also be designed to enhance mRNA stability. Expression regulatory sequences that affect transcriptional and translational stability, such as promoters, and sequences that affect translation, such as Kozak sequences, are known in insect cells. Expression regulatory sequences may be of a nature that modulates the nucleotide sequence to which they are operably linked so that lower or higher expression levels are achieved.
[0055] As used herein, the term “library” refers to a collection of elements or members that are distinct from each other in at least one embodiment. For example, a library of nucleic acid molecules or a library of nucleic acid constructs is a collection of at least two nucleic acid molecules or at least two nucleic acid constructs in which at least one nucleotide is distinct from each other. Similarly, a library of AAV capsid variants is a collection of at least two AAV capsid variants (virions) in which at least one nucleotide and / or at least one amino acid in their capsid proteins is distinct from each other.
[0056] As used herein, the term “targeting” refers to the preferential targeting by a virus (e.g., AAV) by cells of a particular host species or by a particular cell type within a host species. For example, a virus that can infect heart, lung, liver, and muscle cells has broader (i.e., increased) targeting than a virus that can infect only lung and muscle cells. Targeting can also include the dependence of a virus on a particular type of cell surface molecule of its host. For example, some viruses can infect only cells that have surface glycosaminoglycans, while others can infect only cells that have sialic acid (such dependence can be tested using various cell lines that lack a particular class of molecules as potential host cells for viral infection). In some cases, viral affinity describes the relative preference of a virus. For example, a first virus can infect all cell types, but is far more successful in infecting those cells with surface glycosaminoglycans. Even if the absolute transduction efficiency of the second virus is not similar, if the second virus also prefers the same characteristics (for example, the second virus also succeeds by infecting these cells with surface glycosaminoglycans), it can be considered to have similar (or identical) directivity to the first virus. For example, if the second virus may be more efficient than the first virus in infecting any given cell type tested, but the relative priorities are similar (or five are identical), the second virus can still be considered to have similar (or identical) directivity to the first virus. In some embodiments, the directivity of virions containing the target variant AAV capsid protein is unchanged compared to naturally occurring virions. In some embodiments, the directivity of virions containing the target variant AAV capsid protein is expanded (i.e., spreads) compared to naturally occurring virions. In some embodiments, the directivity of virions containing the target variant AAV capsid protein is reduced compared to naturally occurring virions.
[0057] References to nucleotide or amino acid sequences available in public sequence databases herein refer to versions of sequence entries available as of the filing date of this specification.
[0058] [Detailed description of the invention] The inventors investigated means and methods for preparing libraries of AAV capsid variants produced in insect cells. They succeeded in developing methods that minimize the occurrence of cross-packaging and mosaicism in the generated libraries while maintaining a method of sufficient depth and complexity for producing AAV libraries in insect cells that minimizes the occurrence of cross-packaging and mosaicism in the libraries.
[0059] In a first embodiment, a nucleic acid construct is provided. In one embodiment, the nucleic acid construct comprises i) a nucleotide sequence encoding an adeno-associated virus (AAV) capsid variant, operably ligated to a promoter for driving expression in an insect cell; ii) an enhancer element operably ligated to the promoter, dependent on a transcriptional transregulator, wherein the introduction of the transcriptional transregulator into an insect cell containing the nucleic acid construct induces transcription from the promoter; and iii) at least one AAV reverse terminal repeat (ITR) sequence.
[0060] In one embodiment, the nucleic acid construct described herein thus comprises an expression cassette for the expression of an AAV capsid variant in insect cells. "AAV capsid variant" is understood herein as an AAV capsid in which at least one of the capsid proteins differs from the corresponding wild-type AAV capsid protein by at least one amino acid.
[0061] In one embodiment, the nucleic acid construct described herein comprises an expression cassette containing a promoter active in insect cells, the promoter operably ligated to a nucleotide sequence encoding mRNA, the translation of which produces an AAV capsid variant in insect cells. Thus, the nucleotide sequence in the expression cassette encodes mRNA including open reading frame translation, which in insect cells produces the AAV capsid protein necessary for the construction of the AAV capsid variant. Therefore, the nucleotide sequence in the expression cassette encodes mRNA including open reading frame translation which produces at least AAV VP1 and VP3 proteins in insect cells, since AAV VP2 is not strictly required for complete viral particle formation and efficient infectivity (Warrington et al, 2004, J Virol. 78:6595-609). In a preferred embodiment, the nucleotide sequence in the expression cassette encodes mRNA including open reading frame translation which produces all three AAV VP1, VP2, and VP3 proteins in insect cells. Therefore, in one embodiment, the expression cassette encodes mRNA containing a single open reading frame that can be translated into at least AAV VP1 and VP3 capsid proteins by leaky scanning of the VP1 translation start codon and inactivation of the VP2 translation start codon. In a preferred embodiment, the expression cassette encodes mRNA containing a single open reading frame that can be translated into all three AAV VP1, VP2, and VP3 capsid proteins by leaky scanning of the VP1 and VP2 translation start codons.
[0062] Therefore, in one embodiment, the AAV capsid variant expression cassette includes a nucleotide sequence containing a single open reading frame encoding all three parvovirus (AAV) VP1, VP2, and VP3 capsid proteins, and the start codon for translation of the VP1 capsid protein is a suboptimal start codon other than ATG, as described, for example, by Urabe et al. (2002, cited above) and in International Publication No. 2007 / 046703. The suboptimal start codon for the VP1 capsid protein may be as defined above for the Rep78 protein. More preferred suboptimal start codons for the VP1 capsid protein may be selected from ACG, TTG, CTG, and GTG, of which CTG and GTG are most preferred. In alternative embodiments, the expression cassette comprises a nucleotide sequence containing a single open reading frame encoding all three parvovirus (AAV) VP1, VP2, and VP3 capsid proteins, where the start codon for translation of the VP1 capsid protein is ATG, and the mRNA encoding the VP1 capsid protein encoded in the nucleotide sequence contains an alternative start codon containing the VP1 capsid protein outside the frame having an open reading frame (as described in International Publication No. 2019 / 016349). Preferably, the alternative start codon is selected from the group consisting of CTG, ATG, ACG, TTG, GTG, CTC, and CTT, of which ATG is preferred. Preferably, the AAV capsid protein is the AAV5 serotype capsid protein. Preferably, in this embodiment, the nucleotide sequence comprises an alternative open reading frame beginning with an alternative start codon containing the ATG translation start codon for VP1, where preferably the alternative open reading frame following the alternative start codon encodes a peptide of up to 20 amino acids.
[0063] In one embodiment of the AAV capsid variant expression cassette described herein, the promoter for driving expression in insect cells is a baculovirus promoter. In one embodiment, the baculovirus promoter is a late or very late baculovirus promoter, and for example, the promoter is selected from the group consisting of polH, p10, p6.9, and pSel120 promoters. In a preferred embodiment, the late or very late baculovirus promoter is the baculovirus p10 promoter.
[0064] In one embodiment of the AAV capsid variant expression cassette described herein, the expression cassette includes a polyadenylation signal, preferably a polyadenylation signal that results in polyadenylation of mRNA expressed from the expression cassette in insect cells.
[0065] In one embodiment, the nucleic acid construct described herein further comprises an enhancer element operably ligated to a promoter (driving the expression of an AAV cap variant in insect cells). “Enhancer element” or “enhancer” defines a sequence that enhances the activity of a promoter (i.e., increases the transcription rate of a sequence downstream of the promoter), and, in contrast to the promoter, has no promoter activity and can usually function regardless of its position relative to the promoter (i.e., whether it is upstream or downstream of the promoter). Enhancer elements are well known in the art. Non-limiting examples of enhancer elements (or parts thereof) that can be used in the present invention include baculovirus enhancers and enhancer elements found in insect cells. In cells, the enhancer element preferably increases the mRNA expression of the gene to which the promoter is operably ligated by at least 25%, more preferably at least 50%, even more preferably at least 100%, and most preferably at least 200%, compared to the mRNA expression of the gene in the absence of the enhancer element. mRNA expression can be determined, for example, by quantitative RT-PCR.
[0066] In this specification, it is preferable to use enhancer elements to enhance the expression of AAV capsid variants. In one embodiment, at least one enhancer element operably ligated to a promoter in the expression cassette of an AAV capsid variant as defined herein is a transcriptional transregulator-dependent enhancer element. A transcriptional transregulator-dependent enhancer element is understood herein as an enhancer element that, when bound by a transcriptional transregulator protein provided in trans, activates the transcription of the promoter operably ligated to it. Thus, in the absence of the transcriptional transregulator protein, there is little to no transcription by the promoter operably ligated to the enhancer.
[0067] Therefore, in one embodiment, the transcriptional transregulator-dependent enhancer element comprises at least one baculovirus enhancer element and / or at least one ecdysone-responsive element. Preferably, the transcriptional transregulator is a baculovirus pre-initial protein (IE1) or its splice variant (IE0), the transcriptional transregulator-dependent enhancer element is a baculovirus homologous region (hr) enhancer element, and preferably the baculovirus is Autographa californica multicapsid nuclear polyhedron disease virus. IE1 is a highly conserved 67 kDa DNA-binding protein that transcriptionally activates the baculovirus early gene promoter and supports late gene expression in plasmid transfection assays (see, e.g., Olson et al., 2002, J Virol., 76:9505-9515). AcMNPV IE1 has separable domains that contribute to promoter transcriptional activation and DNA binding. The N-terminal half of this 582-residue phosphoprotein contains transcriptional stimulation domains at residues 8–118 and 168–222. IE1 binds to an approximately 28-bp incomplete palindrom (28-mer) that constitutes a repeating sequence within numerous homologous regions (hr) found distributed throughout the AcMNPV genome. The hr28-mer is the minimal sequence motif necessary for IE1-mediated enhancer and origin-specific replication function.
[0068] In one embodiment, the hr enhancer element is selected from the group consisting of hr1, hr2-0.9, hr3, hr4b, and hr5, of which hr4b and hr5 are preferred, and of which hr4b is most preferred. In an alternative embodiment, the hr enhancer element is a variant hr enhancer element, such as a designed element that does not exist in nature. The variant hr enhancer element preferably includes at least one copy of the hr28-mer sequence CTTTACGAGTAGAATTCTACGCGTAAAA (SEQ ID NO: 1), and / or at least 18, 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides are identical to the sequence CTTTACGAGTAGAATTCTACGCGTAAAA (SEQ ID NO: 1), and preferably includes at least one copy of a sequence that binds to the baculovirus IE1 protein, more preferably to the AcMNPV IE1 protein. A variant hr enhancer element is functionally defined such that, more preferably, when the variant element is operably ligated to an expression cassette containing a reporter gene operably ligated to a polH promoter, a) under non-inducible conditions, the expression cassette having the variant element produces less reporter transcript than another identical expression cassette containing the hr2-0.9 element instead of the variant element, or the cassette having the variant element produces less than 1.1, 1.2, 1.5, 2, 5, or 10 times the amount of reporter transcript produced by another identical expression cassette containing the hr4b element instead of the variant element; and b) under inducible conditions, the expression cassette having the variant element produces at least 50, 60, 70, 80, 90, or 100% of the amount of reporter transcript produced by another identical expression cassette containing the hr4b or hr2-0.9 element instead of the variant element. Non-inducible conditions are understood as conditions in which the IE1 protein is absent in the cells being tested with the cassette, while induction conditions are understood as conditions in which sufficient IE1 protein is present to obtain maximum reporter expression with a reference cassette containing the hr4b or hr2-0.9 element.The binding of the variant hr enhancer element to the baculovirus IE1 protein can be assayed using a mobility shift assay, for example, as described in Rodems and Friesen (J Virol. 1995; 69(9): 5368-75).
[0069] In one embodiment, the enhancer element is a hr such as hr4b, which is not leaky but is relatively weak. In a preferred embodiment, the hr4b enhancer is combined with a p10 promoter.
[0070] Therefore, in one embodiment, the nucleic acid construct described herein further comprises at least one AAV reverse terminal repeat (ITR) sequence.
[0071] In this specification, “at least one AAV reverse terminal repeat nucleotide sequence” is understood to mean a palindromic sequence containing nearly complementary and symmetrically arranged sequences, also called “A,” “B,” and “C” regions. The ITR functions as an origin of replication and is a “cis” region in replication, namely the palindrome and recognition sequences of trans-acting replication proteins, such as Rep78 (or Rep68), that recognize specific sequences within the palindrome. One exception to the symmetry of the ITR sequence is the “D” region of the ITR. It is unique (it does not have a complement within a single ITR). Nicking of single-stranded DNA occurs at the junction between the A and D regions. This is the region where new DNA synthesis begins. The D region is usually located on one side of the palindrome and provides directionality to the nucleic acid replication steps. Parvoviruses such as AAV that replicate in mammalian cells typically have two ITR sequences. However, it is possible to manipulate the ITR so that the binding sites on both the A and D regions of the strand are symmetrically located, one on each side of the palindrome. On a double-stranded circular DNA template (e.g., a plasmid), Rep78 or Rep68-assisted nucleic acid replication then proceeds in both directions, and a single ITR is sufficient for parvovirus replication of the circular vector. Therefore, in the present invention, one ITR nucleotide sequence can be used. However, preferably, two or another even number of regular ITRs are used. Most preferably, two ITR sequences are used. A preferred ITR is the AAV2 ITR. For safety reasons, it may be desirable to construct a recombinant parvovirus (rAAV) vector that cannot grow further after initial introduction into cells in the presence of a second AAV. Such a safety mechanism for limiting the amplification of an undesirable vector in the recipient may be provided by using an rAAV having a chimeric ITR, as described in U.S. Patent Application Publication No. 2003148506.
[0072] In one embodiment of the nucleic acid construct, a nucleotide sequence comprising an expression cassette of the AAV capsid variant defined above and an enhancer element operably ligated thereto is adjacent to at least one AAV ITR sequence. In this specification, with respect to sequences adjacent to another element, the term “adjacent” indicates that one or more adjacent elements are present upstream and / or downstream of the sequence, i.e., on the 5' side and / or 3' side. The term “adjacent” is not intended to indicate that the sequences are necessarily contiguous. For example, an intervening sequence may exist between the nucleic acid encoding the transgene and the adjacent element. A sequence in which two other elements (e.g., ITRs) are “adjacent” indicates that one element is located on the 5' side of the sequence and the other on the 3' side of the sequence, but an intervening sequence may exist between them. In a preferred embodiment, the nucleotide sequence of (i) is adjacent to parvovirus reverse terminal repeat nucleotide sequences on both sides.
[0073] In one embodiment of the nucleic acid construct, the expression cassette of the AAV capsid variant defined above and the enhancer element operably linked thereto are adjacent to at least one AAV ITR sequence.
[0074] In one embodiment of the nucleic acid construct, a nucleotide sequence comprising an AAV capsid variant expression cassette and an enhancer element operably linked thereto, adjacent to at least one parvovirus ITR sequence, is preferably incorporated into the genome of a recombinant parvovirus (rAAV) vector produced in insect cells. In one embodiment, the nucleotide sequence comprising the AAV capsid variant expression cassette and the enhancer element operably linked thereto has two AAV ITR nucleotide sequences adjacent to each other, thereby positioning the nucleotide sequence comprising the AAV capsid variant expression cassette and enhancer element between the two AAV ITR nucleotide sequences. In one embodiment, if the nucleotide sequence comprising the AAV capsid variant expression cassette and enhancer element is located between two normal ITRs, or on either side of two D-region manipulated ITRs, it can be incorporated into an rAAV vector produced in insect cells. Therefore, in a preferred embodiment, a nucleic acid construct is provided comprising two AAV ITR nucleotide sequences, wherein a nucleotide sequence comprising an AAV capsid variant expression cassette and an enhancer element operably linked thereto is located between the two AAV ITR nucleotide sequences.
[0075] Typically, nucleic acid constructs comprising an ITR and an AAV capsid variant expression cassette and a nucleotide sequence containing an enhancer element operably ligated thereto are no longer than 5,000 nucleotides (nt) to ensure they do not exceed the maximum AAV packaging limit of 5.5 kbp. However, in other embodiments using nucleic acid constructs that are larger than usual, for example, longer than 5,000 nt or exceeding the maximum AAV packaging limit of 5.5 kbp, the production of rAAV vectors is still possible and therefore not excluded.
[0076] For the generation of recombinant AAV virions, i.e., AAV vectors, in insect cells, AAV sequences that can be used as described herein may be derived from the genome of any AAV serotype. Generally, AAV serotypes have genomic sequences with significant homology at the amino acid and nucleic acid levels, provide the same set of gene functions, produce virions that are substantially physically and functionally equivalent, and replicate and assemble by substantially the same mechanisms. For an overview of the genomic sequences and genomic similarities of various AAV serotypes, see, for example, GenBank accession numbers U89790, J01901, AF043303, AF085716, Chlorini et al. (1997, J.Vir.71:6823-33), Srivastava et al. (1983, J.Vir.45:555-64), Chlorini et al. (1999, J.Vir.73:1309-1319), Rutledge et al. (1998, J.Vir.72:309-319), and Wu et al. (2000, J.Vir.74:8635-47). Any AAV serotype can be used as a source of AAV nucleotide sequences for use in the context of this invention. The AAV ITR sequences for use in connection with the present invention are preferably derived from AAV1, AAV2, AAV4 and / or AAV7. Similarly, the Rep(Rep78 / 68 and Rep52 / 40) coding sequences are preferably derived from AAV1, AAV2, AAV4 and / or AAV7. Sequences encoding AAV capsid proteins for use herein are described in more detail below.
[0077] AAV Rep and ITR sequences are particularly conserved among most serotypes. Rep78 proteins from various AAV serotypes are, for example, over 89% identical, and the total nucleotide sequence identity at the genomic level between AAV2, AAV3A, AAV3B, and AAV6 is approximately 82% (Bantel-Schaal et al., 1999, J. Virol., 73(2):939-947). Furthermore, many AAV serotype Rep and ITR sequences are known to efficiently cross-complement (i.e., functionally substitute) corresponding sequences from other serotypes in the production of AAV particles in mammalian cells. U.S. Patent Application Publication No. 2003148506 also reports that AAV Rep and ITR sequences efficiently cross-complement other AAV Rep and ITR sequences in insect cells. Modified “AAV” sequences can also be used in this context, for example, for the production of rAAV vectors in insect cells. Examples of such modified sequences include sequences having at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 95%, or more nucleotide and / or amino acid sequence identity with AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, or AAV13ITR (for example, sequences with approximately 75-99% nucleotide sequence identity), or Rep can be used instead of the wild-type AAV ITR or Rep sequence.
[0078] While similar in many respects to other AAV serotypes, AAV5 differs from other known human and monkey AAV serotypes more than other known human and monkey AAV serotypes. From this perspective, the production of rAAV5 may differ from the production of other serotypes in insect cells. When using the method of the present invention to produce rAAV5, one or more constructs, or collectively in the case of multiple constructs, preferably comprise a nucleotide sequence containing an AAV5 ITR, wherein the nucleotide sequence comprises an AAV5 Rep coding sequence (i.e., the nucleotide sequence comprises AAV5 Rep78). Such ITR and Rep sequences can be modified as necessary to efficiently produce rAAV5 or pseudotyped rAAV5 vectors in insect cells. For example, the production of rAAV5 vectors in insect cells can be improved by modifying the start codon of the Rep sequence, modifying or removing the VP splice site, and / or modifying the VP1 start codon and nearby nucleotides.
[0079] In further embodiments, the nucleic acid constructs described herein optionally include a reporter gene operably linked to a promoter for driving expression in mammalian cells. As understood herein, the reporter gene preferably encodes a protein that, when expressed in mammalian cells, generates a colorimetric signal, a fluorescent signal, or a luminescent signal that enables cells containing and expressing the reporter gene to be distinguished from surrounding cells that do not contain the reporter gene.
[0080] In one embodiment, the reporter gene encodes a protein that generates a colorimetric signal, such as a gene encoding secreted alkaline phosphatase (SEAP) or β-galactosidase.
[0081] In one embodiment, the reporter gene encodes a protein that generates a fluorescent signal, preferred examples of which include EGFP, mCherry, mCloverS, mRubyS, mApple, iRFP, tdTomato, mVenus, YFP, and RFP.
[0082] In one embodiment, the reporter gene encodes a protein that generates a luminescence signal, preferred examples of which include genes encoding Photinus pyralis (firefly) luciferase, Renilla reniformis (sea coral) luciferase, and nanoluciferase.
[0083] In one embodiment of the nucleic acid construct described herein, the promoter operably ligated to the reporter gene to drive its expression in mammalian cells is a constitutive promoter. In other embodiments, the promoter operably ligated to the reporter gene is a cell-type and / or tissue-specific promoter. As will be understood by those skilled in the art, the cell-type and / or tissue specificity of the promoter depends on the desired cell-type and / or tissue orientation of the resulting AAV capsid variant. In some embodiments, the cell-type and / or tissue-specific promoter for driving the expression of the reporter gene in mammalian cells is a promoter selected from the group consisting of the human synapsin promoter (hSynl), the trans tyretin promoter (TTR), the cytokeratin 18, cytokeratin 19, the unc-45 myosin chaperone B (unc45b) promoter, the cardiac troponin T (cTnT) promoter, the glial fiber acidic protein (GFAP) promoter, the myelin basic protein (MBP) promoter, and the methyl-CpG binding protein 2 (Mecp2) promoter. In a preferred embodiment, the cell type and / or tissue-specific promoter is the human synapsin promoter (hSynl).
[0084] In embodiments comprising a reporter gene operably ligated to a promoter for driving the expression of a nucleic acid construct described herein in mammalian cells, the promoter-ligated reporter gene is flanked by at least one AAV ITR sequence, preferably between two AAV ITR sequences, and the promoter-ligated reporter gene, together with an AAV capsid variant expression cassette and a nucleotide sequence operably ligated thereto, and an optional barcode, is incorporated into the genome of an rAAV vector produced in insect cells.
[0085] In further embodiments, the nucleic acid constructs described herein optionally include a barcode. The barcode is preferably at least one of a sample identifier and a unique molecular identifier (UMI).
[0086] A UMI is a substantially unique, preferably fully unique, sequence or barcode specific to a nucleic acid molecule, i.e., unique to each nucleic acid molecule or construct used herein. A UMI may have random, pseudo-random, partially random, or non-random nucleotide sequences. A UMI can be used to uniquely identify the original molecule from which a sequencing read originates. For example, reads of amplified nucleic acid molecules can be folded from each original nucleic acid molecule into a single consensus sequence. As described above, a UMI may be fully or substantially unique. All nucleic acid molecules or constructs provided herein should be understood as fully unique, containing a unique tag distinct from all other tags contained in further nucleic acid molecules / constructs used herein. Substantially unique should be understood herein to mean that each nucleic acid molecule or construct provided herein contains a random UMI, but these nucleic acid molecules / constructs may contain the same UMI at a low percentage. Preferably, a substantially unique molecular identifier is used when the opportunity to tag exactly the same molecule containing the same sequence with the same UMI is negligible. Preferably, the UMI is completely unique with respect to a particular sequence of nucleic acid molecule or construct, for example, completely unique with respect to a particular capsid variant encoded by a particular sequence of nucleic acid molecule or construct. The UMI is preferably long enough to guarantee this uniqueness. The identifier sequence may be about 2 to 100 nucleotides or longer, preferably about 4 to 18 nucleotides long, and usually about 8 to 12 nucleotides long. Preferably, the identifier sequence does not contain two or more consecutive identical bases. Furthermore, preferably there are differences of at least two, preferably at least three, bases between the individual identifier sequences.
[0087] In embodiments of the nucleic acid construct described herein that optionally include a barcode, the barcode is located adjacent to, preferably between, at least one AAV ITR sequence, and preferably between two AAV ITR sequences, and as a result, the barcode is incorporated into the genome of an rAAV vector produced in insect cells, together with a nucleotide sequence comprising an AAV capsid variant expression cassette and an enhancer element operably linked thereto, and optionally a reporter gene.
[0088] In one embodiment, the nucleic acid construct described herein further comprises an endonuclease recognition sequence, preferably a site-specific endonuclease recognition sequence. Thus, the recognition sequence may be a restriction endonuclease or a programmable endonuclease, such as an RNA-induced CRISPR endonuclease, a zinc finger nuclease, a TAL effector nuclease, or a meganuclease. In a preferred embodiment, the recognition sequence is a protelomerase recognition sequence.
[0089] Protelomerase is a site-specific endonuclease that cleaves (restricts or digests) double-stranded nucleic acid molecules at its recognition sequence, and simultaneously covalently closes the double-stranded ends resulting from the cleavage of the nucleic acid molecule. Suitable protelomerases for use in the method described herein are protelomerases selected from the group consisting of bacteriophage protelomerases, such as phiHAP-1 from Halomonas aquamarina, PY54 from Yersinia enterolytica, phiKO2 from Klebsiella oxytoca, VP882 from Vibrio sp., and N15 (TelN) from Escherichia coli, or any variant thereof. Protelomerases may have amino acid sequences and protelomerase recognition sequences as disclosed in International Publication No. 2010 / 086626 (this document is incorporated herein by reference). Therefore, in one embodiment, the nucleic acid construct described herein includes the recognition sequence of the above-mentioned bacteriophage protelomerase. Preferably, the nucleic acid construct includes the recognition sequence of TeIN protelomerase. In one embodiment, the recognition sequence of TeIN protelomerase includes or consists of a sequence having at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the following sequences: 5'-TATCAGCACACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGTGCTGATA-3'(Sequence ID 2).
[0090] In one embodiment, the recognition sequence of the endonuclease is unique within the nucleic acid construct described herein; that is, there is only one recognition sequence within the nucleic acid construct.
[0091] In one embodiment of the nucleic acid construct described herein, the endonuclease recognition sequence is not adjacent to at least one AAV ITR sequence, preferably not between two AAV ITR sequences, and as a result, the recognition sequence is not incorporated into the genome of the rAAV vector produced in insect cells.
[0092] In one embodiment, the nucleic acid construct described herein comprises, in order from 5' to 3', a) a first AAV ITR sequence; b) an enhancer element; c) an expression cassette comprising a nucleotide sequence encoding an AAV capsid variant, operably linked to a promoter for driving expression in insect cells; and d) a second AAV ITR, wherein a (unique) recognition sequence is located at least one of i) upstream of the first AAV ITR and ii) downstream of the second AAV ITR, and optionally, a reporter gene operably linked to a promoter for driving expression in mammalian cells is located between a) and b), between b) and c), or between c) and d), and optionally, a barcode is located between a) and b), between b) and c), or between c) and d).
[0093] In one embodiment, the nucleic acid construct described herein is a DNA construct, preferably a double-stranded DNA construct.
[0094] In a second embodiment, nucleic acid molecules comprising the nucleic acid constructs described herein are provided. In one embodiment, the nucleic acid molecule is a DNA molecule, preferably a double-stranded or duplex DNA molecule. In one embodiment, the nucleic acid molecule is a cyclic nucleic acid molecule.
[0095] In one embodiment, the nucleic acid molecule described herein is a linear nucleic acid molecule, preferably a linear double-stranded or double-stranded DNA molecule. In one embodiment, the linear nucleic acid molecule is obtained or can be obtained by digestion with an endonucleases that cleave at a recognition sequence in the nucleic acid construct. In one embodiment, the endonucleases are site-specific endonucleases, and the site-specific endonucleases and their recognition sequences are preferably as defined above. Thus, it is understood that the linear nucleic acid molecule is obtained or can be obtained by digesting a “precursor” molecule containing a nucleic acid construct with an endonucleases that cleave at a recognition sequence in the nucleic acid construct. The precursor molecule may be, for example, a cyclic molecule, or the precursor molecule may be a long chain molecule containing many copies of a continuously linked DNA construct that is the product of rolling circle amplification (see the method described below).
[0096] In one embodiment, the nucleic acid molecule described herein is a linear nucleic acid molecule in which at least one end of the linear nucleic acid molecule is protected from exonuclease degradation, preferably both ends of the linear nucleic acid molecule are protected from exonuclease degradation. The ends of the linear nucleic acid molecule can be protected from exonuclease degradation by means and methods known in the art. For example, the ends of the linear nucleic acid molecule can be chemically protected or protected by ligation of an adapter which includes chemical modifications that protect against exonucleases, such as a phosphorothioate-containing skeleton.
[0097] In one embodiment, the nucleic acid molecule described herein is a linear nucleic acid molecule, and at least one end, preferably both ends, of the linear nucleic acid molecule is protected from exonuclease degradation by covalent closure.
[0098] In this specification, a double-stranded nucleic acid end in which the 3' terminal nucleotide of each upper strand is covalently bonded to the 5' terminal nucleotide of each lower strand is referred to as a “closed end.” Similarly, a double-stranded nucleic acid end in which the 5' terminal nucleotide of each upper strand is covalently bonded to the 3' terminal nucleotide of each lower strand is also referred to as a “closed end.” Thus, in this specification, a “closed end” is understood as the end of a double-stranded nucleic acid in which the terminal nucleic acids from the opposite strands are covalently bonded to each other, in contrast to an “open end” in this specification, which is understood as the end of a double-stranded nucleic acid in which the terminal nucleic acids from the opposite strands are not covalently bonded to each other.
[0099] In one embodiment, the nucleic acid molecule described herein is a linear nucleic acid molecule, wherein at least one end, preferably both ends, of the linear nucleic acid molecule is protected from exonuclease degradation by covalent closure, and the covalently closed ends are obtained or can be obtained by digestion with the above-defined protelomerase, preferably by digestion with TeIN protelomerase.
[0100] In one embodiment, the nucleic acid molecule described herein comprises a nucleic acid construct described herein, wherein the position of the recognition sequence of an AAV Rep protein in the insect cell is such that the expression of the AAV Rep protein in the insect cell contains the nucleic acid construct is such that the expression of the AAV Rep protein in the insect cell contains the nucleic acid construct results in the replication and packaging of the DNA fragment containing elements i) to v) into the AAV capsid.
[0101] In a third embodiment, a library is provided comprising a plurality of nucleic acid molecules, each comprising a nucleic acid construct described herein. In one embodiment, the library comprises a plurality of nucleic acid molecules, each comprising a nucleic acid construct described herein, wherein in each member of the library, the nucleotide sequence encoding the AAV capsid variant differs by at least one nucleotide from the nucleotide sequence encoding the AAV capsid variant in another member of the library. In one embodiment, the library comprises a plurality of nucleic acid molecules, each comprising a nucleic acid construct described herein, wherein in each member of the library, the nucleotide sequence encoding the AAV capsid variant differs by at least one nucleotide from the AAV capsid variant encoded by another member of the library.
[0102] In one embodiment, the nucleotide sequence encoding the AAV capsid variant in the library comprises one or more amino acid mutations (e.g., insertions, deletions, and / or substitutions) that encode the AAV capsid variant.
[0103] In one embodiment, the nucleotide sequence encoding the AAV capsid variant in the library encodes an AAV capsid variant that has at least one amino acid difference (mutation) compared to the wild-type AAV capsid protein. The wild-type AAV capsid protein is understood herein as an AAV capsid protein having the amino acid sequence of a naturally occurring AAV isolate, such as a clinical isolate.
[0104] In one embodiment, the nucleotide sequence encoding the AAV capsid variant in the library encodes an AAV capsid variant that, compared to the wild-type AAV capsid protein, contains at least one amino acid difference (mutation) in the hypervariable and / or surface-exposed loop of the capsid protein variant.
[0105] In one embodiment, the nucleotide sequence encoding the AAV capsid variant in the library encodes an AAV capsid variant that, compared to the wild-type AAV capsid protein, contains at least one amino acid difference (mutation) in the VR region of the surface loop of the capsid protein variant, for example, VR-I, VR-II, VR-III, VR-IV, VR-V, VR-VI, VR-VII, VR-VIII, and VR-IX (of which VR-IV and VR-VIII are preferred, and VR-VIII is more preferred).
[0106] In one embodiment, the above-mentioned difference (mutation) of at least one amino acid compared to the wild-type AAV capsid protein includes the insertion of a peptide. The inserted peptide may contain 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more amino acids.
[0107] In one embodiment, the above-mentioned difference (mutation) of at least one amino acid compared to the wild-type AAV capsid protein is introduced into AAV capsid proteins VP1, VP2, or VP3, or into any combination of two capsid proteins, or into all three. In some embodiments, the mutation is introduced into VP1. In some embodiments, the mutation is introduced into VP2. In some embodiments, the mutation is introduced into VP3. In some embodiments, the mutation is introduced into VP1 and VP2. In some embodiments, the mutation is introduced into VP1 and VP3. In some embodiments, the mutation is introduced into VP2 and VP3. In some embodiments, the mutation is introduced into VP1, VP2, and VP3.
[0108] In one embodiment, the above-mentioned difference (mutation) of at least one amino acid compared to the wild-type AAV capsid protein is a single mutation (e.g., an amino acid insertion, deletion, and / or substitution) introduced at a single site in the amino acid sequence of the capsid protein, while in other embodiments, more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 100 or more (including any number between 1 and 100 or more) mutations (e.g., insertions, deletions, and / or substitutions) are introduced into the amino acid sequence of the capsid protein.
[0109] In one embodiment, the nucleotide sequence encoding the AAV capsid variant in the library encodes an AAV capsid variant that has at least one amino acid difference (mutation) compared to the wild-type AAV capsid protein, and the wild-type AAV capsid protein may be a serotype capsid protein selected from the group consisting of AAV1, AAV2, AAV3A, AAV3B, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, and AAVrh10. In further embodiments, the wild-type AAV capsid protein may be the capsid protein of clade A, clade B, clade C, clade D, clade E, or clade FAAV isolate as defined by Gao et al. (2004, J. Virol. 78: 6381-6388), the exemplary sequences of which are disclosed in International Publication No. 2005 / 033321 and are available from the GenBank® database under accession numbers AY530553 to AY530629.
[0110] In a fourth aspect, insect cells are provided that include a nucleic acid molecule containing a nucleic acid construct described herein, or a library containing (a plurality of) such nucleic acid molecules containing a nucleic acid construct described herein.
[0111] In one embodiment, the insect cells may be any cells suitable for the production of heterologous proteins. Preferably, the insect cells can enable the replication of baculovirus vectors and be maintained in a culture, more preferably in a suspension culture. In a preferred embodiment, the insect cells enable the replication of recombinant parvovirus vectors, including (r)AAV vectors. For example, the cell lines used may be derived from cell lines of Spodoptera frugiperda, Drosophila, or mosquitoes, such as Aedes albopictus. Preferred insect cells or cell lines are cells derived from insect species susceptible to baculovirus infection, including, for example, S2 (CRL-1963, ATCC), Se301, SeIZD2109, SeUCR1, Sf9, Sf900+, Sf21, BTI-TN-5B1-4, MG-1, Tn368, HzAm1, Ha2302, Hz2E5, High Five (Invitrogen, CA, USA), and expressSF+ (registered trademark) (U.S. Patent No. 6,103,526; Protein Sciences Corp., CT, USA). Preferred insect cells according to the present invention are insect cells for producing recombinant parvovirus vectors, more specifically recombinant AAV vectors.
[0112] The conditions for the proliferation of insect cells in culture, and the production of heterologous products from insect cells in culture, are well known in the art and are described, for example, in the above-mentioned cited literature on the molecular engineering of insect cells (see also International Publication No. 2007 / 046703).
[0113] In one embodiment, an insect cell comprising a nucleic acid molecule as described herein, or a library as described herein, further comprising a nucleic acid construct comprising at least one expression cassette for the expression of the AAV Rep protein in the insect cell, and a nucleic acid construct comprising at least one expression cassette for the expression of a transcriptional transregulator (e.g., baculovirus pre-initial protein (IE1) or its splice variant (IE0)). In a preferred embodiment, the single nucleic acid construct comprises both an expression cassette for the expression of the AAV Rep protein and an expression cassette for the expression of a transcriptional transregulator. More preferably, since the baculovirus vector already comprises at least one expression cassette for the expression of the baculovirus pre-initial protein (IE1) and / or its splice variant (IE0), the single nucleic acid construct is a baculovirus vector comprising at least one expression cassette for the expression of the AAV Rep protein in the insect cell.
[0114] AAV replicases are non-structural proteins encoded by the rep gene. In wild-type parvovirus, the rep gene, due to its internal P19 promoter, produces two duplicated messenger ribonucleic acids (mRNAs) of different lengths. Each of these mRNAs may or may not be spliced, ultimately resulting in four Rep proteins: Rep78, Rep68, Rep52, and Rep40. Rep78 / 68 and Rep52 / 40 are crucial for ITR-dependent AAV genome or transgene replication and viral particle assembly. Rep78 / 68 acts as a viral replication initiation protein and functions as a replicase of the viral genome (Chejanovsky and Carter, J Virol., 1990, 64:1764-1770; Hong et al., Proc Natl Acad Sci USA, 1992, 89:4673-4677; Ni., et al., J Virol., 1994, 68:1128-1138). Rep52 / 40 proteins are 3' to 5' polar DNA helicases that play a crucial role in packaging viral DNA into empty capsids and are thought to be part of the packaging motor complex (Smith and Kotin, J. Virol., 1998, 4874-4881; King, et al., EMBO J., 2001, 20:3282-3291). The presence of both Rep68 and Rep40 is not essential for AAV production from baculovirus vectors in insect cell platforms (Urabe, et al., 2002).
[0115] A nucleotide sequence encoding one or more AAV Rep proteins is understood herein as a nucleotide sequence encoding at least one of two non-structural Rep proteins, Rep78 and Rep52, which together are necessary and sufficient for parvovirus vector production in insect cells. The AAV nucleotide sequence is preferably derived from human or monkey AAV, most preferably from AAV that typically infects humans (e.g., serotypes 1, 2, 3A, 3B, 4, 5, 6, 8, and 9) or primates (e.g., serotypes 1 and 4). Examples of nucleotide sequences encoding parvovirus Rep proteins are shown in SEQ ID NOs: 3-9.
[0116] It is understood that the exact molecular weights of the Rep78 and Rep52 proteins, as well as the exact location of the translation start codon, may differ among different parvoviruses. However, those skilled in the art know how to identify the corresponding locations in nucleotide sequences derived from other parvoviruses besides AAV-2. Preferably, this nucleotide sequence encodes a parvovirus Rep protein, which is functionally active in the sense that it is a viral replication initiation protein, a replicase of the viral genome, a DNA helicase, and has the necessary activity to package the viral DNA into an empty capsid as described above, and is sufficient for the production of parvovirus vectors in insect cells. In one embodiment, possible false translation start sites in the Rep protein coding sequence other than the Rep78 and Rep52 translation start sites are excluded. In one embodiment, putative splice sites that may be recognized in insect cells are excluded from the Rep protein coding sequence. The exclusion of these sites will be well understood by those skilled in the art.
[0117] In one embodiment, a nucleic acid construct expressing parvovirus Rep proteins comprises a single expression cassette for the expression of at least both parvovirus Rep78 and Rep52 proteins. In one embodiment, the single expression cassette for the expression of at least both parvovirus Rep78 and Rep52 proteins comprises a single open reading frame encoding at least both parvovirus Rep78 and Rep52 proteins and having a suboptimal translation start codon for the Rep78 coding sequence, which results in partial exon skipping so that at least both parvovirus Rep78 and Rep52 proteins are translated in insect cells, for example, as described in U.S. Patent No. 8,512,981, incorporated herein by reference. Suitable suboptimal translation start codons include, for example, ACG, CTG, TTG, and GTG. In another embodiment, a single expression cassette for the expression of at least both parvovirus Rep78 and Rep52 proteins comprises, in order from 5' to 3', (i) a first promoter operably ligated to the 5' portion of a first open reading frame of the parvovirus Rep78 protein, the first open reading frame containing a translation start codon, and (ii) an intron containing a second insect cell promoter, the second promoter operably ligated to the 5' portion of at least one additional open reading frame of the parvovirus Rep52 gene, the at least one additional open reading frame containing at least one additional translation start codon, and an intron overlapping the 3' portion of the first open reading frame (as described, for example, in U.S. Patent No. 8,945,918, incorporated herein by reference).
[0118] In another embodiment, the nucleic acid construct for the expression of parvovirus Rep proteins comprises at least two separate expression cassettes, one for the expression of at least parvovirus Rep78 protein and the other for the expression of at least parvovirus Rep52 protein. Preferably, in this embodiment, the parvovirus Rep78 protein and the parvovirus Rep52 protein contain a common amino acid sequence including the amino acids from the second amino acid to the C-terminal amino acid of the parvovirus Rep52 protein, and the common amino acid sequence of the parvovirus Rep78 protein and the parvovirus Rep52 protein is identical by at least 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100%, and the nucleotide sequence encoding the common amino acid sequence of the parvovirus Rep78 protein and the parvovirus Rep The nucleotide sequences encoding the common amino acid sequence of the 52 proteins are 90%, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, and less than 60% identical (for example, as described in U.S. Patent No. 8,697,417, incorporated herein by reference). In further embodiments, the nucleotide sequences encoding the common amino acid sequence of the parvovirus Rep78 protein exhibit improved cellular codon use bias compared to the nucleotide sequences encoding the common amino acid sequence of the parvovirus Rep52 protein. However, preferably, the nucleotide sequences encoding the common amino acid sequence of the parvovirus Rep52 protein exhibit improved cellular codon use bias compared to the nucleotide sequences encoding the common amino acid sequence of the parvovirus Rep78 protein.Preferably, the difference in codon adaptation index (as defined above) between the nucleotide sequence encoding the common amino acid sequence in the parvovirus Rep78 protein and the parvovirus Rep52 protein is at least 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0, and more preferably, the CAI of the nucleotide sequence encoding the common amino acid sequence in the parvovirus Rep52 protein is at least 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0.
[0119] In one embodiment, the nucleotide sequence encoding the parvovirus Rep78 protein is Sequence ID No. 9, which encodes the wild-type AAV Rep78 protein, and the nucleotide sequence encoding parvovirus Rep52 is selected from one of Sequence ID Nos. 4 to 7, each of which is modified to have a different codon usage than the wild-type Rep78 encoding sequence ID No. 9. In a preferred embodiment, the nucleotide sequence encoding the parvovirus Rep78 protein is Sequence ID No. 9, which is used in combination with Sequence ID No. 7 as the nucleotide sequence encoding parvovirus Rep52, the latter being modified to differ as much as possible from Sequence ID No. 9 in codon usage.
[0120] In one embodiment, two separate expression cassettes for the Rep78 and Rep52 proteins in insect cells are optimized so that the Rep78 and Rep52 proteins in the cell are obtained in a desired molar ratio. Preferably, the combination of the Rep78 and Rep52 expression cassettes in the cell results in a Rep78 to Rep52 molar ratio in the (insect) cell range of 1:10 to 10:1, 1:5 to 5:1, or 1:3 to 3:1. More preferably, the combination of the Rep78 and Rep52 expression cassettes results in a Rep78 to Rep52 molar ratio of at least 1:2, 1:3, 1:5, or 1:10. The molar ratio of Rep78 to Rep52 can be determined by Western blotting, preferably using a monoclonal antibody that recognizes a common epitope of both Rep78 and Rep52, or by using, for example, a mouse anti-Rep antibody (303.9, Progen, Germany; dilution 1:50). The desired molar ratio of Rep78 to Rep52 can be obtained by selecting promoters in Rep78 and Rep52 expression cassettes, respectively, as further described below herein. Alternatively, or in combination, the desired molar ratio of Rep78 to Rep52 can be obtained by using means to reduce the steady-state level of at least one of the parvovirus Rep78 and 52 proteins. Thus, in one embodiment, the nucleotide sequence encoding the mRNA of the parvovirus Rep protein includes modifications that affect the reduction of the steady-state level of the parvovirus Rep protein. The reduction of steady-state conditions can be achieved, for example, by cleaving regulatory elements or upstream promoters (Urabe et al., cited above; Dong et al., cited above), adding proteolytic signal peptides such as PEST or ubiquitinated peptide sequences, substituting a more suboptimal start codon, or introducing artificial introns as described in International Publication No. 2008 / 024998. When using two separate expression cassettes for the Rep78 and Rep52 proteins in insect cells, the promoter in the Rep52 cassette is preferably stronger than the promoter in the Rep78 cassette.In one embodiment, the promoters of the Rep78 cassette and the Rep52 cassette are baculovirus promoters. In one embodiment, the promoters of the Rep78 cassette and the Rep52 cassette are different. In one embodiment, the Rep78 promoter is a delayed early baculovirus promoter such as the 39k promoter. In one embodiment, the Rep52 promoter is a late or very late baculovirus promoter such as the polH, p10, p6.9, and pSel120 promoters. In one embodiment, the late or very late baculovirus promoter used in the Rep52 cassette is a different promoter from the promoter used in the nucleic acid construct defined above, which includes an expression cassette for capsid protein expression.
[0121] In a preferred embodiment, the nucleotide sequence encoding at least one parvovirus Rep protein includes an open reading frame beginning with a suboptimal translation start codon. The suboptimal start codon is preferably a start codon that affects partial exon skipping. Partial exon skipping is understood herein to mean that at least a portion of the ribosome does not initiate translation at the suboptimal start codon of the Rep78 protein but can initiate at a further downstream start codon, preferably the further downstream (first) start codon being the start codon of the Rep52 protein. Alternatively, the nucleotide sequence encoding the parvovirus Rep protein includes an open reading frame beginning with a suboptimal translation start codon and having no further downstream start codons. The suboptimal start codon preferably affects partial exon skipping during the expression of the nucleotide sequence in insect cells. Preferably, the suboptimal start codon influences partial exon skipping in insect cells such that the molar ratio of Rep78 to Rep52 in insect cells is in the range of 1:10–10:1, 1:5–5:1, or 1:3–3:1. The molar ratio of Rep78 to Rep52 can be determined by Western blotting, preferably using a monoclonal antibody that recognizes a common epitope of both Rep78 and Rep52, or by using, for example, a mouse anti-Rep antibody (303.9, Progen, Germany; dilution 1:50).
[0122] In this specification, the term “suboptimal start codon” refers not only to the trinucleotide start codon itself but also to the situation in which it occurs. Therefore, a suboptimal start codon may consist of an “optimal” ATG codon in a suboptimal situation, such as a non-Kozak situation. However, more preferable is a trinucleotide start codon that is suboptimal itself, i.e., a suboptimal start codon that is not ATG. In this specification, suboptimal is understood to mean that the codon has a lower translation initiation efficiency compared to a normal ATG codon under otherwise identical circumstances. Preferably, the efficiency of a suboptimal codon is less than 90, 80, 60, 40, or 20% of the efficiency of a normal ATG codon under otherwise identical circumstances. Methods for comparing the relative efficiency of translation initiation are known to those skilled in the art. Preferred suboptimal start codons may be selected from ACG, TTG, CTG, and GTG. More preferable is ACG. In this specification, the nucleotide sequences encoding parvovirus Rep proteins are understood as nucleotide sequences encoding non-structural Rep proteins, such as Rep78 and Rep52 proteins, that are necessary and sufficient for the production of parvovirus vectors in insect cells.
[0123] In a fifth aspect, a library is provided comprising a plurality of AAV capsid variants, wherein the members of the library comprise nucleic acids replicated and packaged from members of a library of nucleic acid molecules comprising nucleic acid constructs, as described herein, for example, in a third aspect. Thus, the library comprises or consists of a plurality of AAV vector virions comprising a plurality of AAV capsid variants and nucleic acid constructs. Thus, each member of the library comprises an AAV capsid variant comprising one or more amino acid mutations (e.g., insertions, deletions and / or substitutions). Preferably, each member of the library comprises an AAV capsid variant comprising mutations compared to the wild-type AAV capsid protein. Mutations in a capsid variant as defined herein above. In one embodiment, the library comprises at least 10 4 , 10 5 , 10 6, 10 7 , 10 8 or more different AAV capsid variant genes.
[0124] A sixth aspect provides a method for preparing a library of AAV capsid variants. In one embodiment, a method for preparing a library of AAV capsid variants includes: a) providing a library comprising a plurality of nucleic acid constructs as defined above, wherein the amino acid sequence of each member AAV capsid variant encoded in the member construct differs by at least one amino acid from the amino acid sequence of the AAV capsid variant encoded in the other member constructs in the library; b) amplifying the library of nucleic acid constructs, preferably obtained by rolling circle amplification; c) digesting the amplified library of nucleic acid constructs obtained in b) by a site-specific endonuclease that cleaves at recognition sequences in the nucleic acid constructs to produce a library comprising unit-length linear amplified nucleic acid constructs; d) transfecting insect cells with the library comprising the linear amplified nucleic acid constructs obtained in c); e) infecting insect cells with a baculovirus vector comprising at least one expression cassette for expressing AAV Rep protein in the insect cells and culturing the cells; and f) recovering the library of AAV capsid variants produced by the insect cells in e).
[0125] In one embodiment, a method for preparing a library of AAV capsid variants includes providing a library comprising a plurality of nucleic acid constructs as described herein, wherein in each member of the library, the nucleotide sequence encoding an AAV capsid variant encodes a capsid variant that differs from the AAV capsid variant encoded by the other members of the library by at least one amino acid, thereby the AAV capsid variants and their differences may be as described herein.
[0126] In one embodiment, in one embodiment of a method for preparing a library of AAV capsid variants, the library in step a) is provided by i) providing the nucleic acid construct as defined above (as a template or raw material); and ii) replacing a fragment in the construct containing at least a portion of the nucleotide sequence encoding an AAV capsid variant with a library of fragments, wherein the amino acid sequence of the AAV capsid variant of each member of the library is encoded in the fragment of that member and differs by at least one amino acid from the amino acid sequence of the AAV capsid variant encoded in the construct of the other members of the library. In one embodiment of the method, the replacement of the fragment containing at least a portion of the nucleotide sequence encoding an AAV capsid variant in step ii) is performed by Gibson assembly (Gibson et al., 2009, Nat Methods. 6(5):343-5). It is preferable to use an in vitro recombinant DNA cloning method such as Gibson assembly to avoid bias in the cloning of some library members compared to others, as can occur in in vivo biological host systems.
[0127] In one embodiment, a method for preparing a library of AAV capsid variants includes amplifying a library of nucleic acid constructs. As used herein in relation to nucleic acids or nucleic acid reactions, “amplification” is understood to refer to an in vitro method for producing copies or aggregates of specific nucleic acids, such as nucleic acid construct members in a library. Many methods for amplifying nucleic acids are known in the art, and amplification reactions include polymerase chain reaction, ligase chain reaction, strand displacement amplification, rolling circle amplification, transcription-mediated amplification methods such as NASBA (e.g., U.S. Patent No. 5,409,818), loop-mediated amplification methods (e.g., “LAMP” amplification using loop-forming sequences, e.g., U.S. Patent No. 6,410,278), and isothermal amplification. The present method amplifies nucleic acid construct members in a library to obtain a sufficient quantity of nucleic acid construct members for transfection of insect cells. It is preferable to use an in vitro amplification method to avoid bias in the amplification of some library members compared to others, as can occur in an in vivo biological host system. In one embodiment, the library is amplified by rolling circle amplification using Φ29 DNA polymerase. Phi29 DNA polymerase is a replication polymerase derived from the Bacillus subtilis phage phi29 (Φ29), which has exceptional strand substitution and sequential synthesis properties and possesses intrinsic 3'→5' proofreading exonuclease activity. In one embodiment, random primers, such as random hexamer primers, are used for priming the DNA polymerase, such as Φ29 DNA polymerase.
[0128] In one embodiment, a method for preparing a library of AAV capsid variants includes digesting the amplified library of nucleic acid constructs obtained in b) with a site-specific endonuclease that cleaves at a recognition sequence in the nucleic acid construct to produce unit-length linear amplified nucleic acid constructs. The site-specific endonuclease cleaves at the endonuclease recognition sequence in the nucleic acid construct. This recognition sequence is preferably uniquely present in the nucleic acid construct. The site-specific endonuclease, its recognition sequence, and its location in the construct may be as described herein above.
[0129] In one embodiment, a method for preparing a library of AAV capsid variants is a method in which, in step c), the ends of a linear amplification library of unit-length nucleic acid constructs are protected against exonuclease degradation. Protection against exonuclease degradation may be provided herein by means of the methods and means described above.
[0130] In one embodiment, the site-specific endonuclease is a protelomerase that provides a covalently closed end to a linear amplification library of unit-length nucleic acid constructs. The protelomerase is preferably phage N15 TelN protelomerase. Thus, in this embodiment, the nucleic acid construct includes a recognition sequence for the protelomerase in question, which is located in the construct at the position described herein as above.
[0131] A seventh aspect provides a method for identifying AAV capsid variants having desired characteristics. In one embodiment, the method includes: a) providing a library of AAV capsid variants as defined above, or obtained in a method for preparing a library of AAV capsid variants as defined above; b) contacting the library with a culture of cells, organoids, or tissues, or administering the library to a non-human animal; c) transducing the AAV capsid variants in the library into cells, organoids, tissues, or cells in an animal; and d) identifying at least one AAV capsid variant transduced into at least one desired cell type in cells, organoids, tissues, or an animal in b) as an AAV capsid variant having desired characteristics, and optionally recovering the AAV capsid variant having desired characteristics from a desired cell.
[0132] A method for identifying AAV capsid variants having the desired characteristics described herein further comprises step (b) contacting the library with a culture of cells, organoids, or tissues, or administering the library to a non-human animal. In one embodiment, the culture of cells, organoids, or tissues, or the non-human animal, comprises at least one desired cell type, e.g., a cell type from which an AAV capsid variant with improved directivity is identified. In one embodiment, the culture of cells, organoids, or tissues, or the non-human animal, may further comprise at least one undesired cell type, e.g., a cell type in which the transduction efficiency of the identified AAV capsid variant is reduced. In one embodiment, the library is systemically administered to a non-human animal, e.g., by an intravenous route.
[0133] In the next step (c) of the method for identifying AAV capsid variants having the desired features described herein, the AAV capsid variant in the library is transduced into cells, organoids, tissues, or animal cells. Thus, in one embodiment of step (c), the AAV capsid variant in the library is transduced into (desired) cells, organoids, tissues, or animal cells, and sufficient time is allowed to elapse for the reporter gene to be expressed.
[0134] In the next step (d) of the method for identifying an AAV capsid variant having desired features as described herein, at least one AAV capsid variant transduced into at least one cell of a desired cell type in a cell, organoid, tissue, or animal is identified. In one embodiment, the at least one AAV capsid variant is identified by sequencing its UMI. In one embodiment, the at least one AAV capsid variant is identified as an AAV capsid variant having desired features. In one embodiment, step (d) further includes recovering the AAV capsid variant having desired features from a cell of the desired cell type.
[0135] In one embodiment, step (d) includes detecting transduction of a cell of a desired cell type by detecting the expression of a reporter gene in at least one cell of the desired cell type.
[0136] In one embodiment of the method for identifying AAV capsid variants having the desired characteristics described herein, the desired cell type, for example, the cell type to which the improved-targeting AAV capsid variant should be identified, is a cell from an organ or tissue selected from the liver, skeletal muscle, cardiac muscle, diaphragmatic muscle, kidney, brain, stomach, intestine, skin, endothelial cells, and lung. In one embodiment of the method for identifying AAV capsid variants having the desired characteristics described herein, the AAV capsid variant to be identified has a low transduction efficiency of an undesirable cell type, which may be a cell from an organ or tissue selected from the liver, skeletal muscle, cardiac muscle, diaphragmatic muscle, kidney, brain, stomach, intestine, skin, endothelial cells, and lung.
[0137] In one embodiment, a method for identifying an AAV capsid variant having the desirable characteristics described herein is a method in step d) in which, for at least two members in a library, at least one of the mRNA expressed from at least two members and at least one of the genome copy numbers of at least two members in at least one cell of the desired cell type is quantified and identified by sequencing of their barcodes / UMIs, and the member having at least one of the highest mRNA expression level and genome copy number is identified as an AAV capsid variant having the desirable characteristics. In one embodiment of the method, in step d) in cells of two or more cell types including at least the desired cell type, at least one of the mRNA expressed from at least two members in a library and at least one of the genome copy numbers of at least two members is quantified and identified by sequencing of their barcodes / UMIs, and the member showing the most desired distribution across two or more cell types is identified as an AAV capsid variant having the desired characteristics. In one embodiment, the two or more cell types include at least one undesired cell type.
[0138] In one embodiment, a method for identifying an AAV capsid variant having the desired features described herein is a method for identifying a variant having the desired features by performing sequencing of at least one of the following: i) a barcode / UMI, and ii) a region of a DNA construct that codes for a difference of at least one amino acid.
[0139] The present invention has been described above with reference to several exemplary embodiments shown in the drawings. Modifications and alternative implementations of some parts or elements are possible and fall within the scope of protection as defined in the appended claims. [Brief explanation of the drawing]
[0140] [Figure 1] A method for producing DNA scaffolds and AAVs by transfection / baculovirus infection. Replication of the transfected DNA scaffold is facilitated by infection with Rep baculovirus via AAV2 ITR and hr4b enhancer. Furthermore, co-infection with baculovirus transcriptionally activates the P10 promoter, which drives the expression of the capsid VP protein. The capsid is then assembled in the cell, and subsequently, the replicated transgene is packaged into the capsid by the Rep protein to obtain an AAV library. [Figure 2] The transfection / baculovirus infection method results in a low total / full ratio of AAV batches because only cells that accept both components produce AAV capsids, allowing the transgene to be packaged within them. [Figure 3] Transfection / baculovirus infection is used to introduce mutations, amplify libraries, and create templates suitable for AAV library generation, followed by the Grit process (Gibson assembly - rolling circle amplification - TelN digestion). [Figure 4]A) Production of AAV2 / 5 scaffolds by rolling circle / TelN digestion and transfection / infection. AAV2 / 5 scaffolds are amplified from plasmid pIMRAku04 by rolling circle amplification. Next, linear, covalently closed scaffolds are produced by TelN protelomerase digestion. After transfection and baculovirus infection, AAV is recovered and analyzed by Q-PCR and SDS-Page. B) AAV2 / 5 scaffold pIMRaku04. Elements present on the scaffold are: 1. AAV2 ITR, 2. hr4b enhancer, 3. Nanoluc reporter cassette, 4. AAV2 / 5 capsid expression cassette under the control of the P10 promoter, and 5. TelN protelomerase digestion site. [Figure 5] SDS Page gel run with purified AAV batches produced with pIMRaku-04 scaffold or pIMRaku-04 plasmid. The VP123 ratio of AAV2 / 5 material produced using pIMRaku-04 scaffold shows a natural stoichiometry of approximately 1:1:10 for VP123 with respect to AAV. [Figure 6] pIMRAku-04 AAV2 / 5 and pIMRAku-011 are library DNA scaffolds with random 21bp libraries inserted into VR-VIII. The insertion should result in the generation of an AAV2 / 5 or AAV9 capsid library with a heptamer peptide insertion in the VR-VIII surface-exposed loop. The elements present on the scaffold are as described above in Figure 4b. [Figure 7] Viral titers (gc / ml) of AAV2 / 5 (CLB = crude lysate) and AAV9 (purified) libraries, measured by QPCR specific to the hr4b enhancer element. [Figure 8] SDS-Page gels run with purified AAV2 / 5 and AAV9 libraries. The capsid stoichiometry of both libraries shows a natural VP123 ratio of 1:1:10 for AAV. [Figure 9A]Sanger sequencing at the VR-VIII library insertion site of AAV2 / 5. Sequencing of intermediate samples taken during AAV library preparation reveals mixed readings at the insertion site throughout the entire preparation process. [Figure 9B] Sanger sequencing at the insertion site of the VR-VIII library of AAV9. Sequencing of intermediate samples taken during AAV library preparation reveals mixed readings at the insertion site throughout the entire preparation process. [Figure 10] Maintaining complexity during the AAV2 / 5 library preparation process using NGS. Of the initial 1.22 × 10⁷ variants detected in the starter G-Block, 8.1 × 10⁶ variants are amplified by the GRiT process. 3.2 × 10⁶ variants are detected in the purified AAV2 / 5 library preparation. Approximately 25% of the original complexity in the G-Block is maintained in the purified AAV2 / 5 library. [Figure 11] AAV cross-packaging occurs when there is a mismatch during genome and capsid generation. For example, an AAV5 capsid with an AAV9 genome or an AAV9 capsid with an AAV5 genome. [Figure 12] A: A simplified method for determining the cross-packaging rate in two capsid libraries generated in expressSF+ cells or HEK293t cells. AAV2 / 5 and AAV9f are co-produced in SF+ cells and Hek293t cells. After serial purification with serotype-specific affinity resin, genomic identity is established by serotype-specific QPCR. The cross-packaging percentage is calculated by determining the proportion of mismatched genomes in the resulting pool. B: Cross-packaging rate (qPCR) after purification with serotype-specific resin. [Examples]
[0141] [Introduction] AAV capsid libraries have been widely and successfully applied in directed evolutionary studies to select capsids with improved properties. The construction of these AAV libraries utilizes a mammalian production platform in Hek293 cells. Libraries produced on this platform possess sufficient depth and titer to conduct these types of studies. However, their construction requires the transfection of large amounts of plasmid DNA, which makes cross-packaging of the transgenes more likely. Cross-packaging occurs when an AAV genome is packaged into a mismatched capsid. AAV libraries with a high rate of cross-packaging generate a significant amount of noise in downstream analyses of directed evolutionary studies. This is because it is not possible to guarantee that the detected AAV genome originates from a cross-packaged AAV genome or from a correctly packaged AAV genome. Therefore, multiple cycles using in vitro or in vivo models are necessary to increase confidence that the detected genome (with potentially beneficial capsid properties) originates from a correctly packaged AAV and not from a cross-packaged AAV. These cycles can significantly increase the length of such studies, raise their costs, and reduce the feasibility of directed evolution approaches chosen for new capsid development. Therefore, minimizing the degree of cross-packing of AAV libraries is essential.
[0142] The method developed by the inventors relates to the generation of AAV capsid libraries in insect cells. In-house experiments have shown that insect cells are highly sensitive to transfection with high levels of exogenous DNA, resulting in a self-restriction mechanism when delivering large amounts of DNA to the cells. The maximum amount of DNA that can be successfully delivered to insect cells is in the range of 1-2 pg of DNA / cell, far below the lower limit of DNA typically transfected to create AAV libraries in Hek293 (>5 pg DNA / cell). The AAV capsid library construction method developed by the inventors utilizes this observation and combines it with the ability of insect cells to undergo transcriptional activation and transgene replication through co-infection with replicase (Rep) baculovirus.
[0143] The inventors designed a DNA scaffold that promotes the production of high-titer AAV in insect cells using a small amount of DNA (<1 pg / cell) delivered to cells by transfection. Following transfection of the DNA scaffold, the cells are co-infected with Rep-containing baculovirus. In the cells, the Rep-baculovirus promotes the replication of the DNA scaffold (via the AAV2 ITR and hr4b enhancer) and promotes the transcriptional activation of the promoter (P10) that drives the expression of the AAV capsid protein (Figure 1). A further advantage of this method of production is the low percentage of empty capsids in the produced AAV. This is because only cells that have been transfected with the DNA scaffold and infected with the Rep-baculovirus produce AAV (Figure 2).
[0144] Regardless of the production platform (mammalian or insect cells), AAV libraries require a method for introducing mutations into the AAV capsid gene. To develop deeply complex libraries, these introduced mutations need to be maintained throughout different AAV production process steps (e.g., mutation introduction into the scaffold, amplification of the library DNA scaffold, AAV production, and purification). The libraries prepared in these examples introduce complexity into the capsid gene by inserting a randomized 21 bp DNA sequence into the surface-exposed loop VR-VIII of the capsid gene. The inventors have developed a three-step process called GRiT (Gibson Assembly-Rolling Circle Amplification-TelN Proteoromerase Digestion) that can not only amplify the DNA library scaffold but also introduce mutations. The GRiT process can maintain the complexity of the introduced insertion from insertion to amplification of the DNA library scaffold.
[0145] In short, a library of 21 bp-long randomized DNA sequences (encoding a heptamer peptide) is inserted into a linearized DNA scaffold by Gibson assembly. Gibson assembly recirculates the DNA library scaffold, enabling amplification by rolling circle amplification using the phi29 enzyme. A covalently closed single library DNA scaffold is generated from the amplified template by TelN protelomerase digestion to produce a linear cassette. The main advantage of a covalently closed scaffold is that it increases the transfection efficiency of the scaffold compared to covalently open linear DNA. Finally, the library DNA scaffold can be transfected into insect cells to create an AAV capsid library (Figure 3).
[0146] In the first embodiment, the inventors investigate the feasibility of using a rolling circle amplification / Teln digestion AAV2 / 5 DNA scaffold without a library for the production of AAV2 / 5 batches in insect cells. In the second embodiment, the inventors evaluate the GRiT process for constructing an AAV2 / 5 library with randomized heptameric peptide insertions at VR-VIII of AAV2 / 5 in insect cells. AAV2 / 5 is an AAV5 chimera in which VP1 is derived from AAV2 and VP2 / VP3 is derived from AAV5. Furthermore, to evaluate the applicability of the GRiT process / transfection infection method to AAV9 serotypes, the inventors introduced randomized 21bpDNA sequence insertions at VR-VIII of the AAV9 gene and constructed a library in insect cells.
[0147] [1. Example 1: Production of AAV by an AAV2 / 5 DNA scaffold amplified by rolling circle amplification and digested with TelN protelomerase] [1.1 Method] [1.1.1 Generation of AAV5 DNA scaffold by rolling circle amplification and TelN digestion] Plasmid pIMRaku-04 (AAV2 / 5, no library; SEQ ID NO: 10) was amplified by rolling circle amplification (RCA) using the EquiPhi29 enzyme (ThermoFisher). 10 ng of plasmid was amplified at 43°C for 3 hours in the presence of 100 μM exo-resistant random hexamer (ThermoFisher) and 1 mM dNTP (NEB). After amplification, the DNA was precipitated at 21000 × g for 15 minutes in the presence of 0.1 × final volume 3 M potassium acetate (pH=5.2) and 0.7 × final volume 100% isopropanol, followed by washing with 70% ethanol at 21000 × g for 5 minutes. Finally, the DNA pellet was dissolved in 5 mM Tris-HCl (pH=8.5). Next, the long linear chains of the RCA-amplified DNA were digested into covalently closed DNA scaffolds with adjacent ITRs using TelN protelomerase (NEB). Digestion was performed at 30°C using 1 μg of amplified DNA per 1 μl of enzyme. After digestion, short linear DNA was purified by column chromatography using a PCR purification kit (Machery-Nagel). The DNA concentration of the purified DNA scaffold was determined using a Nanodrop100 system (Thermofisher).
[0148] [1.1.2 Production of AAV2 / 5 (pIMRaku-04) in expresSF+ insect cells using transfection / baculovirus infection] Multiple concentrations of RCA amplification / TelN digestion pIMRaku-04 scaffolds ranging from 0.01 μg to 5 μg were transfected using cellfectin II transfection reagent (ThermoFisher) in a 1.5 × 10⁻⁶ ratio. 7100 expresSF+ cells were transfected. A circular plasmid control was also used in the experiment. All conditions were normalized to 15 μg of DNA using UltraPure Herring Sperm carrier DNA (ThermoFisher) to transfect cells with an equal amount of DNA. 72 hours after transfection, fresh baculovirus expressing AAV2 replicase (SEQ ID NO: 14) was inoculated into the cells at 1% of the final volume (40 ml). 72 hours after baculovirus infection, the AAV batch was collected by adding 10× lysis buffer (1.5 M NaCl, 0.5 M Tris-HCl, 1 mM MgCl2, 10% Triton X-100, pH=8.5) and lysing at 28°C for 1 hour. Genomic DNA was digested by benzonase treatment at 37°C for 1 hour. Cellular debris was removed by centrifugation at 1900×g for 15 minutes, and the supernatant containing AAV2 / 5 was then stored at 4°C. AAV2 / 5 particles were purified from crude lysates using AVBSepharose affinity resin (GE Healthcare) via a batch binding protocol. The crude AAV cell lysates were added to the resin washed with 0.2 M HPO4 (pH=7.5) buffer. The samples were then incubated at room temperature for 2 hours with gentle mixing. After binding, the resin was washed with 0.2 M HPO4 (pH=7.5) buffer, and the binding vector was eluted by adding 0.2 M glycine (pH=2.5). The pH of the eluted vector was immediately neutralized by adding 0.5 M Tris-HCl (pH=8.5). The purified AAV batches were stored at -20°C.
[0149] [1.1.3 Analysis of AAV batches by QPCR and SDS-PAGe gel electrophoresis] The titer of the lysed AAV2 / 5 batch was determined using Q-PCR specific to the hr4b enhancer region of the scaffold. AAV was treated with DNAse at 37°C to degrade the exogenous DNA. Then, AAV DNA was released from the particles by short-term (30-minute) heat treatment (37°C) in the presence of 1M NaOH. Next, the alkaline environment was neutralized by adding an equal volume of 1M HCl. The neutralized DNA was diluted 10-fold with WFI 16 ng / μl PolyA, and the sample was then used for QPCR with primers specific to the hr4b element on the transgene: forward CGAGGGATGATGTCATTTGTAGA (SEQ ID NO: 17), reverse ATCCACCGATCTTGCGTTAC (SEQ ID NO: 18), and probe Fam-ACCGAACTCGCTTTACGAGTAGAATTCT-mgb (SEQ ID NO: 19). The VP protein composition of the purified AAV2 / 5 vector was determined by SDS-Page gel electrophoresis. In short, 15 μl of purified AAV2 / 5 was mixed with 5 μl of 4×Leammli loading buffer (Biorad) supplemented with β-mercaptoethanol (Biorad). After denaturing the protein by short-term heat treatment (95°C for 5 minutes), the material was loaded onto a stain-free polyacrylamide gel (Biorad). Next, the samples were separated by electrophoresis at 200 volts for 35 minutes. Following electrophoretic staining (tryptophan-based), the gel was developed under UV light for 5 minutes, and then the VP protein was visualized under UV light using a Chemidoc imaging system (Biorad).
[0150] [1.2 Results] In this embodiment, the inventors investigate whether a DNA scaffold designed to produce an AAV library in insect cells can produce high-titer AAV with a correct 1:1:10 capsid virus protein stoichiometry. A VP123 stoichiometry of approximately 1:1:10 reflects the infectious capsid. To produce AAV, the construct pIMRAku04 (AAV2 / 5; SEQ ID NO: 10) was amplified by rolling circle amplification using EquiPhi29 and linearized with TelN protelomerase. The linearized pIMRaku04 scaffold was transfected into expressSF+ insect cells at concentrations ranging from 0.01 μg to 5 μg / reaction. Three days after transfection, the cells were infected with 1 v / v% AAV2 Rep-expressing baculovirus (SEQ ID NO: 14). After recovery, the AAV2 / 5 particles were purified using a batch coupling protocol with AVB Sepharose (Figure 4a).
[0151] In the experiment, the inventors used the DNA scaffold pIMRAku04, which contains elements necessary for the production of AAV2 / 5 capsids. These elements are as follows: 1. A "packaging" AAV transgene cassette with two adjacent AAV2 ITRs. The ITRs facilitate the replication of the scaffold in insect cells by Rep, as well as its packaging into AAV2 / 5 particles. 2. A Nanoluc reporter gene under the control of a CNS-specific promoter, with a barcode adjacent. The promoter driving reporter expression can be changed depending on the type of tissue the library is developed in (e.g., liver, CNS, or heart-specific). 3. An AAV2 / 5 capsid expression cassette (SEQ ID NO: 11) under the control of a P10 promoter. The activity of the P10 promoter is enhanced by baculovirus infection. 4. An Hr4b enhancer element that not only enhances scaffold replication after entering the cell but also participates in the transcriptional activation of the P10 promoter, promoting the expression of the capsid protein. 5. TelN protelomerase cleavage sites used by the ITR to linearize and covalently close the DNA scaffold outside the adjacent cassette (Figure 4b).
[0152] After production, AAV titers were determined in crude lysated bulk (CLB) using Q-PCR specific to the hr4b element (Table 1). The inventors found measurable AAV titers under all conditions using their pIMRAku04 scaffold. Furthermore, the inventors found that the AAV titer increased with the amount of transfected DNA scaffold. Compared to plasmid transfection, the efficiency of the linearized / covalently closed pIMRAku04 scaffold was approximately 1-logarithm lower (3.29e+11 for 5 μg plasmid, and 3.70e+10 for 5 μg linearized scaffold). SDS-PAGE gel electrophoresis was performed on batch-bound purified material to investigate whether the AAV2 / 5 capsid produced with the pIMRAku04 scaffold contained the correct VP stoichiometry (Figure 6). The VP1 / 2 / 3 ratio of the AAV2 / 5 material produced by this method has a natural ratio of approximately 1:1:10, which reflects the infectious AAV capsid.
[0153] In summary, the DNA scaffold designed by the inventors can produce relatively high-titer AAV batches with a natural VP123 protein ratio of approximately 1:1:10, reflecting the infectious capsid. While the AAV titer produced using the pIMRaku04 scaffold is approximately one logarithm lower than that of the pIMRaku04 plasmid, the amount of virus produced per 1 ml of CLB is sufficient for further library development using this scaffold design. This is because AAV libraries used in directed evolutionary studies in vivo typically have a titer of 5 × 10⁶. 12 This is because it is injected at relatively low concentrations of less than gc / kg. These levels are within the range achievable with the productivity of this AAV production method (when transfecting with a scaffold DNA range of 0.5–2 μg).
[0154] [Table 1]
[0155] [2. Example 2: Preparation of AAV2 / 5 and AAV9 libraries with randomized heptameric peptides inserted using variable loop VR-VIII] [2.1 Method] [2.1.1 Generation of AAV2 / 5 and AAV9 VR-VIII Random 21bp Library Scaffolds by Gibson Assembly, Rolling Circle Amplification, and TelN Digestion (GRiT)] Plasmid pIMRAku-11 (AAV9; containing SEQ ID NO: 12 and AAV9 expression cassette SEQ ID NO: 13) was digested with XagI (NEB) and BamHI (NEB) restriction enzymes. 50 ng of linearized plasmid was combined with 6 ng of G-block fragment EcoNI-Cap9-VR8-7-mer-BamHI (IDT) in a 1:2 scaffold-to-G-block ratio. The amount of DNA used in this reaction was based on the size of the fragments. Next, Gibson assembly was performed at 50°C for 2 hours using the NEBuilder HIFI DNA Assembly Kit (NEB). Plasmid pIMRAku-04 (AAV2 / 5) was digested with NcoI restriction enzyme (NEB). 50 ng of linearized pIMRAku-04 was combined with 6 ng of G-block-7-mer AAV5 VIII (IDT) in a 1:2 scaffold-to-G-block ratio. Gibson assembly was performed at 50°C for 2 hours using the NEBuilder HIFI DNA Assembly Kit (NEB). The resulting circular plasmid, containing an AAV9 or AAV2 / 5 capsid library with a randomized 21 bp insertion in the variable loop VR-VIII of the capsid gene, was used as a template for rolling circle amplification.
[0156] A 10 ng recirculated library plasmid for AAV2 / 5 or AAV9 was amplified at 43°C for 3 hours in the presence of 100 μM exo-resistant random hexamer (ThermoFisher) and 1 mM dNTP (NEB). After amplification, the DNA was precipitated at 21000 × g for 15 minutes in the presence of 0.1 × final volume 3 M potassium acetate (pH=5.2) and 0.7 × final volume 100% isopropanol, followed by washing with 70% ethanol at 21000 × g for 5 minutes. Finally, the DNA pellet was dissolved in 5 mM Tris-HCl (pH=8.5). Next, the long linear chains of the RCA-amplified DNA were digested into covalently closed DNA scaffolds with adjacent ITRs using TelN protelomerase (NEB). Digestion was performed at 30°C using 1 μg of amplified DNA per 1 μl of enzyme. After digestion, short linear DNA was purified using a PCR purification kit (Machery-Nagel) via column chromatography. The DNA concentration of the prepared DNA scaffold was determined using a Nanodrop100 system (Thermofisher).
[0157] [2.1.2 Preparation of AAV2 / 5 and AAV9 Randomized Heptameric Peptide Libraries using Transfection / Baculovirus Infection in expresSF+ Insect Cells] A 1 μg AAV2 / 5 (pIMRAku-04) or AAV9 (pIMRAku-11) DNA scaffold library was transfected into 1.5e7 expressSF+ cells using cellfectin II transfection reagent (ThermoFisher). 15 μg of total DNA was transfected into the cells; that is, 14 μg of UltraPure Herring Sperm carrier DNA (ThermoFisher) was added to the transfection reaction. 72 hours after transfection, fresh baculovirus expressing AAV2 replicase (Bac.VD183) was inoculated into the cells at 1% of the final volume (40 ml). 72 hours after baculovirus infection, AAV batches were collected by adding 10× lysis buffer (1.5M NaCl, 0.5M Tris-HCl, 1mM MgCl2, 10% Triton X-100, pH=8.5) and lysis at 28°C for 1 hour. Genomic DNA was digested by benzoase at 37°C for 1 hour. Cellular debris was removed by centrifugation at 1900×g for 15 minutes, and the supernatant containing the AAV2 / 5 or AAV9 library was then stored at 4°C. AAV2 / 5 or AAV9 library particles were purified from crude lysates using AVB Sepharose (AAV2 / 5 library) affinity resin (GE Healthcare) or AAVx affinity resin (AAV9, ThermoFisher) using a batch coupling protocol. Briefly, AAV crude cell lysates were added to resin washed (with 0.2M HPO4 pH=7.5 buffer). Subsequently, the samples were incubated at room temperature for 2 hours with gentle mixing. After binding, the resin was washed in 0.2 M HPO4 (pH=7.5) buffer, and the binding vector was eluted by adding 0.2 M glycine (pH=2.5). The pH of the eluted vector was immediately neutralized by adding 0.5 M Tris-HCl (pH=8.5). The purified AAV batch was stored at -20°C.
[0158] [2.1.3 Analysis of AAV2 / 5 and AAV9 library batches by QPCR and SDS-PAGe gel electrophoresis] The titers of lysed or purified AAV2 / 5 or AAV9 batches were determined using hr4b enhancer-specific Q-PCR. AAV was treated with DNAse at 37°C to degrade exogenous DNA. AAV DNA was then released from the particles by short-term (30-minute) heat treatment (37°C) in the presence of 1M NaOH. Next, the alkaline environment was neutralized by adding an equal volume of 1M HCl. The neutralized DNA was diluted 10-fold with WFI 16 ng / μl PolyA, and the sample was then used in QPCR with primers specific to the hr4b element on the transgene: forward CGAGGGATGATGTCATTTGTAGA (SEQ ID NO: 17), reverse ATCCACCGATCTTGCGTTAC (SEQ ID NO: 18), and probe Fam-ACCGAACTCGCTTTACGAGTAGAATTCT-mgb (SEQ ID NO: 19). The VP protein composition of purified AAV2 / 5 or AAV9 libraries was determined by SDS-Page gel electrophoresis. Briefly, 15 μl of purified AAV2 / 5 was mixed with 5 μl of 4×Leammli loading buffer (Biorad) supplemented with β-mercaptoethanol (Biorad). After denaturing the proteins by short-term heat treatment (95°C for 5 minutes), the material was loaded onto a stain-free polyacrylamide gel (Biorad). The samples were then separated by electrophoresis at 200 volts for 35 minutes. Following electrophoretic staining (tryptophan-based), the gel was developed under UV light for 5 minutes, and the VP proteins were then visualized under UV light using a Chemidoc imaging system (Biorad).
[0159] [2.1.4 Sanger and whole-genome sequencing of AAV2 / 5 and AAV9 libraries] Using the Purelink viral DNA / RNA mini-kit (ThermoFisher), transviral DNA was isolated from AAV2 / 5 and AAV9 libraries according to the kit protocol. Fragments containing library insertions were amplified by PCR using primers specific to the VR-VIII insertion site (AAV2 / 5 VR-VIII forward ACCAGCGAGAGCGAGAC and reverse CTGCCGGGCACGATTTC, AAV9 VR-VIII forward AAAGCCTGGACCGACTAATG and reverse TCTGTCCTGCCAAACCATAC; SEQ ID NOs. 20-23, respectively). After column purification of the PCR products (Machery-Nagel), the presence of randomized insertions was verified by Sanger sequencing (using the same primers as PCR amplification, Macrogen) and next-generation sequencing (Genomscan). Whole-genome sequencing was performed with the following settings: paired-end 150bp sequencing runs were performed on NovaSeq 6000 (Illumina), obtaining 3Gb of data per library. The DNA was aligned using the barcode detection algorithm https: / / github.com / JonasWeinmann / AAV-barcode-detection-and-normalization.
[0160] [2.1 Results] This example investigates whether AAV2 / 5(pIMRaku-04) and AAV9(pIMRaku-11) scaffolds can be used for the production of AAV libraries in which randomized heptameric peptides are inserted into the variable loop VR-VIII. Furthermore, the inventors investigate the maintenance of complexity introduced into the AAV5 and AAV9 capsid genes using Gibson assembly throughout various process steps of AAV library production.
[0161] Randomized 21 bp DNA libraries were inserted into the variable loop VR-VIII of the AAV2 / 5 and AAV9 capsid genes using the GRiT process described in Figure 3 above. AAV2 / 5 and AAV9 libraries were prepared using the transfection infection method described in Example 1. The DNA scaffolds used in this example are summarized in Figure 7 and contain the same elements as those described in Figure 4b above, differing only in that the capsid expression cassette is either AAV2 / 5 (pIMRAku04) or AAV9 (pIMRaku11). To investigate the productivity of the libraries, both AAV2 / 5 and AAV9 DNA library scaffolds were transfected into insect cells at multiple concentrations. Crude cell lysates were collected following transfection with the library DNA scaffolds and infection with AAV2 Rep baculovirus (SEQ ID NO: 14). The AAV library was purified from a solution containing an affinity resin for AVB Sepharose (AAV2 / 5 library) or AAVx POROS (AAV9 library).
[0162] The viral titers (gc / ml) of crude lysates (CLB, AAV2 / 5 library) or purified libraries (BBNE, AAV9 library) were determined using hr4b enhancer-specific QPCR (Figure 8). For both AAV2 / 5 and AAV9 libraries, the inventors confirmed a decrease in productivity compared to plasmid controls without library inserts. The decrease in productivity for the AAV2 / 5 library was approximately 2 logarithms compared to the plasmid control and 1 logarithm compared to the production of clean scaffolds in Example 1 (Table 1). For the AAV9 library, the decrease in productivity appeared to be approximately 1.5 logarithms compared to the plasmid control. A decrease in titer was expected for the AAV libraries because not all inserts introduced into the AAV2 / 5 or AAV9 capsid DNA library can produce capsids (due to incorrect capsid folding). Furthermore, the inventors observed a decrease in AAV library production when less library DNA was transfected into insect cells. The overall introduction of complexity into the scaffold reduces productivity in both serotypes. However, the library titer remains high enough to be used for directed evolutionary studies in vivo. The production conditions for libraries used in directed evolutionary studies must be balanced based on the AAV library titer required for in vivo experiments and the need to reduce cross-packaging.
[0163] The purified AAV2 / 5 and AAV9 libraries were also separated by SDS-PAGE to determine the capsid stoichiometry and VP size (Figure 8). The capsid stoichiometry in both the AAV2 / 5 and AAV9 libraries is at a native VP123 ratio of approximately 1:1:10 for AAV. Interestingly, the VP proteins in both AAV2 / 5 and AAV9 appear to be slightly heavier than those of the plasmid control (as seen by the slight upward shift in the fragment size of the library on the gel, indicated by the arrow in Figure 8). This can be explained by the insertion of 7aa into the VP protein in the library. This should result in a heavier capsid and therefore a shift in the VP protein in the library on the gel.
[0164] Next, the inventors investigated whether the insertion of randomized 21 bp fragments into their DNA scaffold was successfully introduced and transferred during the AAV library production process. Using primer sets specific to the VR-VIII region of the AAV2 / 5 and AAV9 capsid genes, the region surrounding the library insertion site was amplified by PCR. This PCR was performed on intermediate samples collected while constructing the AAV2 / 5 and AAV9 libraries. The intermediates included: 1. the original library G-Block (AA2 / 5), 2. the Gibson assembly (AA2 / 5), 3. the RCA library amplification reaction product (AAV2 / 5), 4. the unpurified lysate of AAV library production (CLB, AAV2 / 5 and AAV9), and 5. the purified library (BBNE, AAV2 / 5 and AAV9). The amplified PCR products were analyzed for the presence of library insertion by Sanger sequencing. This analysis provides a qualitative assessment of the presence of libraries in the intermediate, as randomized DNA insertions should produce mixed reads in the sequence chromatogram at the insertion site. Sequence analysis confirmed the presence of mixed reads at the insertion site throughout the entire generation process in both AAV2 / 5 and AAV9 libraries (Figure 9). This result demonstrates that our library amplification (GRiT) and AAV generation (transfection / infection) method can introduce and maintain insertions throughout the entire AAV library generation process.
[0165] In the final assessment of the AAV2 / 5 library, the inventors monitored the complexity of the 21 bp insertions and quantitatively tracked their maintenance throughout the AAV library generation process. For the PCR products previously generated for Sanger sequencing, the inventors performed next-generation sequencing analysis. The inventors obtained 3 GB of 150 bp paired-end reads from the following intermediate samples on the NovaSeq 6000 system: 1) the original library G-Block; 2) the RCA library amplification reaction, and 3) the purified library. The resulting sequence reads were then aligned against the AAV2 / 5 capsid cassette, and the number of insertions found in the data pool was quantified using the barcode detection algorithm https: / / github.com / JonasWeinmann / AAV-barcode-detection-and-normalization. The inventors found that out of 1.22×10 7 21 bp insertion variants detected in the starter G-Block, 8.1×10 6 were carried over into the amplified DNA library, and of these, 3.2×10 6 variants were ultimately transferred to the AAV2 / 5 capsid library (Figure 10). Over the entire process, 25% of the original complexity present in the G-Block was transferred to the AAV capsid library, and most of the diversity was lost during AAV production. This loss is thought to be due to a high proportion of AAV capsids misfolded due to non-viable insertions. However, during the library DNA amplification process, the complexity is well maintained, probably due to the rolling circle amplification method used for scaffold generation.
[0166] In summary, the inventors have developed a DNA scaffold suitable for AAV production via a transfection / baculovirus infection-based process in insect cells. The GRiT process allows for the insertion of random 21bp DNA insertions into this scaffold and their amplification without significant loss of complexity. AAV production using these DNA libraries results in a sufficiently concentrated AAV library and the complexity necessary for its use in directed evolutionary studies.
[0167] [3. Example 3: Comparative analysis of genome cross-packaging in AAV libraries produced in Hek293t and expressresSF+ insect cells] [3.1 Method] [3.1.1 Generation of AAV9f and AAV2 / 5 DNA scaffolds and AAV production in HEK293t and insect cells] Linear DNA suitable for the production of AAV9F and AAV2 / 5 capsids in Hek293t or expressSF+ was prepared from plasmids containing AAV9F or AAV2 / 5 capsid expression cassettes (HEK293t AAV2 / 5 (SEQ ID NO: 63), HEK293t AAV9F (SEQ ID NO: 64), SF+ AAV2 / 5 (SEQ ID NO: 65), and SF+ AAV9f (SEQ ID NO: 66). DNA scaffolds were prepared by RCA followed by TELn digestion using the method described in Example 1. In expressSF+ insect cells, AAV was produced by transfecting 1 μg of plasmid into 1.5e7 cells using cellfectin II transfection reagent (ThermoFisher). A 1:1 mixture of scaffolds pVD1943 (AAV2 / 5) and pVD1950 (AAV9F) (1 μg) was transfected into SF+ cells. Infection was performed according to the method described in Example 1. In HEK293t cells, AAV was produced by transfecting 1e7 adherent HEK293t cells with 1 μg of plasmid using PEI (Polysciences). HEK293t cells were transfected with 1 μg of a 1:1 mixture of scaffold pVD2004 (AAV2 / 5) and pVD2006 (AAV9F). 15 μg of total DNA was transfected into the cells. That is, 14 μg of UltraPure Herring Sperm carrier DNA (ThermoFisher) was added to the transfection reaction. Transfection of both cell lines was performed three times. The expressSF+ AAV was harvested 144 hours after transfection and the Hek293t AAV was harvested 72 hours after transfection using the protocol described in Example 1.
[0168] [3.1.2 Continuous purification of AAV2 / 5 and AAV9f particles by coupling affinity batches with serotype-specific resins] A mixture of AAV2 / 5 and AAV9f prepared using expresSF+ or HEK293t was subjected to serial batch coupled purification. This serial purification used affinity resins specific to AAV5 (AVB Sepharose, Cytiva) or AAV9 particles (Poros9, Thermofisher). Crude lysates from the two capsid library products were first purified with AVB Sepharose resin (bound to AAV2 / 5 capsids), and then the flow-through from this purification (containing unbound AAV9 capsids) was used for purification with Poros9 resin (bound to AV9 capsids). Briefly, the purification protocol involved adding washed AAV crude cell lysates (using 0.2M HPO4 (pH=7.5) buffer) to resin (AVB sepharose or Poros9). Subsequently, the samples were incubated at room temperature for 2 hours with gentle mixing. After binding, the resin was washed in 0.2 M HPO4 (pH=7.5) buffer, and the binding vector was eluted by adding 0.2 M glycine (pH=2.5). The pH of the eluted vector was immediately neutralized by adding 0.5 M Tris-HCl (pH=8.5). The purified AAV batch was stored at -20°C.
[0169] [3.1.3 Genome identification by QPCR using serotype-specific primers] The identity of genomes packaged in serially purified particles produced using ExpresSF+ or HEK293t cells was analyzed using AAV serotype-specific QPCR. DNA was isolated from the AAV particles using the method described in Example 1. The primers used for this QPCR are summarized in Table 3 (SEQ ID NOs. 55-60).
[0170] [Table 3]
[0171] [3.1.4 Genome Identification by Next-Generation Sequencing] The identity of the genomes packaged in sequentially purified particles generated using expressresSF+ was also analyzed using next-generation sequencing. Viral genomic DNA was isolated from sequentially purified particles using the Purelink Viral RNA / DNA minikit. Next, the genomic region encoding the capsid was amplified by PCR using the following primers: forward CCTGCTGTTCCGAGTAACCATC and reverse GTCGGCGTGGTTGTACTTGAGG (SEQ ID NOs. 61 and 62). The same primer set was used to amplify the AAV2 / 5 and AAV9F genomes. The PCR samples were sent for Illumina sequencing with a depth of 10 million reads. Genomic identity was analyzed using a barcode count script obtained from github:https: / / github.com / JonasWeinmann / AAV-barcode-detection-and-normalization.
[0172] [3.1 Results] Cross-packaging occurs during AAV library generation when a genome encoding capsid X is packaged into a mismatched capsid Y (Figure 11). In this example, the cross-packaging rate of our novel library generation method in insect cells was investigated and compared with a standard library generation method using HEK293t. We prepared libraries containing a mixture of two serotypes (AAV2 / 5 and AAV9F) using either insect cells or HEK293t. After sequential purification with affinity resins that selectively purify AAV2 / 5 or AAV9F capsids, we identified the origin of the genomes packaged within these AAV capsids using serotype-specific QPCR. The cross-packaging % of the library generation method can be established by determining the proportion of mismatched genomes in the sampled pool after serotype-selective purification (e.g., AAV9F genome in AAV2 / 5 capsid) (Figure 12A). This simplified experiment to evaluate cross-packaging in our library is based entirely on a method previously described by Nonnemmacher in 2015 to evaluate cross-packaging in HEK293t cells.
[0173] Figure 12B shows the percentage of cross-packaging measured by QPCR in the AAV particle pool obtained from two capsid libraries after sequential serotype-specific purification. A significant decrease in the percentage of cross-packaging was observed in the two capsid libraries generated in expressSF+ cells compared to the two capsid libraries generated in HEK293t cells. This was the case for both the AAV9F genome found in the AAV2 / 5 capsid and the AAV2 / 5 genome found in the AAV9F capsid. NGS analysis of the percentage of cross-packaging in samples generated in expressSF+ cells revealed a similar level of cross-packaging as measured by serotype-specific QPCR (Table 4).
[0174] In summary, the inventors measured a significant reduction in the cross-packing rate in two capsid libraries produced by their method in insect cells compared to the industry standard in HEK293t cells. While this method only provides a simplified view of the packaging process, a similar method has been used to establish cross-packing in AAV libraries produced in HEK293t cells. Low levels of cross-packing correlate with an increased rate of targeted evolutionary research aimed at selecting capsids with novel properties. The inventors predict that the reduced cross-packing in their capsid library production method in insect cells will lead to faster selection of novels in future research.
[0175] [Table 4]
Claims
1. i) A nucleotide sequence encoding an adeno-associated virus (AAV) capsid variant, wherein the nucleotide sequence encodes mRNA, and its translation in an insect cell generates the AAV capsid variant, which is operably linked to a promoter for driving its expression in the insect cell; ii) An enhancer element operably coupled to the promoter, wherein the introduction of the transcriptional transregulator into an insect cell containing the nucleic acid construct induces transcription from the promoter; iii) At least one AAV reverse terminal repeat (ITR) sequence; iv) A reporter gene operably linked to a promoter for driving expression in mammalian cells by arbitrary selection; and v) Barcodes (optional selection) A linear nucleic acid molecule containing a nucleic acid construct that includes [the specified component].
2. The nucleic acid molecule according to claim 1, wherein the transcriptional transregulator is a baculovirus pre-initial protein (IE1) or its splice variant (IE0), the transcriptional transregulator-dependent enhancer element is a baculovirus homologous region (hr) enhancer element, and preferably the baculovirus is Autographa californica multicapsid nuclear polyhedron disease virus.
3. The hr enhancer element comprises at least one copy of the hr28-mer sequence CTTTACGAGTAGAGAATTCCTACGCGTAAAAA, and / or at least 20, 21, 22, 23, 24, 25, 26, or 27 nucleotides identical to the sequence CTTTACGAGTAGAGAATTCCTACGCGTAAAAA, and includes at least one copy of a sequence that binds to the baculovirus IE1 protein, wherein the hr enhancer element is operably linked to an expression cassette containing a reporter gene operably linked to a pollH promoter, a) Under non-inducible conditions, the expression cassette having the hr enhancer element produces less reporter transcript than another identical expression cassette containing the hr2-0.9 element, or the cassette having the hr enhancer element produces less than 1.1, 1.2, 1.5, 2, 5, or 10 times the amount of reporter transcript produced by another identical expression cassette containing the hr4b element; b) Under induction conditions, the expression cassette having the hr enhancer element produces at least 50, 60, 70, 80, 90, or 100% of the amount of reporter transcript produced by another identical expression cassette containing the hr4b or hr2-0.9 element. The nucleic acid molecule according to claim 2, more preferably, the hr enhancer element is selected from the group consisting of hr1, hr2-0.9, hr3, hr4b, and hr5, of which hr4b and hr5 are preferred, and of which hr4b is most preferred.
4. The promoter for driving expression in insect cells is a baculovirus promoter, preferably a late or very late baculovirus promoter, more preferably a promoter selected from the group consisting of polH, p10, p6.9, and pSel120 promoters. A nucleic acid molecule according to any one of claims 1 to 3.
5. The nucleic acid molecule according to any one of claims 1 to 4, wherein the linear nucleic acid molecule is obtained by digestion with a site-specific endonuclease that cleaves at a recognition sequence in the nucleic acid construct.
6. The nucleic acid molecule according to any one of claims 1 to 5, wherein at least one end of the linear nucleic acid molecule is protected from exonuclease degradation, preferably both ends of the linear nucleic acid molecule are protected from exonuclease degradation.
7. The nucleic acid molecule according to claim 5 or 6, wherein the endonuclease is a protelomerase, preferably phage N15 TeLN protelomerase, and at least one end of the linear nucleic acid molecule is closed by a covalent bond, or preferably both ends of the linear nucleic acid molecule are closed by covalent bonds.
8. The nucleic acid molecule according to any one of claims 5 to 7, wherein the positions of elements i) to v) and the recognition sequence in the construct are such that the expression of the AAV Rep protein in an insect cell containing the nucleic acid construct results in the replication of the DNA fragment containing elements i) to v) and packaging into an AAV capsid.
9. The nucleic acid construct is arranged sequentially from the 5' side to the 3' side. a) The first AAV ITR sequence; b) Enhancer element; c) an expression cassette comprising a nucleotide sequence encoding the AAV capsid variant, operably linked to a promoter for driving expression in the insect cells; and d) Second AAV ITR Includes, The recognition sequence is located at least one of i) upstream of the first AAV ITR and ii) downstream of the second AAV ITR, and optionally, the reporter gene operably linked to the promoter for driving expression in mammalian cells is located between a) and b), between b) and c), or between c) and d), and optionally, the barcode is located between a) and b), between b) and c), or between c) and d), and preferably, the recognition sequence is a unique recognition sequence. A nucleic acid molecule according to any one of claims 1 to 8.
10. a) The promoter for driving expression in mammalian cells is a cell type and / or tissue-specific promoter, preferably a promoter selected from the group consisting of the human synapsin promoter (hSynl), trans tyretin promoter (TTR), cytokeratin 18, cytokeratin 19, unc-45 myosin chaperone B (unc45B) promoter, cardiac troponin T (cTnT) promoter, glial fiber acidic protein (GFAP) promoter, myelin basic protein (MBP) promoter, and methyl-CpG binding protein 2 (Mecp2) promoter; b) The reporter gene encodes a colorimetric, fluorescent, or bioluminescent reporter protein, preferably the reporter protein is selected from the group consisting of secreted alkaline phosphatase (SEAP), β-galactosidase EGFP, mCherry, mCloverS, mRubyS, mApple, iRFP, tdTomato, mVenus, YFP, RFP, firefly luciferase, sea pansy luciferase, and nanoluciferase; c) The barcode is a nucleotide sequence with a length of 4 to 18 nucleotides. A nucleic acid molecule according to any one of claims 1 to 9, wherein the nucleic acid molecule is at least one of the following.
11. A library comprising a plurality of nucleic acid molecules, each comprising a construct as defined in any one of claims 1 to 10, wherein in each member of the library, the nucleotide sequence encoding the AAV capsid variant differs by at least one nucleotide from the nucleotide sequence encoding the AAV capsid variant in the other members of the library, preferably, in each member of the library, the nucleotide sequence encoding the AAV capsid variant differs by at least one amino acid from the AAV capsid variant encoded by the other members of the library.
12. Insect cells comprising a construct as defined in any one of claims 1 to 10, or a library as defined in claim 11.
13. The insect cell according to claim 12, further comprising a baculovirus vector comprising at least one expression cassette for the expression of AAV Rep in the insect cell.
14. A library comprising multiple AAV capsid variants, wherein the members of the library comprise nucleic acids replicated and packaged from members in a library of nucleic acid constructs as defined in claim 11.
15. A method for preparing a library of AAV capsid variants, a) A step of providing a library comprising a plurality of nucleic acid constructs as defined in any one of claims 1 to 10, wherein the amino acid sequence of the AAV capsid variant encoded in the construct of a member of the library differs by at least one amino acid from the amino acid sequence of the AAV capsid variant encoded in the construct of another member of the library; b) The step of amplifying the library of nucleic acid constructs obtained by rolling circle amplification in a); c) b) A step of digesting the amplified library of nucleic acid constructs obtained by the endonuclease to prepare a library containing linear amplified nucleic acid constructs of unit length; d) A step of transfecting insect cells with the library containing the linearly amplified nucleic acid construct obtained in c); e) Infecting the insect cells with a baculovirus vector containing at least one expression cassette for expressing the AAV Rep protein in the insect cells, and culturing the cells; and f) e) The step of recovering a library of AAV capsid variants produced by the insect cells. A method that includes this.
16. a) The library in the above is i) to provide a nucleic acid construct as defined in any one of claims 1 to 10; and ii) Provided by replacing a fragment in the construct that includes at least a portion of the nucleotide sequence encoding the AAV capsid variant with a library of fragments, wherein the amino acid sequence of the AAV capsid variant of each member of the library differs by at least one amino acid from the amino acid sequence of the AAV capsid variant encoded in the fragment of that member and encoded in the construct of the other members of the library. The method according to claim 15.
17. -In step ii) of claim 16, the substitution is performed by a Gibson assembly; - In step b) of claim 15, the rolling circle amplification is performed by using Φ29 DNA polymerase, preferably in combination with a random primer; and / or - In claim 15c), the terminal of the linear amplification library of the nucleic acid construct of unit length is provided with protection against exonuclease degradation. The method according to claim 16.
18. - In step c) of claim 15, the endonuclease is phage N15 TeLN protelomerase, which provides a covalently closed end to the linear amplification library of the nucleic acid construct of unit length. The method according to any one of claims 15 to 17.
19. A method for identifying AAV capsid variants having desired characteristics, a) To provide a library of AAV capsid variants defined in claim 14 or obtained by any one of claims 15 to 18; b) Contacting the library with a culture of cells, organoids, or tissue, or administering the library to a non-human animal; c) Transduction of the AAV capsid variant in the library into the cells, organoids, tissues, or cells in the animals; and d) Identifying at least one AAV capsid variant obtained by transducing at least one desired cell type in the cells, organoids, tissues, or animals in b) as an AAV capsid variant having desired characteristics, and optionally recovering the AAV capsid variant having the desired characteristics from the desired cells. A method that includes this.
20. The method according to claim 19, wherein step d) includes detecting the expression of the reporter gene in the desired cells, thereby detecting at least the transduction of the desired cells.
21. The method according to claim 19 or 20, wherein the variant having the desired features is identified by sequencing i) the barcode and ii) at least one region of the DNA construct that codes for the difference of at least one amino acid.