Capsid engineered adeno-associated viruses

EP4698637A1Pending Publication Date: 2026-02-25REGENTS OF THE UNIVERSITY OF MINNESOTA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024793506
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-18
Filing Date
2024-04-18
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Current Adeno-associated virus (AAV) based gene therapies face limitations in production yield, DNA packaging capacity, immunogenicity, cell type specificity, and infectivity, despite their clinical potential, due to the constraints of natural AAV serotypes.

Method used

The development of capsid engineered AAVs with saturated programmable insertion engineering (SPINE) to identify and utilize engineerable hotspots for large domain insertions, such as HUH domains and nanobodies, at specific residues on the viral capsid, allowing for targeted cell specificity and enhanced infectivity.

Benefits of technology

This approach reveals a robustness of AAV capsids to accommodate large domain insertions, enabling targeted cell tropism and improved infectivity, facilitating the covalent attachment of molecules like antibodies, thereby enhancing the therapeutic potential of AAV-based gene therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000042_0001
    Figure IMGF000042_0001
  • Figure IMGF000060_0001
    Figure IMGF000060_0001
  • Figure IMGF000060_0002
    Figure IMGF000060_0002
Patent Text Reader

Abstract

Provided herein is an infectious recombinant adeno-associated virus (rAAV), comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises one or more insertions of amino acids that includes one or more single stranded nucleic acid binding domains, one or more small binding proteins or a combination thereof at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1 residues.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CAPSID ENGINEERED ADENO-ASSOCIATED VIRUSES

[0002] STATEMENT OF GOVERNMENT SUPPORT

[0003] This invention was made with government support under GM141152, and DA048742 awarded by the National Institutes of Health. The government has certain rights in the invention. BACKGROUND

[0004] Recombinant Adeno-associated virus (rAAV) has proven to be safe and able to drive longterm expression in dividing and non-dividing human cells. Several of AAV-based therapeutics have been approved by the FDA (1-4) and numerous clinical trials using AAV for the treatment of genetic diseases are underway. Despite the exceptional clinical potential of naturally evolved AAV serotypes, they could be substantially improved with respect to production yield, DNA packaging capacity, immunogenicity, cell type specificity, and infectivity (reviewed in (5)).

[0005] Addressing the drawbacks is facilitated by the relatively simple structural and genetic organization of AAV. The ~4.7kb single-stranded DNA genome comprises two genes, rep and cap, flanked by inverted terminal repeats (ITR). One ORF of the capsid gene cap encodes for three viral proteins VP1 (737 aa, 87 kDa), VP2 (600 aa, 72 kDa), and VP3 (535-503 aa, 62 kDa). It is expressed from the p40-promoter (C-terminus of the rep gene) and translated from overlapping open reading frames in a way that VP2 is lacking the N-terminus of VP1, and VP3 is missing the N-terminal part of VP1 / VP2 (6). Other cap ORFs express the assembly-activating protein (AAP) (7) and membrane-associated accessory protein (MAAP (8)). 60 VP monomers assemble into a viral capsid at an average ratio of 1: 1:10 VP1, VP2, and VP3, respectively (9-12). This ratio is highly divergent and assembly is stochastic, such that every capsid has a unique structural assembly (13). The icosahedral capsid features a cylindrical pore at the 5-fold interface, depressions surrounding the 5-fold pore continuing through the 2-fold axis, as well as protrusions at the 3-fold axis (14). Although the overall topology of AAV capsids is conserved across serotypes, Govindasamy et al. determined variable regions (VR1-9) mapping to surface loops of the capsid (15). The unique N-terminal portion of VP1, also known as VPlu, is located inside the capsid and is indispensable for infection. Upon infection, acidification during endosomal trafficking causes unfolding of the VPlu domain so that it can be externalized through the pore (16-20). A conserved phospholipase A2 domain (PLA2 (21)) and nuclear localization signals (NLS (22)), which are part of VPlu, can then facilitate endosomal escape and nuclear entry (23).

[0006] A variety of capsid engineering approaches have been applied in the past to improve the natural infection efficiency of A A Vs ranging from shuffling of natural AAV serotypes, recovery of ancestral serotypes, and peptide display. These methods resulted in significantly advanced capsid variants, such as AAV-DI (24), Anc80 (25), AAV-PHP.eB (26) or AAV2.7m8 (27). Moreover, there have been a few attempts to incorporate larger, structured protein domains. The first was the fusion of a green fluorescent protein (GFP) N-terminally to VP2, which was used to visualize intracellular trafficking of AAV particles (28). In the same manner, Gaussia Luciferase (29) and an ankyrin repeat protein (DARPin (30)) were successfully incorporated into the capsid. Other studies incorporated domains for a cell type-specific targeting into VR4 of VP1 or VP2, including nanobodies (31), HUH tags (32) or DARPins (33). While the aforementioned domain insertions were either N-terminal fusions or insertions into VR4 and peptide insertions predominantly focused on VR8, few studies attempted to comprehensively survey permissive capsid regions for domain insertions. Judd et al. constructed a random insertion library of mCherry into the VP3 encoding section of VP1 of AAV2, identifying only a single clone, in VR4, that tolerated insertion (34). Thus, there is a paucity of large-scale domain insertional datasets that comprehensively assess (i) the effect on biologically relevant functions, i.e., packaging, binding, and infection, as well as (ii) the effect of inserting domains with different physicochemical properties. In absence of these data, the boundaries of AAV capsid plasticity with respect to accepting domain insertions while maintaining fitness (i.e., assembly, packaging, cell entry, etc.) are yet to be fully understood.

[0007] SUMMARY

[0008] Evolved properties of Adeno- Associated Virus (AAV), such as broad tropism and immunogenicity in humans, are barriers to A AV-based gene therapy. Previous efforts to reengineer these properties have focused on variable regions near AAV’s 3-fold protrusions and capsid protein termini. To comprehensively survey AAV capsids for engineerable hotspots, as disclosed herein, multiple AAV fitness phenotypes were obtained upon insertion of large, structured protein domains into the entire AAV-DJ VP1. The data revealed a surprising robustness of AAV capsids to accommodate large domain insertions. There was positional, domain-type, and fitness phenotype dependence of insertion permissibility, which clustered into correlated structural units that could be linked to distinct roles in AAV assembly, stability, and infectivity. Engineerable hotspots of AAV were identified that facilitate the covalent attachment of binding scaffolds, which may represent an alternative approach to re-direct AAV tropism.

[0009] In particular, Saturated Programmable Insertion Engineering (SPINE) was combined with sequencing-based fitness assays to comprehensively determine multiple AAV fitness phenotypes upon insertion of the FLAG peptide tag as well as several large, structured protein domains into VP1 of AAV-DJ. The data revealed a surprising robustness of AAV viral capsids to accommodate large protein domain insertions. The positional, domain-type, and fitness phenotype dependence of insertion permissibility, which can be mapped to contiguous structural units of the AAV capsid which in turn can be linked to distinct roles in AAV assembly, stability, and infectivity. The engineerable hotspots of AAV accept insertion of 1 protein tags that facilitate the covalent attachment of molecules of interest such as antibodies. These hotspots may enable alternative approaches to redirect AAV tropism.

[0010] The disclosure provides an infectious recombinant adeno-associated virus (rAAV), comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises one or more insertions of amino acids that includes of one or more single stranded nucleic acid binding domains, one or more small binding proteins or a combination thereof at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1, e.g., from 659 to 669 or from 703 to 713, of VP1 (numbering is based on reference VP1 sequence of SEQ ID NOs: 80 and 81). In one embodiment, the insertion is at reside 262, 328, 387, 631, 651, 664, 670, 684 or 708.

[0011] In one embodiment, the insertion is the in a protruding, cylindric pore at the 5-fold interface, a valley extending toward the 2-fold interface around the pore, and protrusions at the 3- fold interface, such as near the 2-fold valley and / or 5-fold interface, including the 5-fold port, 3- fold protrusion, 2-fold valley and / or valley surround the pore, including those sites on the surface of the capsid.

[0012] In one embodiment, the one or more single stranded nucleic acid binding domains are single stranded DNA binding domains. In one embodiment, the one or more single stranded DNA binding domains are HUH domains. In one embodiment, the modification in the viral capsid further comprises a deletion of one or more amino acids of the VP. In one embodiment, the deletion is 2, 3, 4, 5, or 10 amino acids or less than 50 amino acids. In one embodiment, the insertion comprises 5 to 100 amino acids. In one embodiment, the one or more small binding proteins are nanoboides, DARPins, and / or GP2 scaffolds. In one embodiment, VP1 of AAV-DI is modified. In one embodiment, the AAV genome or capsid is AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, or AAVrhlO. In one embodiment, the AAV genome or capsid is AAV1, AAV2, AAV5, AAV6, AAV8, AAV9, AAV-DJ, or AAV10. In one embodiment, the rAAV infects human cells. In one embodiment, the one or more HUH domains have a sequence having at least 80% amino acid sequence identity to a HUH domain in one or more of SEQ ID Nos. 2-10 and 20-21. In one embodiment, the one or more HUH domains have a sequence having at least 90% amino acid sequence identity to a HUH domain in one or more of SEQ ID NOs 2-10 and 20-21. In one embodiment, the one or more insertions are flanked by at least one linker. In one embodiment, at least one insertion is flanked by two linkers. In one embodiment, at least one insertion is flanked by one linker. In one embodiment, up to 10 insertions are in the one or more insertions. In one embodiment, at least two different single stranded nucleic acid binding domains or small binding proteins are in the insertion. In one embodiment, the HUH domains are concatemers. In one embodiment, the rAAV has a plurality of the same insertions. In one embodiment, the insertion is from about 2 kDa up to 300 kDa, about 5 kDa up to 50 kDa, about 10 kDa up to 250 kDa, about 20 kDa up to a 100 kDa, about 10 kDa up to 50 kDa, or about 75 kDa up to about 150 kDa. In one embodiment, the viral genome is a recombinant genome having at least one expression cassette for an exogenous gene product. In one embodiment, the exogenous gene product is a prophylactic or therapeutic gene product. In one embodiment, the exogenous gene product is a cytotoxic gene product. In one embodiment, a targeting molecule is linked to the insertion. In one embodiment, the targeting molecule comprises an antibody or an antigen binding portion thereof. In one embodiment, the targeting molecule comprises one or more albumin-binding domains (ABD), adhirons, adnectin / monobodies, affibodies, affilins, affimers, affitins / nanofitins, alphabodies, anticalins, armadillo repeat proteins, atrimers, avimers, aentyrins, DARPins, fynomers, Kunitz-domains, peptide toxins, peptide ligands, obodies / OB-Fold proteins, pronectins, or repebodies, RNA aptamers, one or more antibodies, or any combination thereof. In one embodiment, the HUH substrate comprises a nucleotide sequence having at least 90% nucleic acid sequence identity to any one of SEQ ID Nos. 1-4 or 11, or to at least 9, 12 or 15 nucleotides thereof that bind one or more of the HUH domains. In one embodiment, the HUH substrate comprises PNA or LNA.

[0013] Further provided is a method of targeting mammalian cells in vivo, comprising: providing a population of infectious rAAV, comprising: rAAV comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises one or more insertions of amino acids that includes of one or more single stranded nucleic acid binding domains, one or more small binding proteins or a combination thereof at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1; providing a single stranded nucleic acid binding substrate for the one or more single stranded nucleic acid binding domains covalently linked to a targeting molecule; combining the population of recombinant virus and the substrate covalently linked to the targeting molecule to form a conjugate; and administering the conjugate to a mammal or administering a recombinant virus comprising an insertion of one or more small binding proteins. In one embodiment, the single stranded nucleic acid binding domain comprises one or more HUH domains. In one embodiment, the insertion is at reside 262, 328, 387, 631, 651, 664, 670, 684 or 708. In one embodiment, the mammal is a non-human mammal. In one embodiment, the mammal is a human. In one embodiment, the targeting molecule is an antibody or an antigen binding portion thereof. In one embodiment, the antibody is an anti-CD3 antibody, anti-CD4 antibody, anti-CD7 antibody, anti- Her2 antibody, anti-CD34 antibody, anti-CD8, anti-CD20 antibody, or anti-CD19 antibody. Also provided is a system comprising: a population of infectious rAAV comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an one or more insertions of amino acids that includes of one or more single stranded nucleic acid binding domains, one or more small binding proteins or a combination thereof at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1, for example, 659 to 669 or from 708 to 713 of VP1; and optionally a substrate for the single stranded nucleic acid binding domain. In one embodiment, the one or more single stranded nucleic acid binding domains comprise HUH domains. In one embodiment, the system further includes a targeting molecule. In one embodiment, the substrate is covalently linked to a targeting molecule. In one embodiment, the insertion is at reside 262, 328, 387, 631, 651, 664, 670, 684 or 708.

[0014] An infectious recombinant adeno-associated virus (rAAV) is provided comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises one or more insertions of amino acids that includes of one or more single stranded nucleic acid binding domains, one or more small binding proteins or a combination thereof at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1, which insertion is optionally bound to the one or more single stranded nucleic acid binding domains which are linked to a molecule. In one embodiment, the molecule is a targeting molecule. In one embodiment, the modified capsid protein has a domain that covalently binds ssDNA in a sequence specific and non-overlapping fashion and by linking different monoclonal antibodies (mAB) to one of the ssDNA, which nucleotides may be modified, e.g., with peptide nucleic acid (PNA) or locked nucleic acid (LNA), and then combining the mAB / ssDNA with virus having the modified capsid, the virus becomes decorated with a selected antibody to form covalent mAB / virus composites. This approach extends to other receptor-specific binding proteins (e.g., nanobodies, Darpin, Affibodies, etc.) or non-antibody-based proteins. In one embodiment, the viral capsid includes an affinity tag, e.g., a His-tag, HA-tag, a FLAG-tag, or a Strep-tag, or other tags useful for affinity purification and isolation, e.g., to purify virions. In one embodiment, the viral genome is a recombinant genome having at least one expression cassette for an exogenous gene product. In one embodiment, the exogenous gene product is a prophylactic or therapeutic gene product, e.g., a cytotoxic gene product.

[0015] Further provided is a method of targeting mammalian cells. In one embodiment, a population of the rAAV comprising an insertion that includes one or more HUH domains that is combined with a HUH substrate for the one or more HUH domains covalently linked to a targeting molecule to form a conjugate, and the conjugate is contacted with mammalian cells, e.g., primate cells including human cells, or bovine, equine, canine, feline ovine, swine or caprine cells, in vitro or in vivo. In one embodiment, the targeting molecule is an antibody or an antigen binding portion thereof.

[0016] BRIEF DESCRIPTION OF THE FIGURES

[0017] Figs. 1A-1L. Design and analysis of AAV domain insertion libraries. (A) Schematic of library design. (B) Plasmids used for the triple transfection in AAV production. (C) Schematic of AAV capsid assembly. (D) Purification of AAV via gradient purification. (E-H) Schematic of workflows analyzing the fitness of AAV domain insertion libraries, i.e., pulldown, binding, uptake, and infectivity assays. (I) Quantification of packaging titer via qPCR. Data are means ± SEM. One way ANOVA test p value 0.19, not significant (ns). (J) Representative Western blot image of AAV domain insertion libraries stained with Al antibody (detecting VP1 subunits). (K) Western blot quantification of VP1, VP2 and VP3 subunits. Data are means (n=3). Two-way ANOVA test p value 0.641, not significant (ns). (L) Tmof AAV libraries obtained by DSF assay. Data are means (n=2).

[0018] Figs. 2A.1, 2A.2, 2A.3 and 2B. Fitness of AAV domain insertion libraries. (A) Heatmap representing the fitness of each motif insertion in each VP1 position for AAV packaging, pulldown, binding, uptake, and infectivity assays. Green indicates higher and magenta lower fitness than AAV-DJ (white). Yellow denotes positions without data. VP1 secondary structure elements and VR1-9 are indicated on top. (B) Fitness distributions of all AAV libraries compared to AAV-DJ (fitness = 0) ± standard error (red dashed lines).

[0019] Figs. 3A-3B. Packaging fitness of AAV domain insertion libraries mapped to the capsid structure. (A) Top left corner: AAV-DJ capsid structure view from the inside (left) and outside (right) radially color-cued. 2-, 3-, and 5-fold axes are indicated. VPlu domain was modeled using RoseTTAFold (98) and manually positioned. All other structures show packaging fitness heatmaps of the indicated domain insertions. Green indicates higher and magenta lower fitness than AAV-DJ (RCSB PDB 7KFR). (B) Zoom of the outside structures from (A). 2-, 3-, and 5- fold axes are outlined.

[0020] Figs. 4A-4B. Pulldown fitness of AAV domain insertion libraries mapped to the capsid structure. (A) Top left corner: AAV-DJ capsid structure view from the inside (left) and outside (right) radially color-cued. 2-, 3-, and 5-fold axes are indicated. VPlu domain was modeled using RoseTTAFold (98) and manually positioned. All other structures show pulldown fitness heatmaps of the indicated domain insertions. Green indicates higher and magenta lower fitness than AAV- DJ (white) (RCSB PDB 7KFR). (B) Zoom of the outside structures from (A). 2-, 3-, and 5-fold axes are outlined. Figs. 5A-5E. Variance of packaging and uptake fitness. (A, B) Variance of packaging fitness (A) and uptake fitness (B) from all domain insertion libraries mapped to the AAV capsid structure. The capsid inside (top) and the capsid outside (bottom) are shown (RCSB PDB 7KFR). (C) Empirical cumulative density insertional fitness of residues within (petrol green) and outside (red) the 2-fold axis (left) and 3-fold axis (right). Significance of distribution differences was tested using a two-sided, two-sample Kolmogorov-Smirnov test. Significance level and p values are shown. (D) Venn diagram showing which of the 335 AAV interface residues are unique and shared among interfaces.

[0021] Figs. 6A-6D. Unbiased clustering of insertion fitness. (A) UMAP cluster analysis of the AAV domain insertional profiling data resulting in five distinct clusters. (B) Cluster map to distinct capsid regions: (1) N-terminus of VP1 and bases of the 3-fold and 5-fold interface in red, (2) the HI-loop, the pore and the inner connecting residue layer in yellow, (3) protrusions of the 3-fold interface in green, (4) 2-fold axis in blue, and (5) the surrounding of the 5-fold pore in light lilac (C) Mean domain insertion fitness of packaging, pulldown, binding, and uptake shown by residues aligned to the UMAP clusters in (A). (D) Distribution of insertion fitness of packaging, pulldown, and uptake for each cluster.

[0022] Figs. 7A-7E. Fitness of WDV insertion variants. (A) Quantification of packaging titer of AAV-DJ and 10 different WDV insertion variants via qPCR. Data are means ± SD. (B) Infection fitness of the WDV insertion variants in (A) quantified by measuring the percentage of tdTomato positive cells 48 hours post transduction at the indicated MOIs. Data are mean ± SEM. (C) Zoom to the 2-fold and 5-fold axes (outlined) of the capsid surface. Positions of N664 and K708 are shown as red spheres. (D) Quantification of packaging titers of AAV-DJ and WDV insertion variants N664 and K708. Data are mean ± SD. (E) Infection fitness of the WDV insertion variants N644 and K708 quantified by measuring the percentage of tdTomato positive cells 48 hours post transduction at an MOI of 1x104vg. If indicated, cells were co-expressing GFP-GPI and / or a ssDNA-anti-GFP antibody was added. Data shown as box plots. Lower and upper hinges of boxes indicate 25thand 75thpercentile, respectively. Mean is indicated by a horizontal bar in each box. Whiskers extend 1.5 * IQR. Two-way repeated measure ANOVA was used to test the significance of variance between means (AAV-DJ, N664, K708) for conditions with and without GFP present. Presence of the mAb was a significant source of variation when GFP-GPI was expressed (p value 0.0063) but not without (p value 0.514). Significance levels and p values for pairwise comparisons using a Bonferroni correction are shown.

[0023] Figs. 8A-8B. Properties of inserted domains. (A) Physical descriptions of domains used in this study. (B) Cartoon representation of domain structures (left) and surface representation with net surface charge shown as a gradient from red (negative), over white (neutral), to blue (positive).

[0024] Fig. 9. Packaging fitness of silent mutation variant compared to AAV-DJ. Crude lysate packaging titers quantified via qPCR. Three replicates with three technical replicates each were performed. Data are means ± SD. No statistical significance (ns) between AAV-DJ and the silent mutations variant by an unpaired, two-sided Student’s t-test (p-value: 0.3327).

[0025] Figs. 10A-10F. Infection fitness assay gating scheme. (A) Whole HEK293FT cells are gated on side (SSC-A) and forward scattering (FSC-A). (B-C) Forward scattering height (FSC- H), forward scattering width (FSC-W), and side scattering width (SSC-W) are used to gate single cells. (D) Cells are further gated using miRFP670nano as an infection marker (representative example). (E-F) Sort statistic for each gated cell population.

[0026] Figs. 11A-11D. Fitness assay biological replicates, data completeness, and depth. (A) Read counts of replicate 1 are plotted against read counts of replicate 2 for the plasmid library, packaging, pulldown, binding, uptake, and infection assays. Data for all seven inserted domains are shown. Linear correlation was calculated (blue line). Pearson correlation coefficient and p values are shown. (B) Contingency plots showing the fraction of insertion positions below (red) and above (grey) the read count quantity cut-off of 50 reads for all seven domains and measured phenotypes. Data are shown for replicate 1 (top) and replicate 2 (bottom). (C) Total read counts of the plasmid library, packaging, pulldown, binding, uptake, and infection assays for each position are represented by insertion position. Data for all seven inserted domains are shown. (D) Cumulative density plots showing the read counts of the plasmid library, packaging, pulldown, binding, uptake, and infection assays for all seven inserted domains for replicate 1 and replicate 2. Cut off for sufficient read quantity was set to 50 reads and is represented by the black dashed lines.

[0027] Figs. 12A-12B. Western blot of AAV domain insertion libraries. (A-B) Representative Western blot image of AAV domain insertion libraries stained with Bl antibody (detecting VP1, VP2 and VP3 subunits) at a short exposure time (A) and long exposure time (B).

[0028] Fig. 13. Thermal profiles of AAV domain insertion libraries. Data of two replicates are shown as "6(Fluorescence) / <)(Temperature)" versus Temperature in °C.

[0029] Figs. 14A-14B. Quantification of empty to full capsid ratio by negative staining transmission electron microscopy. (A) More than 300 capsids were counted manually and grouped into full, empty, and undecided and the percentage of full capsids was calculated. (B) Representative transmission electron microscopy images of the indicated samples.

[0030] Fig. 15. Accuracy of domain insertional profiling data. Standard errors of AAV fitness of all seven domain insertion libraries by insertion position and fitness assay. Secondary structure elements and VR1-9 of the capsid are indicated on top. Boundaries of oligos from VP1 assembly using SPINE are indicated by purple vertical bars.

[0031] Fig. 16. Distribution of pulldown fitness by residue location. Cumulative density function of pulldown fitness for each inserted motif stratified by residue location; within VPlu (blue), buried or exposed inside the capsid and not in VPlu (red), external residues (gray). Significance of distribution differences was tested using a two-sided, two-sample Kolmogorov-Smirnov test. Significance level and p values are shown.

[0032] Figs. 17A-17B. Binding fitness of AAV domain insertion libraries mapped to the capsid structure. (A) Top left corner: AAV-DJ capsid structure view from the inside (left) and outside (right) radially color-cued. 2-, 3-, and 5-fold axes are indicated. VPlu domain was modeled using RoseTTAFold and manually positioned. All other structures show binding fitness heatmaps of the indicated domain insertions. Green indicates higher and magenta lower fitness than AAV-DJ (white) (RCSB PDB 7KFR). (B) Zoom of the outside structures from (A). 2-, 3-, and 5-fold axes are outlined.

[0033] Figs. 18A-18B. Uptake fitness of AAV domain insertion libraries mapped to the capsid structure. (A) Top left corner: AAV-DJ capsid structure view from the inside (left) and outside (right) radially color-cued. 2-, 3-, and 5-fold axes are indicated. VPlu domain was modeled using RoseTTAFold and manually positioned. All other structures show uptake fitness heatmaps of the indicated domain insertions. Green indicates higher and magenta lower fitness than AAV-DJ (white) (RCSB PDB 7KFR). (B) Zoom of the outside structures from (A). 2-, 3-, and 5-fold axes are outlined

[0034] Figs. 19A-19C. Variance of all measured fitness assays. (A) Variance of different fitness measures from all domain insertion libraries mapped to the AAV capsid structure. The capsid inside (left) and the capsid outside (right) are shown (RCSB PDB 7KFR). (B) Empirical cumulative density insertional fitness of residues within (petrol green) and outside (red) at the 2- fold axis (top) and 3-fold axis (bottom). Significance of distribution differences was tested using a two-sided, two-sample Kolmogorov-Smirnov test. Significance levels are shown (**** < 0.0001, *** < 0.001, ** < 0.01, * < 0.05).

[0035] Figs. 20.A-20B. Packaging and infection fitness of cysteine mutants. (A) Crude lysate packaging titers quantified via qPCR. (B) Infection fitness quantified by measuring the percentage of tdTomato positive cells 48 hours post transduction with an MOI of 1E4 vg. Data are means ± SD. Data points of single mutants are colored by cluster membership, missing residues are gray, wildtype AAV-DJ fitness is shown as open circles and horizontal dashed line.

[0036] Figs. 21A-21C. Quantification of VP ratios of WDV insertion variants N664 and K708. (A) Representative Western blot image of AAV domain insertion libraries stained with Bl antibody (detecting VP1, VP2 and VP3 subunits). (B) Representative Western blot image of AAV domain insertion libraries stained with A l antibody (detecting VP1 subunits). (C) Western blot quantification of VP1, VP2, and VP3 subunits. Data are means (n=3).

[0037] Fig. 22. Amino acid residues in VP1 for exemplary serotypes.

[0038] Figs. 23A-23C. FAP nanobody incorporation at different positions of AAV-DJ. (A) Schematic of enhanced infection of FAP receptor-positive cells compared to FAP receptornegative cells by incorporating a FAP nanobody into the AAV capsid. (B) AAV-DJ capsid structure (RCSB PDB 7KFR) from the outside (left) with a zoom-in (right). The eight FAP nanobody insertion positions are highlighted in red. (C) Plasmid design of nanobody (nb) insertion into either VP1 (top) or VP2 (bottom) and the respective trans-complementation plasmids for AAV productions.

[0039] Figs. 24A-24D. FAP nanobody incorporation at various positions enhances infection of FAP receptor-positive cells. (A, B, D) R1 (A), Rl-FAP (B), and SK-MEE-24 (D) cells were transduced at an MOI of 1x103 vg / cell with the indicated capsid variants, followed by a luciferase assay. Photon counts were normalized to the DJ controls. (C) Relative luciferase values from A and B were divided by each other to calculate the specificity gain. The dashed horizontal line highlights the DJ control level. Data are means ± SEM. *p < 0.05, ***p<0.001 by one-way ANOVA and a Dunnett’s post hoc test.

[0040] Figs. 25A-25B. Comparison of infection gain between Rl-FAP and SK-MEE-24 cells. (A) Scatter plot representing the specificity gain of FAP nanobody insertion variants in Rl-FAP cells compared to SK-MEE-24 cells. Data are means ± SEM. (B, C) Averaged infectivity gains determined in Rl-FAP and SK-MEL-24 from both linker types for FAP-nanobody insertion into VP1 (B) and VP2 (C) were mapped onto the AAV-DJ capsid structure (RCSB PDB 7KFR).

[0041] Fig. SI. Protein sequences of FAP (yellow) and GFP (purple) nanobodies with linkers (grey) (SEQ ID NOS: 53-55).

[0042] Fig. S2. Production titers of nanobody insertion variants. Quantification of AAV crude lysate titers by qPCR. The dashed horizontal line highlights the DJ control level. Data are means ± SEM. Differences are not significant for any capsid variant compared to DJ by one-way ANOVA and a Dunnett’s post hoc test. Figs. S3A-S3E. FAP receptor expression profiles. (A) Whole cells (gate Pl) were gated on forward scattering area (FSC-A) and side scattering area (SSC-A). (B) Forward scattering width (FSC-W) and forward scattering area (FSC-A) were used to gate single cells (P2). (C-E) Histograms of unstained or FAP receptor-stained single cell populations of R1 cells (C), Rl-FAP cells (D), and SK-MEL-24 cells (E).

[0043] Figs. S4A-S4B. Specificity gain is dose-independent. (A-B) R1 and Rl-FAP cells were transduced at an MOI of 5x l02(A) or 1x102(B) vg / cell with the indicated capsid variants, followed by a luciferase assay. First, photon counts were normalized to the DJ controls of each cell line and second the Rl-FAP values were divided by the R1 values to calculate the specificity gain. The dashed horizontal line highlights the DJ control level. Data are means ± SEM. *p < 0.05, **p < 0.01, ***p<0.001 by one-way ANOVA and a Dunnett’s post-hoc test.

[0044] Figs. S5A-S5D. Specificity gain of infection is caused specifically by the FAP nanobody incorporated into the capsid. (A) Quantification of AAV crude lysate titers by qPCR. (B-C) R1 cells (B) and Rl-FAP cells. (C) were transduced at an MOI of 1x103vg / cell with the indicated capsid variants, followed by a luciferase assay. Photon counts were normalized to the DJ controls. (D) Relative luciferase values from B and C were divided by each other to calculate the specificity gain. The dashed horizontal line highlights the DJ control level Data are means ± SEM. **p < 0.01, ***p<0.001 by one-way ANOVA and a Dunnett’s post-hoc test.

[0045] DETAILED DESCRIPTION

[0046] Definitions

[0047] A "vector" or "construct" (sometimes referred to as gene delivery or gene transfer "vehicle") refers to a macromolecule or complex of molecules comprising a polynucleotide (such as a plasmid) to be delivered to a host cell, either in vitro or in vivo. The polynucleotide to be delivered may comprise a sequence of interest for gene therapy. Vectors include, for example, viral vectors, e.g., adenovirus including helper-dependent adenovirus vectors, which do not express any adenovirus genes and are immunologically silent to allow for persistent expression, adeno-associated virus (AAV), e.g., an AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV- 7, AAV-8, AAV-9, AAV-10, AAV-11 or AAV-12, and including pseudotyped viruses and nonnatural serotypes such as AAV-DJ (Grimm et al., J. Virol., 82:5887 (2008)), and AAV-PHP.eB or AAV-PHP.S (Chan et al., Nat. Neurosci., 20:1172 (2017)), the disclosures of which are incorporated by reference herein. Vectors can also comprise other components or functionalities that further modulate gene delivery and / or gene expression, or that otherwise provide beneficial properties to the cells to which the vectors will be introduced. Such other components include, for example, components that influence binding or targeting to cells; components that influence uptake of the vector nucleic acid by the cell; components that influence localization of the polynucleotide within the cell after uptake (such as agents mediating nuclear localization); and components that influence expression of the polynucleotide. Such components can be provided as a natural feature of the vector (such as the use of certain viral vectors which have components or functionalities mediating binding and uptake), or vectors can be modified to provide such functionalities. A large variety of such vectors are known in the art and are generally available. When a vector is maintained in a host cell, the vector can either be stably replicated by the cells during mitosis as an autonomous structure, incorporated within the genome of the host cell, or maintained in the host cell's nucleus or cytoplasm.

[0048] A "recombinant viral vector" refers to a viral vector comprising one or more heterologous genes or sequences. Since many viral vectors exhibit size constraints associated with packaging, the heterologous genes or sequences are typically introduced by replacing one or more portions of the viral genome. Such viruses may become replication-defective, requiring the deleted function(s) to be provided in trans during viral replication and encapsidation (by using, e.g., a helper virus or a packaging cell line carrying genes necessary for replication and / or encapsidation). Modified viral vectors in which a polynucleotide to be delivered is carried on the outside of the viral particle have also been described.

[0049] "Gene delivery," "gene transfer," and the like as used herein, are terms referring to the introduction of an exogenous polynucleotide (sometimes referred to as a "transgene"), e.g., via a recombinant virus, into a host cell, irrespective of the method used for the introduction. Such methods include a variety of well-known techniques such as vector-mediated gene transfer (by, e.g., viral infection / transfection, or various other protein-based or lipid-based gene delivery complexes) as well as techniques facilitating the delivery of "naked" polynucleotides (such as electroporation, iontophoresis, "gene gun" delivery and various other techniques used for the introduction of polynucleotides). The introduced polynucleotide may be stably or transiently maintained in the host cell. Stable maintenance typically requires that the introduced polynucleotide either contains an origin of replication compatible with the host cell or integrates into a replicon of the host cell such as an extrachromosomal replicon (e.g., a plasmid) or a nuclear or mitochondrial chromosome. A number of vectors are known to be capable of mediating transfer of genes to mammalian cells, as is known in the art.

[0050] By "transgene" is meant any piece of a nucleic acid molecule (for example, DNA) which is inserted by artifice into a cell either transiently or permanently and becomes part of the cell if integrated into the genome or maintained extrachromosomally. Such a transgene may include a gene which is partly or entirely heterologous (i.e., foreign) to the transgenic host cell or organism, or may represent a gene homologous to an endogenous gene of the host cell or organism.

[0051] “AAV” is adeno-associated virus and may be used to refer to the naturally occurring wildtype virus itself or derivatives thereof. The term covers all subtypes, serotypes and pseudotypes, and both naturally occurring and recombinant forms, except where required otherwise. As used herein, the term “serotype” refers to an AAV which is identified by and distinguished from other AAVs based on capsid protein reactivity with defined antisera, e.g., there are at least eight serotypes of primate AAVs, for example, AAV-1 to AAV-8. For example, serotype AAV2 is used to refer to an AAV which contains capsid proteins encoded from the cap gene of AAV 2 and a genome containing 5' and 3’ ITR sequences from the same AAV2 serotype. The abbreviation “rAAV” refers to recombinant adeno-associated virus, also referred to as a recombinant AAV vector (or “rAAV vector”).

[0052] A cell has been "transformed", "transduced", "transfected" or “genetically modified” by exogenous or heterologous nucleic acids when such nucleic acids have been introduced inside the cell. Transforming DNA may or may not be integrated (covalently linked) with chromosomal DNA making up the genome of the cell. In mammalian cells for example, the transforming DNA may be maintained on an episomal element, such as a plasmid. In a eukaryotic cell, a stably transformed cell is one in which the transforming DNA has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication.

[0053] The term "wild type" with respect to a gene or gene product refers to a gene or gene product that has the characteristics of that gene or gene product when isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designated the "normal" or "wild-type" form of the gene. In contrast, the term "modified" or "mutant" or "variant" refers to a gene or gene product that displays modifications in sequence and or functional properties (i.e., altered characteristics) when compared to the wildtype gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product.

[0054] The term "heterologous" as it relates to nucleic acid sequences such as gene sequences and control sequences, denotes sequences that are not normally joined together, and / or are not normally associated with a particular cell. Thus, a "heterologous" region of a nucleic acid construct or a vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with the other molecule in nature. For example, a heterologous region of a nucleic acid construct could include a coding sequence flanked by sequences not found in association with the coding sequence in nature, i.e., a heterologous promoter. Another example of a heterologous coding sequence is a construct where the coding sequence itself is not found in nature (e.g., synthetic sequences having codons different from the native gene). Similarly, a cell transformed with a construct which is not normally present in the cell would be considered heterologous for purposes of this invention.

[0055] A "gene," "polynucleotide," "coding region," or "sequence" which "encodes" a particular gene product, is a nucleic acid molecule which is transcribed and optionally also translated into a gene product, e.g., a polypeptide, in vitro or in vivo when placed under the control of appropriate regulatory sequences. The coding region may be present in either a cDNA, genomic DNA, or RNA form. When present in a DNA form, the nucleic acid molecule may be single-stranded (i.e., the sense strand) or double-stranded. The boundaries of a coding region are determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxy) terminus. A gene can include, but is not limited to, cDNA from prokaryotic or eukaryotic mRNA, genomic DNA sequences from prokaryotic or eukaryotic DNA, and synthetic DNA sequences. Thus, a gene includes a polynucleotide which may include a full-length open reading frame which encodes a gene product (sense orientation) or a portion thereof (sense orientation) which encodes a gene product with substantially the same activity as the gene product encoded by the full-length open reading frame, the complement of the polynucleotide, e.g., the complement of the full-length open reading frame (antisense orientation) and optionally linked 5’ and / or 3’ noncoding sequence(s) or a portion thereof, e.g., an oligonucleotide, which is useful to inhibit transcription, stability or translation of a corresponding mRNA. A transcription termination sequence will usually be located 3' to the gene sequence.

[0056] The term "control elements" refers collectively to promoter regions, polyadenylation stimulations, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites ("IRES"), enhancers, splice junctions, and the like, which collectively provide for the replication, transcription, post-transcriptional processing and translation of a coding sequence in a recipient cell. Not all of these control elements need always be present so long as the selected coding sequence is capable of being replicated, transcribed and translated in an appropriate host cell.

[0057] The term "promoter region" is used herein in its ordinary sense to refer to a nucleotide region comprising a DNA regulatory sequence, wherein the regulatory sequence is derived from a gene which is capable of binding RNA polymerase and initiating transcription of a downstream (3' direction) coding sequence. Thus, a "promoter," refers to a polynucleotide sequence that controls transcription of a gene or coding sequence to which it is operably linked. A large number of promoters, including constitutive, inducible and repressible promoters, from a variety of different sources, are well known in the art.

[0058] By "enhancer element" is meant a nucleic acid sequence that, when positioned proximate to a promoter, confers increased transcription activity relative to the transcription activity resulting from the promoter in the absence of the enhancer domain. Hence, an "enhancer" includes a polynucleotide sequence that enhances transcription of a gene or coding sequence to which it is operably linked. A large number of enhancers, from a variety of different sources are well known in the art. A number of polynucleotides which have promoter sequences (such as the commonly used CMV promoter) also have enhancer sequences.

[0059] "Operably linked" refers to a juxtaposition, wherein the components so described are in a relationship permitting them to function in their intended manner. By "operably linked" with reference to nucleic acid molecules is meant that two or more nucleic acid molecules (e.g., a nucleic acid molecule to be transcribed, a promoter, and an enhancer element) are connected in such a way as to permit transcription of the nucleic acid molecule. A promoter is operably linked to a coding sequence if the promoter controls transcription of the coding sequence. Although an operably linked promoter is generally located upstream of the coding sequence, it is not necessarily contiguous with it. An enhancer is operably linked to a coding sequence if the enhancer increases transcription of the coding sequence. Operably linked enhancers can be located upstream, within or downstream of coding sequences. A polyadenylation sequence is operably linked to a coding sequence if it is located at the downstream end of the coding sequence such that transcription proceeds through the coding sequence into the polyadenylation sequence.

[0060] "Operably linked" with reference to peptide and / or polypeptide molecules is meant that two or more peptide and / or polypeptide molecules are connected in such a way as to yield a single polypeptide chain, i.e., a fusion polypeptide, having at least one property of each peptide and / or polypeptide component of the fusion. Thus, a cleavable or targeting peptide sequence is operably linked to another protein if the resulting fusion is cleaved into two or more parts as a result of cleavable sequence or is transported into an organelle as a result of the presence of an organelle targeting peptide.

[0061] "Homology" refers to the percent of identity between two polynucleotides or two polypeptides. For example, polypeptides, e.g., viral capsid protein, may have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity to a reference polypeptide sequence, e.g., have at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity to any of Accession Nos. P03135 (AAV-2), AGA39530 (AAV-8), YP_077178.1 (AAV-7, AAS99264.1 (AAV-9), or AAC58045. 1 (AAV-4). The correspondence between one sequence and another can be determined by techniques known in the art. For example, homology can be determined by a direct comparison of the sequence information between two polypeptide molecules by aligning the sequence information and using readily available computer programs. Alternatively, homology can be determined by hybridization of polynucleotides under conditions which form stable duplexes between homologous regions, followed by digestion with single strand-specific nuclease(s), and size determination of the digested fragments. Two DNA, or two polypeptide, sequences are "substantially homologous" to each other when at least about 80%, at least about 90%, or at least about 95% of the nucleotides, or amino acids, respectively match over a defined length of the molecules, as determined using the methods above.

[0062] By "mammal" is meant any member of the class Mammalia including, without limitation, humans and nonhuman primates such as chimpanzees and other apes and monkey species; farm animals such as cattle, sheep, pigs, goats and horses; domestic mammals such as dogs and cats; laboratory animals including rodents such as mice, rats, rabbits and guinea pigs, and the like. An “animal” includes vertebrates such as mammals, avians, amphibians, reptiles and aquatic organisms including fish.

[0063] By "derived from" is meant that a nucleic acid molecule was either made or designed from a parent nucleic acid molecule, the derivative retaining substantially the same functional features of the parent nucleic acid molecule, e.g., encoding a gene product with substantially the same activity as the gene product encoded by the parent nucleic acid molecule from which it was made or designed.

[0064] By "expression construct" or "expression cassette" is meant a nucleic acid molecule that is capable of directing transcription. An expression construct includes, at the least, a promoter. Additional elements, such as an enhancer, and / or a transcription termination stimulation, may also be included.

[0065] The term "exogenous," when used in relation to a protein, gene or nucleic acid, e.g., polynucleotide, in a cell or organism refers to a protein, gene, or nucleic acid which has been introduced into the cell or organism by artificial or natural means, or in relation to a cell refers to a cell which was isolated and subsequently introduced to other cells or to an organism by artificial or natural means (a “donor” cell). An exogenous nucleic acid may be from a different organism or cell, or it may be one or more additional copies of a nucleic acid which occurs naturally within the organism or cell. An exogenous cell may be from a different organism, or it may be from the same organism. By way of a non-limiting example, an exogenous nucleic acid is in a chromosomal location different from that of natural cells or is otherwise flanked by a different nucleic acid sequence than that found in nature.

[0066] The term "isolated" when used in relation to a nucleic acid, peptide, polypeptide or virus refers to a nucleic acid sequence, peptide, polypeptide or virus that is identified and separated from at least one contaminant nucleic acid, polypeptide, virus or other biological component with which it is ordinarily associated in its natural source. Isolated nucleic acid, peptide, polypeptide or virus is present in a form or setting that is different from that in which it is found in nature. For example, a given DNA sequence (e.g., a gene) is found on the host cell chromosome in proximity to neighboring genes; RNA sequences, such as a specific mRNA sequence encoding a specific protein, are found in the cell as a mixture with numerous other mRNAs that encode a multitude of proteins. The isolated nucleic acid molecule may be present in single-stranded or double- stranded form. When an isolated nucleic acid molecule is to be utilized to express a protein, the molecule will contain at a minimum the sense or coding strand (i.e., the molecule may singlestranded), but may contain both the sense and anti-sense strands (i.e., the molecule may be doublestranded).

[0067] The term “peptide”, “polypeptide” and protein” are used interchangeably herein unless otherwise distinguished to refer to polymers of amino acids of any length. These terms also include proteins that are post-translationally modified through reactions that include glycosylation, acetylation and phosphorylation.

[0068] The term "linked" in the context of polypeptide sequences includes a linkage introduced through recombinant means or chemical means.

[0069] The terms "effective amount" or "amount effective to" or "therapeutically effective amount" refers to an amount sufficient to induce a detectable therapeutic response in the subject. Assays for determining therapeutic responses are well known in the art.

[0070] The terms "patient" or "subject" are used interchangeably and refer to a mammalian subject to be treated, for instance, a human patient. In some cases, the methods of the invention find use in experimental animals, in veterinary application, and in the development of animal models for disease, including, but not limited to, rodents including mice, rats, and hamsters, and primates.

[0071] As used herein, "administering" or "delivering" a molecule or treatment to a cell (e.g., a molecule such as a linear or circular nucleic acid optionally in a delivery vehicle) includes contacting the molecule with the cell, e.g., by mixing, fusing, transducing, transfecting, microinjecting, electroporating, or shooting. For instance, for in vivo delivery, a molecule may be delivered via a device such as a catheter, canula or needle. The term “antibody,” as used herein, refers to a full-length immunoglobulin molecule or an immunologically active fragment of an immunoglobulin molecule such as the Fab or F(ab’)2 fragment generated by, for example, cleavage of the antibody with an enzyme such as pepsin or co-expression of an antibody light chain and an antibody heavy chain in, for example, a mammalian cell, or ScFv. The antibody can also be an IgG, IgD, IgA, IgE or IgM antibody. Full- length immunoglobulin "light chains" (about 25 kD or 214 amino acids) are encoded by a variable region gene at the amino-terminus (about 1 10 amino acids) and a kappa or lambda constant region gene at the carboxy-terminus. Full-length immunoglobulin "heavy chains" (about 50 kD or 446 amino acids), are similarly encoded by a variable region gene (about 116 amino acids) and one of the other aforementioned constant region genes, e.g., gamma (encoding about 330 amino acids). Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which in turn define the immunoglobulin classes, IgG, IgM, IgA, IgD and IgE, respectively. An exemplary immunoglobulin (antibody) structural unit comprises a tetramer. Each tetramer is composed of two identical pairs of polypeptide chains, each pair having one "light" (about 25 kD) and one "heavy" chain (about 50-70 kD). The N-terminus of each chain defines a variable region of about 100 to 110 or more amino acids primarily responsible for antigen recognition. The terms variable light chain (VL) and variable heavy chain (VH) refer to these light and heavy chains respectively. In each pair of the tetramer, the light and heavy chain variable regions are together responsible for binding to an antigen, and the constant regions are responsible for the antibody effector functions. In addition to naturally occurring antibodies, immunoglobulins may exist in a variety of other forms including, for example, Fv, ScFv, Fab, and F(ab')2, as well as bifunctional hybrid antibodies (e.g., Lanzavecchia et al. (1987)) and in single chains (e.g., Huston et al. (1988) and Bird et al. (1988), which are incorporated herein by reference). (See, generally, Hood et al., "Immunology", Benjamin, N.Y., 2nded. (1984), and Hunkapiller and Hood (1986), which are incorporated herein by reference). Thus, the term "antibody" includes antigen binding antibody fragments, as are known in the art, including Fab, Fab2, single chain antibodies (scFv for example), chimeric antibodies, etc., either produced by the modification of whole antibodies or those synthesized de novo using recombinant DNA technologies.

[0072] An immunoglobulin light or heavy chain variable region consists of a "framework" region interrupted by three hypervariable regions, also called CDR's. The extent of the framework region and CDR's have been precisely defined (see, "Sequences of Proteins of Immunological Interest," E. Kabat et al., U.S. Department of Health and Human Services, (1983); which is incorporated herein by reference). The sequences of the framework regions of different light or heavy chains are relatively conserved within a species. As used herein, a "human framework region" is a framework region that is substantially identical (about 85% or more, usually 90 to 95% or more) to the framework region of a naturally occurring human immunoglobulin. The framework region of an antibody, that is the combined framework regions of the constituent light and heavy chains, serves to position and align the CDR's. The CDR’s are primarily responsible for binding to an epitope of an antigen.

[0073] Chimeric antibodies are antibodies whose light and heavy chain genes have been constructed, typically by genetic engineering, from immunoglobulin variable and constant region genes belonging to different species. For example, the variable segments of the genes from a mouse monoclonal antibody may be joined to human constant segments, such as gamma 1 and gamma 3. One example of a chimeric antibody is one composed of the variable or antigen-binding domain from a mouse antibody and the constant or effector domain from a human antibody, although other mammalian species may be used.

[0074] As used herein, the term "humanized" immunoglobulin refers to an immunoglobulin having a human framework region and one or more CDR's from a non-human (usually a mouse or rat) immunoglobulin. The non-human immunoglobulin providing the CDR's is called the "donor" and the human immunoglobulin providing the framework is called the "acceptor." Constant regions need not be present, but if they are, they are generally substantially identical to human immunoglobulin constant regions, i.e., at least about 85-90%, or about 95% or more identical. Hence, all parts of a humanized immunoglobulin, except possibly the CDR's, are substantially identical to corresponding parts of natural human immunoglobulin sequences. A "humanized antibody" is an antibody comprising a humanized light chain and a humanized heavy chain immunoglobulin. One says that the donor antibody has been "humanized", by the process of "humanization", because the resultant humanized antibody is expected to bind to the same antigen as the donor antibody that provides the CDR's.

[0075] Thus, humanized forms of non-human (e.g., murine) antibodies are chimeric immunoglobulins, immunoglobulin chains or fragments thereof (such as Fv, Fab, Fab', F(ab")2 or other antigen-binding subsequences of antibodies) which contain minimal sequence derived from non-human immunoglobulin. Humanized antibodies include human immunoglobulins (recipient antibody) in which residues from a complementary determining region (CDR) of the recipient are replaced by residues from a CDR of a non-human species (donor antibody) such as mouse, rat or rabbit having the desired specificity, affinity and capacity. In some instances, Fv framework residues of the human immunoglobulin are replaced by corresponding non-human residues. Humanized antibodies may also comprise residues which are found neither in the recipient antibody nor in the imported CDR or framework sequences. In general, the humanized antibody has substantially all of at least one, and typically two, variable domains, in which all or substantially all of the CDR regions correspond to those of a non-human immunoglobulin and all or substantially all of the framework regions are those of a human immunoglobulin consensus sequence. The humanized antibody optimally also will include at least a portion of an immunoglobulin constant region (Fc), typically that of a human immunoglobulin (Jones et al. (1986); Riechmann et al. (1988); and Presta (1992)). ft is understood that the humanized antibodies may have additional conservative amino acid substitutions which have substantially no effect on antigen binding or other immunoglobulin functions. By conservative substitutions are intended combinations such as gly, ala; val, ile, leu; asp, glu; asn, gin; ser, thr; lys, arg; and phe, tyr.

[0076] Humanized immunoglobulins, including humanized antibodies, have been constructed by means of genetic engineering. Methods for humanizing non-human antibodies are well known in the art. Generally, a humanized antibody has one or more amino acid residues introduced into it from a source which is non-human. These non-human amino acid residues are often referred to as "import" residues, which are typically taken from an "import" variable domain. Humanization can be essentially performed following the method of Winter and co-workers (Jones et al., Nature, 321:522 (1986); Riechmann et al., Nature, 332:323 (1988); Verhoeyen et al., Science, 239:1534 (1988)), by substituting rodent CDRs or CDR sequences for the corresponding sequences of a human antibody. Accordingly, such "humanized" antibodies are chimeric antibodies that have substantially less than an intact human variable domain has been substituted by the corresponding sequence from a non-human species. In practice, humanized antibodies are typically human antibodies in which some CDR residues and possibly some framework residues are substituted by residues from analogous sites in rodent antibodies.

[0077] Human antibodies can also be produced using various techniques known in the art, including phage display libraries (Hoogenboom and Winter, J. Mol. Biol., 227:381 (1991); Marks et al., J. Mol. Biol., 222:581 (1991)). The techniques of Cole et al. and Boerner et al. are also available for the preparation of human monoclonal antibodies (Cole et al., Monoclonal Antibodies and Cancer Therapy, Alan R. Liss, p. 77 (1985) and Boerner et al., J. Immunol., 147:86 (1991)). Similarly, human antibodies can be made by introducing of human immunoglobulin loci into transgenic animals, e.g., mice in which the endogenous immunoglobulin genes have been partially or completely inactivated. Upon challenge, human antibody production is observed, which closely resembles that seen in humans in all respects, including gene rearrangement, assembly, and antibody repertoire. This approach is described, for example, in U.S. Patent Nos. 5,545,807; 5,545,806; 5,569,825; 5,625,126; 5,633,425; 5,661,016, and in the following scientific publications: Marks et al., Bio / Technology 10:779 (1992); Lonberg et al., Nature, 368:856 (1994); Morrison, Nature, 368:812 (1994); Fishwild et al., Nature Biotechnology, 14:845 (1996); Neuberger, Nature Biotechnology, 14:826 (1996); Lonberg and Huszar, Intern. Rev. Immunol., 13:65 (1995). Most humanized immunoglobulins that have been previously described have a framework that is identical to the framework of a particular human immunoglobulin chain and three CDR's from a non-human donor immunoglobulin chain.

[0078] A framework may be one from a particular human immunoglobulin that is unusually homologous to the donor immunoglobulin to be humanized, or a consensus framework derived from many human antibodies. For example, comparison of the sequence of a mouse heavy (or light) chain variable region against human heavy (or light) variable regions in a data bank (for example, the National Biomedical Research Foundation Protein Identification Resource) shows that the extent of homology to different human regions varies greatly, typically from about 40% to about 60-70%. By choosing one of the human heavy (respectively light) chain variable regions that is most homologous to the heavy (respectively light) chain variable region of the other immunoglobulin, fewer amino acids will be changed in going from the one immunoglobulin to the humanized immunoglobulin. The precise overall shape of a humanized antibody having the humanized immunoglobulin chain may more closely resemble the shape of the donor antibody, also reducing the chance of distorting the CDR’s.

[0079] Typically, one of the 3-5 most homologous heavy chain variable region sequences in a representative collection of at least about 10 to 20 distinct human heavy chains is chosen as acceptor to provide the heavy chain framework, and similarly for the light chain. One of the 1 to 3 most homologous variable regions may be used. The selected acceptor immunoglobulin chain may have at least about 65% homology in the framework region to the donor immunoglobulin.

[0080] In many cases, it may be considered desirable to use light and heavy chains from the same human antibody as acceptor sequences, to be sure the humanized light and heavy chains will make favorable contacts with each other. Regardless of how the acceptor immunoglobulin is chosen, higher affinity may be achieved by selecting a small number of amino acids in the framework of the humanized immunoglobulin chain to be the same as the amino acids at those positions in the donor rather than in the acceptor.

[0081] Humanized antibodies generally have advantages over mouse or in some cases chimeric antibodies for use in human therapy: because the effector portion is human, it may interact better with the other parts of the human immune system (e.g., destroy the target cells more efficiently by complement-dependent cytotoxicity (CDC) or antibody-dependent cellular cytotoxicity (ADCC)); the human immune system should not recognize the framework or constant region of the humanized antibody as foreign, and therefore the antibody response against such an antibody should be less than against a totally foreign mouse antibody or a partially foreign chimeric antibody.

[0082] DNA segments having immunoglobulin sequences typically further include an expression control DNA sequence operably linked to the humanized immunoglobulin coding sequences, including naturally associated or heterologous promoter regions. Generally, the expression control sequences will be eukaryotic promoter systems in vectors capable of transforming or transfecting eukaryotic host cells, but control sequences for prokaryotic hosts may also be used. Once the vector has been incorporated into the appropriate host, the host is maintained under conditions suitable for high level expression of the nucleotide sequences, and, as desired, the collection and purification of the humanized light chains, heavy chains, light / heavy chain dimers or intact antibodies, binding fragments or other immunoglobulin forms may follow (see, S. Beychok, Cells of Immunoglobulin Synthesis, Academic Press, New York, (1979), which is incorporated herein by reference).

[0083] Other "substantially homologous" modified immunoglobulins to the native sequences can be readily designed and manufactured utilizing various recombinant DNA techniques well known to those skilled in the art. For example, the framework regions can vary at the primary structure level by several amino acid substitutions, terminal and intermediate additions and deletions, and the like. Moreover, a variety of different human framework regions may be used singly or in combination as a basis for the humanized immunoglobulins of the present invention. In general, modifications of the genes may be readily accomplished by a variety of well-known techniques, such as site-directed mutagenesis (see, Gillman and Smith, Gene, 8:81 (1979) and Roberts et al., Nature, 328:731 (1987), both of which are incorporated herein by reference). Substantially homologous immunoglobulin sequences are those which exhibit at least about 85% homology, usually at least about 90%, or at least about 95% homology with a reference immunoglobulin protein.

[0084] Alternatively, polypeptide fragments comprising only a portion of the primary antibody structure may be produced, which fragments possess one or more immunoglobulin activities (e.g., antigen binding). These polypeptide fragments may be produced by proteolytic cleavage of intact antibodies by methods well known in the art, or by inserting stop codons at the desired locations in vectors known to those skilled in the art, using site-directed mutagenesis.

[0085] As used herein, the term “binds specifically” or “specifically binds,” in reference to an antibody / antigen interaction, means that the antibody binds with a particular antigen without substantially binding to other distinct antigens or to unrelated antigens. The term "peptide" when used with reference to a linker, describes a sequence of 2 to 25 amino acids (e.g., as defined hereinabove) or peptidyl residues. The sequence may be linear or cyclic. For example, a cyclic peptide can be prepared or may result from the formation of disulfide bridges between two cysteine residues in a sequence. A peptide can be linked to another molecule through the carboxy terminus, the amino terminus, or through any other convenient point of attachment, such as, for example, through the sulfur of a cysteine. In one embodiment, a peptide linker may comprise 3 to 25, 5 to 21 , or 5 to 15, or any integer in between, amino acids. Exemplary System for Targeted Gene Delivery

[0086] Viral vectors are a major means of gene delivery with the potential to impact a number of pediatric diseases including inherited genetic disorders and cancer. Naturally evolved properties of many viral vectors are, however, mismatched to clinical delivery needs. In gene therapy, for example, cell type specificity is paramount, as ectopic expression in off-target tissues or cells is undesirable and poses a safety risk. We propose to remove these legacy constraints of natural evolution by functionally separating viral entry (host recognition) and viral replication (gene delivery).

[0087] This is achieved by removing endogenous tropism of the virus, and by introducing into the virus capsid small protein endonuclease domains that covalently bind single strand DNA (ssDNA) in a sequence specific and nonoverlapping (orthogonal) fashion. By linking different targeting molecules, e.g., monoclonal antibodies (mAB), with ssDNA substrates, virus becomes decorated with targeting molecules, e.g., antibody to form covalent mAB / AAV composites (Figure 1A) or by inserting small binding proteins, such as nanobodies, DARPins or GP2 scaffoled (proteins that are 5-10 or 15 kd).

[0088] In one embodiment, to achieve programmable cell type-specific viral gene delivery, at least two components are combined. In one embodiment, as in ADC and CAR-T, antibodies are used to recognize surface markers of a targeted cell type. In one embodiment, as in viral approaches to gene therapy, AAV is employed as a viral vector that is amenable to facile capsid engineering. In one embodiment, antibodies and AAV are combined via a covalently link using, for example, HUH endonuclease domains. HUH domains are small (10-20 kDa) and robustly form covalent bonds with short single strand (ss) DNA in a sequence specific fashion. HUH domains are orthogonal and multiplexable; there is no overlap between sequences that each HUH domain recognizes and binds to. By displaying different HUH domains on different viruses, e.g., different AAVs, delivering different payloads, different ssDNA-conjugated targeting molecules such as antibodies recognizing different cell surface markers are covalently attached to specific viral particles in a programmable fashion. In one embodiment, the present disclosure achieves a separation of function and roles: cell targeting is mediated by targeting molecules, e.g., antibodies, and gene delivery mediated by the virus. By separating these roles into two independently engineerable components, and then combining them as needed, complex interdependency issues are avoided that have hampered previous efforts in viral vector engineering. A functional separation of components also allows rapid target switching without labor intensive repackaging of virus. Thus, a set of template viruses (packaging a battery of therapeutic / cytotoxic payloads) can be prepared at large scale, biobanked, and 'armed' with a target-directed antibody ad hoc.

[0089] In one embodiment, a chimeric AAV that incorporates a ssDNA-binding HUH domain was prepared. The HUH tag is functionally displayed in this context and incorporated into infectious, genome-containing capsids. Without attachment of a nanobody conjugated to the cognate ssDNA of the HUH, the chimeric virus is not infectious. With attachment of an anti-GFP nanobody via the HUH tag, e.g., formation of the mAB / AAV composite, the virus becomes 'armed' and is able to infect cells expressing a surface displayed GFP.

[0090] A chimeric AAV that incorporates a ssDNA-binding HUH domain was prepared. The HUH tag is functionally displayed in this context and incorporated into infectious, genomecontaining capsids. Without attachment of a nanobody conjugated to the cognate ssDNA of the HUH, the chimeric virus is not infectious. With attachment of an anti-GFP nanobody via the HUH tag, i.e., formation of the mAB / AAV composite, the virus becomes 'armed' and is able to infect cells expressing a surface displayed GFP.

[0091] In another embodiment, a chimeric AAV that incorporates one or more small binding proteins was prepared.

[0092] The approach enables simultaneous gene delivery to different cell types with minimal crosstalk or off-target effects. This allows for new kinds of gene therapy approaches that could significantly impact the care of children with genetic disorders or cancer.

[0093] Exemplary Capsid Regions for Modification

[0094] In one embodiment, the VP is an AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAVrhlO, or AAV-9 VP. In one embodiment, regions for modification, e.g., mutation, deletion and / or insertion, in an AAV capsid include positions reside 262, 328, 387, 631, 651, 664, 670, 684 or 708.

[0095] Exemplary Pairs of HUH Domains and ssDNA Substrates Thereof

[0096] Exemplary HUH domains may be obtained from PCV2 (SEQ ID NO:2), phiX174 (SEQ ID NO:3), mMobA (SEQ ID NO:4),TraI36 (SEQ ID NO:5), RepB (SEQ ID NO:6), FBNYV (SEQ ID NO:7), NES (SEQ ID NO:8), TrwC (SEQ ID NO:9), TYLCV (SEQ ID NQ:10), RepBm (SEQ ID NO:20), or DCV (SEQ ID NO:21). Moreover, fragments of those sequences may be employed in viral capsids so long as the ssDNA binding and binding specificity is not substantially altered.

[0097] SEQ ID NO: 2 is SPSKKNGRSG PQPHKRWVFT LNNPSEDERK KIRDLPISLF DYFIVGEEGN EEGRTPHLQG FANFVKKQTF NKVKWYLGAR CHIEKAKGTD QQNKEYCSKE GNLLMEEGAP RSQGQR.

[0098] SEQ ID NO: 3 is KSRRGFAIQR LMNAMRQAHA DGWFIVFDTL TLADDRLEAF YDNPNALRDY FRDIGRMVLA AEGRKANDSH ADCYQYFCVP EYGTANGRLH FHAVHFMRTL PTGSVDPNFG RRVRNRRQLN SLQNTWPYGH SMPIAVRYTQ DAFSRSGWLW PVDAKGEPLK ATSYMAVGFY VAKYVNKKSD MDLAAKGLGA KEWNNSLKTK LSLLPKKLFR IRMSRNFGMK MLTMTNLSTE CLIQLTKLGY DATPFNQILK QNAKREMRLR LGKVTVADVL AAQPVTTNLL KFMRASIKMI GVSNLQSFIA SMTQKLTLSD ISDESKNYLD KAGITTACLR IKSKWTAGGK.

[0099] SEQ ID NO: 4 is MAIYHLTAKT GSRSGGQSAR AKADYIQREG KYARDMDEVL HAESGHMPEF VERPADYWDA ADLYERANGR LFKEVEFALP VELTLDQQKA LASEFAQHLT GAERLPYTLA IHAGGGENPH CHLMISERIN DG1ERPAAQW FKRYNGKTPE KGGAQKTEAL KPKAWLEQTR EAWADHANRA LERAGH.

[0100] SEQ ID NO: 5 is MMSIAQVRSA GSAGNYYTDK DNYYVLGSMG ERWAGRGAEQ LGLQGSVDKD VFTRLLEGRL PDGADLSRMQ DGSNRHRPGY DLTFSAPKSV SMMAMLGGDK RLIDAHNQAV DFAVRQVEAL ASTRVMTDGQ SETVLTGNLV MALFNHDTSR DQEPQLHTHA VVANVTQHNG EWKTLSSDKV GKTGFIENVY ANQIAFGRLY REKLKEQVEA LGYETEVVGK HGMWEMPGVP VEAFSGRSQT IREAVGEDAS LKSRDVAALD TRKSKQHVDP EIKMAEWMQT LKETGFDIRA YRDAADQRAD LRTLTPGPAS QDGPDVQQAV TQAIAGLSER.

[0101] SEQ ID NO: 6 is MAKEKARYFT FLLYPESIPS DWELKLETLG VPMAISPLHD KDKSSIKGQK YKKAHYHVLY IAKNPVTADS VRKKIKLLLG EKSLAMVQVV LNVENMYLYL THESKDAIAK KKHVYDKADI KLINNFDIDR YLE FBNYV.

[0102] SEQ ID NO: 7is MARQVICWCF TLNNPLSPLS LHDSMKYLVY QTEQGEAGNI HFQGYIEMKK RTSLAGMKKL IPGAHFEKRR GTQGEARAYS MKEDTRLEGP WEYGEFVP NES.

[0103] SEQ ID NO: 8 is AMYHFQNKFV SKANGQSATA KSAYNSASRI KDFKENEFKD YSNKQCDYSE ILLPNNADDK FKDREYLWNK VHDVENRKNS QVAREIIIGL PNEFDPNSNI ELAKEFAESL SNEGMIVDLN IHKINEENPH AHLLCTLRGL DKNNEFEPKR KGNDYIRDWN TKEKHNEWRK RWENVQNKHL EKNGFSVRVS ADSYKNQNID LEPTKKEGWK ARKFEDETG. SEQ ID NO: 9 is MLSHMVLTRQ DIGRAASYYE DGADDYYAKD GDASEWQGKG AEELGLSGEV DSKRFRELLA GNIGEGHRIM RSATRQDSKE RIGLDLTFSA PKSVSLQALV AGDAEIIKAH DRAVARTLEQ AEARAQARQK IQGKTRIETT GNLVIGKFRH ETSRERDPQL HTHAVILNMT KRSDGQWRAL KNDEIVKATR YLGAVYNAEL AHELQKLGYQ LRYGKDGNFD LAHIDRQQIE GFSKRTEQIA EWYAARGLDP NSVSLEQKQA AKVLSRAKKT SVDREALRAE WQATAKELGI DFS TLYCV.

[0104] SEQ ID NO: 10 is MPRLFKIYAK NYFLTYPNCS LSKEEALSQL KKLETPTNKK YIKVCKELHE NGEPHLHVLI QFEGKYQCKN QRFFDLVSPN RSAHFHPNIQ AAKSSTDVKT YVEKDGNFID FGVSQIDGRS.

[0105] SEQ ID NO: 20 is MSEKKEIVKG RDWTFLVYPE SAPENWRTIL DETFMRWVES PLHDKDVNAD GEIKKPHWHI LLSSDGPITQ TAVQKIIGPL NCPNAQKVGS AKGLVRYMVH LDNPEKYQYS LDEIVGHNGA DVASYFELTA.

[0106] SEQ ID NO: 21 is MAKSGNYSYK RWVFTINNPT FEDYVHVLEF CTLDNCKFAI VGEEKGANGT PHLQGFLNLR SNARAAALEE SLGGRAWLSR ARGSDEDNEE YCAKESTYLR VGEPVSKGRS S.

[0107] In one embodiment, the HUH domain has at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, or 99% amino acid sequence identity to one of SEQ ID Nos. 2-10 and 20-21.

[0108] In one embodiment, the HUH domain is a fragment of one of SEQ ID Nos. 2-10 or 20-21, e.g., one having a deletion of 1, 2, 3, 4, 5, or more, e.g., 10, 15, 20, 25, 30, 35, 40, 45 or 50 residues, that has at least 80%, 85%, 90%, 95%, 98%, 99% or 100% the activity of SEQ ID Nos. 2-10 or 20-21.

[0109] In one embodiment, the HUH substrate has at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, or 99% nucleic acid sequence identity to one of SEQ ID Nos. 1-4 or 11. In one embodiment, the HUH substrate is a fragment of one of SEQ ID Nos. 1-4 or 11, e.g., one having a deletion of 1, 2, 3, 4, 5, or more, e.g., 10, 15, 20, 25, 30, 35, 40, 45 or 50 residues, that has at least 80%, 85%, 90%, 95%, 98%, 99% or 100% the activity of SEQ ID Nos.1-4 or 11.

[0110] Exemplary Linkers

[0111] Exemplary linkers include but are not limited to (GGGGS)3(SEQ ID NO: 22), (Gly)s (SEQ ID NO: 23), (Gly)6(SEQ ID NO: 24), (EAAAK)3(SEQ ID NO: 25), (EAAAK)„ (n=l-3) (SEQ ID NO: 26), A(EAAAK)4ALEA(EAAAK)4A (SEQ ID NO: 27), A(EAAAK)4ALEA(EAAAK)4A (SEQ ID NO: 28), (GGGGS)3(SEQ ID NO: 29), A(EAAAK)4ALEA(EAAAK)4A (SEQ ID NO: 30), A(EAAAK)4ALEA(EAAAK)4A (SEQ ID NO: 21), GGGGS (SEQ ID NO: 32), PAPAP (SEQ ID NO: 33), AEAAAKEAAAKA (SEQ ID NO: 34), (GGGGS)n (n=l, 2, 4) (SEQ ID NO: 35), (Ala-Pro)n (10 - 34 aa) (SEQ ID NO: 36), VSQTSKLTR1AETVFPDV (SEQ ID NO: 37), PEG LWA (SEQ ID NO: 38), RVLJ.AEA (SEQ ID NO: 39); EDVVCCjSMSY (SEQ ID NO: 40); GGIEGRiGSc(SEQ ID NO: 41), TRHRQPRJ.GWE (SEQ ID NO: 42); AGNRVRR^SVG (SEQ ID NO: 43; RRRRRRR;R|R (SEQ ID NO: 44), GFLGJ, (SEQ ID NO: 45), A(EAAAK)4ALEA(EAAAK)4A (SEQ ID NO: 46), (GlyGlyGlyGlySerh- GGGGSGGGGS (SEQ ID NO: 47), LE, or GGSSGGGSGG (SEQ ID NO: 48).

[0112] Exemplary Targeting Molecules

[0113] In one embodiment, the target molecule binding protein is an antibody or a portion thereof, e.g., a scFV or a single domain antibody (sdAb) that is based on the recombinant variable heavy domains from the heavy chain only antibodies found in Camelids and sharks. Other binding proteins include intrabodies and nanobodies.

[0114] The targeting molecules may be specific for any selected antigen, e.g., any cell surface marker including any post-translational modificiation (CD52, HER2, CD19, CD20, CD30, CD3, 0X40, VEGFR, EGER, CD205, CD64, CD4, CD8), any membrane protein including any post- translational modificiation (Navi.8, Clc3, NMDA receptors, DI & D2 muscarinic receptors, ASIC1 / 3), any synthetic or otherwise engineered proteins targeted to the cell surface (GFP-GPI, GFP-PDGR, mGRASP), and any glycosylated lipids (cerebroside, galactoside, lipopolysaccharide).

[0115] Exemplary Genes for Delivery

[0116] In one embodiment, the gene product is a therapeutic gene product, e.g., GM-CSF, CD40L, IL-2, CD80, MDA-7, or TNF-alpha. In one embodiment, the gene product is a prophylactic gene product, e.g. kallikrein or is pathogen-derived protein fragments (Pestisvirus E2, C.psittaci MOMP, S.aureus FnBP, S. aureus ClfA, Avian paramyxovirus HN, B. melitensis OMP-31, .S'. japonicum Sj23). In one embodiment, the gene product is a catalytic RNA. In one embodiment the gene product is a guide RNA. In one embodiment the gene is a donor for homologous recombination. In one embodiment the gene product is a nuclease suitable for genome editing, e.g. TALENS, 5. pyogenes Cas9, S.aureus Cas9. In one embodiment, the gene product is a cytotoxic gene product, e.g., suicide genes such as rexin-G or HSVtk in combination with ganciclovir, an apoptosis inducer, e.g., p53, p27Kipl, p21Wafl, p!6INK4A, Ad5IkB, or cyclin- dependent kinase inhibitors, or an angiogenesis inhibitor. In one embodiment, the gene product is a chimeric T-cell receptor (anti-CD19 scfv / CD28 / CD3^ CAR, anti-BCMA scfv / 4-1 BB / CD3^ CAR). In one embodiment, the gene to be delivered includes but is not limited to cystic fibrosis transmembrane conductance regulator, a-antitrypsin, p-globin, y-globin, tyrosine hydroxylase, glucocerebrosidase, aryl sulfatase A, factor VIII, dystrophin or erythropoietin, a viral, bacterial, tumor or fungal antigen, or an immune response modulator, e.g., a cytokine including but not limited to IFN-alpha, IFN-gamma, TNF, IL-1, IL-17, or IL-6.

[0117] Exemplary rAAV Genomes

[0118] An AAV vector typically comprises a polynucleotide that is heterologous to AAV. The polynucleotide is typically of interest because of a capacity to provide a function to a target cell in the context of gene therapy, such as up- or down-regulation of the expression of a certain phenotype. Such a heterologous polynucleotide or “transgene,” generally is of sufficient length to provide the desired function or encoding sequence.

[0119] Where transcription of the heterologous polynucleotide is desired in the intended target cell, it can be operably linked to its own or to a heterologous promoter, depending for example on the desired level and / or specificity of transcription within the target cell, as is known in the art. Various types of promoters and enhancers are suitable for use in this context. Constitutive promoters provide an ongoing level of gene transcription and may be preferred when it is desired that the therapeutic or prophylactic polynucleotide be expressed on an ongoing basis. Inducible promoters generally exhibit low activity in the absence of the inducer and are up-regulated in the presence of the inducer. They may be preferred when expression is desired only at certain times or at certain locations, or when it is desirable to titrate the level of expression using an inducing agent. Promoters and enhancers may also be tissue-specific: that is, they exhibit their activity only in certain cell types, presumably due to gene regulatory elements found uniquely in those cells.

[0120] Illustrative examples of promoters are the SV40 late promoter from simian virus 40, the Baculovirus polyhedron enhancer / promoter element, Herpes Simplex Virus thymidine kinase (HSV tk), the immediate early promoter from cytomegalovirus (CMV) and various retroviral promoters including LTR elements. Inducible promoters include heavy metal ion inducible promoters (such as the mouse mammary tumor virus (mMTV) promoter or various growth hormone promoters), and the promoters from T7 phage which are active in the presence of T7 RNA polymerase. By way of illustration, examples of tissue-specific promoters include various surfactin promoters (for expression in the lung), myosin promoters (for expression in muscle), and albumin promoters (for expression in the liver). A large variety of other promoters are known and generally available in the art, and the sequences of many such promoters are available in sequence databases such as the GenBank database. Where translation is also desired in the intended target cell, the heterologous polynucleotide will preferably also comprise control elements that facilitate translation (such as a ribosome binding site or “RBS” and a polyadenylation signal). Accordingly, the heterologous polynucleotide generally comprises at least one coding region operatively linked to a suitable promoter, and may also comprise, for example, an operatively linked enhancer, ribosome binding site and poly-A signal. The heterologous polynucleotide may comprise one encoding region, or more than one encoding regions under the control of the same or different promoters. The entire unit, containing a combination of control elements and encoding region, is often referred to as an expression cassette.

[0121] The heterologous polynucleotide is integrated by recombinant techniques into or in place of the AAV genomic coding region (i.e., in place of the AAV rep and cap genes) but is generally flanked on either side by AAV inverted terminal repeat (ITR) regions. This means that an ITR appears both upstream and downstream from the coding sequence, either in direct juxtaposition, e.g., (although not necessarily) without any intervening sequence of AAV origin in order to reduce the likelihood of recombination that might regenerate a replication-competent AAV genome. However, a single ITR may be sufficient to carry out the functions normally associated with configurations comprising two ITRs (see, for example, WO 94 / 13788), and vector constructs with only one ITR can thus be employed in conjunction with the packaging and production methods of the present invention.

[0122] The native promoters for rep are self-regulating and can limit the amount of AAV particles produced. The rep gene can also be operably linked to a heterologous promoter, whether rep is provided as part of the vector construct, or separately. Any heterologous promoter that is not strongly down-regulated by rep gene expression is suitable; but inducible promoters may be preferred because constitutive expression of the rep gene can have a negative impact on the host cell. A large variety of inducible promoters are known in the art; including, by way of illustration, heavy metal ion inducible promoters (such as metallothionein promoters); steroid hormone inducible promoters (such as the MMTV promoter or growth hormone promoters); and promoters such as those from T7 phage which are active in the presence of T7 RNA polymerase. One subclass of inducible promoters are those that are induced by the helper virus that is used to complement the replication and packaging of the rAAV vector. A number of helper-virus- inducible promoters have also been described, including the adenovirus early gene promoter which is inducible by adenovirus El A protein; the adenovirus major late promoter; the herpesvirus promoter which is inducible by herpesvirus proteins such as VP 16 or 1CP4; as well as vaccinia or poxvirus inducible promoters. Methods for identifying and testing helper-virus-inducible promoters have been described (see, e.g., WO 96 / 17947). Thus, methods are known in the art to determine whether or not candidate promoters are helper-virus-inducible, and whether or not they will be useful in the generation of high efficiency packaging cells. Briefly, one such method involves replacing the p5 promoter of the AAV rep gene with the putative helper-virus-inducible promoter (either known in the art or identified using well-known techniques such as linkage to promoter-less “reporter” genes). The AAV rep-cap genes (with p5 replaced), e.g., linked to a positive selectable marker such as an antibiotic resistance gene, are then stably integrated into a suitable host cell (such as the HeLa or A549 cells exemplified below). Cells that are able to grow relatively well under selection conditions (e.g., in the presence of the antibiotic) are then tested for their ability to express the rep and cap genes upon addition of a helper virus. As an initial test for rep and / or cap expression, cells can be readily screened using immunofluorescence to detect Rep and / or Cap proteins. Confirmation of packaging capabilities and efficiencies can then be determined by functional tests for replication and packaging of incoming rAAV vectors. Using this methodology, a helper-virus-inducible promoter derived from the mouse metallothionein gene has been identified as a suitable replacement for the p5 promoter and used for producing high titers of rAAV particles (as described in WO 96 / 17947).

[0123] Removal of one or more AAV genes is in any case desirable, to reduce the likelihood of generating replication competent AAV (“RCA”). Accordingly, encoding or promoter sequences for rep, cap, or both, may be removed, since the functions provided by these genes can be provided in trans.

[0124] The resultant vector is referred to as being “defective” in these functions. In order to replicate and package the vector, the missing functions are complemented with a packaging gene, or a plurality thereof, which together encode the necessary functions for the various missing rep and / or cap gene products. The packaging genes or gene cassettes are in one embodiment not flanked by AAV ITRs and in one embodiment do not share any substantial homology with the rAAV genome. Thus, in order to minimize homologous recombination during replication between the vector sequence and separately provided packaging genes, it is desirable to avoid overlap of the two polynucleotide sequences. The level of homology and corresponding frequency of recombination increase with increasing length of homologous sequences and with their level of shared identity. The level of homology that will pose a concern in a given system can be determined theoretically and confirmed experimentally, as is known in the art. Typically, however, recombination can be substantially reduced or eliminated if the overlapping sequence is less than about a 25 nucleotide sequence if it is at least 80% identical over its entire length, or less than about a 50 nucleotide sequence if it is at least 70% identical over its entire length. Of course, even lower levels of homology are preferable since they will further reduce the likelihood of recombination. It appears that, even without any overlapping homology, there is some residual frequency of generating RCA. Even further reductions in the frequency of generating RCA (e.g., by nonhomologous recombination) can be obtained by “splitting” the replication and encapsidation functions of AAV, as described by Allen et al., WO 98 / 27204).

[0125] The rAAV vector construct, and the complementary packaging gene constructs can be implemented in this invention in a number of different forms. Viral particles, plasmids, and stably transformed host cells can all be used to introduce such constructs into the packaging cell, either transiently or stably.

[0126] In certain embodiments of this invention, the AAV vector and complementary packaging gene(s), if any, are provided in the form of bacterial plasmids, AAV particles, or any combination thereof. In other embodiments, either the AAV vector sequence, the packaging gene(s), or both, are provided in the form of genetically altered (preferably inheritably altered) eukaryotic cells. The development of host cells inheritably altered to express the AAV vector sequence, AAV packaging genes, or both, provides an established source of the material that is expressed at a reliable level.

[0127] A variety of different genetically altered cells can thus be used in the context of this invention. By way of illustration, a mammalian host cell may be used with at least one intact copy of a stably integrated rAAV vector. An AAV packaging plasmid comprising at least an AAV rep gene operably linked to a promoter can be used to supply replication functions (as described in U.S. Patent 5,658,776). Alternatively, a stable mammalian cell line with an AAV rep gene operably linked to a promoter can be used to supply replication functions (see, e.g., Trempe et al., WO 95 / 13392); Burstein et al. (WO 98 / 23018); and Johnson et al. (U.S. No. 5,656,785). The AAV cap gene, providing the encapsidation proteins as described above, can be provided together with an AAV rep gene or separately (see, e.g., the above-referenced applications and patents as well as Allen et al. (WO 98 / 27204). Other combinations are possible and included within the scope of this invention.

[0128] Uses

[0129] The virus can be used for administration to an individual for purposes of gene therapy or vaccination. Suitable diseases for therapy include but are not limited to those induced by viral, bacterial, or parasitic infections, various malignancies and hyperproliferative conditions, autoimmune conditions, and congenital deficiencies.

[0130] Gene therapy can be conducted to enhance the level of expression of a particular protein either within or secreted by the cell. Vectors may be used to genetically alter cells either for gene marking, replacement of a missing or defective gene, or insertion of a therapeutic gene. Alternatively, a polynucleotide may be provided to the cell that decreases the level of expression. This may be used for the suppression of an undesirable phenotype, such as the product of a gene amplified or overexpressed during the course of a malignancy, or a gene introduced or overexpressed during the course of a microbial infection. Expression levels may be decreased by supplying a therapeutic or prophylactic polynucleotide comprising a sequence capable, for example, of forming a stable hybrid with either the target gene or RNA transcript (antisense therapy), capable of acting as a ribozyme to cleave the relevant mRNA or capable of acting as a decoy for a product of the target gene.

[0131] Vaccination can be conducted to protect cells from infection by infectious pathogens. As the traditional vaccine methods, vectors of this invention may be used to deliver transgenes encoding viral, bacterial, tumor or fungal antigen and their subsequent expression in host cells. The antigens, which expose to the immune system to evoke an immune response, can be in the form of virus-like particle vaccines or subunit vaccines of virus-coding proteins. Alternatively, as the method of passive immunization, vectors of this invention might be used to deliver genes encoding neutralizing antibodies and their subsequent expression in host non-hematopoietic tissues. The vaccine-like protection against pathogen infection can be conducted through direct provision of neutralizing antibody from vector-mediated transgene expression, bypassing the reliance on the natural immune system for mounting desired humoral immune responses.

[0132] The introduction of the virus to an animal may involve use of any number of delivery techniques (both surgical and non-surgical) which are available and well known in the art. Such delivery techniques, for example, include vascular catheterization, cannulization, injection, inhalation, endotracheal, subcutaneous, inunction, topical, oral, percutaneous, intra-arterial, intravenous, and / or intraperitoneal administrations. Vectors can also be introduced by way of bioprostheses, including, by way of illustration, vascular grafts (PTFE and dacron), heart valves, intravascular stents, intravascular paving as well as other non-vascular prostheses. General techniques regarding delivery, frequency, composition and dosage ranges of vector solutions are within the skill of the art.

[0133] In particular, for delivery of a vector of the invention to a tissue, any physical or biological method that will introduce the vector to a host animal can be employed. Vector means both a bare recombinant vector and vector DNA packaged into viral coat proteins, as is well known for administration. There are no known restrictions on the carriers or other components that can be co-administered with the vector (although compositions that degrade DNA should be avoided in the normal manner with vectors). Pharmaceutical compositions can be prepared as injectable formulations or as topical formulations to be delivered to the muscles by transdermal transport. Numerous formulations for both intramuscular injection and transdermal transport have been previously developed and can be used in the practice of the invention. The vectors can be used with any pharmaceutically acceptable carrier for ease of administration and handling.

[0134] For purposes of intramuscular injection, solutions in an adjuvant such as sesame or peanut oil or in aqueous propylene glycol can be employed, as well as sterile aqueous solutions. Such aqueous solutions can be buffered, if desired, and the liquid diluent first rendered isotonic with saline or glucose. Solutions of the virus as a free acid (DNA contains acidic phosphate groups) or a pharmacologically acceptable salt can be prepared in water suitably mixed with a surfactant such as hydroxypropylcellulose. A dispersion of viral particles can also be prepared in glycerol, liquid polyethylene glycols and mixtures thereof and in oils. Under ordinary conditions of storage and use, these preparations contain a preservative to prevent the growth of microorganisms. In this connection, the sterile aqueous media employed are all readily obtainable by standard techniques well-known to those skilled in the art.

[0135] The pharmaceutical forms suitable for injectable use include sterile aqueous solutions or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. In all cases the form must be sterile and must be fluid to the extent that easy syringability exists. It must be stable under the conditions of manufacture and storage and must be preserved against the contaminating action of microorganisms such as bacteria and fungi. The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, liquid polyethylene glycol and the like), suitable mixtures thereof, and vegetable oils. The proper fluidity can be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of a dispersion and by the use of surfactants. The prevention of the action of microorganisms can be brought about by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid, thimerosal and the like. In many cases it will be preferable to include isotonic agents, for example, sugars or sodium chloride. Prolonged absorption of the injectable compositions can be brought about by use of agents delaying absorption, for example, aluminum monostearate and gelatin.

[0136] Sterile injectable solutions are prepared by incorporating the virus in the required amount in the appropriate solvent with various of the other ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incorporating the sterilized active ingredient into a sterile vehicle which contains the basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the methods of preparation include but are not limited to vacuum drying and the freeze-drying technique which yield a powder of the active ingredient plus any additional desired ingredient from the previously sterile-filtered solution thereof.

[0137] For purposes of topical administration, dilute sterile, aqueous solutions (usually in about 0.1% to 5% concentration), otherwise similar to the above parenteral solutions, are prepared in containers suitable for incorporation into a transdermal patch, and can include known carriers, such as pharmaceutical grade dimethylsulfoxide (DMSO).

[0138] Compositions may be used in vivo as well as ex vivo. In vivo gene therapy comprises administering the vectors directly to a subject. Pharmaceutical compositions can be supplied as liquid solutions or suspensions, as emulsions, or as solid forms suitable for dissolution or suspension in liquid prior to use. For administration into the respiratory tract, one mode of administration is by aerosol, using a composition that provides either a solid or liquid aerosol when used with an appropriate aerosolubilizer device. Another mode of administration into the respiratory tract is using a flexible fiberoptic bronchoscope to instill the vectors. Typically, the viral vectors are in a pharmaceutically suitable pyrogen-free buffer such as Ringer’s balanced salt solution (pH 7.4). Although not required, pharmaceutical compositions may optionally be supplied in unit dosage form suitable for administration of a precise amount.

[0139] An effective amount of virus is administered, depending on the objectives of treatment. An effective amount may be given in single or divided doses. Where a low percentage of transduction can cure a genetic deficiency, then the objective of treatment is generally to meet or exceed this level of transduction. In some instances, this level of transduction can be achieved by transduction of only about 1 to 5% of the target cells, but is more typically 20% of the cells of the desired tissue type, usually at least about 50%, at least about 80%, at least about 95%, or at least about 99% of the cells of the desired tissue type. As a guide, the number of vector particles present in a single dose given by bronchoscopy will generally be at least about 1 x 1012, e.g., about 1 x 1013, 1 x 1014, 1 x 1015or 1 x 1016particles, including both DNAse-resistant and DNAse- susceptible particles. In terms of DNAse-resistant particles, the dose will generally be between 1 x 1012and 1 x 1016particles, more generally between about 1 x 1012and 1 x IO13particles. The treatment can be repeated as often as every two or three weeks, as required, although treatment once in 180 days may be sufficient.

[0140] The decision of whether to use in vivo or ex vivo therapy, and the selection of a particular composition, dose, and route of administration will depend on a number of different factors, including but not limited to features of the condition and the subject being treated. The assessment of such features and the design of an appropriate therapeutic or prophylactic regimen is ultimately the responsibility of the prescribing physician.

[0141] It is understood that variations may be applied to these methods by those of skill in this art without departing from the spirit of this invention.

[0142] Dosages, Formulations and Routes of Administration

[0143] Administration of the recombinant viruses may be continuous or intermittent, depending, for example, upon the recipient’s physiological condition, whether the purpose of the administration is therapeutic or prophylactic, and other factors known to skilled practitioners. The administration of the recombinant viruses may be essentially continuous over a preselected period of time or may be in a series of spaced doses. Both local and systemic administration is contemplated. When the recombinant viruses are employed for prophylactic purposes, recombinant viruses are amenable to chronic use, e.g., by systemic administration.

[0144] One or more suitable unit dosage forms comprising the recombinant viruses, can be administered by a variety of routes including oral, or parenteral, including by rectal, transdermal, subcutaneous, intravenous, intramuscular, intraperitoneal, intrathoracic, intrapulmonary and intranasal routes. For example, for administration to the liver, intravenous administration may be preferred. For administration to the lung, airway administration may be preferred. The formulations may, where appropriate, be conveniently presented in discrete unit dosage forms and may be prepared by any of the methods well known to pharmacy. Such methods may include the step of bringing into association the recombinant viruses with liquid carriers, solid matrices, semisolid carriers, finely divided solid carriers or combinations thereof, and then, if necessary, introducing or shaping the product into the desired delivery system.

[0145] When the recombinant viruses are prepared for oral administration, they may be combined with a pharmaceutically acceptable carrier, diluent or excipient to form a pharmaceutical formulation, or unit dosage form. The total active ingredients in such formulations comprise from 0.1 to 99.9% by weight of the formulation. By “pharmaceutically acceptable” it is meant the carrier, diluent, excipient, and / or salt must be compatible with the other ingredients of the formulation, and not deleterious to the recipient thereof. The active ingredient for oral administration may be present as a powder or as granules; as a solution, a suspension or an emulsion; or in achievable base such as a synthetic resin for ingestion of the active ingredients from a chewing gum. The active ingredient may also be presented as a bolus, electuary or paste.

[0146] Pharmaceutical formulations containing the recombinant viruses can be prepared by procedures known in the art using well known and readily available ingredients. For example, the recombinant viruses can be formulated with common excipients, diluents, or carriers, and formed into tablets, capsules, suspensions, powders, and the like. Examples of excipients, diluents, and carriers that are suitable for such formulations include the following fillers and extenders such as starch, sugars, mannitol, and silicic derivatives; binding agents such as carboxymethyl cellulose, HPMC and other cellulose derivatives, alginates, gelatin, and polyvinyl-pyrrolidone; moisturizing agents such as glycerol; disintegrating agents such as calcium carbonate and sodium bicarbonate; agents for retarding dissolution such as paraffin; resorption accelerators such as quaternary ammonium compounds; surface active agents such as cetyl alcohol, glycerol monostearate; adsorptive carriers such as kaolin and bentonite; and lubricants such as talc, calcium and magnesium stearate, and solid polyethyl glycols.

[0147] The recombinant viruses can also be formulated as elixirs or solutions for convenient oral administration or as solutions appropriate for parenteral administration, for instance by intramuscular, subcutaneous or intravenous routes.

[0148] The pharmaceutical formulations of the recombinant viruses can also take the form of an aqueous or anhydrous solution or dispersion, or alternatively the form of an emulsion or suspension.

[0149] Thus, the recombinant viruses may be formulated for parenteral administration (e.g., by injection, for example, bolus injection or continuous infusion) and may be presented in unit dose form in ampules, pre-filled syringes, small volume infusion containers or in multi-dose containers with an added preservative. The active ingredients may take such forms as suspensions, solutions, or emulsions in oily or aqueous vehicles, and may contain formulatory agents such as suspending, stabilizing and / or dispersing agents. Alternatively, the active ingredients may be in powder form, obtained by aseptic isolation of sterile solid or by lyophilization from solution, for constitution with a suitable vehicle, e.g., sterile, pyrogen-free water, before use.

[0150] These formulations can contain pharmaceutically acceptable vehicles and adjuvants which are well known in the prior art. It is possible, for example, to prepare solutions using one or more organic solvent(s) that is / are acceptable from the physiological standpoint, chosen, in addition to water, from solvents such as acetone, ethanol, isopropyl alcohol, glycol ethers such as the products sold under the name “Dowanol”, polyglycols and polyethylene glycols, C1-C4 alkyl esters of shortchain acids, e.g., ethyl or isopropyl lactate, fatty acid triglycerides such as the products marketed under the name “Miglyol”, isopropyl myristate, animal, mineral and vegetable oils and polysiloxanes.

[0151] For administration to the upper (nasal) or lower respiratory tract by inhalation, the recombinant viruses are conveniently delivered from an insufflator, nebulizer or a pressurized pack or other convenient means of delivering an aerosol spray. Pressurized packs may comprise a suitable propellant such as dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoroethane, carbon dioxide or other suitable gas. In the case of a pressurized aerosol, the dosage unit may be determined by providing a valve to deliver a metered amount.

[0152] Alternatively, for administration by inhalation or insufflation, the composition may take the form of a dry powder, for example, a powder mix of the agent and a suitable powder base such as lactose or starch. The powder composition may be presented in unit dosage form in, for example, capsules or cartridges, or, e.g., gelatine or blister packs from which the powder may be administered with the aid of an inhalator, insufflator or a metered-dose inhaler.

[0153] For intra-nasal administration, the recombinant viruses may be administered via nose drops, a liquid spray, such as via a plastic bottle atomizer or metered-dose inhaler. Typical of atomizers are the Mistometer (Wintrop) and the Medihaler (Riker).

[0154] The local delivery of the recombinant viruses can also be by a variety of techniques which administer the agent at or near the site of disease. Examples of site-specific or targeted local delivery techniques are not intended to be limiting but to be illustrative of the techniques available. Examples include local delivery catheters, such as an infusion or indwelling catheter, e.g., a needle infusion catheter, shunts and stents or other implantable devices, site specific carriers, direct injection, or direct applications.

[0155] Drops, such as eye drops or nose drops, may be formulated with an aqueous or nonaqueous base also comprising one or more dispersing agents, solubilizing agents or suspending agents. Liquid sprays are conveniently delivered from pressurized packs. Drops can be delivered via a simple eye dropper-capped bottle, or via a plastic bottle adapted to deliver liquid contents dropwise, via a specially shaped closure.

[0156] The formulations and compositions described herein may also contain other ingredients such as antimicrobial agents, or preservatives. Furthermore, the active ingredients may also be used in combination with other agents, for example, bronchodilators.

[0157] The recombinant viruses may be administered to a mammal alone or in combination with pharmaceutically acceptable carriers. As noted above, the relative proportions of active ingredient and carrier are determined by the solubility and chemical nature of the compound, chosen route of administration and standard pharmaceutical practice.

[0158] The dosage of the recombinant viruses will vary with the form of administration, the particular compound chosen and the physiological characteristics of the particular patient under treatment. Generally, small dosages will be used initially and, if necessary, will be increased by small increments until the optimum effect under the circumstances is reached.

[0159] The invention will be further described by the following non-limiting examples. Example 1

[0160] Results

[0161] In the context of recombinant AAV as a gene therapy vector, capsid structural features impinge on a multitude of distinct virion functions. First, AAV must be produced recombinantly, which entails a stochastic oligomerization process forming an empty capsid in the nucleus that is governed by ordered interaction between capsid protein monomers (13, 36, 37), followed by packaging of a single-strand DNA payload through the 5-fold pore that is braced by complex subunit interactions (16, 38-40). The resulting virion must be sufficiently stable to survive biochemical purification after lysis of the producer cells (41). Next, virions must attach to (co- )receptors on the surface of a target cell (AAV-DJ: heparan sulfate (42, 43)), followed by endocytosis (23). After endocytosis, virions must interact with the recently identified universal AAV receptor (AAVR) to ensure proper trafficking to the trans-Golgi network (23, 44, 45), undergo the required conformational changes (autoproteolysis, externalization of the N-terminal PLA2 domain (19, 21, 46)) that together mediate virion escape into the cytosol. Once in the cytosol, several intracellular trafficking events remain: capsids interact with the nuclear pore complex, enter the nucleus, and are forwarded to the nucleolus, where the genetic payload is released (20, 23).

[0162] The goal in this study was to survey as many proxies (‘AAV fitness phenotypes’) for these distinct virion functions as possible. We posit that the resulting comprehensive multiparametric AAV fitness datasets can facilitate the optimization of engineering AAV along multiple axes (8, 47, 48). Furthermore, based on our previous studies in ion channels (49, 50), we hypothesized that systematic domain insertion (i.e., perturbation scanning) across different measured phenotypes may uncover the underlying topological organization of AAV capsids and provide insight into capsid determinants for assembly, stability, and dynamics of virus capsids.

[0163] Domain insertional profiling in AAV-DJ VP1

[0164] AAV-DJ (24) was the testbed for inserting a peptide tag (FLAG) and six different protein domains (Fig. 8) in between every two residues of VP1. The rationale for focusing on VP1, as opposed to the more abundant VP3 or the non-essential VP2, was as follows: As the least abundant VP isoform there are, on average, between 1-5 copies of VP1 incorporated per capsid (9-12). Note that this is an average copy number based on bulk measurements (9-12); because assembly is a stochastic process, there are many particles with copy numbers at the extreme tails (i.e., 0 copies or > 10 copies) (13). It was reasoned that keeping a low number of VPs carrying inserted domains, which are potentially very disruptive, is more likely to result in assembled and functional virions. This is akin to applying a low or intermediate amount of selection pressure in directed protein evolution experiments, which can reveal more facetted fitness phenotypes (51). Furthermore, it had been observed in a prior study that inserting large domains into VR4 (part of the VP common region) completely abolished AAV production unless it was limited to VP1 (or VP2) only (32). Most importantly, focusing on VP1 allows us to interrogate the effect of domain insertion on AAV cell entry; VP1 is required for AAV infectivity as it contains the PLA2 domain that mediates endosomal escape (19, 21, 46).

[0165] To generate separate expression constructs for the VP1 domain insertion library and cap expressing only VP2 & VP3, the AAV-DJ cap gene was duplicated, and mutations introduced (MIK; T138A / M203K / M211L / M235L) to suppress expression of VP1 or VP2 / VP3, respectively (Fig. 1A, table 1). The VP1 heparin binding domain (HBD: residues 587-590) was replaced with an HA tag. The VPl-only cap gene was then subjected to SPINE (35), which resulted in a library of VP1 variants with a peptide tag or protein domain, flanked by short linkers, inserted in between every two residues. One motif we inserted was the FLAG peptide tag (DYKDDDDK; SEQ ID NO: 56), because of its similar size to prior AAV peptide insertions (52-54). With an eye towards redirecting AAV tropism, the focus was on protein domains that themselves have retargeting abilities (nanobody binding to GFP (55)) or that enable covalent linkage of retargeting moieties (SNAP-tag to link 06-benzylguanine derivatives (56); SpyCatcher to link SpyTag fusions (57)), and three different HUH tags (WDV, DCV, mMobA) to covalently link ssDNA-conjugated molecules in a sequence specific manner (58-61).

[0166] Table 1

[0167] Plasmids and libraries used in this study.

[0168]

[0169] Since insertions can affect any of the steps in AAV packaging and infection, it was attempted to independently assay different insertion variant fitness phenotypes. The relatively large insertion variant library size (744 AAV positions (including HA tag) x 7 motifs = 5,208 variants for each assay) necessitated a high throughput format. We therefore devised assays in which the different fitness phenotypes of an insertion variant were assessed by NextGen sequencing variant populations before and after a fitness test (62). A requirement for this approach was a stringent linkage between genotype (the insertion variant) and the measured phenotype (determined by properties of the capsid into which this VP 1 -variant is assembled). This was achieved by flanking the VP 1 variant library with ITRs such that the gene encoding a specific VP1 variant became packaged into the capsid that incorporated this variant during assembly. The potential pitfail of cross-packaging (a mismatch between packaged VP1 variant gene and VP1 variant protein that incorporated into the capsid) was avoided by transfecting producer cells at very low MOI. This has been demonstrated to reduce cross-packaging (63).

[0170] These VP1 variant input libraries were used, stratified by inserted motif, for helper-free virus production followed by gradient purification (Fig. 1B-D). Using NextGen sequencing (NGS) of packaged genomes, ‘AAV packaging fitness’ was assayed by counting the frequency of a given VP1 variant ( / ) after packaging Cv) relative to the frequency of that variant in the input library (w), normalized to wildtype AAV (wt):

[0171] Absolute wildtype fitness was measured by spiking in an AAV-DJ ITR flanked VP1 only cap gene containing 10 synonymous mutations into the VPl-library mix (table 1). Wildtype AAV-DJ and AAV-DJ with silent mutations showed no difference in production titers (Fig. 9). Using a similar approach to count VP1 variants before and after selection, assays were established that determine pulldown fitness (using affinity purification material), cell binding, cell uptake, and infectivity fitness (Fig. 1E-H). Infectivity assays were based on transduction of HEK293FT cells with miRFP670nano (64) expressed from the AAV payload (Fig. 1A). This enabled flow sorting of infectious variants (enriched in miRFP670nanohlghcells) (Fig. 10).

[0172] VP1 insertion library quality and completeness

[0173] Overall, coverage for all phenotype assays and inserted motifs was excellent (median dataset completeness is 98.5%; see table 2 for sequencing statistics). Biological replicates were highly correlated for plasmid library, packaging, and pulldown assays (Pearson correlation coefficient: 0.67-0.99, Fig. 11 A). The infection assay was very noisy for some domains (Pearson correlation coefficient 0.38-0.88) likely related to small number of cells collected for the miRFP670nanohlghcell pool (Fig. 10). For all tested motifs, the majority of missing insertion positions (no data in either replicate, for example positions 428-445) was missing in all phenotypes and the input library, suggesting that the dropout rate was related to library construction (Fig. 11B, C). Since replicate 2 had a higher read depth and completeness overall, we used this dataset for further analysis. Median read depth across tested motifs and phenotype assays was 635 reads per position (Fig. 1 ID, Table 2).

[0174] Domain insertions do no alter bulk properties

[0175] After helper-free production and iodixanol gradient purification, virus titers were measured by qPCR. While titers appeared somewhat lower for insertion libraries, this difference was not significant (Fig. II, one-way ANOVA p-value 0.19). Western blot with the Al antibody (which recognizes a VP 1 -unique epitope) confirmed that VP1 was incorporated into virions for all libraries (Fig. 1 J) . We next used the B 1 antibody densitometry to determine bulk VP1 / VP2 / VP3 ratios (Fig. IK, Fig. 12). It was found that wildtype AAV-DJ virion contains on average three VP1 copies per 60-mer capsid, in line with prior studies (12, 13). There was no significant difference in VP composition for the different insertion libraries with respect to VP1 (two-way ANOVA p- value: 0.641). As another bulk characteristic, we measured capsid melting temperatures using differential scanning fluorimetry. While SNAP and FLAG libraries had slightly elevated melting points (Fig. IL, Fig. 13), all libraries were overall remarkably similar and within 1°C range of AAV-DJ, suggesting that domain insertions in VP1 did not significantly impact capsid stability. Negative stain electron microscopy was used to compare full / empty capsid ratios of wildtype AAV-DJ, SNAP-, and nanobody libraries and found that full particles comprised between 80-90% of all samples (Fig. 14).

[0176] Taken together, these data suggest that, on average (i.e., considering the entire insertion library), bulk properties are not altered by domain insertions.

[0177] High resolution AAV fitness profiles across different phenotypes

[0178] It might be expected that different insertion types at different sites do have different impact on AAV fitness, so attention was turned to insertion motif- and region-specific differences. These data are summarized in Fig. 2, showing a heatmap of insertion fitness of all seven motifs in all 744 VP1 positions, segregated by measured fitness phenotypes (packaging, pulldown, binding, uptake, and infectivity). Fitness values are mapped from magenta over white to green, corresponding to lower to higher than wildtype AAV-DJ fitness (white). Note that fitness measurements are derived from AAV capsid with a variable copy number of VP1 located at random faces, which will affect phenotype penetrance. Poisson errors were generally low (Fig. 15) and only slightly elevated at the junctions between the 14 different fragments used in the SPINE-mediated assembly of the AAV insertion library. Median error with respect to dynamic range of each fitness assay was (0.102 / 2) log units = 5.1%. Overall, there was a strong dependence on insertion position and type of inserted domain for most assays, with the notable exception of uptake and infectivity.

[0179] Effects of domain insertions in VPlu

[0180] Focusing on packaging fitness, distinct increase in fitness was observed when motifs (in particular: nanobody, SNAP tag, or DCV HUH tag) were inserted into the PLA2 domain of VPlu (Fig. 3). This suggests that perturbing the PLA2 domain at the assembly stage resulted in increased packaging efficiency. One possible explanation is that an insertion causes such a dramatic conformational change that VP 1 becomes incompatible with trimer or pentamer formation, the initial steps of capsid assembly (36, 37). Malformed VP1 may ‘poison’ trimer assembly in the cytosol, may interfere with transport to the nucleus, or may interfere with trimer addition to the growing capsid. In this scenario, VP1 carrying motifs inserted in the PLA2 are not incorporated at all in the assembling capsid thus leaving more room for genome packaging, because the steric hindrance of internalized VPlu is removed. An alternative explanation for the increased packaging fitness is that the inserted motif forces VPlu to remain external, thus removing steric barrier to genome packaging. Our data provide support for the latter: pulldown with the respective affinity materials (e.g., GFP-agarose beads) showed significant enrichment if nanobody, mMobA, DCV, and SNAP tag were inserted into VPlu (residues 1:160), all of which improved also packaging fitness (Fig. 4, Fig. 16). This means that these VP variants are incorporated in the capsid such the inserted motif is accessible on the capsid exterior. FLAG tag insertions into VPlu, which presumably did not interfere with internalization, did not increase pulldown fitness. Neither did WDV insertions, which is consistent with their deleterious impact on packaging fitness. Furthermore, it was found that overall binding fitness was neutral (peaked around wildtype fitness, Fig. 2 and Fig. 17), which was expected as most VP subunits (VP2 & VP3) contain the wildtype determinants of proteoglycan and AAVR binding (45, 65). However, a notable drop for binding fitness for VPlu insertions of motifs that package well (e.g., nanobody; Pearson correlation coefficient -0.641) was observed, possibly due to steric hindrance of virus binding when it carries large external motifs. The only motifs that were not impaired for binding upon VPlu insertions are the FLAG tag, which can be explained by its small size, as well as SpyCatcher and WDV, which are motifs that our pulldown data suggest were not compatible with virus assembly when inserted into VPlu. Consistent with the role and required timing of VPlu in virus trafficking upon endocytosis, uptake fitness (assay measures presence of viral genome in any compartment inside the cell (66)) was impaired for all motifs inserted into the N-terminus of VP1 (Fig. 2 and Fig. 18). It has previously been demonstrated that premature exposition of VPlu decreases infectivity (67), meaning that variants with pre-externalized VPlu escape less efficiently from the endosome and are degraded in the lysosomal compartment (23). The same was observed for infectivity fitness, with the notable exceptions of FLAG and WDV.

[0181] Effects of domain insertions near AAV’s 3 -fold axis

[0182] Insertion into protrusion near the 3 -fold axis (residues 420-620; including VR4-8) generally impaired virus packaging (Fig. 2 and Fig. 3), which is consistent with the highly interdigitated structure of this region and its role in early assembly of VP trimers that are then forwarded to the nucleus as capsid building blocks (20, 36, 37). VP1 with insertions in this region may interfere with efficient trimer assembly (through a kinetic mechanism or by promoting off- pathway products), which would lower the overall capsid assembly efficiency. Nevertheless, several lines of evidence suggested that there appears to be some plasticity in trimer assembly to accommodate VP1 insertion variants so that they can be incorporated. For one, the overall hit to packaging fitness depended on the specific inserted motif. As expected, FLAG peptide insertions were relatively benign, but so were insertions of two HUH tags (DCV and WDV), and SpyCatcher (Fig. 2 and Fig. 3). There was no clear correlation with motif size hinting at more complex determinants for insertion fitness in this region. Second, despite impaired packaging fitness, pulldown fitness was greater than wildtype for all motifs, suggesting that they did become incorporated into purified AAV capsids, albeit at a lower overall efficiency (Fig. 2 and Fig. 4).

[0183] As for the N-terminus of VP1, binding fitness was not impaired for insertions in the 3-fold protrusion (Fig. 2 and Fig. 17). In fact, the three HUH tags showed increased binding, which could be due to their generally higher cationic surface charge (Fig. 8B) aiding interaction with negatively charged components of the extracellular matrix, such as proteoglycans. Consistent with higher surface binding, higher uptake fitness than wildtype AAV-DJ in the case of DCV and WDV was found, when inserted into VR5 or VR8 (Fig. 2 and Fig. 18). It was noted that binding and uptake fitness measured in this high throughput assay for insertion of mMobA into VR4 match our results from previous engineering of this region (32). Despite higher binding and uptake, none of the insertions in this region could achieve wildtype infection efficiency (Fig. 2), suggesting that domain insertion impacted later steps of virus trafficking to the nucleus.

[0184] Effects of domain insertions near AAV’s 2-fold valleys The neighborhood near the 2-fold symmetry center, which includes VR9, emerged as another region with distinct motif-specific phenotypes. This is consistent with earlier studies showing that dynamics of the 2-fold regions are essential for genome packaging (39) and AAV infectivity (68). Here, it was observed both strongly deleterious fitness (e.g., SNAP tag) and strongly beneficial fitness (SpyCatcher and WDV). In all cases pulldown fitness was positive (Fig. 4), suggesting the incorporation of at least one VP1 insertion variant that contains a motif insertion in this region. Interestingly, binding, uptake, and infectivity were impaired for most motifs. Unbiased clustering of insertion fitness reveals contiguous functional units in AAV capsids

[0185] Taken the fitness measurements across all motifs, all insertion positions, and all measured phenotypes in aggregate, patterns in fitness variance that appear correlated in contiguous regions of the AAV capsid were noticed. For example, packaging fitness varied predominantly in the PLA2 domain of VPlu and the 2-fold symmetry axis (Fig. 5A). Focusing on residues unique to this interface, it was found that packaging fitness distributions of interface and non-interface residues were not significantly different when DCV or FLAG peptide were inserted. However, fitness was significantly improved for mMobA, SpyCatcher, or WDV insertions into interface residues (Fig. 5C). SNAP tag insertions were strongly deleterious. The variance at the 3-fold symmetry axis was markedly different. As described above, all motifs except WDV lowered packaging fitness, which resulted in lower overall fitness variance at this interface (Fig. 5D). Remarkably, uptake fitness variance was greatest in the protrusion around the 3 -fold axis and still considerably high along the 2-fold axis (Fig. 5B). For all measured phenotypes, variance was relatively low around the 5-fold symmetry axis (Fig. 19).

[0186] If different inserted motifs are thought of as different degrees of perturbation (e.g., weak for FLAG peptide insertion, strong for large nanobody insertion), then the positional and domaintype dependence of insertion permissibility that was observed in the data suggests that we were measuring spatially resolved information of how different aspects of AAV fitness responds to these different degrees of perturbation.

[0187] To link insertion permissibility phenotypes to mechanistic structure / function relationships, an unbiased clustering approach (UMAP) (69) was used. This resulted in five robust clusters that map to regions of the capsid with distinct roles in AAV biology (Fig. 6A). Importantly, these clusters mapped to structurally contiguous (not interspersed) regions of the AAV capsid (Fig. 6B, C).

[0188] Cluster 1 contains VPlu in addition to residues lining the bases of the 3-fold and 5-fold axes. Cluster 2 forms an extended network that comprises the HI loop and connects to the 3-fold axis protrusions. Cluster 3 represents protrusion at the 3-fold axis and residues on the external turns of the DE loop that line the pore at the 5 -fold symmetry axis. Cluster 4 maps to the depression near the 2-fold symmetry axis and buried regions, which are part of the 3-fold axis. Cluster 5 predominantly maps to residues that line the capsid interior and the 5-fold pore, or that interdigitate HI loop, external residues of the 2-fold valley, and 3-fold protrusions (Fig. 6B).

[0189] To understand the underlying mechanisms that drive clustering, fitness phenotype distributions were segregated by cluster identity (Fig. 6D). Considering packaging fitness, it was found that insertion into two clusters (1 and 5; containing VPlu and the network of residues that connect HI loop to the base of the 3-fold protrusion) were associated with improved packaging fitness, while clusters 3 and 4, which represent the interdigitated external region 3-fold- and 2- fold axes, were associated with poor packaging fitness. Insertions into capsid lining regions (cluster 2) were neutral with respect to packaging. With different measured phenotypes, these association patterns change: Considering pulldown fitness, we found that clusters representing buried residues or those lining the capsid interior (i.e., clusters 1, 2 & 5) were associated with poor fitness compared to those in externally accessible regions (cluster 3 & 4). For uptake, cluster 1 (which contains VPlu) had the worst fitness and cluster 3 (the 3-fold protrusion) was closest to wildtype fitness.

[0190] Using existing knowledge about structure and function of AAV, we can begin to make association between clusters and their specific roles in AAV packaging and infection. The strong association between poor pulldown fitness, uptake fitness, and cluster 1 is obvious considering it being comprised of mostly buried or internal residues and containing VPlu, which encodes the required PLA2 domain and nuclear localization signals (19, 21, 22, 46). Conversely, association with higher packaging fitness would be compatible with motif insertion promoting externalization of VPlu, thus decreasing steric hindrance with the packaged genome. The sensitivity of cluster 3 with respect to packaging efficiency is consistent with trimer formation and stability, which are key determinants of capsid assembly (36, 37). Most binding sites for cellular receptors (e.g., proteoglycan, AAVR) are located near the 3-fold axis, as well (14, 42-45, 65, 70). Flexibility of the 2-fold interface has previously been linked not only to AAV infectivity (68), but also genome packaging (39), which would explain how insertions may drive poor packaging fitness. The relatively neutral fitness of cluster 2 (e.g., packaging, uptake) is consistent with most insertions ending up on the capsid interior, not interfering with assembly or cellular uptake. While the 5-fold pore is part of this cluster, a numeric simulation of VP1 copy numbers ranging from 1-10 suggests that around half of the twelve 5-fold pores are assembled from non-VPl only, and thus available as an alternative pathway for Rep-mediated genome packaging. Cluster 5 has an intricate structure, with the HI loops that surround the 5 -fold pore like the blades of an aperture and connecting to the base of the 3-fold axis. Prior studies that have suggested a link between conformational changes at the 3-fold protrusion upon binding cell surface proteoglycan are communicated to conformational change at the 5-fold pore, priming the release of VPlu (71). Disulfide crosslinking to probe conformational flexibility

[0191] If the correlated conformational plasticity of clusters plays a role in packaging and / or infectivity, it would be expected that changing conformational dynamics would impact these functions. One way to test this idea is by replacing two proximal residues by cysteines, such that a cystine disulfide link is formed once AAV is exposed to an oxidizing environment (i.e., after release from producer cells). If the two mutated residues are part of the same cluster of correlated conformational plasticity, and assuming that the individual cysteine substitution are benign, we expect a less significant effect on infectivity compared to if the two residues belong to different clusters. Our choices of residue pairs are summarized in fig. S13A. Unlike domain insertions, which were only done in VP1, cysteine substitutions were introduced into all VPs. All single- and double-mutants produced near to wildtype titers (Fig. 20A; one-way ANOVA n.s.; Dunnett’s test with wildtype AAV-DJ as control: n.s.).

[0192] Some single cysteine mutants had dominant deleterious effects on infectivity (e.g., W608C, H292C). For double-mutants in the ‘within cluster’ set that had wild-type single mutant infectivity, infectivity was comparable to wildtype (e.g., H625C / Y426C, Fig. 20B), suggesting the minimal disruption of conformational dynamics. For pairs that belong to different clusters, all but one (H643C / Y350C) showed effects on infectivity that differed from what was predicted based on the individual single mutants (Fig. 20B). Several of these pairs involve interfaces that undergo conformational dynamics during infection.

[0193] For example, mutating F671C in the HI loop was strongly deleterious to infection, but this phenotype was rescued in the background of H255C, which by itself was benign. Prior studies have shown that heparin binding near the 3-fold and 2-fold axes induces an HI loop rearrangement and an iris-like opening of the channel located at the 5-fold axis (71). HI loop deletions, substitutions, and insertions have shown that this loop, while flexible in amino acid composition and length, is critical for proper VP1 incorporation and infectivity (72). The same study showed that interaction of the HI loop with the underlying EF loop is mediated by hydrophobic pi-stacking interactions (F661 / P373; in AAV2 numbering) and that disrupting this interaction lowers infectivity by preventing VP1 incorporating into assembled capsid. Our results are reminiscent of this mechanism. Positions F671 / H255, which can form NH- -n hydrogen bonds, are both conserved across AAV serotypes. H225C may be benign as this supports formation of an aromatic-thiol 71 hydrogen bond, thus allowing VP1 to incorporate or retaining structural rearrangement after receptor binding. Conversely, F671C is disruptive as it removes the aromatic component of 71-stacking interactions, prevents VP1 incorporation and / or disrupts these rearrangements. The H255C / F671C double mutant may rescue infectivity by forming a disulfide bond to substitute as a stand in for n interactions.

[0194] In another example, H423C (at the base of the 3-fold axis) and V613C (with the 2-fold interface) individually had little effect, but together strongly impaired infectivity. This trend held true for adjacent pairs (H360C / 437C; H428CL737C) that similar linked clusters comprising the 3-fold protrusion and 2-fold axis. Interestingly, several prior studies have linked AAV infectivity and conformational dynamics at the 3-fold and 2-fold axes. In addition to structural rearrangements in the HI loops, heparin binding to AAV2 causes significant rearrangement of 3- fold protrusions and the 2-fold valleys (71). Selective oxidation of tyrosine residue at the 2-fold dimer interface lowered infectivity (68). A mutation (R432A in AAV2) remodeled intramolecular and intermolecular hydrogen bond networks propagating from the 3-fold to both 2-fold and 5-fold axes (38, 40).

[0195] Engineerable hotspots near the 2-fold axis and in the HI loop

[0196] The HUH tag mMobA was used in AAV-DJ VR4 to covalently link targeting scaffolds to the AAV capsid, which redirected AAV tropism in vitro (32). Here, the entire AAV capsid for suitable HUH tag insertion hotspots was interrogated. Comparing fitness maps for all three HUH tags, there were many differences amongst tags, which is likely related to their different biophysical properties (Fig. 8). It was noticed that WDV was remarkably different two all other inserted domains in several regards. For one, packaging fitness was improved over wildtype when this domain was inserted into HI loops or along the 2-fold axis (Fig. 3). Binding fitness was generally strong, but insertions into the 2-fold axis were deleterious (Fig. 17); this was the only insertion type for which we saw a deleterious phenotype for this assay. For uptake fitness, it was observed a strong segregation in fitness between 3-fold protrusion and 2-fold and 5 -fold axes (Fig. 18). Given that most previous studies have investigated VRs in the 3-fold protrusion for capsid engineering, attention was turned to insertion sites near the 2-fold and 5-fold axes that had near wildtype packaging fitness in the NGS-based assay. Ten VP1 WDV insertion variants individually were produced as crude cell lysates and measured titers, which all were comparable to wildtype (Fig. 7A, one-way ANOVA n.s.). Testing each crudely enriched variant for the ability to infect HEK293FT cells, it was found that several WDV variants inserted into surface exposed sites retained infection potency (Fig. 7B). Among those were insertions into the 3-fold protrusions (S268), DE loop of the 5-fold pore (T331), HI loop (N664), and two sites along the 2-fold axis (Y702, K708). All sites had positive fitness in the NGS-based pulldown assay suggesting the WDV-VP1 does become incorporated into AAV capsid (Fig. 2 and Fig. 4). For two variants, N664 and K708 (Fig, 7C), they were produced as iodixanol-gradient purified virus, which both trended to produce at higher titer compared to wildtype (Fig, 7D) and incorporated VP1 as confirmed by Western blot (Fig. 21). As previously shown (32), HUH tags mediate the attachment of ssDNA antibodies to AAV, which in turn increased infectivity in cells that express, on the cell surface, the antigen recognized by the antibody. Using surface-expressed GFP (GFP-GPI) as a test case, we tested infectivity of the two purified WDV variants and wildtype AAV-DJ with and without conjugation to an ssDNA-linked anti-GFP antibody. Note that expression of GFP-GPI alone reduced cell health likely related to ER stress. Whereas infectivity of WDV variants was unaffected by conjugation to anti-GFP just like for wildtype AAV-DJ (expected as it was a mock conjugation since it does not contain a HUH tag), a boost to infectivity was observed for both N664-WDV and K708-WDV upon co-expression of surface expressed GFP (Fig. 7E), suggesting that anti-GFP became conjugated to WDV and then enhanced infectivity by directing AAV toward surface expressed GFP as a binding receptor (two-way repeated measure ANOVA p-value 0.0063).

[0197] Discussion

[0198] Viral vectors are an essential component for gene delivery in therapies treating inherited disorders and cancer. Adeno-associated virus (AAV), in particular, is widely used in both approved therapies and in ongoing clinical trials because of its good safety in humans and ability to drive long-term expression in both dividing and non- dividing cells. Unfortunately, several of the evolved properties of AAV are mismatched to clinical needs (e.g., broad tropism, limited payload capacity, existing serum-immunity in most of the human population) or they pose biomanufacturing challenges (e.g., scale up of helper virus free production, yield of full virions to maximize potency).

[0199] Motivated by these challenges, there have been extensive efforts to re-engineer AAV properties in the past including directed evolution approaches, such as repeated mutagenesis, capsid shuffling (73), viral display of short peptides (52-54), and adding larger, structured targeting scaffolds, such as antibodies, nanobodies, DARPins, or affibodies (30-33, 74). Most of these studies have focused on regions in the 3-fold protrusion, commonly VR8, VR4, or VP termini. Recently, deep mutagenesis of the entire capsid protein, combined with machine learning (8, 47, 75), demonstrated that learned sequence & function relationships can aid the prediction of sequence variation to improve desired AAV traits. However, deep mutagenesis in AAV has so far been limited to amino acid substitutions. Here, the concept of deep mutagenesis was combined with scaffold insertion to systematically measure the fitness of AAV containing VP 1 with seven different motifs inserted in between every two residues. This comprehensive analysis quantitatively links where different structured motifs can be inserted into VP1 to retain compatibility with AAV virion assembly and packaging, cell binding, and uptake. The fitness profiles show that there is a strong dependence on insertion position and type of inserted motif, with several regions showing diverging fitness for different measured phenotypes (e.g., VPlu for packaging vs. cell uptake, Fig. 2). This highlights that the outcome of sequence variation can have multi-facetted impact on AAV properties, and it calls for integration of assays across several clinically relevant AAV attributes to safeguard against inadvertent optimization for undesired traits. In recent years, we have gained broad access to precision variant library engineering (35, 76, 77), Next-Gen sequencing (enabling counting number of sequence variants in a highly diverse population before and after applying a test for fitness (62)), unified analytical frameworks to interpret these large dataset (78), and machine learning approaches for deciphering sequence / function relationships (79). Taken together, generating and interrogating large AAV variant libraries has become feasible. The focus can now shift towards what these datasets tell about AAV biology and how they can guide and accelerate viral vector engineering.

[0200] For example, while fitness of many insertion sites is consistent with known AAV structure and function (e.g., the importance of trimer assembly along 3-fold axis in capsid assembly (36)), the high packaging fitness when motifs are inserted into the PLA2 domain of VPlu was surprising. Several studies have elucidated VPlu and VP2 dynamics as part of events after virus uptake by the cell, in which the internalization of VPlu PLA2 is a required step for endosomal escape (21, 46). VPlu internalization during maturation of AAV virions is less well understood but appears to be coordinated with genome packaging (39). The existence of such mechanism would reconcile recently proposed models of stochastic assembly (13) (which should result in both internalized and externalized VPlu) and the well-known requirement for heating or pH lowering to induce VPlu externalization before it can be detected by VPl-specific antibodies (17, 19, 72, 80). It is plausible that domain insertions interfere with this process, leaving VPlu externalized, which would leave more room inside the AAV capsid for genome packaging by Rep78 due to lack of steric hindrance (one VP1 occupies 85x103A3or l / 35thof the available space inside the capsid) but at the loss of infectivity. Interestingly, VPlu and VP2 of the related parvovirus B19V have acquired an receptor-binding domain insertion just upstream of PLA2 and are always external (81-84), supporting the idea that addition of extra domains into VPlu leave it externalized. Given that precise timing of VPlu externalization in the correct endosomal compartment is required for maximum infectivity (premature externalization reduced infectivity (67)), this may point to an evolutionary mechanism that balances virion packaging efficiency with infectivity. If the full / empty ratio is fundamentally constrained by a packaging / infectivity balance, this would have implications for biomanufacturing of AAV in which one of the major ongoing efforts is to find ways to enrich full particles. One prediction of our hypothesis is that full particles may contain fewer VP1 , on average, compared to empty particles, and this is negatively impacting infection potency. The observation that overexpression of VP1 inhibits rAAV packaging is consistent with this idea (85). There may exist a Goldilocks regimen, just the right copy number of VP1, that maximizes both - genome content and endosomal escape. Further research is required to fully test this hypothesis, including measuring VP stoichiometry and genome content at the single capsid level.

[0201] Unbiased clustering of insertion fitness across several phenotypes also revealed a topological organization of AAV into regions that can be linked to correlates in AAV capsid assembly, genome packaging, and infectivity. Given that many of these roles have a basis in distinct regimes of capsid stability and dynamics, we hypothesize that inserting different domains (which represent different degrees of perturbation) probes conformational plasticity at or near the insertion site. Put simply, it probes whether the insertion site is conformationally rigid (allowing no insertions) or conformationally flexible (allowing some or all insertions). Clusters emerge because conformational plasticity of different capsid regions impinges differently with different measured phenotypes (e.g., capsid flexing required for efficient genome packaging (39) vs. flexing during externalization of VPlu / PLA2 after cell uptake, which only happens in full, but not empty capsids (17)).

[0202] Similar ideas of spatially contiguous protein regions linked to specific functions have been proposed in the past, including protein “sectors” mapped through measuring amino acid coevolution (86, 87), regional conformational flexibility mapped by circular permutation profiling (88) and domain insertion (49, 50, 76), or revealing the functional architecture of an enzyme from high-throughput enzyme variant kinetics (89). At their core, all these approaches use mutations to perturb sequence / function relationships. Similarly, by perturbing VP1 through domain insertion and measuring how it responds (in terms of assembly, infectivity, etc.), we learn how AAV structure intersects with AAV function. Going forward, domain insertional profiling in the background of different genetic backgrounds (i.e., serotypes) may further separate general principles of AAV assembly, function and serotype-specific properties (e.g., stability, immune evasion, tropism). The systematic domain insertion approach also revealed new opportunities for viral engineering. Following the intuition that engineering non-conserved, surface exposed, and tropism-determining loops is the likeliest path to change AAV properties, much of AAV engineering so far has focused on N-termini of capsid proteins, or variable loops of the 3-fold protrusions (30, 32, 34, 74, 90-94). However, systematic studies we and others conducted (50, 88), suggest that this intuition may be misleading. It was found that permissibility to domain insertion is not correlated with conservation, surface exposure, or other static, structural features. Instead, dynamic features, such as regional flexibility, are predictive as to where a domain insertion is tolerated. In this study, we identified two new regions near the 2-fold and 5-fold axes that can tolerate the insertion of HUH tags (15kDa), which in turn enable the covalent linkage of antibodies (150kDa). In a proof-of-principle, it was found that insertions have little impact on production levels and provide a modest boost to infecting cells that express the antibody’s cognate antigen. Further research and engineering are required to fully leverage the potential of these new engineerable hotspots. They represent an exciting opportunity to sidestep the constraint of directed evolution of targeting the region near the 3-fold axis. As this region is important for receptor binding, it is also the most antigenic region. In fact, the binding sites for several proteoglycans, AAVR, and neutralizing antibodies (42, 43, 45, 70) overlap. Approaches that shuffle the sequence of this region must apply selection pressure to selectively remove antibody binding whilst retaining the mode of cell binding and uptake. By providing an alternative site to which targeting scaffold can be linked, it may be possible to address this challenge more effectively.

[0203] Materials and Methods

[0204] Cloning and library generation

[0205] All plasmids and libraries used in this study are listed in table 1 and were generated either by classical restriction enzyme cloning or Golden Gate Assembly (95). Restriction enzymes were obtained from NEB, standard oligos, labeled oligos as well as gBlocks from IDT, and oligo pools for library cloning from Agilent Technologies. For amplification of DNA sequences for cloning and NGS the PrimeSTAR Max DNA polymerase (Takara Bio Inc.) and for colony PCRs the OneTaq Quick-Load Master Mix Polymerase was used (NEB). PCR products were analyzed on 1 % TAE agarose gels, cut out and purified using the Zymoclean Gel DNA Extraction Kit (Zymo Research) by following the manufacturer’s instructions. Post cloning, plasmids were transformed into NEB® Stable Competent E. coli cells and libraries into MegaX DH10B T1R Electrocomp Cells (Thermo Fisher), before plated on LB plates containing either carbenicillin (100 pg / ml) alone or a combination with chloramphenicol (25 jig / ml) depending on the selection marker(s) on the plasmids and libraries. To assess coverage of libraries a small amount from the transformed cells was taken, serial dilutions prepared, and plated on LB plates with the respective selection marker(s). The next day, colonies were counted to estimate coverage and colony PCRs were run to verify library diversity. Plasmids were isolated using the Zyppy Plasmid Miniprep Kit, ZymoPURE II Midiprep Kit or the ZymoPURE II Maxiprep Kit (all Zymo Research).

[0206] For the generation of AAV-DJ insertion libraries an altered cap-DJ gene sequence was used with mutated start sites for VP2 (T138A) and VP3 (M203K, M211L, and M235L), a T176A mutation eliminating the BsmBI cutting site, and a replacement of the HBD domain (R587-R590) by an HA-tag (AYPYDVPDYAA; SEQ ID NO: 57). The libraries were created using SPINE as previously published by us (35). In brief, the cap DJ-VP1 sequence was split up into 14 fragments, and oligos containing a genetic handle behind every amino acid position were designed. Every oligo contained barcodes for amplification, matching BsmBI restriction sites to assemble the 14 fragments, and Bsal restriction sites to swap out the handle. The genetic handle was first replaced by a chloramphenicol expression cassette flanked by BsmBI cutting sites to insert a selection marker for library presence. Next, the library was transferred into an AAV plasmid backbone encoding an EFS-driven miRFP670nano sequence terminated by a SV40-polyA, a p40 promoter with BsmBI sites for the library insertion and flanked by ITRs. Last, the chloramphenicol was replaced by different domains (nanobody, SpyCatcher, SNAP, mMobA, WDV or DCV) or a FLAG-tag. While the domains had 5aa SGGGG-domain-GGGGS linkers (SEQ ID NOS: 57 and 58), the FLAG-tag was flanked by short SG-FLAG-GS linkers only. The DJ-VP1 silent mutation plasmid, that was used as a reference, was designed by introducing ten silent mutations into the VPl-only DNA sequence of AAV-DJ. To this end, codons for either arginine, serine or leucine were altered by two nucleotides each at positions that were 200-300 bp apart from each other. The silent mutation VP1 coding sequence was ordered as a gBlock and cloned into the same backbone as the plasmid libraries, i.e., miRFP670nano expression cassette and a p40 promoter to drive VP1 expression and flanked by ITRs. Cysteine point mutations were introduced by site-directed mutagenesis of the wildtype rep2-capDJ plasmid. WDV insertion variants were generated by inserting the domain into the above mentioned p40-driven and altered AAV-DJ-VPl-only sequence by Golden Gate Cloning.

[0207] Tissue culture

[0208] HEK293FT cells (Invitrogen) and 293AAV cells (Cell Biolabs) were maintained in DMEM (Gibco) containing 4.5 g / L D-glucose, L-glutamine, l lO mg / L sodium pyruvate, and supplemented with 10 % fetal bovine serum (Gibco) and 100 U per mL penicillin / 100 pg per mL streptomycin (Gibco). Cells were kept in a humidified cell culture incubator at 5 % CO2 and 37 °C and passaged every 2-3 days when reaching 70-90 % confluency. For experiments with HEK293FT cells, plates were pre-coated with growth factor reduced basement membrane matrix Matrigel (Corning) prior to seeding. If applicable, HEK293FT cells were transfected with a plasmid encoding GFP-GPI using Turbofect (Invitrogen) while seeding and according to the manufacturer's protocol. The amounts of DNA used for transfection are further specified in the sections of the different assays. For AAV productions, 293 AAV cells were used only till reaching passage 10.

[0209] AAV crude lysate production

[0210] 293AAV cells were seeded into 6-well plates at a density of 500,000 cells per well. The next day, cells were transfected with 2.5 pg DNA using PEI and an equimolar ratio of the plasmids necessary for the respective AAV production. Three days post transfection, cells were harvested by flushing off the cells by pipetting. Cells were washed with PBS once and then subjected to five freeze and thaw cycles by alternating between liquid nitrogen and a 37 °C water bath. Cell debris was pelleted by centrifugation at 18,000 rpm at 4 °C for 10 minutes and the supernatant containing the AAV particles, was stored at -20 °C until use.

[0211] Purified AAV production

[0212] Large scale productions of AAV were either done by the University of Minnesota Viral Vector and Cloning Core using a sucrose gradient or by ourselves following published iodixanol gradient density protocols (41, 96). In brief, four million 293 AAV cells were seeded into 15 cm dishes and transfected using PEI and 47 pg total DNA per dish 48 hours post seeding. For the AAV-DJ library control, an equimolar triple-transfection was used composed of an Adeno-helper plasmid, a plasmid encoding the rep2 and capDJ genes, and a transgene plasmid encoding an EFS promoter-driven miRFP670nano sequence flanked by ITRs. For AAV-DJ insertion library productions, a plasmid ratio of 1:0. 1:0.1 of an Adeno-helper plasmid, a plasmid encoding rep2 and only VP1 of capDJ, as well as the respective AAV-DJ insertion library in which the AAV-DJ silent mutation variant was spiked in was transfected. To top up to 47 pg total DNA, a pUC19 stuffer plasmid was added. 72 hours post transfection, cells were detached with a cell lifter and cells pelleted by centrifugation at 400 xg for 15 minutes. Cell pellet was washed once with PBS and resuspended in a buffer containing 2 mM MgC12, 0.15 M NaCl and 50 mM Tris-HCl at pH 8.5. Cells were cracked open using five freeze and thaw cycles. Free genomic and plasmid DNA was digested with a Benzonase Nuclease (Sigma- Aldrich). Cell debris was removed by centrifugation and, subsequently, the lysate transferred into ultracentrifugation tubes (Beckman Coulter). The iodixanol discontinuous gradient (15%, 25%, 40%, and 60% iodixanol concentration) was layered underneath the cell lysate. Density gradient centrifugation was done at 50,000 rpm for two hours at 4 °C using a 70.1 Ti rotor (Beckman Coulter). Post centrifugation, the 40% iodixanol phase containing the AAV was isolated, aliquoted, and stored at -80 °C until use. For pulldown, binding and uptake, as well as the DSF assays, AAV samples were dialyzed to PBS supplemented with 5% glycerol using 10 kDa Amicon Ultra-15 Centrifugal Filter Units (MilliporeS igma) . qPCR

[0213] To determine the titer of crude lysate AAV samples, 1-5 pl of the crude lysate were mixed with PBS supplemented with 2 mM MgCb to a final volume of 50 pl. Then, 0.1 pl ultrapure Benzonase Nuclease (Sigma-Aldrich) was added. Samples were incubated for 30 minutes at 37 °C to digest non-encapsidated DNA. Next, 5 pl of a 1 Ox Proteinase K buffer (100 nM Tris-HCl, pH 8.0, 10 mM EDTA, and 10% SDS) and 1 pl Proteinase K (20 mg / ml; Zymo Research) were added to stop the DNA digest and start the protein digest to free the ssDNA from the AAV particles. Samples were incubated for 20 minutes at 50 °C, followed by a heat inactivation of the enzymes for 5 minutes at 95 °C. The DNA was purified using the DNA Clean & Concentrator-5 Kit (Zymo Research) according to the manufacturer’s instructions for ssDNA purification. For purified AAV samples, the Benzonase digest step was skipped and only the Proteinase K and DNA purification steps were performed. All samples were diluted 1: 1,000 in H2O prior to qPCR. The viral genome (vg) quantification was done on a QuantStudio5 Real-Time PCR System (Applied Biosystems) using the PowerUp SYBR Green Master Mix (Applied Biosystems) and following the manufacturer’s instructions. To calculate the viral titer (vg / ml) a plasmid standard with a known concentration of plasmid copies was used. Primer sets binding either the CMV- enhancer or the p40 promoter of the AAV genomes, as well as in the plasmid standard were selected (table S3).

[0214] Pulldown assay of libraries

[0215] Between le9 and lelO vg were used as input material to bind to different magnetic beads for pulldown assays: SNAP-Capture Magnetic Beads (NEB) for SNAP tag insertion; Pierce Anti- DYKDDDDK (SEQ ID NO: 59) Magnetic Agarose (Thermo Scientific) for FLAG tag insertion; Streptavidin Magnetic Beads (NEB) for nanobody, SpyCatcher, and HUH tag insertions. 80 pl bead slurry for SNAP pulldowns and 50 pl bead slurry for FLAG pulldowns were washed three times with 300 pl wash buffer I (0.15 M NaCl, 20 mM Tris-HCl pH 7.5, 1 mM EDTA). Then AAV libraries were mixed with wash buffer I and the beads to a final volume of 300 pl, before incubated on a slow shaker for 30 minutes at room temperature. Beads were washed again three times with 300 pl wash buffer I for SNAP beads and with PBS (pH 7.4, Gibco) for FLAG beads to remove unbound AAV particles, and finally resuspended in 50 pl PBS. For pulldown assays with streptavidin beads, 100 pl bead slurry was washed three times with 300 pl wash buffer II (0.15 M NaCl, 20 mM Tris-HCl pH 7.5, 1 mM EDTA). For nanobody and SpyCatcher binding, beads were pre-incubated with either 320 pmol biotinylated superfolder-GFP or 1,000 pmol biotinylated SpyTag, respectively, for 30 minutes on a slow shaker at room temperature in a total volume of 300 pl in wash buffer II. Afterwards, unbound superfolder-GFP and SpyTag were removed by washing the beads three times with 300 pl wash buffer II. For HUH tag pulldowns, AAV libraries were first reacted with 1 nmol biotinylated ssDNA oligos (sequences are given in table S3) in PBS supplemented with 0.05 % v / v salmon sperm DNA (Invitrogen), 1 mM MgCE, and 1 mM MnCI? for 15 minutes at 37 °C. Next, beads were incubated with AAV libraries for 30 minutes on a slow shaker at room temperature, before unbound AAV particles were removed by washing three times with 300 pl wash buffer II. Beads with bound AAV particles were resuspended in 50 pl PBS. Viral genomes of samples after the pulldown were purified using the Quick-DNA Microprep Plus Kit (Zymo research) and by following the manufacturer’ s protocol. Binding and uptake assays of libraries

[0216] The binding and uptake assays were performed as previously described in Berry et al. 2017 (66), but using HEK293FT cells. In brief, 375,000 HEK293FT cells were seeded into 6-well plates using 2 ml culturing media per well. The next day, cells were incubated for 30 minutes at 4 °C. Afterwards, the media was aspirated and 200 pl cold DMEM containing AAV particles at an MOI of le4 added. The cells were further incubated for one hour at 4 °C. Next, cells were washed three times with ice-cold PBS to eliminate unbound AAV particles. For the binding assay, 150 pl ice- cold PBS was added, the cells detached with a cell scraper, and the cell suspension transferred to a microcentrifuge tube. For the uptake assay, 1 ml of pre-warmed culturing media was added immediately after the PBS wash and the cells incubated for two hours in a cell culture incubator to allow for uptake of the AAV particles. Then, the cells were detached by trypsinization and collected in a microcentrifuge tube, before washed three times with 200 pl PBS. The DNA from binding and uptake samples was purified using the Quick-DNA Microprep Plus Kit (Zymo research) and by following the manufacturer’s protocol.

[0217] Infectivity assay of libraries

[0218] 150,000 HEK293FT cells were seeded into 12-well plates using 1 ml culturing media per well. The next day, cells were transduced with purified AAV (in PBS supplemented with 5 % glycerol) at an MOI of 2e5 vg / cell. To this end, the media was aspirated, cells were washed with 500 pl PBS once. Purified AAV were mixed with DMEM without supplements to a final volume of 500 pl and added onto the cells. After two hours of incubation in the cell culture incubator, 1.5 ml culturing media (with supplements) was added. 24 hours post transduction, the temperature was reduced to 33 °C to promote protein expression rather than cell growth (97). 72 hours post transduction cells were prepared for cell sorting as follows. Media was aspirated, cells were washed with 500 pl PBS, 500 pl Accutase solution (Sigma-Aldrich) was added and incubated at room temperature until all cells detached. Cell suspension was transferred to a microcentrifuge tube and centrifuged for three minutes at 400 xg to pellet cells. Cells were washed two times with 500 pl PBS, before resuspended in 650 pl cell sorting buffer (PBS supplemented with 5 mM EDTA and 2.5 % FBS) and passed through a 35 pm cell strainer to avoid cell clumps. Cell sorting was performed by the University of Minnesota Flow Cytometry Resource (UFCR) on a FACS Aria II instrument (BD Biosciences) with a 85 pm nozzle by sorting miRFP670nano positive (excitation 640 nm laser, emission 670 nm / 30nm bandpass filter). Post sorting, the DNA was extracted from the cells using the Quick-DNA Microprep Plus Kit (Zymo research) and by following the manufacturer’s protocol.

[0219] NGS preparation

[0220] Purified DNA samples (library plasmid DNA, ssDNA from purified AAV, and DNA extracted after pulldown, binding, uptake, and infectivity assays) were amplified using primers binding 50 base pairs up and down stream of the VP1 coding sequence (table 3). For amplification, the PrimeSTAR Max DNA Polymerase (Takara Bio Inc.) was used according to the manufacturer’s recommendation with a 25 pl reaction volume, an annealing temperature of 62 °C and an elongation time of 15 seconds. At least five reactions were pooled for each DNA sample, whereas the cycle number was kept at a minimum to obtain >50 ng per sample. PCR products were purified using the DNA Clean & Concentrator-5 kit (Zymo Research) and by following the manufacturer’ s instructions for dsDNA purification. The DNA was eluted in a 10 mM TRIS buffer with pH 8.0. Prior to sequencing, the DNA was quantified using the Qubit IX dsDNA HS assay kit and a Qubit 4 Fluorometer (both Invitrogen), and the DNA of the insertion libraries from the same assay was pooled in an equimolar ratio. >50 ng of each DNA pool was submitted to the University of Minnesota Genomics Center, where Nextera XT libraries were created, and samples sequenced using a NovaSeq SPrime 150 paired end run.

[0221] Table 2.

[0222] Sequencing statistics.  Table 3

[0223] DNA oligos used in this study(SEQ ID NOS: 60-69).

[0224] Sequencing data analysis and enrichment calculation

[0225] Forward and reverse reads were aligned individually using a DIP-seq pipeline (76), slightly modified for SPINE compatibility and for updated python packages. The code for handling data from domain insertion library sequencing is available at: https: / / github.com / SavageLab / dipseq. This pipeline results in .csv spreadsheets (available as processed data, along scripts to reproduce manuscript figures, which are available at https: / / github / com / Schmidt-lab / AAV_Insertion_Profiling) indicating insertion position, direction, and whether it was in frame. Fitness was calculated from the frequency of a given VP1 variant (;) after packaging (5) relative to the frequency of that variant in the input library (w), normalized to wildtype AAV (wt):

[0226] Fitness standard error for each variant was calculated assuming a Poisson distribution.

[0227] 1 1 1 1

[0228] Counti s+ 0.5 Counti u+ 0.5 Countwt.s+ 0.5 Countwt u+ 0.5

[0229] Western blot

[0230] For VP protein analysis 3x109- 1x1010vg of purified AAV were mixed with 12.5 pl 4x Laemmli Sample Buffer (Bio-Rad, supplemented with 10% 2-mercaptoethanol) and topped up to a final volume of 50 pl with PBS. Samples were denatured for 10 minutes at 95 °C and afterwards chilled on ice. Protein samples, and 5 pl of the Precision Plus Protein Dual Color Standard (BioRad), were separated by molecular weight on a 7.5 % precast polyacrylamide gel (Bio-Rad) in Tris / Glycine / SDS Electrophoresis Buffer (Bio-Rad) for 85 minutes at 120 V. Next, proteins were transferred to a nitrocellulose membrane (pore size 0.45 pm; Thermo Scientific) in an ice-cold blotting buffer (25 mM Tris Base, 96 mM glycine, 20 % methanol) for 80 minutes at 110 V. The membrane was washed once in TBS-T (20 mM Tris Base, 137 mM NaCl, pH 7.6, 0.05% Tween- 20) and incubated in 5 % skim milk solution in TBS-T for two hours at room temperature on a slow shaker to block non-specific binding. A primary antibody detecting either all three VP proteins (anti- AAV VP1 / VP2 / VP3 mouse monoclonal, Bl, supernatant, Progen) or only VP1 (anti- AAV VP1 mouse monoclonal, A, lyophilized, purified, Progen) was diluted 1:250 in 5 % skim milk solution in TBS-T and incubated overnight at 4 °C. The next day, the membrane was washed four times for 5 minutes in TBS-T on a shaker, before the secondary antibody was added (anti-mouse IgG-peroxidase antibody produced in goat, Sigma- Aldrich). The secondary antibody was diluted 1:50,000 in 5% skim milk solution in TBS-T and incubated for two hours at room temperature on a slow shaker. Afterwards, the membrane was washed again four times for 5 minutes in TBS-T at room temperature to remove unbound antibodies, before the SuperSignal West Dura Extended Duration Substrate kit solution (Thermo Scientific) was applied and incubated for two minutes at room temperature. The chemiluminescence signal was detected with an Amersham Imager 600 (GE Healthcare) using exposure times between one second and ten minutes, depending on the signal intensities. Quantification of VP expression was done using ImageJ.

[0231] Differential Scanning Fluorimetry (DSF)

[0232] First, 5,000X SYPRO Orange dye (Invitrogen) was diluted 1:100 in PBS with 5 % glycerol. Then, each sample was prepared by mixing 50X SYPRO Orange dye 1:10 with >5x109vg of purified AAV samples in PBS with 5 % glycerol to a final volume of 25 pl or 50 pl. Samples were mixed by pipetting up and down and pipetted into a 0.1 ml MicroAmp Fast Optical 96-Well Reaction Plate (Applied Biosystems). The plate was sealed with an optical adhesive cover (Applied Biosystems) and spun down for two minutes at 1,000 xg. A melt curve experiment was run on a QuantStudio5 Real-Time PCR instrument (Applied Biosystems) using the xl-m4 filter set (excitation filter: 470 nm / 15nm, emission filter: 623 nm / 14nm) and the following settings: 30 °C for two minutes, temperature increase from 30 °C to 99 °C in 0.5 °C and 30 seconds increments, and a final incubation step of two minutes at 99 °C. Lysozyme at a concentration of 0. 1 mg / ml with a determined melting temperature of 70 °C was used as a reference control within each run. Post processing of the melt-curve data was done in MATLAB R2021a (The Mathworks Inc., Natick, Massachusetts). A smoothing spline (smoothing parameter p=0.9) was fitted to the data before calculating the numerical gradient 6(Fluorescence) / 6(Temperature).

[0233] Transmission electron microscopy

[0234] To quantify the empty to full capsid ratio, negative staining and transmission electron microscopy was performed by the Characterization Facility, University of Minnesota. At least 300 AAV particles were manually counted per sample.

[0235] Infectivity assay of cysteine mutants and WDV variants

[0236] 75,000 HEK293FT cells were seeded per well of a 24-well plate using 0.5 ml culturing media. If applicable, cells were transfected with 100 ng GFP-GPI plasmid while seeding. The next day, media was exchanged, and cells transduced with crude lysates or purified AAV at the indicated MOIs. 48 hours post transduction cells were prepared for flow cytometry as follows. Media was aspirated, cells were washed with 500 pl PBS, 250 pl Accutase solution (Sigma- Aldrich) was added and incubated at room temperature until all cells detached. Cell suspension was transferred into a microcentrifuge tube and centrifuged for three minutes at 400 xg to pellet cells. Cells were washed two times with 300 pl PBS, before resuspended in 600 pl flow cytometry buffer (PBS supplemented with 5 mM EDTA and 2.5 % FBS) and passed through a 35 pm cell strainer to avoid cell clumps. Flow cytometry was performed either on a LSRFortessa X-20 or a FACSymphony A3 Cell Analyzer (both BD Biosciences) equipped with 561 nm and 488 nm lasers to detect tdTomato and GFP positive cells, respectively. Minimum 10,000 single cell events were recorded per sample. Data analysis was performed using the Flowlo 10.8.0 software (BD Biosciences).

[0237] AADV-DJ fitness (green is residue better)

[0238] 618 0.339840045

[0239] 619 0.359591273

[0240] 621 0.326996723

[0241] 631 0.479188539

[0242] 632 0.357827682

[0243] 634 0.368064318

[0244] 635 0.323942369

[0245] 637

[0246] 639

[0247] 640 0.3 15300684

[0248] 641 0.433638403

[0249] 642 0.529220168

[0250] 643 0.302692676 0.352844371

[0251] 0.320206932 702 0.307600768

[0252] 703 0.40429721 1

[0253] 704 0.349797292

[0254] 705 0.418279786

[0255] 706 0.327908168

[0256] 707

[0257] 708 0.305703207

[0258] 709

[0259] 710

[0260] 711

[0261] 715

[0262] 723

[0263] 724

[0264] 725

[0265] 726

[0266] 727

[0267] An exemplary parental capsid sequence is (AAV-DJ) (SEQ ID NO: 70):

[0268] MAADGYLPDWLEDTLSEGIRQWWKLKPGPPPPKPAERHKDDSRGLVLPGYKYLGPFN

[0269] GLDKGEPVNEADAAALEHDKAYDRQLDSGDNPYLKYNHADAEFQERLKEDTSFGGNL

[0270] GRAVFQAKKRLLEPLGLVEEAAKAAPGKKRPVEHSPVEPDSSSGTGKAGQQPARKRLN

[0271] FGQTGDADSVPDPQP1GEPPAAPSGVGSLTKAAGGGAPLADNNEGADGVGNSSGNWH

[0272] CDSTWLGDRVITTSTRTWALPTYNNHLYKQISNSTSGGSSNDNAYFGYSTPWGYFDFN

[0273] RFHCHFSPRDWQRLINNNWGFRPKRLSFKLFNIQVKEVTQNEGTKTIANNLTSTIQVFTD

[0274] SEYQLPYVLGSAHQGCLPPFPADVFMIPQYGYLTLNNGSQAVGRSSFYCLEYFPSQMLR

[0275] TGNNFQFTYTFEDVPFHSSYAHSQSLDRLMNPLIDQYLYYLSRTQTTGGTTNTQTLGFS

[0276] QGGPNTMANQAKNWLPGPCYRQQRVSKTSADNNNSEYSWTGATKYHLNGRDSLVNP

[0277] GPAMASHKDDEEKFFPQSGVLIFGKQGSEKTNVDIEKVMITDEEEIRTTNPVATEQYGS

[0278] VSTNLQRGNRQAATADVNTQGVLPGMVWQDRDVYLQGPIWAKIPHTDGHFHPSPLMG

[0279] GFGLKHPPPQILIKNTPVPADPPTTFNQSKLNSFITQYSTGQVSVEIEWELQKENSKRWNP

[0280] EIQYTSNYYKSTSVDFAVNTEGVYSEPRPIGTRYLTRNL

[0281] Exemplary VP1 (AAV-DJ) sequences: atggctgccgatggttatcttccagattggctcgaggacactctctctgaaggaataagacagtggtggaagctcaaacctggcccaccac caccaaagcccgcagagcggcataaggacgacagcaggggtcttgtgcttcctgggtacaagtacctcggacccttcaacggactcgac aagggagagccggtcaacgaggcagacgccgcggccctcgagcacgacaaagcctacgaccggcagctcgacagcggagacaacc cgtacctcaagtacaaccacgccgacgccgagttccaggagcggctcaaagaagatacgtcttttgggggcaacctcgggcgagcagtc ttccaggccaaaaagaggcttcttgaacctcttggtctggttgaggaagcggctaagacggctcctggaaagaagaggcctgtagagcact ctcctgtggagccagactcctcctcgggaaccggaaaggcgggccagcagcctgcaagaaaaagattgaattttggtcagactggagac gcagactcagtcccagaccctcaaccaatcggagaacctcccgcagccccctcaggtgtgggatctcttacaatggctgcaggcggtggc gcaccaatggcagacaataacgagggcgccgacggagtgggtaattcctcgggaaattggcattgcgattccacatggatgggcgacag agtcatcaccaccagcacccgaacctgggccctgcccacctacaacaaccacctctacaagcaaatctccaacagcacatctggaggatc ttcaaatgacaacgcctacttcggctacagcaccccctgggggtattttgactttaacagattccactgccacttttcaccacgtgactggcag cgactcatcaacaacaactggggattccggcccaagagactcagcttcaagctcttcaacatccaggtcaaggaggtcacgcagaatgaa ggcaccaagaccatcgccaataacctcaccagcaccatccaggtgtttacggactcggagtaccagctgccgtacgttctcggctctgccc accagggctgcctgcctccgttcccggcggacgtgttcatgattccccagtacggctacctaacactcaacaacggtagtcaggccgtggg acgctcctccttctactgcctggaatactttccttcgcagatgctgagaaccggcaacaacttccagtttacttacaccttcgaggacgtgcctt tccacagcagctacgcccacagccagagcttggaccggctgatgaatcctctgattgaccagtacctgtactacttgtctcggactcaaaca acaggaggcacgacaaatacgcagactctgggcttcagccaaggtgggcctaatacaatggccaatcaggcaaagaactggctgccag gaccctgttaccgccagcagcgagtatcaaagacatctgcggataacaacaacagtgaatactcgtggactggagctaccaagtaccacc tcaatggcagagactctctggtgaatccgggcccggccatggcaagccacaaggacgatgaagaaaagttttttcctcagagcggggttct catctttgggaagcaaggctcagagaaaacaaatgtggacattgaaaaggtcatgattacagacgaagaggaaatcaggacaaccaatcc cgtggctacggagcagtatggttctgtatctaccaacctccagagaggcaacagacaagcagctaccgcagatgtcaacacacaaggcgt tcttccaggcatggtctggcaggacagagatgtgtaccttcaggggcccatctgggcaaagattccacacacggacggacattttcacccc tctcccctcatgggtggattcggacttaaacaccctccgcctcagatcctgatcaagaacacgcctgtacctgcggatcctccgaccaccttc aaccagtcaaagctgaactctttcatcacccagtattctactggccaagtcagcgtggagatcgagtgggagctgcagaaggaaaacagc aagcgctggaaccccgagatccagtacacctccaactactacaaatctacaagtgtggactttgctgttaatacagaaggcgtgtactctga accccgccccattggcacccgttacctcacccgtaatctgtaa (SEQ ID NO: 80)

[0282] MAADGYLPDWLEDTLSEGIRQWWKLKPGPPPPKPAERHKDDSRGLVLPGYKYLGPFN GLDKGEPVNEADAAALEHDKAYDRQLDSGDNPYLKYNHADAEFQERLKEDTSFGGNL GRAVFQAKKRLLEPLGLVEEAAKTAPGKKRPVEHSPVEPDSSSGTGKAGQQPARKRLN FGQTGDADSVPDPQPIGEPPAAPSGVGSLTMAAGGGAPMADNNEGADGVGNSSGNWH CDSTWMGDRVITTSTRTWALPTYNNHLYKQISNSTSGGSSNDNAYFGYSTPWGYFDFN RFHCHFSPRDWQRLINNNWGFRPKRLSFKLFNIQVKEVTQNEGTKTIANNLTSTIQVFTD SEYQLPYVLGSAHQGCLPPFPADVFMIPQYGYLTLNNGSQAVGRSSFYCLEYFPSQMLR TGNNFQFTYTFEDVPFHSSYAHSQSLDRLMNPLIDQYLYYLSRTQTTGGTTNTQTLGFS QGGPNTMANQAKNWLPGPCYRQQRVSKTSADNNNSEYSWTGATKYHLNGRDSLVNP GPAMASHKDDEEKFFPQSGVLIFGKQGSEKTNVDIEKVMITDEEEIRTTNPVATEQYGS VSTNLQRGNRQAATADVNTQGVLPGMVWQDRDVYLQGPIWAKIPHTDGHFHPSPLMG GFGLKHPPPQILIKNTPVPADPPTTFNQSKLNSFITQYSTGQVSVEIEWELQKENSKRWNP EIQYTSNYYKSTSVDFAVNTEGVYSEPRPIGTRYLTRNL* (SEQ ID NO: 81)

[0283] References

[0284] 1. S. Russell, J. Bennett, J. A. Wellman, D.C. Chung, Z.F. Yu, A. Tillman, J. Wittes, J.

[0285] Pappas, O. Elci, S. McCague, D. Cross, K.A. Marshall, J. Walshire, T.L. Kehoe, H. Reichert,

[0286] M. Davis, L. Raffini, L.A. George, F.P. Hudson, L. Dingfield, X. Zhu, J.A. Haller, E.H. Sohn, V.B. Mahajan, W. Pfeifer, M. Weckmann, C. Johnson, D. Gewaily, A. Drack, E. Stone, K.

[0287] Wachtel, F. Simonelli, B.P. Leroy, J.F. Wright, K.A. High, and A.M. Maguire. Efficacy and safety of voretigene neparvovec (AAV2-hRPE65v2) in patients with RPE65-mediated inherited retinal dystrophy: a randomised, controlled, open-label, phase 3 trial. Lancet. 390, 849-860 (2017). 2. S.M. Hoy. Onasemnogene Abeparvovec: First Global Approval. Drugs. 79, 1255-1262 (2019).

[0288] 3. S. Yla-Herttuala. Endgame: glybera finally recommended for approval as the first gene therapy drug in the European union. Mol Ther. 20, 1831-1832 (2012).

[0289] 4. D. Wang, P.W.L. Tai, and G. Gao. Adeno-associated virus vector as a platform for gene therapy delivery. Nat Rev Drug Discov. 18, 358-378 (2019).

[0290] 5. C. Li, and R.J. Samulski. Engineering adeno-associated virus vectors for gene therapy. Nat Rev Genet. 21, 255-272 (2020).

[0291] 6. F.T. Jay, C.A. Laughlin, and B.J. Carter. Eukaryotic translational control: adeno- associated virus protein synthesis is affected by a mutation in the adenovirus DNA-binding protein. Proc Natl Acad Sci U S A. 78, 2927-2931 (1981).

[0292] 7. F. Sonntag, K. Schmidt, and J. A. Kleinschmidt. A viral assembly factor promotes AAV2 capsid formation in the nucleolus. Proc Natl Acad Sci U S A. 107, 10220-10225 (2010).

[0293] 8. P.J. Ogden, E.D. Kelsic, S. Sinai, and G.M. Church. Comprehensive AAV capsid fitness landscape reveals a viral gene and enables machine-guided design. Science. 366, 1139-1143 (2019).

[0294] 9. R.M. Buller, and J.A. Rose. Characterization of adenovirus-associated virus-induced polypeptides in KB cells. J Virol. 25, 331-338 (1978).

[0295] 10. F.B. Johnson, H.L. Ozer, and M.D. Hoggan. Structural proteins of adenovirus-associated virus type 3. J Virol. 8, 860-863 (1971).

[0296] 11. J.A. Rose, J.V. Maizel, J.K. Inman, and A.J. Shatkin. Structural proteins of adenovirus- associated viruses. J Virol. 8, 766-770 (1971).

[0297] 12. J. Snijder, M. van de Waterbeemd, E. Damoc, E. Denisov, D. Grinfeld, A. Bennett, M. Agbandje-McKenna, A. Makarov, and A.J. Heck. Defining the stoichiometry and cargo load of viral and bacterial nanoparticles by Orbitrap mass spectrometry. J Am Chem Soc. 136, 7295-7299 (2014).

[0298] 13. T.P. Worner, A. Bennett, S. Habka, J. Snijder, O. Friese, T. Powers, M. Agbandje- McKenna, and A.J.R. Heck. Adeno-associated virus capsid assembly is divergent and stochastic. Nat Commun. 12, 1642 (2021).

[0299] 14. Q. Xie, W. Bu, S. Bhatia, J. Hare, T. Somasundaram, A. Azzi, and M.S. Chapman. The atomic structure of adeno-associated virus (AAV-2), a vector for human gene therapy. Proc Natl Acad Sci U S A. 99, 10405-10410 (2002). 15. L. Govindasamy, E. Padron, R. McKenna, N. Muzyczka, N. Kaludov, J.A. Chiorini, and M. Agbandje-McKenna. Structurally mapping the diverse phenotype of adeno-associated virus serotype 4. J Virol. 80, 11556-11570 (2006).

[0300] 16. S. Bieker, F. Sonntag, and J.A. Kleinschmidt. Mutational analysis of narrow pores at the fivefold symmetry axes of adeno-associated virus type 2 capsids reveals a dual role in genome packaging and activation of phospholipase A2 activity. J Virol. 79, 2528-2540 (2005).

[0301] 17. S. Kronenherg, B. Bottcher, C.W. von der Lieth, S. Bieker, and J.A. Kleinschmidt. A conformational change in the adeno-associated virus type 2 capsid leads to the exposure of hidden VP1 N termini. J Virol. 79, 5296-5303 (2005).

[0302] 18. H.J. Nam, B.L. Gurda, R. McKenna, M. Potter, B. Byrne, M. Salganik, N. Muzyczka, and M. Agbandje-McKenna. Structural studies of adeno-associated virus serotype 8 capsid transitions associated with endosomal trafficking. J Virol. 85, 11791-11799 (2011).

[0303] 19. B. Venkatakrishnan, J. Yarbrough, J. Domsic, A. Bennett, B. Bothner, O.G. Kozyreva, R.J. Samulski, N. Muzyczka, R. McKenna, and M. Agbandje-McKenna. Structure and dynamics of adeno-associated virus serotype 1 VPl-unique N-terminal domain and its role in capsid trafficking. J Virol. 87, 4974-4984 (2013).

[0304] 20. S.F. Cotmore, and P. Tattersail. Parvoviruses: Small Does Not Mean Simple. Annu Rev Virol. 1, 517-537 (2014).

[0305] 21. Z. Zadori, J. Szelei, M.C. Lacoste, Y. Li, S. Gariepy, P. Raymond, M. Allaire, I.R. Nabi, and P. Tijssen. A viral phospholipase A2 is required for parvovirus infectivity. Dev Cell. 1, 291-302 (2001).

[0306] 22. J.C. Grieger, S. Snowdy, and R.J. Samulski. Separate basic region motifs within the adeno- associated virus capsid proteins are essential for infectivity and assembly. J Virol. 80, 5199- 5210 (2006).

[0307] 23. J.M. Riyad, and T. Weber. Intracellular trafficking of adeno-associated virus (AAV) vectors: challenges and future directions. Gene Ther. 28, 683-696 (2021).

[0308] 24. D. Grimm, J.S. Lee, L. Wang, T. Desai, B. Akache, T.A. Storm, and M.A. Kay. In vitro and in vivo gene therapy vector evolution via multispecies interbreeding and retargeting of adeno-associated viruses. J Virol. 82, 5887-5911 (2008).

[0309] 25. E. Zinn, S. Pacouret, V. Khaychuk, H.T. Turunen, L.S. Carvalho, E. Andres-Mateo s, S. Shah, R. Shelke, A.C. Maurer, E. Plovie, R. Xiao, and L.H. Vandenberghe. In Silico Reconstruction of the Viral Evolutionary Lineage Yields a Potent Gene Therapy Vector. Cell Rep. 12, 1056-1068 (2015). 26. K.Y. Chan, M.J. Jang, B.B. Yoo, A. Greenbaum, N. Ravi, W.L. Wu, L. Sanchez- Guardado, C. Lois, S.K. Mazmanian, B.E. Deverman, and V. Gradinaru. Engineered AAVs for efficient noninvasive gene delivery to the central and peripheral nervous systems. Nat Neurosci. 20, 1172-1179 (2017).

[0310] 27. D. Dalkara, L.C. Byrne, R.R. Klimczak, M. Visel, L. Yin, W.H. Merigan, J.G. Flannery, and D.V. Schaffer. In vivo-directed evolution of a new adeno-associated virus for therapeutic outer retinal gene delivery from the vitreous. Sci Trans! Med. 5, 189ra76 (2013).

[0311] 28. K.H. Warrington, O.S. Gorbatyuk, J.K. Harrison, S.R. Opie, S. Zolotukhin, and N. Muzyczka. Adeno-associated virus type 2 VP2 capsid protein is nonessential and can tolerate large peptide insertions at its N terminus. J Virol. 78, 6595-6609 (2004).

[0312] 29. A. Asokan, J.S. Johnson, C. Li, and R.J. Samulski. Bioluminescent virion shells: new tools for quantitation of AAV vector dynamics in cells and live animals. Gene Ther. 15, 1618-1622 (2008).

[0313] 30. R.C. Miinch, H. Janicki, I. Volker, A. Rasbach, M. Hallek, H. Biining, and C.J. Buchholz. Displaying high-affinity ligands on adeno-associated viral vectors enables tumor cell-specific and safe gene transfer. Mol Ther. 21, 109-118 (2013).

[0314] 31. A.M. Eichhoff, K. Borner, B. Albrecht, W. Schafer, N. Baum, F. Haag, J. Korbelin, M. Trepel, I. Braren, D. Grimm, S. Adriouch, and F. Koch-Nolte. Nanobody -Enhanced Targeting of AAV Gene Therapy Vectors. Mol Ther Methods Clin Dev. 15, 211-220 (2019).

[0315] 32. A.C. Zdechlik, Y. He, E.J. Aird, W.R. Gordon, and D. Schmidt. Programmable Assembly of Adeno-Associated Virus-Antibody Composites for Receptor-Mediated Gene Delivery. Bioconjug Chem. 31, 1093-1106 (2020).

[0316] 33. A. Michels, A.M. Frank, D.M. Gunther, M. Mataei, K. Borner, D. Grimm, J. Hartmann, and C.J. Buchholz. Lentiviral and adeno-associated vectors efficiently transduce mouse T lymphocytes when targeted to murine CD8. Mol Ther Methods Clin Dev. 23, 334-347 (2021).

[0317] 34. J. Judd, F. Wei, P.Q. Nguyen, L.J. Tartaglia, M. Agbandje-McKenna, J.J. Silberg, and J. Suh. Random Insertion of mCherry Into VP3 Domain of Adeno-associated Virus Yields Fluorescent Capsids With no Loss of Infectivity. Mol Ther Nucleic Acids. 1, e54 (2012).

[0318] 35. W. Coyote-Maestas, D. Nedrud, S. Okorafor, Y. He, and D. Schmidt. Targeted insertional mutagenesis libraries for deep domain insertion profiling. Nucleic Acids Res. 48, ell (2020).

[0319] 36. M.G. Mateu. Assembly, stability and dynamics of virus capsids. Arch Biochem Biophys. 531, 65-79 (2013).

[0320] 37. M. Medrano, M.A. Fuertes, A. Valbuena, P.J. Carrillo, A. Rodriguez-Huete, and M.G. Mateu. Imaging and Quantitation of a Succession of Transient Intermediates Reveal the Reversible Self-Assembly Pathway of a Simple Icosahedral Virus Capsid. J Am Chem Soc. 138, 15385-15396 (2016).

[0321] 38. S. Bieker, M. Pawlita, and J.A. Kleinschmidt. Impact of capsid conformation and Repcapsid interactions on adeno-associated virus type 2 genome packaging. J Virol. 80, 810-820 (2006).

[0322] 39. B. Gerlach, J.A. Kleinschmidt, and B. Bbttcher. Conformational changes in adeno- associated virus type 1 induced by genome packaging. J Mol Biol. 409, 427-438 (201 1).

[0323] 40. L.M. Drouin, B. Lins, M. Janssen, A. Bennett, P. Chipman, R. McKenna, W. Chen, N. Muzyczka, G. Cardone, T.S. Baker, and M. Agbandje-McKenna. Cryo-electron Microscopy Reconstruction and Stability Studies of the Wild Type and the R432A Variant of Adeno- associated Virus Type 2 Reveal that Capsid Structural Stability Is a Major Factor in Genome Packaging. J Virol. 90, 8542-8551 (2016).

[0324] 41. J.C. Grieger, V.W. Choi, and R.J. Samulski. Production and characterization of adeno- associated viral vectors. Nat Protoc. 1, 1412-1428 (2006).

[0325] 42. T.F. Lerch, J.K. O’Donnell, N.L. Meyer, Q. Xie, K.A. Taylor, S.M. Stagg, and M.S. Chapman. Structure of AAV-DJ, a retargeted gene therapy vector: cryo-electron microscopy at 4.5 A resolution. Structure. 20, 1310-1320 (2012).

[0326] 43. N.L. Meyer, and M.S. Chapman. Adeno-associated virus (AAV) cell entry: structural insights. Trends Microbiol. 30, 432-451 (2022).

[0327] 44. S. Pillay, N.L. Meyer, A.S. Puschnik, O. Davulcu, J. Diep, Y. Ishikawa, L.T. Jae, J.E. Wosen, C.M. Nagamine, M.S. Chapman, and J.E. Carette. An essential receptor for adeno- associated virus infection. Nature. 530, 108-112 (2016).

[0328] 45. N.L. Meyer, G. Hu, O. Davulcu, Q. Xie, A.J. Noble, C. Yoshioka, D.S. Gingerich, A. Trzynka, L. David, S.M. Stagg, and M.S. Chapman. Structure of the gene therapy vector, adeno-associated virus with its cell receptor, AAVR. Elife. 8, e44707 (2019).

[0329] 46. A. Girod, C.E. Wobus, Z. Zadori, M. Ried, K. Leike, P. Tijssen, J.A. Kleinschmidt, and M. Hallek. The VP1 capsid protein of adeno-associated virus type 2 is carrying a phospholipase A2 domain required for virus infectivity. J Gen Virol. 83, 973-978 (2002).

[0330] 47. D.H. Bryant, A. Bashir, S. Sinai, N.K. Jain, P.J. Ogden, P.F. Riley, G.M. Church, L.J. Colwell, and E.D. Kelsic. Deep diversification of an AAV capsid protein by machine learning. Nat Biotechnol. 39, 691-696 (2021).

[0331] 48. S. Ravindra Kumar, T.F. Miles, X. Chen, D. Brown, T. Dobreva, Q. Huang, X. Ding, Y. Luo, P.H. Einarsson, A. Greenbaum, M.J. Jang, B.E. Deverman, and V. Gradinaru. Multiplexed Cre-dependent selection yields systemic AAVs for targeting distinct brain cell types. Nat Methods. 17, 541-550 (2020).

[0332] 49. W. Coyote-Maestas, Y. He, C.L. Myers, and D. Schmidt. Domain insertion permissibility- guided engineering of allostery in ion channels. Nat Commun. 10, 290 (2019).

[0333] 50. W. Coyote-Maestas, D. Nedrud, A. Suma, Y. He, K.A. Matreyek, D.M. Fowler, V. Carnevale, C.L. Myers, and D. Schmidt. Probing ion channel functional architecture and domain recombination compatibility by massively parallel domain insertion profiling. Nat Commun. 12, 7114 (2021).

[0334] 51. S. Bershtein, M. Segal, R. Bekerman, N. Tokuriki, and D.S. Tawfik. Robustness-epistasis link shapes the fitness landscape of a randomly drifting protein. Nature. 444, 929-932 (2006).

[0335] 52. L. Perabo, H. Biining, D.M. Koller, M.U. Ried, A. Girod, C.M. Wendtner, J. Enssle, and M. Hallek. In vitro selection of viral vectors with modified tropism: the adeno-associated virus display. Mol Then 8, 151-157 (2003).

[0336] 53. K. Varadi, S. Michelfelder, T. Korff, M. Hecker, M. Trepel, H.A. Katus, J. A. Kleinschmidt, and O.J. Muller. Novel random peptide libraries displayed on AAV serotype 9 for selection of endothelial cell-directed gene transfer vectors. Gene Ther. 19, 800-809 (2012).

[0337] 54. D.A. Waterkamp, O.J. Muller, Y. Ying, M. Trepel, and J.A. Kleinschmidt. Isolation of targeted AAV2 vectors from novel virus display libraries. J Gene Med. 8, 1307-1319 (2006).

[0338] 55. U. Rothbauer, K. Zolghadr, S. Tillib, D. Nowak, L. Schermelleh, A. Gahl, N. Backmann, K. Conrath, S. Muyldermans, M.C. Cardoso, and H. Leonhardt. Targeting and tracing antigens in live cells with fluorescent nanobodies. Nat Methods. 3, 887-889 (2006).

[0339] 56. A. Juillerat, T. Gronemeyer, A. Keppler, S. Gendreizig, H. Pick, H. Vogel, and K. Johnsson. Directed evolution of O6-alkylguanine-DNA alkyltransferase for efficient labeling of fusion proteins with small molecules in vivo. Chem Biol. 10, 313-317 (2003).

[0340] 57. B. Zakeri, J.O. Fierer, E. Celik, E.C. Chittock, U. Schwarz-Linek, V.T. Moy, and M. Howarth. Peptide tag forming a rapid covalent bond to a protein, through engineering a bacterial adhesin. Proc Natl Acad Sci U S A. 109, E690-7 (2012).

[0341] 58. B.A. Everett, L.A. Litzau, K. Tompkins, K. Shi, A. Nelson, H. Aihara, R.L. Evans lii, and W.R. Gordon. Crystal structure of the Wheat dwarf virus Rep domain. Acta Crystallogr F Struct Biol Commun. 75, 744-749 (2019).

[0342] 59. K.N. Lovendahl, A.N. Hayward, and W.R. Gordon. Sequence-Directed Covalent Protein- DNA Linkages in a Single Step Using HUH-Tags. J Am Chem Soc. 139, 7030-7035 (2017). 60. A.T. Smiley, K.J. Tompkins, M.R. Pawlak, A.J. Krueger, R.L. Evans, K. Shi, H. Aihara, and W.R. Gordon. Watson-Crick Base-Pairing Requirements for ssDNA Recognition and Processing in Replication-Initiating HUH Endonucleases. mBio. 14, e0258722 (2023).

[0343] 61. K.J. Tompkins, M. Houtti, L.A. Litzau, E.J. Aird, B.A. Everett, A.T. Nelson, L. Pornschloegl, L.K. Limon-Swanson, R.L. Evans, K. Evans, K. Shi, H. Aihara, and W.R. Gordon. Molecular underpinnings of ssDNA specificity by Rep HUH-endonucleases and implications for HUH-tag multiplexing and engineering. Nucleic Acids Res. 49, 1046-1064 (2021).

[0344] 62. D.M. Fowler, and S. Fields. Deep mutational scanning: a new style of protein science. Nat Methods. 11, 801-807 (2014).

[0345] 63. P.F. Schmit, S. Pacouret, E. Zinn, E. Telford, F. Nicolaou, F. Broucque, E. Andres-Mateos, R. Xiao, M. Penaud-Budloo, M. Bouzelha, N. Jaulin, O. Adjali, E. Ayuso, and L.H. Vandenberghe. Cross-Packaging and Capsid Mosaic Formation in Multiplexed AAV Libraries. Mol Ther Methods Clin Dev. 17, 107-121 (2020).

[0346] 64. O.S. Oliinyk, A.A. Shemetov, S. Pletnev, D.M. Shcherbakova, and V.V. Verkhusha. Smallest near-infrared fluorescent protein evolved from cyanobacteriochrome as versatile tag for spectral multiplexing. Nat Commun. 10, 279 (2019).

[0347] 65. S.M. Stagg, C. Yoshioka, O. Davulcu, and M.S. Chapman. Cryo-electron Microscopy of Adeno-associated Virus. Chem Rev. 122, 14018-14054 (2022).

[0348] 66. G.E. Berry, and L.V. Tse. Virus Binding and Internalization Assay for Adeno-associated Virus. Bio Protoc. 7, e2110 (2017).

[0349] 67. F. Sonntag, S. Bieker, B. Leuchs, R. Fischer, and J.A. Kleinschmidt. Adeno-associated virus type 2 capsids with externalized VP1 / VP2 trafficking domains are generated prior to passage through the cytoplasm and are maintained until uncoating occurs in the nucleus. J Virol. 80, 11040-11054 (2006).

[0350] 68. E.D. Horowitz, M.G. Finn, and A. Asokan. Tyrosine cross-linking reveals interfacial dynamics in adeno-associated viral capsids during infection. ACS Chem Biol. 7, 1059-1066 (2012).

[0351] 69. L. Mclnnes, J. Healy, and J. Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv: 1802.03426. (2018).

[0352] 70. E.E. Large, and M.S. Chapman. Adeno-associated virus receptor complexes and implications for adeno-associated virus immune neutralization. Front Microbiol. 14, 1116896 (2023). 71. H.C. Levy, V.D. Bowman, L. Govindasamy, R. McKenna, K. Nash, K. Warrington, W. Chen, N. Muzyczka, X. Yan, T.S. Baker, and M. Agbandje-McKenna. Heparin binding induces conformational changes in Adeno-associated virus serotype 2. J Struct Biol. 165, 146- 156 (2009).

[0353] 72. N. DiPrimio, A. Asokan, L. Govindasamy, M. Agbandje-McKenna, and R.J. Samulski. Surface loop dynamics in adeno-associated virus capsid assembly. J Virol. 82, 5178-5189 (2008).

[0354] 73. P. Wu, W. Xiao, T. Conlon, J. Hughes, M. Agbandje-McKenna, T. Ferkol, T. Flotte, and N. Muzyczka. Mutational analysis of the adeno-associated virus type 2 (AAV2) capsid gene and construction of AAV2 vectors with altered tropism. J Virol. 74, 8635-8647 (2000).

[0355] 74. S. Hagen, T. Baumann, H.J. Wagner, V. Morath, B. Kaufmann, A. Fischer, S. Bergmann, P. Schindler, K.M. Arndt, and K.M. Muller. Modular adeno-associated virus (rAAV) vectors used for cellular virus-directed enzyme prodrug therapy. Sci Rep. 4, 3759 (2014).

[0356] 75. J. Becker, J. Fakhiri, and D. Grimm. Fantastic AAV Gene Therapy Vectors and How to Find Them-Random Diversification, Rational Design and Machine Learning. Pathogens. 11, 756 (2022).

[0357] 76. D.C. Nadler, S.A. Morgan, A. Flamholz, K.E. Kortright, and D.F. Savage. Rapid construction of metabolite biosensors using domain-insertion profiling. Nat Commun. 7, 12266 (2016).

[0358] 77. C. Plesa, A.M. Sidore, N.B. Lubock, D. Zhang, and S. Kosuri. Multiplexed gene synthesis in emulsions for exploring protein functional landscapes. Science. 359, 343-347 (2018).

[0359] 78. A.F. Rubin, H. Gelman, N. Lucas, S.M. Bajjalieh, A.T. Papenfuss, T.P. Speed, and D.M. Fowler. A statistical framework for analyzing deep mutational scanning data. Genome Biol. 18, 150 (2017).

[0360] 79. K.K. Yang, Z. Wu, and F.H. Arnold. Machine-learning-guided directed evolution for protein engineering. Nat Methods. 16, 687-694 (2019).

[0361] 80. C.E. Wobus, B. Hiigle-Dbrr, A. Girod, G. Petersen, M. Hallek, and J.A. Kleinschmidt. Monoclonal antibodies against the adeno-associated virus type 2 (AAV-2) capsid: epitope mapping and identification of capsid domains involved in AAV-2-cell interaction and neutralization of AAV-2 infection. J Virol. 74, 9281-9293 (2000).

[0362] 81. S.F. Cotmore, V.C. McKie, L.J. Anderson, C.R. Astell, and P. Tattersail. Identification of the major structural and nonstructural proteins encoded by human parvovirus B19 and mapping of their genes by procaryotic expression of isolated genomic fragments. J Virol. 60, 548-557 (1986). 82. B. Kaufmann, A. A. Simpson, and M.G. Rossmann. The structure of human parvovirus Bl 9. Proc Natl Acad Sci U S A. 101, 11628-11633 (2004).

[0363] 83. M. Kawase, M. Momoeda, N.S. Young, and S. Kajigaya. Most of the VP1 unique region of B19 parvovirus is on the capsid surface. Virology. 211, 359-366 (1995).

[0364] 84. C. Ros, M. Gerber, and C. Kempf. Conformational changes in the VPl-unique region of native human parvovirus B19 lead to exposure of internal sequences that play a role in virus neutralization and infectivity. J Virol. 80, 12017-12024 (2006).

[0365] 85. Q. Wang, Z. Wu, J. Zhang, J. Firrman, H. Wei, Z. Zhuang, L. Liu, L. Miao, Y. Hu, D. Li, Y. Diao, and W. Xiao. A Robust System for Production of Superabundant VP1 Recombinant AAV Vectors. Mol Ther Methods Clin Dev. 7, 146-156 (2017).

[0366] 86. R.N. McLaughlin, F.J. Poelwijk, A. Raman, W.S. Gosal, and R. Ranganathan. The spatial architecture of protein function and adaptation. Nature. 491, 138-142 (2012).

[0367] 87. O. Rivoire, K.A. Reynolds, and R. Ranganathan. Evolution-Based Functional Decomposition of Proteins. PLoS Comput Biol. 12, el004817 (2016).

[0368] 88. J.T. Atkinson, A.M. Jones, Q. Zhou, and J.J. Silberg. Circular permutation profiling by deep sequencing libraries created using transposon mutagenesis. Nucleic Acids Res. 46, e76 (2018).

[0369] 89. C.J. Markin, D.A. Mokhtari, F. Sunden, M.J. Appel, E. Akiva, S.A. Longwell, C. Sabatti, D. Herschlag, and P.M. Fordyce. Revealing enzyme functional architecture via high- throughput microfluidic enzyme kinetics. Science. 373, eabf8761 (2021).

[0370] 90. B.E. Deverman, P.L. Pravdo, B.P. Simpson, S.R. Kumar, K.Y. Chan, A. Banerjee, W.L. Wu, B. Yang, N. Huber, S.P. Pasca, and V. Gradinaru. Cre-dependent selection yields AAV variants for widespread gene transfer to the adult brain. Nat Biotechnol. 34, 204-209 (2016).

[0371] 91. A. Girod, M. Ried, C. Wobus, H. Lahm, K. Leike, J. Kleinschmidt, G. Deleage, and M. Hallek. Genetic capsid modifications allow efficient re-targeting of adeno-associated virus type 2. Nat Med. 5, 1052-1056 (1999).

[0372] 92. S. Michelfelder, M.K. Lee, E. deLima-Hahn, T. Wilmes, F. Kaul, O. Muller, J. A. Kleinschmidt, and M. Trepel. Vectors selected from adeno-associated viral display peptide libraries for leukemia cell-targeted cytotoxic gene therapy. Exp Hematol. 35, 1766-1776 (2007).

[0373] 93. O.J. Muller, F. Kaul, M.D. Weitzman, R. Pasqualini, W. Arap, J.A. Kleinschmidt, and M. Trepel. Random peptide libraries displayed on adeno-associated virus to select for targeted gene therapy vectors. Nat Biotechnol. 21, 1040-1046 (2003). 94. R.C. Munch, A. Muth, A. Muik, T. Friedel, J. Schmatz, B. Dreier, A. Trkola, A. Pliickthun, H. Biining, and C.J. Buchholz. Off-target-free gene delivery by affinity-purified receptor- targeted viral vectors. Nat Commun. 6, 6246 (2015).

[0374] 95. C. Engler, R. Kandzia, and S. Marillonnet. A one pot, one step, precision cloning method with high throughput capability. PLoS One. 3, e3647 (2008).

[0375] 96. S. Zolotukhin, B.J. Byrne, E. Mason, I. Zolotukhin, M. Potter, K. Chesnut, C. Summerford, R.J. Samulski, and N. Muzycz.ka. Recombinant adeno-associated virus purification using novel methods improves infectious titer and yield. Gene Ther. 6, 973-985 (1999).

[0376] 97. C.Y. Lin, Z. Huang, W. Wen, A. Wu, C. Wang, and L. Niu. Enhancing Protein Expression in HEK-293 Cells by Lowering Culture Temperature. PLoS One. 10, e0123562 (2015).

[0377] 98. M. Baek, F. DiMaio, I. Anishchenko, J. Dauparas, S. Ovchinnikov, G.R. Lee, J. Wang, Q. Cong, L.N. Kinch, R.D. Schaeffer, C. Millan, H. Park, C. Adams, C.R. Glassman, A. DeGiovanni, J.H. Pereira, A.V. Rodrigues, A.A. van Dijk, A.C. Ebrecht, D.J. Opperman, T. Sagmeister, C. Buhlheller, T. Pavkov-Keller, M.K. Rathinaswamy, U. Dalwadi, C.K. Yip, J.E. Burke, K.C. Garcia, N. V. Grishin, P.D. Adams, R.J. Read, and D. Baker. Accurate prediction of protein structures and interactions using a three-track neural network. Science. 373, 871- 876 (2021).

[0378] Example 2

[0379] Unlocking Precision Gene Therapy: Harnessing AAV Tropism with Nanobody Swapping at Capsid Hotspots

[0380] Introduction

[0381] Adeno-associated virus (AAV) is a compact 25 nm virus known for its favorable clinical characteristics, such as low pathogenicity and the ability to induce long-term expression in both dividing and non-dividing cells. These features make it a promising candidate for applications in gene and cell therapy (reviewed in (1)). Its single-stranded DNA genome, spanning 4.7 kb, encodes two genes, rep and cap, flanked by inverted terminal repeats. The cap gene produces three viral proteins (VP1, VP2, and VP3) from the same open reading frame. VP2 is a N-terminal truncated version of VP1, and VP3 is a further truncated version of VP2 (2, 3). The capsid, formed by 60 VP monomers, exhibits an icosahedral structure with an average ratio of 1:1:10 for VP1, VP2, and VP3, respectively (4-6).

[0382] The capsid’s distinctive features include a protruding, cylindric pore at the 5-fold interface, a valley extending toward the 2-fold interface around the pore, and protrusions at the 3-fold interface, which play a role in mediating target cell receptor binding (7). Despite high conservation in the overall structure and topology across serotypes, structural analyses have identified nine variable regions (VR1-9) on the capsid surface (8). VR4 and VR8 have been particularly targeted in AAV capsid engineering to evade neutralizing antibody binding (9) and re-direct viral tropism (10-15).

[0383] While VR8 has been predominantly utilized for peptide insertions, a library screen by Judd et al. (16) revealed that VR4 can accommodate the fluorescent protein mCherry. Subsequent to this discovery, various protein domains with re-targeting capabilities, including DARPins (17), HUH-tags (18), and nanobodies (19, 20), were successfully incorporated into VR4. Nanobodies, originating from camelids and characterized by their small size (15 kDa), specificity, stability, and ability to serve as targeting ligands for chemotherapy drugs, radionuclides, or toxins (21), stand out among these options. Therefore, there is considerable interest in optimizing the incorporation of nanobodies into the AAV capsid to enhance cell type specific AAV targeting.

[0384] We have had performed a domain insertion library screen by incorporating domains with re-targeting abilities, including a GFP nanobody, in between every two amino acid residues of VP1 protein of AAV-DJ (22). We demonstrated that nanobody insertions are tolerated not only in the tip of the 3-fold protrusion (VR4), but also several positions lining the 2-fold valley as well as the 5-fold interface of an AAV capsid. While insertion of a GFP nanobody into AAV increased infectivity towards cells expressing GFP on their cell surface, it was unclear if this nanobody- mediated re-targeting generalizes to different nanobodies, implying modularity, or whether any optimization is required to achieve high infection specificity.

[0385] To address these questions, we selected seven positions near the 2-fold valley and 5-fold interface alongside the benchmark insertion position in VR4. Insertions were made into either the VP1 or VP2 protein of AAV-DJ with two different types of linkers. With an eye toward clinically relevant AAV re-targeting, we chose a nanobody targeting fibroblast activating protein (FAP), which has emerged as a promising cancer target in recent years (23-25). While most of the FAP nanobody insertion variants have weaker infection efficacy on cells lacking the FAP receptor (i.e., off-target cells), most of the variants (six out of the eight positions tested) surpassed the DJ control in FAP receptor-positive (‘on-target’) cells with a specificity gain of up to 18-fold for the best variant: VP2-N262 with asymmetric linkers. These findings reinforce the feasibility of nanobody- mediated AAV targeting for cell and gene therapy applications.

[0386] Materials and Methods

[0387] Cloning

[0388] Oligos and gBlocks were obtained from IDT, restriction enzymes and T4 DNA ligase from NEB and PCRs were done using the PrimeSTAR Max DNA polymerase from Takara Bio. PCR products were purified using the DNA Clean & Concentrator- 5 Kit (Zymo Research) or the Zymoclean Gel DNA Extraction Kit (Zymo Research) if PCR products were analyzed on 1 %TAE agarose gels. Post cloning, plasmids were transformed into NEB® Stable Competent E. coli cells, before plated on LB plates containing carbenicillin at a concentration of 100 pg / ml. Plasmids were isolated using the Zyppy Plasmid Miniprep Kit (Zymo Research), according to the manufacturer’s instructions. All plasmids used in this study are listed in Table SI.

[0389] Table St For the cloning of the FAP nanobody insertion variants, the BsmBI restriction site was eliminated from the plasmid containing the rep2-capDJ (AAV-DJ) gene sequences by introducing a silent mutation using mutagenesis PCR at position D178 of the cap-DJ gene. The plasmids DJ-VP1 and VP2 were obtained by mutating the start codons for VP2 / 3 (T138A, M203K, M211L, M235L) and VP1 / 3 (MIK, M203K, M211L, M235L), respectively. For the AAV productions obligatory plasmids complementing the missing VPs, DJ-VP2 / 3 and DJ-VP1 / 3, were cloned similarly by mutating the start codons for VP1 (MI K) and VP2 (T138A), respectively. FAP or GFP nanobody insertion plasmids were generated using golden gate assembly (26) and the BsmBI restriction enzyme. FAP (Figure S5) and GFP (27) nanobody sequences were human codon-optimized and ordered as gBlocks with either symmetric or asymmetric linkers. Nanobody and linker sequences are given in Figure SI. Post cloning, sequences were verified by Plasmid-EZ sequencing (Azenta Life Sciences).

[0390] Tissue Culture

[0391] The cell lines CWR-Rl-enzalutamide resistant / luciferase+(stably expressing a firefly luciferase), and CWR-Rl-enzalutamide resistant / luciferase plus FAP (additionally expressing a human FAP receptor) were provided by the LeBeau lab from the University of Wisconsin (28). Both cell lines and 293 AAV cells (Cell Biolabs) were cultured in DMEM (Gibco) supplemented with 10 % fetal bovine serum (Gibco), 4.5 g / L D-glucose, L-glutamine, 110 mg / L sodium pyruvate, and 100 U per mL penicillin / 100 pg per mL streptomycin (Gibco). Media for Rl-FAP cells was additionally supplemented with 3 pg / mL puromycin (ApexBio Technology). Cells were kept in a humidified incubator at 37 °C and 5 % CO2 and passaged every 2-4 days when reaching a confluency of 70-80 %.

[0392] AAV crude lysate production

[0393] 293AAV cells were seeded into 6-well plates at a density of 500,000 cells per well. 24 hours later, cells were transfected with 2.5 pg DNA using PEI and an equimolar ratio of the plasmids necessary for the respective AAV production: (i) an Adenohelper plasmid, (ii) a nanoLuciferase encoding plasmid, (iii) a plasmid encoding rep2 and capDJ with start codons of either only VP1 or only VP2, but with a FAP-nanobody insertion, and (iv) a plasmid encoding rep2 and capDJ encoding the VP proteins needed for complementation. Three days post transfection, cells were harvested by flushing off the cells by pipetting and spun down for 5 min at 400 xg. Cells were resuspended in PBS (pH 7.4, Gibco) and then subjected to five freeze and thaw cycles by alternating between liquid nitrogen and a 37 °C water bath. Cell debris was pelleted by centrifugation at 17,000 xg at 4 °C for 10 min. The supernatant containing the AAV particles was stored at -20 °C until use. qPCR

[0394] Production titers of crude lysate samples were determined as follows. 2 pl of the crude lysates were mixed with PBS, supplemented with 2mM MgC12, and 0.1 pl ultrapure Benzonase Nuclease (Sigma-Aldrich) was added. Samples were incubated at 37 °C for 30 min to digest DNA that was not protected by AAV capsids. Next, 5 pl lOx Proteinase K buffer (100 nM Tris-HCl, pH 8.0, 10 mM EDTA, and 10 % SDS) and 1 pl Proteinase K (20 mg / ml; Zymo Research) were added inhibiting the Benzonase and digesting proteins including the AAV capsid to free the ssDNA. Afterwards, samples incubated for 20 min at 50 °C, followed by heat inactivation of the Proteinase K for 5 min at 95 °C. The viral ssDNA was purified using the DNA Clean & Concentrator-5 Kit (Zymo Research) according to the manufacturer’s instructions for ssDNA purification. All samples were diluted 1:500 in H2O prior to qPCR, which was run using a QuantStudio5 Real- Time PCR System (Applied Biosystems) and by using the PowerUp SYBR Green Master Mix (Applied Biosystems), following the manufacturer’s instructions. A primer set binding within the CMV-enhancer (forward: AACGCCAATAGGGACTTTCC (SEQ ID NO: 71), reverse: GGGCGTACTTGGCATATGAT (SEQ ID NO 72) (29) of the transgene expression cassette was used. To calculate the viral titer in vg / ml, a plasmid standard at a known concentration also containing a CMV-enhancer was used.

[0395] Luciferase assay

[0396] R1 and Rl-FAP cells were seeded into 96-well plates at a density of 12,500 cells per well, while SK-MEL-24 cells were plated at a density of 10,000 cells per well. The next day, the media was replaced, and cells were transduced at the indicated multiplicity of infection (MOI) with AAV crude lysates. 48h post transduction, the media was aspirated, cells washed with 100 pl PBS (pH 7.4, Gibco) per well and then 25 pl PBS and 25 pl of the Nano-Gio® Luciferase Assay System reagent (Promega) were added. Cells were incubated for 15 min at room temperature on a shaker at 600 rpm. Afterwards, the suspension was mixed by pipetting up and down before 20 pl were transferred into a white 96-well F-bottom plate (Corning). Luminescence was measured using an Infinite F200 PRO plate reader (Tecan) by using an integration time of 100 ms. Luciferase assays were conducted with three technical replicates per sample. Flow cytometry

[0397] To analyze FAP expression two million Rl, Rl-FAP or SK-MEL-24 cells were detached with Accutase solution (Sigma- Aldrich), collected in a 15 ml conical tube and spun down at 400 xg for 3 min at 4 °C. The cell pellets were resuspended in 1 ml cold flow buffer (PBS supplemented with 5 % FBS and 0.1 % sodium azide). Next, the cell suspensions were split in two halves and transferred into cold microcentrifuge tubes and washed two more times with cold 500 pl flow buffer. The unstained samples remained in flow buffer, while the stained samples were resuspended in flow buffer supplemented with the primary antibody (anti-FAP human B12, (24) at a dilution of 1:500. Incubation was done for one hour at 4 °C on an end-over-end rotator. Both cell batches, unstained and stained, were washed three times in cold flow buffer and then the secondary antibody (Goat anti-Human IgG H+L Cross-Adsorbed Secondary Antibody, Alexa Fluor™ 488, ThermoFisher Scientific) at a dilution of 1:500 was added. The incubation of the secondary antibody was done for one hour at 4°C on an end-over-end rotator. Cells were washed another three times in cold flow buffer before passed through a 35 pm cell strainer to avoid cell clumps. Flow cytometry was done on a SONY SH800 flow sorter equipped with a 488 nm laser. Gates for the whole cell population and single cell population were adjusted to the different cell types tested. At least 20,000 single cell events were recorded for each sample.

[0398] Statistics qPCR values were obtained from three independent crude lysate productions. The luciferase data shown were attained by three biological replicates with three technical replicates each. All error bars indicate the standard error of the mean (SEM). For all qPCR and luciferase assay data shown, the differences between DJ and all other capsid variants were tested for statistical significance by one-way ANOVA analysis of variance followed by a Dunnett’s post- hoc test, p-values < 0.05 were considered statistically significant (*p<0.05; **p < 0.01; ***p < 0.001). All p-values are listed in Table S2-5. Statistical analysis was performed in R (version 4.3.2).

[0399] Table S2 p-vaiues of one-way ANOvA with DunneWs post -hoc test from data in Figure 2A-D (as.: not significant).

[0400] Table S3 p-vaius$ of one-way ANOVA with Dunnett's post-hoc test front data in Figure 3A (n.$.: not Significant). Tabie S4 p-vai ues of one-way ANQVA with Dunnett’s post-hoc test front data in Figure S2A-8 (n.s.: not significant). Table S5 p-values of one-way ANOVA with Dunnetfs post-hoc test from data in Figure 3A-D (n.s.. not significant).

[0401] Results

[0402] Vector design of nanobody insertions into the AAV-DJ capsid for FAP-mediated cell targeting

[0403] To enhance the infection efficacy of AAV specifically for FAP receptor-positive cells, but not FAP receptor-negative cells, we inserted a FAP nanobody into the capsid of AAV-DJ (Figure 23 A). We chose in total eight different insertion positions on the capsid surface: three positions in the 2-fold valley (Q259, N262, and Q387), two positions in the protruding pore (N328 and N337), two positions in the depression around the pore (N664 and S670), as well as the previously proven insertion position T456 at the tip of the 3-fold protrusion (Figure 23B; (19, 20)). Since VP3 makes up the majority of the AAV capsid, we restricted the nanobody embeddings to VP1 or VP2, thereby avoiding potential steric hindrance of capsid formation by too many nanobodies on a single capsid. For insertions into VP1 only, the start codons for VP2 (T138A) and VP3 (M203K, M21 IL, and M235L) were mutated prior to introducing the FAP nanobody. Further, we generated a rep2-capDJ plasmid with a mutated start codon of VP1 (MIK), complementing VP2 and VP3 expression during AAV production. Likewise, a VP2 only plasmid by mutating start codons of VP1 and VP3 was generated, as well as a plasmid providing only VP1 and VP3 by mutating the VP2 start codon (Figure 23C). For the nanobody insertions, we chose two different types of linkers: one short, symmetric linker pair (SGGGG (SEQ ID NO: 73 on both sides) and an asymmetric linker pair, with a long N-terminal 5xSGGGG (SEQ ID NO: 74) linker and a short C-terminal GGGGS (SEQ ID NO: 75) linker, which was previously used by Eichhoff et al. for their nanobody insertions into VR4 (19). To interrogate AAV re-targeting in the context of adding targeting nanobodies to the AAV capsid alone, we did not mutate AAV-DJ’ s heparin binding domain (R587-590), which is often done to attenuate capsid binding to heparan sulfate proteoglycans. Altogether, a total of 32 FAP nanobody variants were generated: 8 positions x 2 different VPs x 2 linker sets.

[0404] FAP nanobody insertions boost transduction in FAP receptor-positive cells. All variants, as well as a DJ control and controls for the split VP expression plasmids (VP1 and VP2 / 3 or VP2 and VP 1 / 3 expressed from separate plasmids), were produced as crude lysates packaging a CAG promoter driven NanoLuc payload. Post-production, titers were assessed by qPCR. We found that none of the variants had titers statistically significant different to the DJ control. However, we noted a trend that almost invariably all VP1 insertion variants resulted in higher titers than DJ (Figure S2). Next, we used these crude lysates to infect the human prostate cancer cell line CWR-R1 -enzalutamide resistant / Fluciferase+(Rl ) at an MOI of I xlO3vg / cell, which is FAP receptor-negative as verified by flow cytometry (Figure S3A-C). 48 hours post transduction, a luciferase assay using Furimazine substrate for the NanoLuc was conducted and photon counts normalized to the DJ control (Figure 24A). Note, the firefly luciferase (stably expressed by the cell line) and the NanoLuc (delivered as transgene by the AAV) use fully orthogonal substrates, D-luciferin and Furimazine, respectively (30), meaning that the integrated firefly luciferase does not contribute to measured signal. All VP2 insertion variants and nearly all VP1 insertion variants infected Rl cells significantly less efficient than the DJ parent (Figure 24 A), indicating that the nanobody insertions to some extend negatively impact the uptake and intracellular processing of the AAV. To test our hypothesis that FAP nanobody insertions into the AAV capsid can boost infection of FAP receptor positive cells, the same AAV samples were used to infected Rl-FAP cells, which stably expressed the human FAP receptor (Figure S3D). We found that insertions into positions N262, N328, Q387, T456, N664, and S670 surpassed the infection potency of DJ, independent of the linker type and whether the FAP nanobody insertion was made into VP1 or VP2. The best variant, VP2-N262-asymmetric, infected Rl-FAP cells even ~8-fold better than DJ. Conversely, insertions into positions Q259 and N337 did not enhance infection and showed an equal infection reduction as seen for the assay with the FAP receptor negative Rl cells (Figure 24B). We calculated a FAP nanobody-mediated specificity gain as the ratios of off- to on-target infection efficacy (Rl-FAP I Rl infectivity normalized to AAV-DJ). By this metric, our data reveals an up to - 18-fold improved infection in Rl than in Rl-FAP cells for the VP2-N262-asymmetric variant. Even all other variants, that were able to mediate a FAP nanobody-specific infection in Rl-FAP cells, showed a specificity gain of >4-fold (Figure 24C). Lowering the dose of transduced AAV to 5x102and 1x102vg / cell showed comparable results, demonstrating that the specificity gain is dose independent (Figure S4). Since the Rl-FAP cell line is engineered to overexpress FAP, we turned our attention to a cell line with lower endogenous expression (compared to Rl-FAP), such as the melanoma cell line SK-MEL-24 (Figure S3E). For the infection of SK-MEL-24 cells, the same trend of infection gain was observed. The symmetric VP1 insertion at position T456 stands out with a 5-fold infection increase, but also VP2 insertions at positions N262, N328, T456, and N664 surpassed the DJ infection potency by at least 2-fold (Figure 24D). To test if the observed specificity gains are FAP nanobody-specific, we conducted an isotype control experiment with a GFP nanobody. To this end, two variants (VP1-T456 and VP1-N664, both with symmetric linkers) were produced and tested in a one-on-one comparison to their FAP nanobody counterparts. As expected, there was no GFP nanobody-specific infection gain as seen for the FAP nanobody (Figure S5). Comparison of infection gain from different cell types reveals the most robust nanobody insertion positions.

[0405] The differences and similarities between the two different cell lines tested, Rl-FAP and SK-MEL-24 are illustrated in Figure 25A. Overall, FAP nanobody embeddings into VP2 yielded a higher infection rate on average than into VP1, except for position T456, the previously published benchmark (19). This position seemed to tolerate nanobody insertions irrespective of the linker. Similarly good as position T456 functioned the variants with insertion at positions N262, N328, and N664, showing the most robust infection gains (Figure 25A). For embeddings into positions Q259 and N337 an infection gain was not observed in either cell line. Mapping of the averaged infection gains for each insertion position into VP1 or VP2 onto the capsid structure revealed no consistent tolerability scheme with regards to the different interfaces of the AAV capsid. For example, all three positions located close to each other within the 2-fold valley (Q259, N262, and Q387) reached very different nanobody-mediated infection gains (Figure 25B). Discussion

[0406] AAV has been proven to be a suitable vector to efficiently deliver transgenes for cell and gene therapies (1). Nonetheless, broad tissue tropism of natural serotypes hampers its applicability whenever a cell type-specific targeting is of interest. Different capsid engineering approaches have been applied to tackle this issue, including peptide display (9-15), the recovery of AAV ancestors (31), and the insertion of domains with re-targeting abilities (17, 18, 32). In particular, the embedding of nanobodies into the AAV capsid has been strikingly successful in boosting a cellspecific transduction (19, 20).

[0407] Nanobodies naturally come with outstanding properties. They are small (15kDa), stable, have antibody-like binding affinities, and, on top, a low immunogenicity (33). Previous studies had successfully incorporated ARTC2.2, P2X7, CD38, CD4, and GFP nanobodies into VR4 of VP1 and subsequently showed a nanobody-specific uptake by cells expressing the cognate receptor (18-20). Based on our prior domain insertional profiling screen (22), we postulated that nanobodies can be embedded throughout the capsid’s surface, instead of only into the 3-fold protrusion (VR4). Additional options for nanobody insertions may potentially synergize existing VR4 engineering approaches. In this study, we incorporated a FAP nanobody into eight different positions, covering different surface areas of AAV-DJ. The benchmark position T456, which is part of VR4 was included as a reference. As hypothesized, we could show that AAV capsids can harbor nanobody insertions at various locations, including the 2-fold valley and the protruding 5-fold pore. None of the tested variants showed a reduction in production titer (Figure S2) and six out of the eight chosen positions were able to mediate a FAP receptor-specific infection boost of up to ~ 18-fold (Figure 24). The two different linker pairs tested had very minor influence on whether an insertion variant boosted infection or not (Figure 24), which is surprising regarding the fact that the two termini of a nanobody are ~40 Angstrom apart and the short symmetric linker could cause sheering forces to the capsid. The choice of engineered VP protein made a bigger difference. We observed increased production titers for almost all VP1 variants, whereas VP2 insertions produced similar to the DJ parent (Figure S2). Despite higher production titers, nanobody incorporations into VP1 had a lower transduction efficacy compared to VP2 insertions (Figure 24, 25). We speculate that this effect is caused by a reduced incorporation rate of VPl-nanobody monomers. It has been shown that VP1 is indispensable for the infection process. The unique N -terminus, also known as VPlu, lays inside the capsid and comprises a phospholipase domain (PLA2) as well as a nuclear localization signal. Both play a role in endosomal escape and nuclear entry, respectively (34-37). On the one hand, if the capsid contains fewer VP1 monomers, fewer VPlu overhangs occupy the inner cavity of the capsid, possibly facilitating the more efficient packaging of the DNA cargo. On the other hand, fewer VP1 monomers also result in a reduced infection potency. Consequently, the VP2 monomer, for which no crucial functions are known, appears to be the better choice for nanobody insertions. While Eichoff et al. only tested a single insertion into VP1, Hamann et al. also tested N-terminal fusions of a nanobody to VP2, which was less effective than the known VR4 insertion position of VP1 (19, 20). Supporting the notion that engineering VP2 is the more promising route, a previous study by us also tested VP2 insertions into VR4 and found that this insertion was superior compared to the VP1 equivalent (18).

[0408] Overall, our FAP nanobody insertion into six out of the eight positions (N262, N328, Q387, T456, N664, and S670) resulted in an infection boost in FAP receptor-positive cells (Figure 24). All these positions have in common that they are on the surface of the capsid, but they are located at very different regions, i.e., at the 5-fold pore (N328), 3-fold protrusion (T456), 2-fold valley (N262 and Q387), and valley surrounding the pore (N664 and S670; Figure 25B, C). One could argue that four of these are part of known VRs and therefore more likely to tolerate insertions (N262 is part of VR1, N328 is part of VR2, Q387 is part of VR3, and T456 is part of VR4), but N664 and S670 also showed a receptor-specific infection, and these two positions are neither part of a VR nor a general protrusion (8). Further, three of the position amenable to nanobody insertions (N262, Q387, T456) are known to play a role in the binding of the ubiquitously used AAVR receptor (38, 39). We speculate that insertions right behind these residues are likely to block the AAVR binding and instead promote the FAP receptor binding. Another important receptor used by AAV is heparan sulfate proteoglycan, which is bound by the heparin-binding domain (HBD, R587-R590) of the capsid. Previous studies had removed this binding domain prior to inserting the nanobody to achieve a null-tropism capsid parent to which the infection potency was compared (18-20). We could show here, that even with leaving this known receptor binding site in place, a nanobody -mediated infection boost occurred.

[0409] Two insertion positions that we tested, N337 and Q259, did not result in any infection gain, although their packaging was not impaired (Figure S2). N337 is partly hidden within the pore and an insertion there could potentially block the externalization of VPlu through the pore during infection. Q259 lays within the 2-fold valley closely abutting a neighboring VP subunit (unlike the nearby N262). Prior work has shown that interfacial dynamics at the twofold axis play a role in externalization of VPlu during infection, which may explain the lack of infectivity of insertion variant at this position (40).

[0410] The FAP nanobody used in this study can be used to direct AAV to cancer or cancer- associated cells with a high FAP expression profile (23). Notably, the straightforward replacement of the GFP nanobody in our previous study (22) with a FAP nanobody here suggests that nanobody swapping at the 2-fold valley and 5-fold axis hotspots is a viable strategy to rapidly diversify AAV tropism. With more and more nanobodies being developed, the herein demonstrated tolerability of nanobody insertions at various locations within the AAV capsid can further expand the applicability of nanobody -directed cell-specific targeting.

[0411] Bibliography

[0412] 1. Wang,D., Tai,P.W.L. and Gao,G. (2019) Adeno-associated virus vector as a platform for gene therapy delivery. Nat Rev Drug Discov 18, 358-378.

[0413] 2. Jay,F.T., Laughlin, C.A. and Carter, B.J. (1981) Eukaryotic translational control: adeno- associated virus protein synthesis is affected by a mutation in the adenovirus DNA-binding protein. Proc Natl Acad Sci U SA 78, 2927-2931.

[0414] 3. Srivastava,A., Lusby, E.W. and Berns, K.I. (1983) Nucleotide sequence and organization of the adeno-associated virus 2 genome. J Virol 45, 555-564.

[0415] 4. Johnson, F.B., Ozer,H.L. and Hoggan,M.D. (1971) Structural proteins of adenovirus-associated virus type 3. J Virol 8, 860-863. 5. Rose,J.A., Maizel,J.V., Inman, J.K. and Shatkin,A.J. (1971) Structural proteins of adenovirus- associated viruses. J Virol 8, 766-770.

[0416] 6. SnijderJ., van de Waterbeemd,M., Damoc,E., Denisov, E., Grinfeld,D., Bennett, A., Agbandje- McKenna,M., Makarov, A. and Heck,A.J. (2014) Defining the stoichiometry and cargo load of viral and bacterial nanoparticles by Orbitrap mass spectrometry. J Am Chem Soc 136, 7295-7299.

[0417] 7. Xie,Q., Bu,W., Bhatia, S., Hare, J., Somasundaram,T., Azzi,A. and Chapman, M.S. (2002) The atomic structure of adeno-associated virus (AAV-2), a vector for human gene therapy. Proc Natl Acad Sci U SA 99, 10405-10410.

[0418] 8. Govindasamy,L. , Padron,E., McKenna, R., Muzyczka,N., Kaludov,N., Chiorini,J.A. and Agbandje-McKenna,M. (2006) Structurally mapping the diverse phenotype of adeno-associated virus serotype 4. J Virol 80, 11556-11570.

[0419] 9. Tse,L.V., Kline, K.A., Madigan, V.J., Castellanos Rivera, R.M., Wells, L.F., Havlik,L.P., Smith, J.K., Agbandje-McKenna,M. and Asokan,A. (2017) Structure-guided evolution of antigenically distinct adeno-associated virus variants for immune evasion. Proc Nall Acad Sci U S A 114, E4812-E4821.

[0420] 10. Chan,K.Y., Jang,M.J., Yoo,B.B., Greenbaum,A., Ravi,N., Wu,W.L., Sanchez-Guardado, L., Lois,C., Mazmanian,S.K., Deverman, B.E. and Gradinaru,V. (2017) Engineered AAVs for efficient noninvasive gene delivery to the central and peripheral nervous systems. Nat Neurosci 20, 1172-1179.

[0421] 11. Dalkara,D., Byrne, L.C., Klimczak,R.R., Visel,M., Yin,L., Merigan,W.H., Flannery, J.G. and Schaffer, D.V. (2013) In vivo-directed evolution of a new adeno-associated virus for therapeutic outer retinal gene delivery from the vitreous. Sci Transl Med 5, 189ra76.

[0422] 12. Deverman, B.E., Pravdo,P.L., Simpson, B.P., Kumar, S.R., Chan,K.Y., Banerjee, A., Wu,W.L., Yang,B., Huber, N., Pasca,S.P. and Gradinaru,V. (2016) Cre-dependent selection yields AAV variants for widespread gene transfer to the adult brain. Nat Biotechnol 34, 204-209.

[0423] 13. KbrbelinJ., Sieber, T., Michelfelder,S., Funding, L., Spies, E., Hunger, A., Alawi,M., Rapti,K., Indenbirken,D., Muller, O.J., Pasqualini,R., Arap,W., Kleinschmidt, J.A. and Trepel,M. (2016) Pulmonary Targeting of Adeno-associated Viral Vectors by Next-generation Sequencing-guided Screening of Random Capsid Displayed Peptide Libraries. Mol Ther 24, 1050-1061.

[0424] 14. Tabebordbar,M., Lagerborg,K.A., Stanton, A., King,E.M., Ye,S., Tellez, L., Krunnfusz,A., Tavakoli,S., Widrick,J.J., Messemer,K.A., Troiano,E.C., Moghadaszadeh,B., Peacker,B.L., Leacock, K.A., Horwitz, N., Beggs, A.H., Wagers, A.J. and Sabeti,P.C. (2021) Directed evolution of a family of AAV capsid variants enabling potent muscle-directed gene delivery across species. Cell 184, 4919-4938. e22. 15. Weinmann, J., Weis,S., SippelJ., Tulalamba,W., Remes, A., El Andari,J., Herrmann,A.K., Pham,Q.H., Borowski, C., Hille, S., Schdnberger,T. , Frey,N., Lenter,M., VandenDriessche,T., Muller, O.J., Chuah,M.K., Lamia, T. and Grimm, D. (2020) Identification of a myotropic AAV by massively parallel in vivo evaluation of barcoded capsid variants. Nat Commun 11, 5432.

[0425] 16. Judd, J., Wei,F., Nguyen, P.Q., Tartaglia,L.J., Agbandje-McKenna,M., Silberg,J.J. and Suh, J. (2012) Random Insertion of mCherry Into VP3 Domain of Adeno-associated Virus Yields Fluorescent Capsids With no Loss of Infectivity. Mol Ther Nucleic Acids 1, e54.

[0426] 17. Michels, A., Frank, A.M., Gunther, D.M., Mataei,M., Borner, K., Grimm, D., Hartmann, J. and Buchholz, C. J. (2021) Lentiviral and adeno-associated vectors efficiently transduce mouse T lymphocytes when targeted to murine CD8. Mol Ther Methods Clin Dev 23, 334-347.

[0427] 18. Zdechlik,A.C., He,Y., Aird,E.J., Gordon, W.R. and Schmidt, D. (2020) Programmable Assembly of Adeno-Associated Virus-Antibody Composites for Receptor-Mediated Gene Delivery. Bioconjug Chern 31, 1093-1106.

[0428] 19. Eichhoff,A.M., Borner, K., Albrecht, B., Schafer, W., Baum,N., Haag,F., K6rbelin,J., Trepel,M., Braren,!., Grimm, D., Adriouch,S. and Koch-Nolte, F. (2019) Nanobody-Enhanced Targeting of AAV Gene Therapy Vectors. Mol Ther Methods Clin Dev 15, 211-220.

[0429] 20. Hamann, M.V., Beschorner,N., Vu,X.K., Hauber,!., Lange, U.C., Traenkle,B., Kaiser, P.D., Foth,D., Schneider, C., Biining,H., Rothbauer,U. and Hauber, J. (2021) Improved targeting of human CD4+ T cells by nanobody-modified AAV2 gene therapy vectors. PLoS One 16, eO261269.

[0430] 21. Bao,G., Tang,M., Zhao, J. and Zhu,X. (2021) Nanobody: a promising toolkit for molecular imaging and disease therapy. EJNMMI Res 11, 6.

[0431] 22. Hoffmann,M.D., Zdechlik,A.C., He,Y., Nedrud,D., Aslanidi,G., Gordon, W. and Schmidt, D. (2023) Multiparametric domain insertional profiling of adeno-associated virus VP1. Mol Ther Methods Clin Dev 31, 101143.

[0432] 23. Garin-Chesa,P., Old,L.J. and Rettig, W. J. (1990) Cell surface glycoprotein of reactive stromal fibroblasts as a potential antibody target in human epithelial cancers. Proc Natl Acad Sci U S A 87, 7235-7239.

[0433] 24. Hintz, H.M., Cowan, A.E., Shapovalova, M. and LeBeau,A.M. (2019) Development of a CrossReactive Monoclonal Antibody for Detecting the Tumor Stroma. Bioconjug Chem 30, 1466-1476.

[0434] 25. Xin,L„ Gao, J., Zheng, Z„ Chen,Y„ Lv,S., Zhao,Z„ Yu,C., Yang,X. and Zhang, R. (2021) Fibroblast Activation Protein-a as a Target in the Bench-to-Bedside Diagnosis and Treatment of Tumors: A Narrative Review. Front Oncol 11, 648187. 26. Engler, C., Kandzia,R. and Marillonnet,S. (2008) A one pot, one step, precision cloning method with high throughput capability. PLoS One 3, e3647.

[0435] 27. Rothbauer,U., Zolghadr,K., Tillib,S., Nowak, D., Schermelleh,L., Gahl,A., Backmann,N., Conrath, K., Muy Hermans, S., Cardoso, M.C. and Leonhardt, H. (2006) Targeting and tracing antigens in live cells with fluorescent nanobodies. Nat Methods 3, 887-889.

[0436] 28. Hintz, H.M., Gallant, J. P., Vander Griend,D.J., Coleman, I.M., Nelson, P.S. and LeBeau,A.M. (2020) Imaging Fibroblast Activation Protein Alpha Improves Diagnosis of Metastatic Prostate Cancer with Positron Emission Tomography. Clin Cancer Res 26, 4882-4891.

[0437] 29. Negrete, A. and Kotin, R.M. (2007) Production of recombinant adeno-associated vectors using two bioreactor configurations at different scales. J Virol Methods 145, 155-161.

[0438] 30. Hall,M.P„ Unch,J., Binkowski,B.F., Valley, M.P., Butler, B.L., Wood,M.G„ Otto,P„ Zimmerman, K., Vidugiris,G., Machleidt,T., Robers,M.B., Benink,H.A., Eggers, C.T., Slater, M.R., Meisenheimer,P.L., Klaubert,D.H., Fan,F., Encell,L.P. and Wood,K.V. (2012) Engineered luciferase reporter from a deep sea shrimp utilizing a novel imidazopyrazinone substrate. ACS Chem Biol 7, 1848-1857.

[0439] 31. Zinn,E., Pacouret,S., Khaychuk,V., Turunen,H.T., Carvalho, L.S., Andres-Mateos,E., Shah,S., Shelke,R., Maurer, A.C., Plovie,E., Xiao,R. and Vandenberghe,L.H. (2015) In Silico Reconstruction of the Viral Evolutionary Lineage Yields a Potent Gene Therapy Vector. Cell Rep 12, 1056-1068.

[0440] 32. Munch, R.C., Muth, A., Muik,A., Friedel, T., Schmatz,J., Dreier,B., Trkola,A., Pliickthun,A., Buning,H. and Buchholz, C.J. (2015) Off-target-free gene delivery by affinity-purified receptor- targeted viral vectors. Nat Commun 6, 6246.

[0441] 33. Muyldermans,S. (2013) Nanobodies: natural single-domain antibodies. Anna Rev Biochem 82, 775-797.

[0442] 34. Grieger, J.C., Snowdy, S. and Samulski,R.J. (2006) Separate basic region motifs within the adeno-associated virus capsid proteins are essential for infectivity and assembly. J Virol 80, 5199- 5210.

[0443] 35. Kronenberg,S., B6ttcher,B., von der Lieth,C.W., Bieker, S. and Kleinschmidt, J. A. (2005) A conformational change in the adeno-associated virus type 2 capsid leads to the exposure of hidden VP1 N termini. J Virol 79, 5296-5303.

[0444] 36. Nam,H.J., Gurda,B.L., McKenna, R., Potter, M., Byrne, B., Salganik,M., Muzyczka,N. and Agbandje-McKenna,M. (2011) Structural studies of adeno-associated virus serotype 8 capsid transitions associated with endosomal trafficking. J Virol 85, 11791-11799. 37. Venkatakrishnan,B., Yarbrough, J., Domsic,J., Bennett, A., Bothner,B., Kozyreva, O.G., Samulski,R.J., Muzyczka,N., McKenna, R. and Agbandje-McKenna,M. (2013) Structure and dynamics of adeno-associated virus serotype 1 VPl-unique N-terminal domain and its role in capsid trafficking. J Virol 81, 4974-4984.

[0445] 38. Xu,G., Zhang, R., Li,H., Yin,K., Ma,X. and Lou,Z. (2022) Structural basis for the neurotropic AAV9 and the engineered AAVPHP.eB recognition with cellular receptors. Mol Ther Methods Clin Dev 26, 52-60.

[0446] 39. Zhang, R„ Cao,L„ Cui,M„ Sun,Z„ Hu,M., Zhang, R„ Stuart, W„ Zhao,X„ Yang,Z., Li,X„ Sun,Y., Li,S., Ding,W., Lou,Z. and Rao,Z. (2019) Adeno-associated virus 2 bound to its cellular receptor AAVR. Nat Microbiol 4, 675-682.

[0447] 40. Horowitz, E.D., Finn,M.G. and Asokan,A. (2012) Tyrosine cross-linking reveals interfacial dynamics in adeno-associated viral capsids during infection. ACS Chem Biol 7, 1059-1066.

[0448] All publications, patents and patent applications are incorporated herein by reference. While in the foregoing specification, this invention has been described in relation to certain preferred embodiments thereof, and many details have been set forth for purposes of illustration, it will be apparent to those skilled in the art that the invention is susceptible to additional embodiments and that certain of the details herein may be varied considerably without departing from the basic principles of the invention.

Claims

WHAT IS CLAIMED IS:

1. An infectious recombinant adeno-associated virus (rAAV), comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an insertion of amino acids that includes one or more single stranded nucleic acid binding domains at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1 residues.

2. The rAAV of claim 1 , wherein the one or more single stranded nucleic acid binding domains are single stranded DNA binding domains.

3. The rAAV of claim 2, wherein the one or more single stranded DNA binding domains are HUH domains.

4. The rAAV of any one of claims 1 to 3, wherein the modification in the viral capsid further comprises a deletion of one or more amino acids of the VP.

5. The rAAV of claim 4, wherein the deletion is 2, 3, 4, 5, or 10 amino acids or less than 50 amino acids.

6. The rAAV of any one of claims 1 to 5, wherein the insertion comprises 5 to 100 amino acids.

7. The rAAV of any one of claims 1 to 6, wherein VP1 of AAV-DJ is modified.

8. The rAAV of any one of claims 1 to 7, wherein the AAV genome or capsid is AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, or AAVrhlO.

9. The rAAV of any one of claims 1 to 7, wherein the AAV genome or capsid is AAV1, AAV2, AAV5, AAV6, AAV8, AAV9, AAV-DJ, or AAV10.

10. The rAAV of claim any one of claims 1 to 6, which infects human cells.

11. The rAAV of any one of claims 3 to 10, wherein the one or more HUH domains have a sequence having at least 80% amino acid sequence identity to a HUH domain in one or more of SEQ ID Nos. 2-10 and 20-21.

12. The rAAV of any one of claims 3 to 10, wherein the one or more HUH domains have a sequence having at least 90% amino acid sequence identity to a HUH domain in one or more of SEQ ID NOs 2-10 and 20-21 .

13. The rAAV of any one of claims 3 to 12, wherein the one or more HUH domains are flanked by at least one linker.

14. The rAAV of claim 13, wherein at least one HUH domain is flanked by two linkers.

15. The rAAV of claim 13 wherein at least one HUH domain is flanked by one linker.

16. The rAAV of any one of claims 3 to 15, wherein up to 10 HUH domains are in the insertion.

17. The rAAV of claim 16, wherein at least two different HUH domains are in the insertion.

18. The rAAV of claim 16, wherein the HUH domains are concatemers.

19. The rAAV of any one of claims 3 to 18, which has a plurality of the same HUH domain.

20. The rAAV of any one of claims 1 to 19, wherein the insertion is from about 2 kDa up to300 kDa, about 5 kDa up to 50 kDa, about 10 kDa up to 250 kDa, about 20 kDa up to a 100 kDa, about 10 kDa up to 50 kDa, or about 75 kDa up to about 150 kDa.

21. The rAAV of any one of claims 1 to 20, wherein the viral genome is a recombinant genome having at least one expression cassette for an exogenous gene product.

22. The rAAV of claim 21, wherein the exogenous gene product is a prophylactic or therapeutic gene product.

23. The rAAV of claim 21, wherein the exogenous gene product is a cytotoxic gene product.

24. The rAAV of any one of claims 1 to 23, further comprising a targeting molecule linked to the insertion.

25. The rAAV of claim 24, wherein the targeting molecule comprises an antibody or an antigen binding portion thereof.

26. The rAAV of claim 24 or 25, wherein the targeting molecule comprises one or more albumin-binding domains (ABD), adhirons, adnectin / monobodies, affibodies, affilins, affimers, affitins / nanofitins, alphabodies, anticalins, armadillo repeat proteins, atrimers, avimers, aentyrins, DARPins, fynomers, Kunitz-domains, peptide toxins, peptide ligands, obodies / OB- Fold proteins, pronectins, or repebodies, RNA aptamers, one or more antibodies, or any combination thereof.

27. The rAAV of any one of claims 3 to 30, wherein the HUH substrate comprises a nucleotide sequence having at least 90% nucleic acid sequence identity to any one of SEQ ID Nos. 1-4 or 11, or to at least 9, 12 or 15 nucleotides thereof that bind one or more of the HUH domains.

28. The rAAV of any one of claims 3 to 27, wherein the HUH substrate comprises PNA or LNA.

29. A method of targeting mammalian cells in vivo, comprising: a) providing a population of infectious rAAV, comprising: rAAV comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an insertion of amino acids that includes one or more single stranded nucleic acid binding domains at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1 residues; b) providing a single stranded nucleic acid binding substrate for the one or more single stranded nucleic acid binding domains covalently linked to a targeting molecule; c) combining the population of recombinant virus and the substrate covalently linked to the targeting molecule to form a conjugate; and d) administering the conjugate to a mammal.

30. The method of claim 29, wherein the single stranded nucleic acid binding domain comprises one or more HUH domains.

31. The method of claim 29 or 30, wherein the mammal is a non-human mammal.

32. The method of claim 29, 30 or 31, wherein the mammal is a human.

33. The method of any one of claims 29 to 32, wherein the targeting molecule is an antibody or an antigen binding portion thereof.

34. The method of claim 33, wherein the antibody is an anti-CD3 antibody, anti-CD4 antibody, anti-CD7 antibody, anti-Her2 antibody, anti-CD34 antibody, anti-CD8, anti-CD20 antibody, or anti-CD19 antibody.

35. A system comprising: a population of infectious rAAV comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an insertion of amino acids that includes one or more single stranded nucleic acid binding domains at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1; and a substrate for the single stranded nucleic acid binding domain.

36. The system of claim 35, wherein the one or more single stranded nucleic acid binding domains comprise HUH domains.

37. The system of claim 35 or 36, further comprising a targeting molecule.

38. The system of claim 37, wherein the substrate is covalently linked to a targeting molecule.

39. An infectious recombinant adeno-associated virus (rAAV), comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an insertion of amino acids that includes one or more single stranded nucleic acid binding domains at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1, which insertion is bound to the one or more single stranded nucleic acid binding domains which are linked to a molecule.

40. The rAAV of claim 39 wherein the molecule is a targeting molecule.

41. A vector comprising a modified recombinant adeno-associated virus (rAAV) capsid, wherein the modified viral capsid comprises an insertion of amino acids that includes one or more small binding proteins at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1 residues.

42. The vector of claim 41, wherein the modification in the viral capsid further comprises a deletion of one or more amino acids of the VP.

43. The vector of claim 42, wherein the deletion is 2, 3, 4, 5, or 10 amino acids or less than 50 amino acids.

44. The vector of any one of claims 41 to 43, wherein the insertion comprises 5 to 100 amino acids.

45. The vector of any one of claims 41 to 44, wherein the one or more small binding proteins are nanobodies, a DARPins, GP2, adhirons, adnectin / monobodies, affibodies, affilins, affimers, atrimers, fynomers, kunitz domains, OBodies, peptide toxins, or peptide ligands.

46. The vector of any one of claims 41 to 45, wherein VP1 of AAV-DJ is modified.

47. The vector of any one of claims 41 to 46, wherein the capsid is AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, or AAVrhlO.

48. The vector of any one of claims 41 to 46, wherein the AAV genome or capsid is AAV1, AAV2, AAV5, AAV6, AAV8, AAV9, AAV-DJ, or AAV10.

49. The vector of any one of claims 41 to 48, wherein the one or more small binding proteins are flanked by at least one linker.

50. The vector of any one of claims 41 to 49, wherein at least one small binding protein is flanked by two linkers.

51. The vector of any one of claims 41 to 48, wherein at least one small binding protein is flanked by one linker.

52. The vector of any one of claims 41 to 51, wherein up to 10 small binding proteins are in the insertion.

53. The vector of any one of claims 41 to 52, wherein at least two different small binding proteins are in the insertion.

54. The vector of any one of claims 41 to 53, which has a plurality of the same small binding proteins.

55. The vector of any one of claims 41 to 54, wherein the insertion is from about 2 kDa up to 300 kDa, about 5 kDa up to 50 kDa, about 10 kDa up to 250 kDa, about 20 kDa up to a 100 kDa, about 10 kDa up to 50 kDa, or about 75 kDa up to about 150 kDa.

56. An infectious recombinant adeno-associated virus (rAAV), comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an insertion of amino acids that includes one or small binding proteins at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1 residues.

57. The rAAV of claim 56, wherein the modification in the viral capsid further comprises a deletion of one or more amino acids of the VP.

58. The rAAV of claim 56 or 57, wherein the one or more small binding proteins are nanobodies, a DARPins, GP2, adhirons, adnectin / monobodies, affibodies, affilins, affimers, atrimers, fynomers, kunitz domains, OBodies, peptide toxins, or peptide ligands.

59. The rAAV of any one of claims 56 to 58, wherein VP1 of AAV-DJ is modified.

60. The rAAV of any one of claims 56 to 59, wherein the AAV genome or capsid is AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, or AAVrhlO.

61. The rAAV of any one of claims 56 to 59, wherein the AAV genome or capsid is AAV1, AAV2, AAV5, AAV6, AAV8, AAV9, AAV-DJ, or AAV10.

62. The rAAV of any one of claims 56 to 61, which infects human cells.

63. The rAAV of any one of claims 56 to 62, wherein the one or more insertions are flanked by at least one linker.

64. The rAAV of any one of claims 56 to 63, wherein up to 10 small binding proteins are in the insertion.

65. The rAAV of any one of claims 56 to 64, wherein at least two different small binding proteins are in the insertion.

66. The rAAV of any one of claims 56 to 65, wherein the viral genome is a recombinant genome having at least one expression cassette for an exogenous gene product.

67. A method of targeting mammalian cells in vivo, comprising: providing a population of infectious rAAV, comprising: rAAV comprising a modified viral capsid and a viral genome, wherein the modified viral capsid comprises an insertion of amino acids that includes one or more small binding proteins at a residue from 260 to 270, 328 to 338, 370 to 400 or 618 to 727 of VP1 residues and administering the recombinant virus comprising the one or more small binding proteins to a mammal.

68. The method of claim 67, wherein the one or more small binding proteins are nanobodies, a DARPins, GP2, adhirons, adnectin / monobodies, affibodies, affilins, affimers, atrimers, fynomers, kunitz domains, OBodies, peptide toxins, or peptide ligands.

69. The method of claim 67 or 68, wherein the mammal is a non-human mammal.

70. The method of claim 67, 68 or 69, wherein the mammal is a human.