Reprogramming tropism via displayed peptide tiling receptor ligands

By inserting ligand peptides into AAV capsids to reprogram tropism, the method addresses inefficient delivery challenges, enabling targeted and efficient delivery of therapeutic agents to specific tissues.

JP2025531824APending Publication Date: 2025-09-25RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025514363
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2023-09-08
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

The challenge in nucleic acid and protein-based therapeutics is the inefficient, unsafe, and non-targeted delivery of viral and non-viral formulations, particularly in modulating tropism for specific tissues.

Method used

Engineering adeno-associated viruses (AAV) by inserting ligand peptides into surface-exposed loops of the capsid, systematically evaluating variants for packaging capacity and tropism, and using recombinant vectors with desired tropism or immuno-orthogonality to target specific tissues like pancreas, heart, brain, lung, liver, kidney, muscle, or intestine.

Benefits of technology

Achieves targeted and efficient delivery of therapeutic agents to specific tissues by reprogramming viral tropism, enhancing tissue specificity and immune evasion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531824000001_ABST
    Figure 2025531824000001_ABST
Patent Text Reader

Abstract

Described herein are methods for engineering proteins and viruses to improve tropism, as well as proteins and viruses produced by using the methods.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority under 35 U.S.C. § 119 of Provisional Application No. 63 / 405,360, filed September 9, 2022, the disclosure of which is incorporated herein by reference in its entirety.

[0002] (Statement of Federally Funded Research) This invention was made with government support under Grant Nos. OD032742, CA222826, and GM123313 awarded by the National Institutes of Health, and Grant No. W81XWH-22-1-0401 awarded by the U.S. Department of Defense. The government has certain rights in this invention.

[0003] FIELD OF THE INVENTION Described herein are methods for engineering proteins and viruses to improve tropism, as well as proteins and viruses produced by using the methods.

[0004] (Incorporated by reference to the sequence listing) Attached to this application is a Sequence Listing entitled "00015-417WO1_SL.xml," created on September 8, 2023, and containing 872,924 bytes of data, machine-formatted on an IBM-PC, MS-Windows operating system. The Sequence Listing is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0005] Nucleic acid and protein-based therapeutics for modulating healthy and disease states are poised to enable the next frontier of human medicine. However, their successful deployment is contingent on the ability to deliver them efficiently, safely, and in a targeted manner. In this regard, a host of viral and non-viral delivery formulations have been developed, but the ability to directly modulate their tropism remains a challenge. Summary of the Invention

[0006] The present disclosure provides a method for improving the tropism of a virus or other delivery agent, comprising identifying ligand protein sequences derived from all known receptor-interacting ligands, systematically tiling the ligand peptides into 5-50 or 10-20 or 20 amino acid peptides that are inserted into surface-exposed loops of AAV capsids, and evaluating the engineered capsids for their packaging capacity, in vivo tropism, and enhanced protein interactions. In one embodiment, the virus is an adeno-associated virus (AAV). In a further embodiment, the AAV is selected from the group consisting of AAV1, AAV2, AAV5, AAV6, AAV7, AAV8, and AAV9. In yet a further embodiment, the AAV is AAV5 or AAV9. In another embodiment, peptide sequences were generated via pooled oligonucleotide synthesis and inserted into four distinct loop regions: AAV5-loop1 (N443), AAV5-loop2 (S576), AAV9-loop1 (Q456), and AAV9-loop2 (A587) to generate over one million AAV variants. In another embodiment, a 20'mer ligand peptide is inserted into one or both of two surface-exposed loops of AAV5 (SEQ ID NO:2) or AAV9 (SEQ ID NO:4).

[0007] The present disclosure provides recombinant vectors comprising a capsid protein comprising one or two peptides inserted into one or both of two surface-exposed loops in the capsid protein, wherein the vector has desired tropism or immuno-orthogonality, and the peptides are independently selected from the group consisting of SEQ ID NOs:5-820 and 870-911. In one embodiment, the virus is an adeno-associated virus (AAV). In a further embodiment, the AAV vector is an AAV5 serotype. In a still further embodiment, the AAV5 comprises a capsid protein having the sequence set forth in SEQ ID NO:2. In another embodiment, the AAV vector is an AAV9 serotype. In a further embodiment, the AAV9 comprises a capsid protein having the sequence set forth in SEQ ID NO:4. In another embodiment, the vector has tropism for the pancreas, heart, brain, lung, liver, kidney, muscle, spleen, or intestine. In another or further embodiment, the peptide of SEQ ID NO: 530-820, 870-910, or 911 is inserted into loop 1 and / or loop 2 of an AAV5 or AAV9 capsid. In another embodiment, the vector is immunoorthogonal. In a further embodiment, the vector is an adeno-associated virus (AAV). In yet a further embodiment, the AAV comprises a capsid protein of any one of SEQ ID NOs: 822, 824, 826, 828, 830, 832, 834, 836, 838, 840, 842, 844, 846, 848, 850, 852, 854, 856, 858, 860, 862, 864, or 866, or a sequence at least 85% to 99% identical to any of the foregoing sequences.

[0008] The present disclosure provides a method for producing a delivery vehicle or vector with a desired tropism, the method comprising: selecting a peptide sequence from any one of SEQ ID NOS:5-820 or 870-911; and (i) cloning a nucleic acid sequence encoding the peptide into a coding sequence for a capsid protein at an exposed loop site to obtain a recombinant capsid coding sequence and producing a vector using the recombinant capsid coding sequence; or (ii) inserting the peptide into an exposed surface of the delivery vehicle. In one embodiment, the virus is an adeno-associated virus (AAV). In a further embodiment, the AAV is selected from the group consisting of AAV1, AAV2, AAV5, AAV6, AAV7, AAV8, and AAV9. In yet a further embodiment, the AAV is AAV5 or AAV9. In a further embodiment, the AAV5 capsid coding sequence comprises SEQ ID NO:1. In another embodiment, the AAV9 capsid coding sequence comprises SEQ ID NO:2. The present disclosure also provides an AAV vector comprising a capsid protein modified according to any of the foregoing embodiments.

[0009] The present disclosure also provides an AAV vector comprising a capsid protein, wherein the capsid protein expresses a peptide of any one of SEQ ID NOs: 5-820 or 870-911 in surface-exposed loop 1 and / or loop 2 of the capsid protein. In one embodiment, the wild-type capsid protein sequence comprises SEQ ID NO: 2 or 4. In another embodiment, the AAV vector has an AAV5 or AAV9 serotype. In yet another embodiment, the vector has a desired tropism. In a further embodiment, the AAV vector has tropism for the pancreas, heart, brain, lung, liver, kidney, muscle, spleen, or intestine.

[0010] The present disclosure also provides a viral vector having a capsid protein, the capsid protein comprising a heterologous targeting peptide of 10 to 30 amino acids in length inserted into a surface-exposed portion of the capsid protein, wherein the targeting peptide is set forth in any one of SEQ ID NOS: 5-820 or 870-911. In one embodiment, the heterologous targeting peptide is about 15 to 25 amino acids in length. In a further embodiment, the heterologous targeting peptide is about 20 amino acids in length. In another embodiment, the viral vector is an adeno-associated virus (AAV). In yet another embodiment, the viral vector is a lentiviral vector. In another embodiment, the capsid protein is a VP1 capsid protein. In yet another embodiment, the capsid protein is a VP2 capsid protein. In yet another embodiment, the capsid protein is a VP3 capsid protein. In another embodiment, the heterologous targeting peptide is inserted into the AAV capsid protein at loop 1 and / or loop 2. In yet another embodiment, the viral vector is AAV5. In another embodiment, the viral vector is AAV9. In yet another embodiment, the heterologous targeting peptide is flanked by linker peptides at the N-terminus and C-terminus of the heterologous targeting peptide. In another embodiment, the heterologous targeting peptide targets the viral vector to hepatocytes or liver tissue. In yet another embodiment, the heterologous targeting peptide targets the viral vector to neuronal cells or brain tissue. In yet another embodiment, the heterologous targeting peptide targets the viral vector to pancreatic cells or pancreatic tissue. In yet another embodiment, the heterologous targeting peptide targets the viral vector to cardiac cells or cardiac tissue. In another embodiment, the heterologous targeting peptide targets the viral vector to lung tissue. In yet another embodiment, the heterologous targeting peptide targets the viral vector to intestinal tissue. In yet another embodiment, the heterologous targeting peptide targets the viral vector to spleen tissue. In yet another embodiment, the heterologous targeting peptide targets the viral vector to kidney cells or kidney tissue. In another embodiment, the heterologous targeting peptide targets the viral vector to muscle cells or tissue.

[0011] The present disclosure also provides an adeno-associated virus (AAV) capsid protein comprising a heterologous targeting peptide cloned into loop 1 and / or loop 2 of the capsid protein, wherein the heterologous targeting peptide is approximately 10-30 amino acids in length and is contained within or comprises any one of the peptides set forth in SEQ ID NOS: 5-820 or 870-911. In one embodiment, the capsid protein is a VP1 capsid protein. In another embodiment, the capsid protein is a VP2 capsid protein. In yet another embodiment, the capsid protein is a VP3 capsid protein. In another embodiment, the heterologous targeting peptide is approximately 15-25 amino acids in length. In yet another embodiment, the heterologous targeting peptide is approximately 20 amino acids in length. In yet another embodiment, the heterologous targeting peptide is flanked by linker peptides at the N- and C-termini of the heterologous targeting peptide. In another embodiment, the heterologous targeting peptide targets hepatocytes or liver tissue. In yet another embodiment, the heterologous targeting peptide targets nerve cells or brain tissue. In another embodiment, the heterologous targeting peptide targets pancreatic cells or tissue. In yet another embodiment, the heterologous targeting peptide targets cardiac cells or tissue. In another embodiment, the heterologous targeting peptide targets lung tissue. In yet another embodiment, the heterologous targeting peptide targets intestinal tissue. In another embodiment, the heterologous targeting peptide targets spleen tissue. In another embodiment, the heterologous targeting peptide targets kidney cells or tissue. In yet another embodiment, the heterologous targeting peptide targets muscle cells or tissue. The present disclosure also provides a recombinant AAV (rAAV) comprising the capsid protein of any of the foregoing embodiments.

[0012] The present disclosure provides a recombinant AAV (rAAV) comprising a capsid protein having a targeting peptide in loop 1 and / or loop 2, wherein the targeting peptide is independently selected from SEQ ID NOs: 5-820 or 870-911. In one embodiment, the targeting peptide is present in both loop 1 and loop 2. In a further embodiment, the targeting peptides have the same tropism. In another embodiment, the recombinant AAV further comprises a heterologous polynucleotide for gene delivery. In a further embodiment, the heterologous polynucleotide is a therapeutic gene. In yet another or further embodiment of any of the foregoing embodiments, the rAAV is present in a pharmaceutical composition.

[0013] The present disclosure also provides a method for delivering a transgene to a subject, comprising administering to the subject a recombinant AAV (rAAV), wherein the rAAV comprises (i) a capsid protein of the present disclosure and (ii) at least one transgene, wherein the rAAV infects cells of a target tissue of the subject. In one embodiment, the at least one transgene encodes a protein. In a further embodiment, the protein is an immunoglobulin heavy chain or light chain or a fragment thereof. In another embodiment, the at least one transgene encodes a small interfering nucleic acid. In a further embodiment, the small interfering nucleic acid is an miRNA. In another embodiment, the small interfering nucleic acid is an miRNA sponge or TuD RNA that inhibits the activity of at least one miRNA in a subject or animal. In yet another embodiment, the miRNA is expressed in cells of the target tissue. In yet another embodiment, the target tissue is skeletal muscle, heart, liver, pancreas, brain, or lung. In another embodiment, the transgene expresses a transcript containing at least one binding site for an miRNA, and the miRNA inhibits transgene activity in tissues other than the target tissue by hybridizing to the binding site. In yet another embodiment, at least one transgene encodes a gene product that mediates genome editing. In another embodiment, the transgene comprises a tissue-specific or inducible promoter. In a further embodiment, the tissue-specific promoter is a liver-specific thyroxin binding globulin (TBG) promoter, an insulin promoter, a glucagon promoter, a somatostatin promoter, a pancreatic polypeptide (PPY) promoter, a synapsin-1 (Syn) promoter, a creatine kinase (MCK) promoter, a mammalian desmin (DES) promoter, an α-myosin heavy chain (α-MHC) promoter, or a cardiac troponin T (cTnT) promoter. In another embodiment, the rAAV is administered intravenously, intravascularly, transdermally, intraocularly, intrathecally, orally, intramuscularly, subcutaneously, intranasally, or by inhalation.In yet another embodiment, the subject is selected from a mouse, a rat, a rabbit, a dog, a cat, a sheep, a pig, and a non-human primate. In yet another embodiment, the subject is a human.

[0014] The present disclosure also provides an isolated nucleic acid encoding an AAV capsid protein comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 5-820 and 870-911.

[0015] The present disclosure also provides a delivery vehicle for delivery of a small molecule drug or biological agent having a desired tropism, the delivery vehicle comprising a peptide or peptide fragment of at least 10-20 amino acids of any one of SEQ ID NOS: 5-820 or 870-911. In one embodiment, the delivery vehicle is selected from the group consisting of a liposome, a nanoparticle, a bacterium, a bacteriophage, a virus-like particle (VLP), an erythrocyte ghost, and an exosome. In another embodiment, the biological agent comprises an siRNA, an antisense molecule, a protein or polypeptide, insulin, a vaccine, or an antibody. In yet another embodiment, the small molecule drug comprises a chemotherapeutic agent, an anti-inflammatory agent, a steroid, and an antibiotic.

[0016] The present disclosure also provides a biological agent having a desired tropism, wherein the biological agent is linked to a peptide or peptide fragment of at least 10-20 amino acids of any one of SEQ ID NOS: 5-820 or 870-911. In one embodiment, the biological agent is a nucleic acid, protein, polypeptide, peptide, antibody, antibody fragment, non-immunoglobulin binding agent, or enzyme. [Brief explanation of the drawings]

[0017] [Figure 1A]Design of an AAV library displaying receptor-ligand tiling peptides. (a) Schematic of the approach for rationally engineering and characterizing AAV variants. Ligand protein sequences derived from all known receptor-interacting ligands are systematically tiled into 20-amino acid peptides that are inserted into surface-exposed loops of the AAV capsid. These engineered variants were then evaluated for their packaging capacity, in vivo tropism, and enhanced protein interactions. [Figure 1B] Design of an AAV library displaying receptor-ligand tiling peptides. (b) Protein class distribution of the tiled ligands used to construct the screening library. These include known receptor-interacting ligands (orange), cell membrane-permeable proteins (green), and protein domains (blue), as well as a stop codon-containing negative control (purple). Peptide sequences were generated via pooled oligonucleotide synthesis and inserted into four distinct loop regions: AAV5-loop1 (N443), AAV5-loop2 (S576), AAV9-loop1 (Q456), and AAV9-loop2 (A587) to generate over one million AAV variants. Capsid surface residues are colored according to their distance from the capsid core, and insertion sites adjacent to the residues are shown in red. [Figure 2A] AAV library packaging analysis reveals biophysical features that contribute to capsid fitness. (a) Schematic showing recombinant production of pooled AAV libraries in HEK293T cells. [Figure 2B] AAV library packaging analysis reveals biophysical features that contribute to capsid fitness. (b) Normalized abundance of each inserted peptide in the plasmid library relative to DNA isolated from a recombinantly produced AAV variant capsid library. The dotted red line indicates where plasmid abundance equals capsid abundance. [Figure 2C]AAV library packaging analysis reveals biophysical features that contribute to capsid fitness. (c) Normalized abundance of a negative control peptide containing a stop codon in both the plasmid library and the recombinantly produced AAV capsid library. [Figure 2D] AAV library packaging analysis reveals biophysical features that contribute to capsid fitness. (d) Distribution of peptide biophysical parameters associated with packaging. Peptide charge, alpha-helical content, flexibility, and hydrophobicity distribution are shown for peptides enriched in capsid pools ("packaged") and depleted ("non-packaged") peptides. Statistical significance between groups was calculated by t-test (****p<0.0001). [Figure 2E] AAV library packaging analysis reveals biophysical features that contribute to capsid fitness. (e) Using the biophysical parameters of the inserted peptides as features, we trained a support vector machine classifier to predict which AAV variants will be successfully packaged into capsids. The receiver operating characteristic curve for the resulting model is shown, with an area under the curve of 0.89. [Figure 2F] AAV library packaging analysis reveals biophysical features that contribute to capsid fitness. (f) UMAP embedding for each AAV variant, colored by packaging state. Inserted peptide charge, alpha-helical content, flexibility, and hydrophobicity were used as input features for embedding. [Figure 3A] In vivo screening analysis enables predictive computational modeling of tropism. (a) Overview of the in vivo screening methodology. Four AAV variant libraries were injected retro-orbitally into C57 / BL6 mice in duplicate. Two weeks after injection, nine organs were harvested from each mouse, and the inserted peptide-containing region of the AAV capsid was amplified and subjected to next-generation sequencing. [Figure 3B]In vivo screening analysis allows for predictive computational models of tropism. (b) Repeat Pearson correlations for log2 fold change (log2FC) values ​​for each library and organ (organ vs. capsid). [Figure 3C] In vivo screening analysis enables predictive computational models of tropism. (c) Summary of screening results. AAV variants significantly enriched in a particular organ are defined as those with a log2FC>1 and an FDR-adjusted p-value <0.05. Bar plots show the number of significantly enriched variants detected per organ for both AAV5 and AAV9, as well as a comparison of peptide hits for each loop insertion site. [Figure 3D] In vivo screening analysis enables predictive computational models of tropism. (d) Overview of the classification model predicting AAV tissue tropism from peptide sequence alone. The inserted peptide sequence was converted to binary one-hot encoding (for each peptide, 20 rows correspond to position and 20 columns correspond to the presence of specific amino acids). This one-hot encoding scheme was then used as input to a convolutional neural network (CNN) to predict organ targeting. Model performance was evaluated separately for each organ via accuracy, area under the receiver operator characteristic curve (AUROC), F1 score, and Matthews correlation coefficient (MCC). The model was trained on two-thirds of the data, and the remaining one-third was excluded as a validation dataset to evaluate performance. The figure discloses SEQ ID NOs: 867-869, respectively, in order of appearance. [Figure 4A] AAV variants identified by in vivo screening exhibit reprogrammed tropism. (a) Heatmap showing the log2FC values ​​for each AAV variant significantly enriched in at least one organ. Rows are individual variants, and columns are organs (n=2 per organ). [Figure 4B] AAV variants identified by in vivo screening exhibit reprogrammed tropism. (b) UMAP embedding of significantly enriched AAV variants. Each point represents a variant colored by the organ with the highest log2FC. [Figure 4C] AAV variants identified by in vivo screening exhibit reprogrammed tropism. (c) AAVs were selected for validation from the pool of significant hits based on their tissue specificity, broad tropism, and / or internal consistency. Internal consistency was quantified by counting the number of similar (>50% homologous) inserted peptides that were also detected as hits for a given organ. AAV variants were characterized structurally via transmission electron microscopy and functionally via in vivo delivery of the mCherry transgene. Heat maps show all variants selected for validation (n = 21). The left heat map shows Z-normalized log2FC values ​​from the pooled in vivo screening (n = 2), and the right heat map shows Z-normalized mCherry expression quantified by RT-qPCR for AAV9 (n = 2). [Figure 4D] AAV variants identified by in vivo screening exhibit reprogrammed tropism. (d) Liver RT-qPCR quantification of mCherry delivery compared with protein level quantification via fluorescence microscopy. [Figure 5A]Characterization and mechanistic exploration of AAV variants with displayed ligand peptides. (a) Full characterization experiment for variant AAV9.DKK1. The upper left shows lung screening counts for all AAV9 loop 2 variants with inserted DKK1-derived peptides. The x-axis indicates the position in the DKK1 structure where a given peptide begins. Lung counts are shown in blue, and capsid counts are shown in orange. The red arrow indicates the position of the inserted peptide in AAV9.DKK1. The bar plot shows RT-qPCR quantification of mCherry transgene expression across individual validations (n ​​= 2) normalized to that of AAV9. Electron microscopy images of AAV variant capsids and fluorescence microscopy of mCherry protein expression levels in the lung and liver are also shown. [Figure 5B] Characterization and mechanistic exploration of AAV variants with displayed ligand peptides. (b) Knockout experiments to verify AAV9.DKK1 receptor dependency. Cas9-containing lentivirus was produced using either a non-targeting control (NTC) or an LRP6-targeting sgRNA (n = 2 each). HEK293T cells were then transduced with the lentivirus and positively selected with puromycin. Lentiviral-transduced cells were then seeded into 24-well plates and transduced with either wild-type AAV9 (4 × 10 viral genomes) or AAV9.DKK1 virus (1 × 10 viral genomes). The mean fluorescence intensity (MFI) of infected cells was then quantified using flow cytometry. Shown on the right is the crystal structure of the 7-mer DKK1 peptide (contained within AAV9.DKK1) in complex with LRP6 (58). [Figure 5C]Characterization and mechanistic exploration of AAV variants with displayed ligand peptides. (c) Full characterization experiment for variant AAV9.PDGFC. The upper left shows muscle screening counts for all AAV9 loop 1 variants with inserted PDGFC-derived peptides. The x-axis indicates the position in the PDGFC structure where a given peptide starts. Blue indicates muscle counts, and orange indicates capsid counts. The red line indicates the position of the peptide inserted into AAV9.PDGFC. The bar plot shows RT-qPCR quantification of mCherry transgene expression across individual validations (n ​​= 2) normalized to that of AAV9. Also shown are electron micrographs of variant capsids and fluorescence microscopy of mCherry protein expression levels in heart and muscle. [Figure 5D] Characterization and mechanistic exploration of AAV variants with displayed ligand peptides. (d) Receptor overexpression studies for variant AAV9.PDGFC. A known receptor for the PDGFC ligand was cloned into an overexpression plasmid and transfected into HEK293T cells in 24-well plates. 24 hours later, cells were transduced with either AAV9 (4 × 10 viral genomes) or AAV9.PDGFC (4 × 10 viral genomes). 24 hours after transduction, cells were harvested and mCherry expression was quantified by flow cytometry. Bar plots show MFI normalized to the mean MFI of AAV9-transduced cells with overexpressed empty vector. Statistical significance between groups was calculated by t-test (*p<0.05, **p<0.01, ***p<0.001, ****p<0.0001). [Figure 6A]The inserted peptides promote efficient liver detargeting across multiple AAV scaffolds in a mouse strain-independent manner. (a) AAV variants significantly enriched across all capsids are projected in two dimensions via UMAP. Distances between points are explicitly calculated via the Levenshtein distance of the inserted peptide's amino acid sequence, then embedded via UMAP. AAV variants are colored by their log2 fold change in the brain. Select variant clusters highly enriched in the brain are highlighted. AAV variants are also colored by their capsid insertion site in the embedding on the right. [Figure 6B] The inserted peptide promotes efficient liver detargeting across multiple AAV scaffolds in a mouse strain-independent manner. (b) Brain and capsid counts for all AAV5 loop 2 and AAV9 loop 1 variants with inserted APOA1-derived peptides. Blue indicates brain counts, and orange indicates capsid counts. Red arrows indicate the location of the inserted peptide in AAV5.APOA1 and AAV9.APOA1, respectively. Electron micrographs for AAV5.APOA1 and AAV9.APOA1 are also shown. [Figure 6C] The inserted peptide promotes efficient liver detargeting across multiple AAV scaffolds in a mouse strain-independent manner. (c) RT-qPCR values ​​for individual in vivo validation of AAV5.APOA1 and AAV9.APOA1 in C57 / BL6 mice (transduced against AAV9) confirming liver detargeting. [Figure 6D] The inserted peptide promotes efficient liver detargeting across multiple AAV scaffolds in a mouse strain-independent manner. (d) Performance of AAV5.APOA1 in BALB / c mice. The bar graph shows RT-qPCR values ​​for AAV5.APOA1 across eight organs (transduction relative to AAV9 in BALB / c), confirming that liver detargeting is strain-independent. [Figure 7A]Mining, characterization, and tropism reprogramming of novel immuno-orthogonal AAV serotypes via displayed ligand peptides. (a) Schematic of the computational pipeline used to identify novel AAV serotypes for testing. Using the Basic local alignment search tool (BLAST), we identified 687 initial capsids with sequence homology to the AAV2 cap gene. This initial list was then filtered to exclude truncated genomes, redundant samples, human and non-mammalian serotypes, and close orthologs. The final list included 23 AAV capsid sequences for further investigation. [Figure 7B] Mining, characterization, and tropism reprogramming of novel immuno-orthogonal AAV serotypes via displayed ligand peptides. (b) Hierarchical clustering dendrogram of AAV capsid sequences. Previously identified AAV serotypes currently in use are shown in red. Novel AAVs identified and used for downstream testing are shown in black. [Figure 7C] Mining, characterization, and tropism reprogramming of novel immuno-orthogonal AAV serotypes via displayed ligand peptides. (c) Two-step filtering of novel AAVs. AAVs were first evaluated by measuring their ability to package and then by their ability to transduce the liver in vivo (using the mCherry transgene). All values ​​are shown relative to the ortholog wild-type AAV5. [Figure 7D] Mining, characterization, and tropism reprogramming of novel immuno-orthogonal AAV serotypes via displayed ligand peptides. (d) Four novel AAVs capable of infecting the liver (AAV MM2, AAV MG2, AAV MG1, and AAV CH1) were tested for immune cross-reactivity with AAV8. Mice were immunized with the indicated AAVs and then tested for antibody cross-reactivity via ELISA 3 weeks after injection. [Figure 7E]Mining, characterization, and tropism reprogramming of novel immuno-orthogonal AAV serotypes via displayed ligand peptides. (e) The PDGFC peptide from AAV9.PDGFC was inserted into loop 1 of AAV MG2 to yield AAV MG2.PDGFC. AAV MG2 and AAV MG2.PDGFC were injected into C57BL / 6 mice, and muscle transduction was quantified via RT-qPCR 3 weeks later. Bar plots show muscle transduction relative to wild-type AAV MG2. Statistical significance between groups was calculated by t-test (*p<0.05, **p<0.01, ***p<0.001, ****p<0.0001). DETAILED DESCRIPTION OF THE INVENTION

[0018] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to a "cell" includes a plurality of such cells, and a reference to a "fragment" includes a reference to one or more fragments and equivalents thereof known to those skilled in the art.

[0019] Additionally, the use of "or" means "and / or" unless otherwise stated. Similarly, "comprise," "comprises," "comprising," "include," "includes," and "including" are interchangeable and are not intended to be limiting.

[0020] It should be further understood that where the descriptions of various embodiments use the term "comprising," those skilled in the art will understand that in some specific instances, an embodiment may alternatively be described using the language "consisting essentially of" or "consisting of."

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although many methods and reagents are similar or equivalent to those described herein, exemplary methods and materials are disclosed herein.

[0022] All publications mentioned herein are incorporated by reference in their entirety for the purpose of describing and disclosing methodologies that might be used in connection with the description herein. Furthermore, with respect to any terms presented in one or more publications that are similar or identical to terms expressly defined in this disclosure, the definition of the term expressly provided in this disclosure shall control in all respects.

[0023] It is to be understood that this disclosure is not limited to the particular methodology, protocols, and reagents, etc., described herein and as such may vary. The terminology used herein is for the purpose of describing particular embodiments or aspects only and is not intended to limit the scope of the present disclosure.

[0024] Except in the examples or unless otherwise indicated, all numbers expressing quantities of ingredients or reaction conditions used herein should be understood to be modified in all instances by the term "about." When used to describe the invention in connection with percentages, the term "about" means ±1%. As used herein, the term "about" can mean within an acceptable range of error for a particular value as determined by one of ordinary skill in the art, which may depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. Alternatively, "about" can mean within ±20%, ±10%, ±5%, or ±1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, within five-fold, or within two-fold of a value. Where specific values ​​are described in the present application and claims, unless otherwise specified, the term "about" can be assumed to mean within an acceptable range of error for the particular value. Additionally, when ranges and / or subranges of values ​​are provided, the ranges and / or subranges may include the endpoints of the ranges and / or subranges. In some cases, variations may include amounts or concentrations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.

[0025] For the recitation of numerical ranges herein, each intervening number is expressly contemplated to the same degree of precision. For example, for the range 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.

[0026] As used herein, the term "adeno-associated virus" or "AAV" refers to a member of the class of viruses related to this name and belonging to the genus Dependoparvovirus, family Parvoviridae. Multiple serotypes of this virus are known to be suitable for gene delivery, and all known serotypes can infect cells from a variety of tissue types. Non-limiting exemplary serotypes useful in the methods disclosed herein include any of the 11 or 12 serotypes, e.g., AAV2, AAV5, and AAV8, or variant serotypes such as AAV-DJ. AAV structural particles are composed of 60 protein molecules composed of VP1, VP2, and VP3. Each particle contains approximately five VP1 proteins, five VP2 proteins, and 50 VP3 proteins arranged in an icosahedral structure. Non-limiting exemplary VP1 sequences useful in the methods disclosed herein are provided below.

[0027] The term "amino acid" refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function similarly to naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified (e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine). In some embodiments, an amino acid analog refers to a compound that has the same basic chemical structure as a naturally occurring amino acid, i.e., a carbon attached to a hydrogen, a carboxyl group, an amino group, and an R group, such as homoserine, norleucine, methionine sulfoxide, or methionine methylsulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. In some embodiments, an amino acid mimetic refers to a compound that has a structure that differs from the general chemical structure of an amino acid but functions in a manner similar to a naturally occurring amino acid. The terms "non-naturally occurring amino acid" and "unnatural amino acid" refer to amino acid analogs, synthetic amino acids, and amino acid mimetics that are not found in nature. In certain instances, one or more D-amino acids may be used in the various peptide compositions of the present disclosure. The present disclosure provides various peptides useful for treating various diseases and infectious diseases. These peptides may contain naturally occurring amino acids. In other embodiments, the peptides may contain unnatural amino acids. The use of unnatural amino acids can improve peptide stability, reduce degradation, and / or improve biological activity. For example, in some embodiments, one or more D-amino acids. In other embodiments, retro-inverso peptides are contemplated using various amino acid configurations.

[0028] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Similarly, nucleotides may be referred to by their commonly accepted single-letter codes.

[0029] The term "Cas9" refers to CRISPR-associated RNA-guided endonucleases, such as Streptococcus pyogenes Cas9 (spCas9; see Accession No. Q99ZW2.1, the sequence of which is incorporated herein by reference), as well as their orthologs and biological equivalents. Biological equivalents of Cas9 include, but are not limited to, C2c1 from Alicyclobacillus acideterrestris and Cpf1 from various bacterial species, including Acidaminococcus spp. and Francisella novicida U112, which perform cutting / cleaving functions similar to Cas9. Cas9 can refer to endonucleases that create double-strand breaks in DNA, nickase variants such as RuvC or HNH mutants that create single-strand breaks in DNA, and other variants such as deadCas-9 ("dCas9"), which lack endonuclease activity. Cas9 may also refer to "split-Cas9," in which Cas9 is split into two halves: C-terminal Cas9 (C-Cas9) and N-terminal Cas-9 (N-Cas9), which can be fused to two intein moieties. See, e.g., U.S. Pat. No. 9,074,199 (B1); Zetsche et al. (2015) Nat Biotechnol. 33(2):139-42; Wright et al. (2015) PNAS 112(10)2984-89. Non-limiting examples of commercially available sources of SpCas9, including plasmids, can be found under the following AddGene reference numbers: 42230:PX330;SpCas9 and single guide RNA; 48138:PX458;SpCas9-2A-EGFP and single guide RNA; 62988:PX459;SpCas9-2A-Puro and single guide RNA; 48873:PX460; SpCas9n (D10A nickase) and single guide RNA; 48140:PX461; SpCas9n-2A-EGFP (D10A nickase) and single guide RNA; 62987:PX462; SpCas9n-2A-Puro (D10A nickase) and single guide RNA; and 48137:PX165;SpCas9; all of which are incorporated herein by reference.

[0030] As used herein, the term "CRISPR" refers to clustered regularly interspaced short palindromic repeats (CRISPR). CRISPR may also refer to a sequence-specific gene manipulation technology or system that relies on the CRISPR pathway. A CRISPR recombinant expression system can be programmed to cleave a target polynucleotide using a CRISPR endonuclease and a guide RNA. The CRISPR system can be used to create double-stranded or single-stranded breaks in a target polynucleotide. The CRISPR system can also be used to recruit proteins or label a target polynucleotide. In some embodiments, CRISPR-mediated gene editing utilizes the pathways of nonhomologous end-joining (NHEJ) or homologous recombination to perform editing. These applications of CRISPR technology are known in the art and have been widely implemented. See, for example, U.S. Patent No. 8,697,359 and Hsu et al. (2014) Cell 156(6):1262-1278.

[0031] As used herein, the term "delivery vehicle" refers to a composition useful for delivering a payload to a cell, tissue, or subject. Delivery vehicles can deliver a variety of payloads, including biological agents and small molecule drugs. Exemplary, but non-limiting, delivery vehicles include liposomes, nanoparticles, bacteria, bacteriophages, virus-like particles (VLPs), erythrocyte ghosts, and exosomes.

[0032] As used herein, the term "domain" may refer to a specific region of a larger molecule (e.g., a specific region of a protein or polypeptide) that may be associated with a particular function. For example, a "cognate binding domain" may refer to a domain of a protein that binds to one or more receptors or other protein moieties. Similarly, the corresponding coding sequence for a particular polypeptide domain may be referred to as a polynucleotide domain.

[0033] The term "encoding" as applied to a polynucleotide can refer to a polynucleotide that is said to "encode" a polypeptide when, in its natural state or when manipulated by methods well known to those of skill in the art, it can be transcribed and / or translated to produce mRNA for the polypeptide and / or fragments thereof. In some cases, the antisense strand is the complement of such a nucleic acid, and the coding sequence can be deduced therefrom.

[0034] The terms "equivalent" or "biological equivalent" are used interchangeably when referring to a particular molecule, biological substance, or cellular substance, and are intended to have minimal homology while still maintaining the desired structure or functionality.

[0035] As used herein, "expression" can refer to the process by which a polynucleotide is transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.

[0036] As used herein, the term "functional" can be used to modify any molecule, biological substance, or cellular substance with the intent of achieving a particular, specified effect.

[0037] As used herein, the term "gRNA" or "guide RNA" refers to a guide RNA sequence used to target a specific gene for correction using CRISPR technology. Techniques for designing gRNAs and donor therapeutic polynucleotides for target specificity are well known in the art. For example, see Doench, J., et al. Nature biotechnology 2014;32(12):1262-7, Mohr, S. et al. (2016) FEBS Journal 283:3232-38, and Graham, D., et al. Genome Biol. 2015;16:260. A gRNA may comprise, alternatively consist essentially of, or even consist of a fusion polynucleotide comprising CRISPR RNA (CRISPR RNA, crRNA) and trans-activating CRIPSPR RNA (trans-activating CRIPSPR RNA, tracrRNA), or a polynucleotide comprising CRISPR RNA (crRNA) and trans-activating CRIPSPR RNA (tracrRNA). In some embodiments, the gRNA is synthetic (Kelley, M. et al. J of Biotechnology 233 (2016) 74-83).

[0038] "Homology" or "identity" or "similarity" can refer to sequence similarity between two peptides or two nucleic acid molecules. Homology can be determined by comparing a position in each sequence, which can be aligned for purposes of comparison. For example, if a position in the compared sequences is occupied by the same base or amino acid, the molecules are homologous at that position. The degree of homology between sequences is a function of the number of matching or homologous positions shared by the sequences. An "unrelated" or "non-homologous" sequence shares less than 40% identity, or less than 25% identity, with one of the sequences of the present disclosure.

[0039] Homology refers to the percent (%) identity of a sequence to a reference sequence. In practice, any specific sequence may be at least 50%, 60%, 70%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98%, or 99% identical to any sequence described herein. Whether such a specific peptide, polypeptide, or nucleic acid sequence has a specific identity / homology can be conventionally determined using known computer programs such as the Bestfit program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, 575 Science Drive, Madison, Wis. 53711). When using Bestfit or any other sequence alignment program to determine whether a specific sequence is, for example, 95% identical to a reference sequence, parameters can be set so that the percentage of identity is calculated over the entire length of the reference sequence and allows for a homology gap of up to 5% of the entire reference sequence. Several sequences are provided herein, and it is contemplated that sequences having at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, and 100% identity to any one of the sequences herein will find use in any of the compositions and methods described herein.

[0040] For example, in certain embodiments, the identity between a reference sequence (query sequence, i.e., a sequence of the present disclosure) and a subject sequence, also referred to as a global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)). In some cases, identity is interpreted narrowly, and parameters for certain embodiments used in FASTDB amino acid alignment can include the following: scoring scheme = PAM (allowed mutation rate) 0, k-tuple = 2, mismatch penalty = 1, joining penalty = 20, randomization group length = 0, cutoff score = 1, window size = sequence length, gap penalty = 5, gap size penalty = 0.05, window size = 500, or the length of the subject sequence, whichever is shorter. According to this embodiment, if the subject sequence is shorter than the query sequence due to N- or C-terminal deletions rather than internal deletions, a manual correction can be made to the results to account for the fact that the FASTDB program does not consider N- and C-terminal truncations of the subject sequence when calculating the overall percent identity. For N- and C-terminally truncated subject sequences, the percent identity can be corrected by calculating the number of query sequence residues flanking the N- and C-termini of the subject sequence that are not matched / aligned with the corresponding subject residues as a percentage of the total bases of the query sequence. The determination of whether a residue is matched / aligned can be determined by the results of a FASTDB sequence alignment. This percentage can then be subtracted from the percent identity calculated by the FASTDB program using specific parameters to arrive at a final percent identity score. This final percent identity score can be used for purposes of this embodiment. In some cases, only residues to the N- and C-termini of the subject sequence that are not matched / aligned with the query sequence are considered for purposes of manually adjusting the percent identity score. That is, only query residue positions outside the farthest N- and C-terminal residues of the subject sequence are considered for this manual correction.For example, a 90-residue subject sequence can be aligned with a 100-residue query sequence to determine percent identity. The deletion occurs at the N-terminus of the subject sequence, so the FASTDB alignment does not show a match / alignment of the first 10 residues at the N-terminus. Because the 10 mismatched residues represent 10% of the sequence (number of residues at the N-terminus and C-terminus that are not matched / total number of residues in the query sequence), 10% is subtracted from the percent identity score calculated by the FASTDB program. If the remaining 90 residues were perfectly matched, the final percent identity would be 90%. In another example, a 90-residue subject sequence is compared with a 100-residue query sequence. Because the deletion is internal this time, there are no residues at the N-terminus or C-terminus of the subject sequence that do not match / align with the query. In this case, the percent identity calculated by FASTDB is not manually corrected. Again, only residue positions outside the N- and C-termini of the subject sequence, as shown in the FASTDB alignment, that are not matched / aligned with the query sequence are manually corrected for.

[0041] "Hybridization" can refer to a reaction in which one or more polynucleotides react to form a complex stabilized through hydrogen bonding between the bases of nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific manner. The complex can include two strands forming a double-stranded structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction can constitute a step in a more extensive process, such as the initiation of a PCR reaction or the enzymatic cleavage of a polynucleotide by a ribozyme.

[0042] Examples of stringent hybridization conditions include an incubation temperature of about 25°C to about 37°C, a hybridization buffer concentration of about 6xSSC to about 10xSSC, a formamide concentration of about 0% to about 25%, and a wash solution of about 4xSSC to about 8xSSC. Examples of moderate hybridization conditions include an incubation temperature of about 40°C to about 50°C, a buffer concentration of about 9xSSC to about 2xSSC, a formamide concentration of about 30% to about 50%, and a wash solution of about 5xSSC to about 2xSSC. Examples of high stringency conditions include an incubation temperature of about 55°C to about 68°C, a buffer concentration of about 1xSSC to about 0.1xSSC, a formamide concentration of about 55% to about 75%, and a wash solution of about 1xSSC, 0.1xSSC, or deionized water. Generally, hybridization incubation times range from 5 minutes to 24 hours, with one, two, or more wash steps, with wash incubation times of about 1, 2, or 15 minutes. SSC is a 0.15 M NaCl and 15 mM citrate buffer. It is understood that equivalents of SSC using other buffer systems can be used.

[0043] As used herein, the term "immuno-orthogonal" refers to the lack of immune cross-reactivity between two or more antigens. In some embodiments, the antigen is a protein (e.g., Cas9). In some embodiments, the antigen is a viral antigen associated with a particular viral vector (e.g., AAV). As recognized in the art, antigens typically comprise antigenic determinants that have a specific sequence of three-dimensional structure. Furthermore, antigenic determinants may comprise domains or subsequences of a larger polypeptide or molecular sequence. In some embodiments, immuno-orthogonal antigens do not share an amino acid sequence of more than 5, more than 6, more than 7, more than 8, more than 9, more than 10, more than 11, more than 12, more than 13, more than 14, more than 15, or more than 16 consecutive amino acids. In some embodiments, immuno-orthogonal antigens do not share any highly immunogenic peptides. In some embodiments, immuno-orthogonal antigens do not share affinity for a major histocompatibility complex (e.g., MHC class I or class II). Antigens that are immuno-orthogonal are amenable to sequential administration to evade the host immune system.

[0044] The term "immunosilent" refers to an epitope or foreign peptide, polypeptide, or protein that does not elicit an immune response from a host upon administration. In some embodiments, the peptide, polypeptide, or protein does not elicit an adaptive immune response. In some embodiments, the peptide, polypeptide, or protein does not elicit an innate immune response. In some embodiments, the peptide, polypeptide, or protein elicits neither an adaptive nor an innate immune response. In some embodiments, an immunosilent peptide, polypeptide, or protein has reduced immunogenicity.

[0045] As used herein, the term "isolated" can refer to a molecule or biological or cellular material that is substantially free of other substances. In one aspect, the term "isolated" can refer to a nucleic acid, such as DNA or RNA, or a protein or polypeptide (e.g., an antibody or derivative thereof), or a cell or organelle, or a tissue or organ, that has been separated from other DNA or RNA, or proteins or polypeptides, or cells or organelles, or tissues or organs, respectively, present in the natural source. The term "isolated" can also refer to a nucleic acid or peptide that, when produced by recombinant DNA technology, is substantially free of cellular material, viral material, or culture medium, or, when chemically synthesized, is substantially free of chemical precursors or other chemicals. Furthermore, "isolated nucleic acid" is meant to include nucleic acid fragments that are not naturally occurring as fragments and would not be found in the natural state. In some cases, the term "isolated" is also used herein to refer to a polypeptide that has been isolated from other cellular proteins, and is meant to encompass both purified and recombinant polypeptides. In some cases, the term "isolated" is also used herein to refer to cells or tissues that are isolated from other cells or tissues, and is meant to encompass both cultured and engineered cells or tissues.

[0046] "Messenger RNA" or "mRNA" is a nucleic acid molecule that is transcribed from DNA and then processed to remove non-coding portions known as introns. In some cases, the resulting mRNA is transported from the nucleus (or another locus where the DNA resides) and translated into protein. The term "pre-mRNA" can refer to the strand before it is processed to remove the non-coding portions. mRNA has a "U" instead of a "T" in the cDNA coding sequence.

[0047] The term "ortholog" is used in reference to another gene or protein and refers to a homolog of that gene or protein that evolved from the same ancestral source or that is artificially evolved using molecular biology and genetic engineering. An ortholog may or may not retain the same function as the gene or protein to which it is orthologous. Non-limiting examples of Cas9 orthologs include S. aureus Cas9 ("spCas9"), S. thermophiles Cas9, L. pneumophilia Cas9, N. lactamica Cas9, N. meningitides Cas9, B. longum Cas9, A. muciniphila Cas9, and O. laneus Cas9.

[0048] The term "payload" refers to therapeutic and diagnostic agents that can be loaded into or onto a delivery vehicle. Such payloads include biological entities and small molecule entities. Exemplary payload agents include small molecule drugs, biomolecules, viruses, therapeutic agents, prodrugs, gene silencing agents, chemotherapeutic agents, diagnostic agents, and / or components of gene editing systems. Examples of biomolecules include, but are not limited to, nucleic acids (e.g., DNA, RNA, mRNA, modified mRNA, small RNA, siRNA, miRNA, genes, and transgenes), peptides / proteins (including antibodies, enzymes, transcription factors, etc.), viruses, hormones, carbohydrates, lipids, and vitamins. Examples of gene silencing agents include siRNA, chRNA, miR, ribozymes, morpholinos, and esiRNA. Examples of gene editing systems include, but are not limited to, CRISPR-Cas systems, zinc finger nucleases, and TALENs. Examples of diagnostic agents include, but are not limited to, dyes and stains, radioactive tracers, and imaging agents. Examples of anti-cancer and chemotherapeutic agents that may be used with or loaded into the delivery vehicle include, but are not limited to, alkylating agents such as thiotepa and CYTOXAN® cyclophosphamide; alkylsulfonates such as busulfan, improsulfan, and piposulfan; aziridines such as benzodopa, carboquone, metadopa, and uredopa; ethyleneimines and methylameramines (altretamine, triethylenemelamine, triethylenephosphoramide, triethylenethiophosphoramide, and thimerosulfan); tiimethylolomelamine); acetogenins (e.g., bullatacin and bullatacinone); camptothecins (including the synthetic analog topotecan); bryostatin; kallistatin; CC-1065 (including its adozelesin, carzelesin, and bizelesin synthetic analogs); cryptophycins (especially cryptophycin 1 and cryptophycin 8); dolastatin; duocarmycins (including synthetic analogs KW-2189 and CB1-TM1); eleutherobin; pancratistatin; sarcodictin; spongistatin;Nitrogen mustards, such as chlorambucil, chlornaphazine, clofosfamide, estramustine, ifosfamide, mechlorethamine, mechlorethamine oxide hydrochloride, melphalan, nobembicine, phenesterine, prednimustine, trofosfamide, uracil mustard; nitrosoureas, such as carmustine, chlorozotocin, fotemustine, lomustine, nimustine, and ranimustine; vinca alkaloids; epipodophyllotoxins; antibiotics, such as enediyne antibiotics (e.g., calicheamicin, especially calicheamicin gamma II and calicheamicin omega II; L-asparaginase; anthracenedione-substituted ureas; methylhydrazine derivatives; dynemicins, including dynemicin A; bisphosphonates, e.g., clodronate; esperamicin; and neocarzinostatin chromophores and related chromoprotein enediyne antibiotic chromophores), aclacinomycin, actinomycin, ausramycin, azaserine, bleomycin, cactinomycin, carabicin, carminomycin, carzinophilin, chromomycin, dactinomycin, daunorubicin, detol Bicine, 6-diazo-5-oxo-L-norleucine, ADRIAMYCIN® doxorubicin (including morpholino-doxorubicin, cyanomorpholino-doxorubicin, 2-pyrrolino-doxorubicin, and deoxydoxorubicin), epirubicin, esorubicin, idarubicin, marcelomycin, mitomycins such as mitomycin C, mycophenolic acid, nogalamycin, olivomycin, peplomycin, potfilomycin, puromycin, chelamycin, rodorubicin, streptonigrin, streptozocin, tubercidin, ubenimex, zinostatin, zorubicin; antimetabolites such as methotrexate and 5-fluorouracil (5-FU); folic acid analogs such as denopterin, methotrexate, pteropterin, trimetrexate; purine analogs such as fludarabine, 6-mercaptopurine, thiamiprine, thioguanine; pyrimidine analogs such as ancitabine, azacitidine, 6-azauridine, carmofur, cytarabine, dideoxyuridine, doxifluridine, enocitabine, floxuridine;androgens, e.g., calsterone, dromostanolone propionate, epithiostanol, mepitiostane, testolactone; antiadrenal agents, e.g., aminoglutethimide, mitotane, trilostane; folate supplements, e.g., florinic acid acid); aceglatone; aldophosphamide glycoside; aminolevulinic acid; eniluracil; amsacrine; bestravcil; bisantrene; edatraxate; defofamine; demecolcine; diaziquone; elflornithine; elliptinium acetate; epothilone; etoglucide; gallium nitrate; hydroxyurea; lentinan; lonidynin; maytansinoids, such as maytansine and ansamitocins; mitoguazone; mitoxantrone; mopidanol; nithiaerin; pentostatin; phenamet; pirarubicin; losoxanthion; podophyllic acid; 2-ethylhydrazide; procarbazine; PSK® polysaccharide complex (JHS Natural Products, Eugene, Oreg.); razoxane; rhizoxin; schizofiran; spirogermanium; tenuazonic acid; triazicon; 2,2,2"-trichlorotiiethylamine; trichothecenes (especially T-2 toxin, verulaculin A, roridin A, and anguidine); urethane; vindesine; dacarbazine; mannomustine; mitobronitol; mitolactol; pipobroman; gacytosine; arabinoside ("Ara-C"); cyclophosphamide; thiotepa; taxoids, such as TAXOL® paclitaxel (Bristol-Myers Squibb Oncology, Princeton, NJ), ABRAXANE® cremophor-free albumin-engineered nanoparticle formulation of paclitaxel (American Pharmaceutical Partners, Schaumberg, Ill.), and TAXOTERE® (docetaxel) (Rhone-Poulenc Rorer, Antony, France); chlorambucil; GEMZAR® (gemcitabine); 6-thioguanine; mercaptopurine; methotrexate; platinum coordination complexes, such as cisplatin, oxaliplatin, and carboplatin; vinblastine; platinum;Examples of inhibitors include etoposide (VP-16); ifosfamide; mitoxantrone; vincristine; NAVELBINE® vinorelbine; novantrone; teniposide; edatrexate; daunomycin; aminopterin; xeloda; ibandronate; irinotecan (e.g., CPT-11); the topoisomerase inhibitor RFS 2000; difluoromethylornithine (DFMO); retinoids, e.g., retinoic acid; capecitabine; leucovorin (LV); irinotecan; adrenocortical suppressants; corticosteroids; progestins; estrogens; androgens; gonadotropin-releasing hormone analogs; and pharmaceutically acceptable salts, acids, or derivatives of any of the above. Also included are anti-cancer drugs such as anti-estrogens and selective estrogen receptor modulators, including, for example, tamoxifen (including NOLVADEX® tamoxifen), raloxifene, droloxifene, 4-hydroxytamoxifen, trioxifene, ketoxifene, LY117018, onapristone, and FARESTON-toremifene. antihormonal agents that act to regulate or inhibit hormone action on tumors, such as hormone receptor modulators (SERMs); aromatase inhibitors that inhibit the enzyme aromatase, which regulates estrogen production in the adrenal glands, such as 4(5)-imidazole, aminoglutethimide, MEGASE® megestrol acetate, AROMASL® exemestane, formestane, fadrozole, RIVISOR® vorozole, FEMARA® letrozole, and ARTIMIDEX® anastrozole; and antiandrogens such as flutamide, nilutamide, bicalutamide, leuprolide, and goserelin; and troxacitabine (a 1,3-dioxolane nucleoside cytosine analog); antisense oligonucleotides, particularly those that inhibit the expression of genes in signaling pathways involved in abnormal cell proliferation, such as PKC-alpha, Ralf, and H-Ras;Ribozymes, such as VEGF-A expression inhibitors (e.g., ANGIOZYME® ribozyme) and HER2 expression inhibitors; vaccines, such as gene therapy vaccines, e.g., ALLOVECTIN® vaccine, LEUVECTIN® vaccine, and VAXID® vaccine; PROLEUKIN® rJL-2; LURTOTECAN® topoisomerase 1 inhibitors; ABARELLX® rmRH; antibodies such as trastuzumab, and pharmaceutically acceptable salts, acids, or derivatives of any of the above. Examples of protein payloads include mammalian proteins such as growth hormone (GH), including, for example, human growth hormone, bovine growth hormone, and other members of the GH supergene family; growth hormone-releasing factor; parathyroid hormone; thyroid-stimulating hormone; lipoproteins; alpha-1-antitrypsin; insulin A chain; insulin B chain; proinsulin; follicle-stimulating hormone; calcitonin; luteinizing hormone; glucagon; clotting factors, such as factor VIIIC, factor IX, tissue factor, and von Willebrand factor; anticoagulants, such as protein C; atrial natriuretic factor; pulmonary surfactant; plasminogen activators, such as urokinase or tissue-type plasminogen activator. activator, t-PA); bombadin; thrombin; alpha-tumor necrosis factor, beta-tumor necrosis factor; enkephalinase; RANTES (regulated upon activation, normally expressed and secreted by T cells); human macrophage inflammatory protein (MIP-1-alpha); serum albumin, e.g., human serum albumin; Müllerian inhibitory substance; relaxin A chain; relaxin B chain; prorelaxin; mouse gonadotropin-related peptide; DNase; inhibin; activin; vascular endothelial growth factor (VEGF); hormone or growth factor receptor; integrin; protein A or D; rheumatoid factor;neurotrophic factors, such as bone-derived neurotrophic factor (BDNF), neurotrophin-3, -4, -5, or -6 (NT-3, NT-4, NT-5, or NT-6), or nerve growth factors, such as NGF-beta; platelet-derived growth factor (PDGF); fibroblast growth factors, such as aFGF and bFGF; epidermal growth factor (EGF); transforming growth factors (TGF), such as TGF-alpha and TGF-beta, including TGF-beta 1, TGF-beta 2, TGF-beta 3, TGF-beta 4, or TGF-beta 5; insulin-like growth factor-I and -II (IGF-I and IGF-II); des(1-3)-IGF-I (brain IGF-D); insulin-like growth factor binding proteins; CD proteins, such as CD3, CD4, CD8, CD19, and CD20; osteoinductive factors; immunotoxins; bone morphogenetic proteins (BMPs); T-cell receptors; surface membrane proteins; decay accelerating factors (DAFs); viral antigens, such as, for example, portions of the AIDS envelope; transport proteins; homing receptors; addressins; regulatory proteins; immunoadhesins; antibodies; and biologically active fragments or variants of any of the above-listed polypeptides.

[0049] Members of the GH supergene family include growth hormone, prolactin, placental lactogen, erythropoietin, thrombopoietin, interleukin-2, interleukin-3, interleukin-4, interleukin-5, interleukin-6, interleukin-7, interleukin-9, interleukin-10, interleukin-11, interleukin-12 (p35 subunit), interleukin-13, interleukin-15, oncostatin M, ciliary neurotrophic factor, leukocyte inhibitory factor, alpha interferon, beta interferon, gamma interferon, omega interferon, tau interferon, granulocyte colony-stimulating factor, granulocyte-macrophage colony-stimulating factor, macrophage colony-stimulating factor, cardiotrophin-1, and other proteins identified and classified as members of this family.

[0050] Other payload drugs that can be incorporated into the delivery vehicle include gastrointestinal therapeutic agents such as aluminum hydroxide, calcium carbonate, magnesium carbonate, sodium carbonate, and the like; nonsteroidal contraceptives; parasympathomimetics; psychotherapeutic agents; primary tranquilizers such as chlorpromazine hydrochloride, clozapine, mesoridazine, metiapine, reserpine, thioridazine, and the like; secondary tranquilizers such as chlordiazepoxide, diazepam, meprobamate, temazepam, and the like; nasal decongestants; and sedative-hypnotics. other steroids, such as testosterone and testosterone propionate; sulfonamides; sympathomimetics; vaccines; vitamins and nutrients, such as essential amino acids, essential fats; antimalarials, such as 4-aminoquinolines, 8-aminoquinolines, pyrimethamine; antimigraine medications, such as mazindol, phentermine; antiparkinsonian medications, such as L-dopa; Anticonvulsants, such as atropine, methscopolamine bromide, etc.; anticonvulsants and anticholinergics, such as bile therapy, digestive aids, enzymes, etc.; antitussives, such as dextromethorphan, noscapine, etc.; bronchodilators; cardiovascular agents, such as antihypertensive compounds, rauwolfia alkaloids, coronary vasodilators, nitroglycerin, organic nitrates, pentaerythritol tetranitrate, etc.; electrolyte substitutes, such as potassium chloride; ergot alkaloids, such as ergotamine with and without caffeine, hydrogenated ergot These include horn alkaloids, dihydroergocristine methanesulfonate, dihydroergocomine methanesulfonate, dihydroergocloptine methanesulfonate, and combinations thereof; alkaloids, such as atropine sulfate, belladonna, hyoscine hydrobromide, and the like; analgesics; narcotics, such as codeine, dihydrocodienone, meperidine, morphine, and the like; and non-narcotics, such as salicylates, aspirin, acetaminophen, d-propoxyphene, and the like.

[0051] Antibiotic payloads include, for example, cephalosporins, chlorarnphenical, gentamicin, kanamycin A, kanamycin B, penicillin, ampicillin, streptomycin A, antimycin A, chloropamtheniol, metronidazole, oxytetracycline penicillin G, tetracycline, and the like.

[0052] Other payload agents may include vaccines or antigenic agents, for example, payload antigens derived from microorganisms such as Neisseria gonorrhea, Mycobacterium tuberculosis, herpesvirus (humonis, types 1 and 2), Candida albicans, Candida tropicalis, Trichomonas vaginalis, Haemophilus vaginalis, Group B Streptococcus sp., Microplasma hominis, Hemophilus ducreyi, Granuloma inguinale, Lymphopathia venereum, Treponema pallidum, Brucella abortus, and the like. Brucella melitensis, Brucella suis, Brucella canis, Campylobacter fetus, Campylobacter fetus intestinalis, Leptospira pomona, Listeria monocytogenes, Brucella ovis, equine herpesvirus 1, equine arteritis virus, IBR-IBP virus, BVD-MB virus, Chlamydia psittaci, Trichomonas foetus, Toxoplasma gondii, Escherichia coli, Actinobacillus equuli, Salmonella abortus ovis, Salmonella abortus equi, Pseudomonas aeruginosa, Corynebacterium equi, Corynebacterium pyogenes, Actinobaccilus seminis, Mycoplasma bovigenitalium, Aspergillus fumigatus, Absidia ramosa, Trypanosoma equiperdum, Babesia The payload may be loaded with Clostridium caballi, Clostridium tetani, Clostridium botulinum, etc. In other embodiments, the payload may comprise neutralizing antibodies against the above microorganisms.

[0053] In other embodiments, the payload may comprise an enzyme such as ribonuclease, neuramidinase, trypsin, glycogen phosphorylase, sperm lactate dehydrogenase, sperm hyaluronidase, adenosine triphosphatase, alkaline phosphatase, alkaline phosphatase esterase, aminopeptidase, trypsin chymotrypsin, amylase, muramidase, acrosomal proteinase, diesterase, glutamate dehydrogenase, succinate dehydrogenase, beta-glycophosphatase, lipase, ATPase alpha-peptate gamma-glutamylotranspeptidase, sterol-3-beta-ol-dehydrogenase, DPN-di-asprolase, and the like.

[0054] Peptide-payload conjugates are also encompassed by the present disclosure, wherein the peptides of the present disclosure are directly linked or fused to the payload molecules set forth herein (without being packaged in a delivery vehicle) such that the peptide-payload conjugates can be delivered directly to target desired tissues based on the tropism of the peptides.

[0055] The term "promoter," as used herein, refers to any sequence that regulates the expression of a coding sequence, such as a gene. Promoters can be, for example, constitutive, inducible, repressible, or tissue-specific. A "promoter" is a regulatory sequence that is a region of a polynucleotide sequence at which the initiation and rate of transcription are controlled. It can include genetic elements to which regulatory proteins and molecules, such as RNA polymerase and other transcription factors, can bind. Non-limiting exemplary promoters include the CMV promoter and the U6 promoter.

[0056] The terms "protein," "peptide," and "polypeptide" are used interchangeably and, in their broadest sense, refer to a compound of two or more subunit amino acids, amino acid analogs, or peptidomimetics. The subunits may be linked by peptide bonds. In alternative embodiments, the subunits may be linked by other bonds, such as esters, ethers, etc. A protein or peptide can include at least two amino acids, with no limit on the maximum number of amino acids that a protein or peptide sequence can contain. As noted above, the term "amino acid" can refer to any natural and / or unnatural or synthetic amino acid, including glycine and both D and L optical isomers, amino acid analogs, and peptidomimetics. As used herein, the term "fusion protein" can generally refer to a protein composed of domains from two or more naturally occurring or recombinantly produced proteins, with each domain performing a different function. In this regard, the term "linker" can refer to a peptide fragment used to link these domains together and, optionally, to preserve the conformation of the fusion protein domains and / or prevent undesirable interactions between the fusion protein domains that may impair their respective functions.

[0057] The terms "polynucleotide" and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotides may have any three-dimensional structure and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, ESTs, or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, RNAi, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polynucleotide. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by conjugation with a labeling component. The term can also refer to both double-stranded and single-stranded molecules. Unless otherwise specified or required, any embodiment of the present disclosure that is a polynucleotide can encompass both the double-stranded form and each of the two complementary single-stranded forms that are known or predicted to make up the double-stranded form.

[0058] The term "polynucleotide sequence" may refer to an alphabetical representation of a polynucleotide molecule that can be input into a data system of a computer with a central processing unit and used in bioinformatics applications such as functional genomics and homology searching.

[0059] Similarly, the terms "polypeptide sequence," "peptide sequence," or "protein sequence" may be an alphabetical representation of a polypeptide molecule that can be input into a data system of a computer with a central processing unit and used in bioinformatics applications such as functional proteomics and homology searching.

[0060] As used herein, the term "recombinant expression system" refers to a genetic construct for the expression of specific genetic material formed by recombinant means.

[0061] As used herein, the term "recombinant protein" can refer to a polypeptide or peptide produced by recombinant DNA techniques; generally, DNA encoding the polypeptide or peptide is inserted into a suitable expression vector, which is then used to transform a host cell to produce the heterologous polypeptide or peptide.

[0062] The term "sequencing" as used herein may include bisulfite-free sequencing, bisulfite sequencing, TET-assisted bisulfite (TAB) sequencing, ACE sequencing, high-throughput sequencing, Maxam-Gilbert sequencing, massively parallel signature sequencing, Polony sequencing, 454 pyrosequencing, Sanger sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real time (SMRT) sequencing, nanopore sequencing, shotgun sequencing, RNA sequencing, Enigma sequencing, or any combination thereof.

[0063] As used herein, the term "subject" is intended to mean any animal. In some embodiments, the subject may be a mammal, and in further embodiments, the subject may be a cow, horse, cat, mouse, pig, dog, human, or rat.

[0064] As used herein, the terms "transformation" and "transfection" are intended to refer to various art-recognized techniques for introducing exogenous nucleic acid into a host cell, including calcium phosphate or calcium chloride co-precipitation, DEAE-dextran-mediated transfection, lipofection (e.g., using commercially available reagents such as LIPOFECTIN® (Invitrogen Corp., San Diego, CA), LIPOFECTAMINE® (Invitrogen), FUGENE® (Roche Applied Science, Basel, Switzerland), JETPEI™ (Polyplus-transfection Inc., New York, NY), EFFECTENE® (Qiagen, Valencia, CA), DREAMFECT™ (OZ Biosciences, France), or electroporation. Suitable methods for transforming or transfecting host cells can be found in Sambrook, et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), and other laboratory manuals. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art and are described in Sambrook, J., Fritsch, E. F., and Maniatis, T., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor, NY, 1989, and other laboratory manuals. nded.; Cold Spring Harbor Laboratory: Cold Spring Harbor, NY, (1989) and Silhavy, TJ, Bennan, M L and Enquist, LW, Experiments with Gene Fusions; Cold Spring Harbor Laboratory: Cold Spring Harbor, NY, (1984), and Ausubel, F M et. al., Current Protocols in Molecular Biology, Greene Publishing and Wiley-Interscience (1987), each of which is incorporated herein by reference in its entirety. Additional useful methods are described in manuals including Advanced Bacterial Genetics (Davis, Roth and Botstein, Cold Spring Harbor Laboratory, 1980), Experiments with Gene Fusions (Silhavy, Berman and Enquist, Cold Spring Harbor Laboratory, 1984), Experiments in Molecular Genetics (Miller, Cold Spring Harbor Laboratory, 1972), Experimental Techniques in Bacterial Genetics (Maloy, in Jones and Bartlett, 1990), and A Short Course in Bacterial Genetics (Miller, Cold Spring Harbor Laboratory 1992), each of which is incorporated herein by reference in its entirety.

[0065] The terms "treat," "treating," and "treatment," as used herein, refer to ameliorating symptoms associated with a disease or disorder (e.g., cancer, Covid-19, etc.), including preventing or delaying the onset of symptoms of the disease or disorder and / or reducing the severity or frequency of symptoms of the disease or disorder.

[0066] As used herein, the term "vector" can refer to a nucleic acid construct designed for transfer between different hosts, including, but not limited to, a plasmid, a virus, a cosmid, a phage, a BAC, a YAC, etc. In some embodiments, a "viral vector" is defined as a recombinantly produced virus or viral particle containing a polynucleotide that is delivered to a host cell either in vivo, ex vivo, or in vitro. In some embodiments, a plasmid vector can be prepared from a commercially available vector. In other embodiments, a viral vector can be produced from a baculovirus, a retrovirus, an adenovirus, an AAV, etc., according to techniques known in the art. In one embodiment, the viral vector is a lentiviral vector. Examples of viral vectors include retroviral vectors, adenoviral vectors, adeno-associated viral vectors, alphaviral vectors, etc. Infectious tobacco mosaic virus (TMV)-based vectors can be used to produce proteins and have been reported to express Griffithsin in tobacco leaves (O'Keefe et al. (2009) Proc. Nat. Acad. Sci. USA 106(15):6099-6104). Alphavirus vectors, such as Semliki Forest virus-based vectors and Sindbis virus-based vectors, have also been developed for use in gene therapy and immunotherapy. See Schlesinger & Dubensky (1999) Curr. Opin. Biotechnol. 5:434-439 and Ying et al. (1999) Nat. Med. 5(7):823-827. In embodiments where gene transfer is mediated by a retroviral vector, the vector construct can refer to a polynucleotide comprising the retroviral genome or a portion thereof and a gene of interest.Further details regarding modern methods of vectors for use in gene transfer can be found, for example, in Kotterman et al. (2015) Viral Vectors for Gene Therapy: Translational and Clinical Outlook Annual Review of Biomedical Engineering 17. Vectors containing both a promoter and a cloning site to which a polynucleotide can be operably linked are well known in the art. Such vectors are capable of transcribing RNA in vitro or in vivo and are commercially available from sources such as Agilent Technologies (Santa Clara, Calif.) and Promega Biotech (Madison, Wis.). In one aspect, the promoter is a pol III promoter.

[0067] Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." In general, expression vectors useful in recombinant DNA techniques are often in the form of plasmids. As used herein, "plasmid" and "vector" can be used interchangeably. However, the present disclosure is intended to include such other forms of expression vectors, such as viral vectors (e.g., replication-defective retroviruses, adenoviruses and adeno-associated viruses), which serve equivalent functions. Typically, a vector or plasmid contains sequences directing the transcription and translation of the relevant gene, a selectable marker, and sequences enabling autonomous replication or chromosomal integration. Suitable vectors contain a region 5' of the gene harboring transcription initiation controls and a region 3' of the DNA segment controlling transcription termination. Both control regions may be derived from genes homologous to the transformed host cell, although it should be understood that such control regions may also be derived from genes that are not native to the species selected as the production host.

[0068] Typically, a vector or plasmid contains sequences that direct the transcription and translation of the gene fragment, a selectable marker, and sequences enabling autonomous replication or chromosomal integration. Suitable vectors contain a region 5' of the gene harboring transcription initiation controls and a region 3' of the DNA fragment that controls transcription termination. Both control regions can be derived from genes homologous to the transformed host cell, although it should be understood that such control regions can also be derived from genes that are not native to the species selected as the production host.

[0069] Initiation control regions or promoters that are useful for driving expression of the relevant pathway coding region in the desired host cell are numerous and familiar to those of skill in the art. Virtually any promoter capable of driving these genetic elements is suitable for the present invention, including, but not limited to, lac, ara, tet, trp, IPL, IPR, T7, tac, and trc (useful for expression in Escherichia coli and Pseudomonas), the amy, apr, npr promoters and various phage promoters useful for expression in Bacillus subtilis and Bacillus licheniformis, nisA (useful for expression in Gram-positive bacteria; Eichenbaum et al. Appl. Environ. Microbiol. 64(8):2763-2769 (1998)), and the synthetic P11 promoter (useful for expression in Lactobacillus plantarum; Rud et al., Microbiology 152:1011-1019 (2006)). Termination control regions may also be derived from various genes native to the preferred hosts.

[0070] Adeno-associated viruses (AAVs) are common gene therapy vectors; however, their efficacy is hindered by poor target tissue transduction and off-target delivery. We hypothesized that naturally occurring receptor-ligand interactions could be repurposed to engineer tropism. This disclosure provides a method in which all annotated protein ligands known to bind to human receptors were fragmented into tiling 20-mer peptides and displayed at two sites on the surface loops of AAV5 and AAV9 capsids. The resulting capsid library, containing >1 million AAV variants, was screened across nine tissues in C57BL / 6 mice. Tracking variant abundance identified >250,000 variants packaged into capsids, and >15,000 variants that efficiently transduced at least one mouse organ. Twenty-one AAV variants were validated, accurately recapitulating 74.3% of organ tropism predictions, confirming the overall effectiveness of the screening. Systematic ligand tiling, which enabled prediction of putative AAV-receptor interactions, was successfully validated by targeted genetic perturbation. Comprehensive peptide tiling also enabled testing of homologous peptide activity. Interestingly, observed functional peptides tended to originate from specific domains on the ligand. Notably, specific peptides also showed consistent activity across mouse strains, capsid insertion contexts, and capsid serotypes, including novel immunoorthogonal serotypes. Further analysis of the displayed peptides revealed that biophysical attributes were highly predictive of AAV variant packaging and that machine-learnable relationships existed between peptide sequence and tissue tropism. The disclosed comprehensive ligand-peptide tiling and display approach enables manipulation of tropism across diverse viral, virus-like, and non-viral delivery platforms, revealing fundamental receptor-ligand biology.

[0071] The present disclosure provides a number of peptides useful for targeting desired tissues. The present disclosure also provides immunosilent peptides. The present disclosure also provides vectors, delivery vehicles, and peptide-payload conjugates comprising one or more of the peptides of the present disclosure with desired tropism.

[0072] Rational screening strategies have enormous potential to expand the molecular tools available for clinical gene therapy applications. While AAV engineering efforts have spanned more than a decade, advances in DNA synthesis have enabled us to create a data-driven library of AAV variants that leverage existing functional biomolecules from nature. Using natural biomolecules as a defined source of inserted peptides has several advantages over random hexamers (and similar methods). First, natural biomolecules have already been filtered for biological functionality by thousands of years of evolutionary selective pressure. Second, defined libraries allow for robust quantification of the fitness of each AAV variant, allowing for easy stratification of AAV variants by their infectivity across organs of interest. While this method has been primarily applied to engineering AAV, mining properties for functional biomolecules has applications in a wide range of protein engineering challenges, such as engineering orthogonal viral (including lentiviral) and nonviral (including lipid nanoparticle) vectors or identifying biological inhibitors of important protein / protein interactions.

[0073] Furthermore, it is believed that the rational ligand tiling approach described herein may lead to important insights into fundamental AAV receptor binding and transduction more generally. Previous efforts have utilized either large-scale genetic screens or unique observations of the transduction behavior of specific AAV variants. Because the AAV variants described herein display peptides derived from natural receptor-interacting ligands, it is believed that the AAV variants may retain some of this binding ability, which may partially explain their altered transduction profiles.

[0074] This disclosure focuses on adeno-associated viruses (AAVs), which have emerged as a leading vector for gene delivery in clinical applications. While several AAV-mediated therapies have received regulatory approval, efficient delivery of therapy to target tissues is challenging with systemic injections. To overcome this, high viral titers are often used in treatments, which in turn have been associated with potential liver toxicity in clinical trials. Local injections are also problematic, often requiring invasive procedures with potential organ damage and long recovery times. Due to these delivery challenges, some gene therapy approaches have chosen to pursue ex vivo therapeutic designs to overcome targeting issues, where feasible, but this in turn can result in reliance on complex laboratory procedures and high manufacturing costs.

[0075] Therefore, there are numerous ongoing research efforts to improve in vivo therapeutic targeting, and AAV variants have been engineered to specifically target tissues such as the brain and muscle. This has primarily been achieved using strategies that iteratively screen random peptides inserted into the AAV capsid, capsid shuffling, random mutagenesis of the capsid sequence as a whole, or direct chemical manipulation. While random oligomer-mediated AAV capsid mutagenesis has yielded functional capsids with novel properties, stochastic mutation screening strategies limit the ability to predict future functional variants. Therefore, rational and programmable manipulation of viral phenotypes remains an elusive goal.

[0076] Toward rational manipulation of viral function, deep mutation libraries and associated functional screens have enabled systematic mapping of capsid mutational fitness, providing important information that can be used to predict future variant activity. Furthermore, defined libraries of pooled oligonucleotides have been used to insert gene fragments derived from proteins with known affinity for synapses into AAV capsids with the goal of improving retrograde axonal transport. While these methodologies have provided important insights for AAV engineering, much remains unknown regarding how AAV genotype influences packaging and tissue transduction. Therefore, systematic datasets mapping AAV genotype to clinically relevant characteristics, such as organ specificity, are greatly needed. Given the clinical risk of hepatotoxicity and other efficacy issues related to off-target transduction, leveraging screening techniques to obtain highly specific AAV capsids would be of great value to the medical and scientific communities.

[0077] With the goal of developing rational and potentially programmable strategies to enable tissue targeting, it was hypothesized that receptor-ligand interactions, which mediate a spectrum of naturally occurring cell-protein and cell-cell interactions, could be repurposed to engineer AAV tropism. However, neither the interaction interface nor the cell-type specificity profile of receptor-ligand interactions have been fully mapped to enable comprehensive exploration and manipulation. To investigate and address this issue, all annotated protein ligands known to bind to human receptors were fragmented into short peptide tilings and displayed on the surface loops of the AAV capsid (an approach termed "AAV-PepTile"). Short peptides grafted onto a stabilizing molecular scaffold can recapitulate local protein domain structures, and protein-protein interface sites are typically located within 1200–2000 Å. 2Because specific peptide hot loops ranging from 4 to 8 amino acids (AA) contribute maximally to the binding energy of protein-protein interactions, it was hypothesized that inserting a 20-amino acid peptide into the AAV capsid could dramatically alter its binding capacity and, therefore, its transduction profile in vivo. While inserting longer peptides could, in principle, more faithfully mimic the local ligand structure, AAV generally does not tolerate insertions larger than 20-25 amino acids without significantly impairing capsid packaging efficiency. Based on this, experiments proceeded using 20-amino acid peptide plus adjacent two-amino acid linker insertions. In total, a library of >1 million AAV variants was constructed by inserting synthetic gene fragments encoding potential receptor ligands and cell membrane-permeable proteins into one of two surface loops on AAV5 and AAV9. Unlike random peptide libraries, ligand tiling enabled robust quantification of tissue transduction rates for all variants screened. Furthermore, systematic testing of the activity of similar peptides generates predictions of putative receptor interactions that drive AAV variant tropism. Quantifying transduction rates across nine organs identified highly specific variants targeting the brain and lung, as well as muscle- and heart-targeted variants with broader organ transduction. The resulting data linking AAV variant genotype to packaging efficiency and tissue specificity expands our understanding of the AAV fitness landscape and provides a unique resource upon which further data-driven engineering efforts can be built.

[0078] A general schematic of the AAV screening is outlined in Figure 3a. Briefly, in the primary screen, a 20 amino acid peptide library consisting of tiled fragments from all known receptor-interacting ligands, along with other interesting protein classes flanked by glycine-serine linkers, was displayed in two surface-exposed loops of two clinically relevant AAV serotypes, AAV5 and AAV9, with their sequences and specific amino acid insertion sites shown in Table 1.

[0079] [Table 1-1]

[0080] [Table 1-2]

[0081] These libraries were used to generate AAV capsid particles, which were then injected systemically into mice in duplicate. Two weeks later, the heart, lungs, liver, small intestine, spleen, pancreas, kidneys, brain, and skeletal muscle were harvested. Genomic DNA was isolated from these tissues, and the variable peptide regions of interest were amplified and subjected to next-generation sequencing (NGS). AAV abundance in each tissue was quantified by measuring the log2 fold change (log2FC) values ​​in each organ relative to the injected pool.

[0082] A list of candidate "hit peptides" in the AAV9 loop1 and AAV9 loop2 scaffolds was synthesized and subjected to deep mutagenesis to obtain an AAV library with millions of variants, which was then profiled in a secondary screen. As in the primary screen, AAV capsid particles were produced from these libraries and injected into mice, where tropism was quantified. Furthermore, to further evaluate the species-dependent nature of the engineered viral particles, the AAV capsid libraries were pooled from the primary and secondary screens and injected into rhesus macaques, an important preclinical non-human primate (NHP) model. Similar to the mouse experiments, various organs were isolated from the NHPs 3 weeks after injection, and variable peptide regions were amplified and isolated. AAV abundance was quantified in the liver, lung, and four lobes of the brain. Furthermore, to apply additional selection pressure, variable peptide regions from the heart, lung, pancreas, intestine, brain, and muscle were amplified and isolated, cloned, and returned to the AAV9 loop1 and AAV9 loop2 scaffolds. This tertiary screen was then subjected to in vivo selection pressure in mice for 2 weeks, after which tropism was quantified by NGS. For all screens, a "hit" was defined as an engineered AAV variant that exhibited a log2FC for capsids >1 in both replications and a p-value <0.1 in at least one organ.

[0083] Table 2 shows peptides present in the starting library of the primary screen that showed significant in vivo activity across all four scaffolds into which they were grafted. In addition to this list, there is a set of peptides that were consistently identified as hits across both the primary, secondary, tertiary, and NHP screens. These screen-independent hits for both AAV9 loop 1 and AAV9 loop 2 are shown in Table 3.

[0084] [Table 2]

[0085]

Table 3-1

[0086]

Table 3-2

[0087]

Table 3-3

[0088]

Table 3-4

[0089]

Table 3-5

[0090]

Table 3-6

[0091]

Table 3-7

[0092] Tables 4A-4B show the top-performing peptides across all of these described screens. These include mutant versions of peptides originally identified in the primary screen. Information includes the peptide's amino acid sequence, the AAV scaffold into which the peptide was inserted, the organ in which the peptide was a hit, the log2FC in that organ, and the screen in which the peptide was identified. While the peptides presented herein were highlighted for their ability to transduce specific organs, many also exhibit other interesting properties, such as their increased ability to form functional capsids (high titers), their significant detargeting away from the liver, and their ability to simultaneously target unique combinations of organs. These peptides and their variants are expected to have the ability to modulate the tropism of AAV capsids, such as AAV5, AAV9, and 23 immunoorthogonal capsids (shown in Table 5), as well as the tropism of other delivery vehicles and peptide-payload conjugates.

[0093] [Table 4-1]

[0094] [Table 4-2]

[0095] [Table 4-3]

[0096] [Table 4-4]

[0097] [Table 4-5]

[0098]

Table 4-6

[0099]

Table 4-7

[0100]

Table 4-8

[0101]

Table 4-9

[0102]

Table 4-10

[0103]

Table 4-11

[0104]

Table 4-12

[0105]

Table 4-13

[0106]

Table 5-1

[0107]

Table 5-2

[0108]

Table 6-1

[0109]

Table 6-2

[0110]

Table 6-3

[0111]

Table 6-4

[0112]

Table 6-5

[0113]

Table 6-6

[0114]

Table 6-7

[0115]

Table 6-8

[0116]

Table 6-9

[0117]

Table 6-10

[0118]

Table 6-11

[0119]

Table 6-12

[0120]

Table 6-13

[0121]

Table 6-14

[0122]

Table 6-15

[0123]

Table 6-16

[0124]

Table 6-17

[0125]

Table 6-18

[0126]

Table 6-19

[0127]

Table 6-20

[0128]

Table 6-21

[0129]

Table 6-22

[0130]

Table 6-23

[0131]

Table 6-24

[0132]

Table 6-25

[0133]

Table 6-26

[0134]

Table 6-27

[0135]

Table 6-28

[0136]

Table 6-29

[0137]

Table 6-30

[0138] This disclosure demonstrates that one of the discovered variants (AAV9.DKK1) exhibits non-preferred transduction when the cognate receptor for the ligand from which the peptide was derived is knocked out (Figure 5a). Furthermore, due to the systematic tiling nature of the peptide library, it was possible to generate full-length transduction maps spanning the entire residue space of the ligand utilized in this study (Figures 5a and 5c). This not only increases confidence in the screening hits, in that homologous peptides behave similarly, but can also provide insight into key structural domains mediating virus-receptor interactions and, more broadly, provide a platform for further expanding the known protein interactome, especially when the dataset is integrated with the cell surfaceome.

[0139] Furthermore, developing predictive models for AAV infectivity has attracted significant interest from multiple research groups in recent years. The application of machine learning to AAV engineering parallels major advances in machine learning across multiple areas of protein science, such as structure prediction, enzyme activity prediction, and antibody binding optimization. While deep learning and similar black-box methodologies have rapidly become mature technologies, the application of these methodologies to AAV engineering remains severely limited by the lack of available training data. The AAV screening data is an ideal training dataset for several reasons: (1) the experimental design features a large, defined library of variants, meaning that all variants have reliable quantification of packaging infectivity (Figure 1); (2) each variant was screened across a panel of nine major organs to map infectivity across diverse tissue types; (3) the library was inserted across two AAV serotypes and two surface loops, providing important functional information about the capsid context in which the peptides reside; and (4) a large cohort of variants was validated, demonstrating that the screening data are reliable (Figure 4). (5) To illustrate the usefulness of the dataset as training data, the packaging efficiency of AAV variants can be accurately predicted from the biophysical characteristics of the inserted peptides (Figure 2e), and the peptide amino acid sequence is directly predictive of tissue tropism across multiple capsids and insertion sites (Figure 3d). Therefore, the screening data will have great utility for the machine learning and computational biology communities.

[0140] Variants identified through pooled screening have tissue transduction exceeding that of AAV9 in many organs (Figures 5a and 5c), and some variants exhibit dramatic liver detargeting (Figures 4d and 6c), but could be further engineered to enhance potency and specificity. Thus, peptides with at least 85% identity to any one of the peptide sequences provided herein are also contemplated. This may be particularly relevant for specific peptides derived from ligands that bind to receptors expressed on multiple cell types or have promiscuous binding activity. In validation experiments, a standard promoter (CMV) was used to drive expression of the mCherry transgene. Alternatively, tissue-specific promoters could be used to increase the specificity and magnitude of transgene expression in the organ of interest. Furthermore, hit capsids identified here could be further engineered for increased activity. Existing hits could serve as scaffolds for additional rounds of targeted mutagenesis and screening, or peptides could be inserted into both loop 1 and loop 2 of the AAV capsid to increase the valency of the displayed ligand. Furthermore, recently developed direct chemical engineering or peptide display strategies and / or alternative peptide linkers can be utilized to enhance peptide-mediated transduction, and scRNAseq can be used to screen hit variants for more specific cell types within the organ of interest.

[0141] While screening strategies address the obstacles of organ-targeting gene therapy, issues with pre-existing AAV immunity may also limit clinical translation. Therefore, experiments computationally identified and fully characterized a set of four immunoorthogonal AAV capsids (Figures 7a-d) and demonstrated that insertion of a retargeting peptide identified through screening could alter the in vivo tropism of one of these identified AAVs (Figure 7e). This not only increases the number of known serotypes with interesting immunological properties to be explored alongside the current AAV set, but also provides confidence that peptide insertions identified through screening methodologies can be grafted onto diverse biological scaffolds to achieve unique targeting across a wide range of delivery contexts.

[0142] In summary, this disclosure presents a large-scale functional screen of engineered AAV variants spanning over one million total variants derived from two capsids and multiple insertional mutagenesis sites. Using this screening data, validation of 21 AAV variants identified AAV capsids with increased organ transduction across multiple organs (heart, muscle, and lung for AAV9.PDGFC, Figure 4c) as well as highly specific AAV capsids (AAV5.APOA1, AAV9.DKK1, Figure 5a). Improved, broadly targeted AAV variants hold great promise for genetic disorders such as hemophilia A, where overall factor VIII expression levels are paramount. At the same time, highly specific AAV capsids such as AAV5.APOA1 (which has less than 1% of the liver infectivity of WT AAV9) have great utility for neurodegenerative disorders, where maximizing transgene expression in the brain is essential. In addition to the novel variants identified herein, the bulk screening data itself is of high value. Given the size, reliability, and translational relevance of the screening dataset, the dataset can serve as a foundation for future computational engineering of designer AAV capsids.

[0143] The following examples are intended to illustrate but not limit the disclosure, and while they are typical of those that might be used, other procedures known to those skilled in the art may alternatively be used. [Example]

[0144] mouse All animal care and experimental procedures were performed in accordance with the University of California Institutional Animal Care and Use Committee. Male C57Bl / 6J (JAX, #000664) and Balb / cJ (JAX, #000651) mice, 6-8 weeks old, were purchased from Jackson Laboratories and administered systemic retro-orbital injections with either AAV or PBS.

[0145] Cell lines and culture conditions All cells were cultured in a humidified incubator at 37°C with 5% CO. HEK293T cells were cultured in DMEM medium supplemented with 10% FBS, GlutaMAX (1x) (GIBCO), and penicillin-streptomycin (100 U / mL) (GIBCO).

[0146] Visualization of AAV capsid structure To obtain structural representations of AAV capsids, AAV5 and AAV9 structural files were downloaded from the Protein Data Bank and visualized using PyMOL Molecular Graphics System, Version 2.0 (Schroedinger, LLC), with the ramp_new function utilized to color the surface capsid representation.

[0147] Design of a displayed peptide library Each AAV library consisted of 275,298 peptides derived from 6,465 proteins. These protein sources were mined from various protein families, including all protein ligands cataloged in the Guide to Pharmacology database, toxins, nuclear localization signals (NLSs), viral receptor binding domains, albumin and Fc binding domains, transmembrane domains, histones, granzymes, and predicted cell-penetration motifs. In addition to peptides encoding functional biomolecules, 444 control peptides encoding FLAG tags with premature stop codons were included. For all human proteins, the cDNA encoding each protein was in silico fragmented to generate DNA encoding all possible 20-mer peptides. For viral proteins, cell-penetration motifs, and FLAG stop codon controls, the protein sequences were reverse-translated into DNA using the most abundant human codon for each amino acid.

[0148] Oligonucleotide array synthesis, amplification, and cloning Oligonucleotide libraries were synthesized by GenScript as three 91,766-element pools. Each oligonucleotide library was amplified using KAPA Hifi Hotstart Readymix, with the manufacturer's recommended cycling conditions of an annealing temperature of 60°C and a 30-second extension time. The number of PCR cycles was optimized to avoid overamplification of the peptide library. After amplifying each oligonucleotide pool and verifying amplicon size on an agarose gel, the amplified sublibraries were pooled to yield a total of 275,298-element peptide libraries.

[0149] These pools were cloned into loop 1 or loop 2 sites to generate pAAV5L1screen, pAAV5L2screen, pAAV9L1screen, and pAAV9L2screen. AAV rep and cap were flanked by AAV inverted terminal repeat (ITR) sequences to facilitate packaging of the cap gene into recombinant AAV particles.

[0150] Recombinant AAV production Using the above library plasmid pools (AAV5-loop1, AAV5-loop2, AAV9-loop1, and AAV9-loop2), each AAV capsid library was produced by transfecting HEK293T cells in 40 15 cm dishes with the plasmid library pool (diluted 1:100 with pUC19 filler DNA to prevent capsid cross-packaging) and an adenovirus helper plasmid (pHelper). Titers were determined via qPCR using iTaq Universal SYBR green supermix and primers binding to the AAV ITR region. To prepare capsid particles as templates for qPCR, 2 μL of virus was added to 50 μL of alkaline digestion buffer (25 mM NaOH, 0.2 mM EDTA) and boiled for 8 minutes. After this, 50 μL of neutralization buffer (40 mM Tris-HCl, 0.05% Tween-20, pH 5) was added to each sample.

[0151] In vivo evaluation of AAV display libraries Each AAV capsid library was administered retroorbitally to mice in duplicate at a dose of 2E12 vg / mouse for the AAV9-based library or 1E12 vg / mouse for the AAV5-based library. Two weeks after injection, the heart, lungs, liver, intestine, spleen, pancreas, kidney, brain, and gastrocnemius muscle were harvested and placed in RNAlater storage solution. Total DNA was extracted from all mouse tissues using TRIzol reagent and the TNES-6U back-extraction method. The resulting precipitated DNA was centrifuged at 18,000 g for 15 minutes, and the supernatant was discarded. The DNA pellet was then washed three times with 70% ethanol. Finally, the pellet was air-dried and resuspended in 300 μL of EB.

[0152] Preparation of plasmid and capsid DNA for next-generation sequencing To sequence the plasmid libraries (AAV5 / 9 and loop 1 / 2 peptide inserts), 50 ng of plasmid was used as template for a 50 μL KAPA Hifi Hotstart Readymix PCR reaction with an annealing temperature of 60°C and a 30-second extension time. Primers were designed to amplify the peptide-coding region from each sublibrary. The number of cycles was optimized to avoid overamplification. The PCR reaction was purified using a QIAquick PCR Purification Kit according to the manufacturer's protocol. 50 ng of the PCR amplicon was then used as template for a secondary 50 μL KAPA Hifi Hotstart Readymix PCR reaction, to which Illumina-compatible adapters and indexes (NEBNext, catalog number E7600S) were added. The PCR reaction was performed with an annealing temperature of 60°C and a 30-second extension time. To sequence the capsid library, a similar protocol was performed using a modified template amount in step 1 PCR. To prepare capsid particles as templates for PCR, 2 μL of virus was added to 50 μL of alkaline digestion buffer (25 mM NaOH, 0.2 mM EDTA) and boiled for 8 minutes. After this, 50 μL of neutralization buffer (40 mM Tris-HCl, 0.05% Tween-20, pH 5) was added to each sample. 1 μL of this digested capsid mixture was then used as template for a 50 μL PCR reaction. For each sample, the number of cycles was optimized to avoid overamplification, followed by a secondary PCR to add Illumina-compatible adapters and indexes. After generating Illumina-compatible libraries, the plasmid and capsid samples were sequenced on a NovaSeq 6000 using an S4 flow cell generating 100-bp paired-end reads.

[0153] Preparation of tissue DNA for next-generation sequencing To sequence the AAV cap genes from each tissue for pooled screening, a two-step PCR-based library preparation method was used, similar to the plasmid / capsid library. For each organ and replicate, a 300 μL PCR reaction was performed using 120 μL of genomic DNA as template. For each tissue, the number of cycles was optimized via initial qPCR to avoid overamplification of the library. All other parameters, such as primers and melting temperature, were identical to those used for the plasmid library PCR. Following this initial PCR, a secondary PCR was performed as described above, adding Illumina-compatible adapters and indexes. The libraries were then sequenced on a NovaSeq 6000 using an S4 flow cell to generate 100-bp paired-end reads.

[0154] In vivo validation of AAV variants Mice were systemically administered in duplicate with either saline or AAV-variant-mCherry, AAV9-mCherry, or AAV5-mCherry capsids at a dose of 5E11vg / mouse. Three weeks after injection, lungs were inflated with PBS / OCT solution, and the lungs, heart, liver, intestine, spleen, pancreas, kidney, brain, and gastrocnemius muscle were harvested. Each organ was divided, one half placed in RNAlater, and the other half embedded in an OCT block and flash-frozen in a dry ice / ethanol slurry. Total RNA was then isolated from all mouse tissues using an RNA isolation kit (Zymo, catalog no. R2072) with TRIzol reagent and on-column DNase treatment. cDNA synthesis was performed using random primers from the Protoscript cDNA synthesis kit (NEB, catalog no. E6560S). Transgene expression was then quantified via qPCR using iTaq Universal SYBR green supermix and primers binding to the mCherry transcript. GAPDH-specific primers were used to normalize mCherry transgene expression to that of an internal GAPDH control. For histological examination, OCT frozen blocks were cryosectioned at approximately 10 μm thickness, and the tissue slides were then imaged using an Olympus SlideScanner S200. Exposure times between 5 and 1000 ms were used, with the same exposure time used for all samples of a given tissue type. mCherry expression from the tissue sections was then quantified using Olympus OlyVIA software to calculate the average pixel intensity across the entire organ section.

[0155] Quantifying AAV variant abundance from NGS data Starting from the FASTQ sequencing files, a count matrix describing AAV abundance in each sample (plasmid / capsid / tissue) was generated using the MAGeCK (94) 'count' function. Following this, the count matrix was normalized for each sample (via multiplication with a constant size factor) to account for non-identical read depth. Sequencing counts were then transformed by taking the base 2 logarithm of the raw counts after adding pseudocounts. Variants with no counts across all experimental samples were excluded from the analysis.

[0156] Biophysical analysis of AAV capsids Biophysical features of inserted peptides were calculated using the "ProteinAnalysis" module in the Biopython Python package. Variants were considered successful packagers if they had higher abundance in capsid particles compared to the plasmid pool. Support vector machine training and visualization were achieved via the "svm" module in the sklearn Python package (96). UMAP projections of peptide biophysical features were achieved via the "plot" function in the UMAP Python package. All default parameters were used for visualization. Boxplots and hexbin plots were generated using the matplotlib and seaborn Python packages.

[0157] Identification of variants significantly enriched in each tissue To identify variants that successfully transduced each tissue, a one-sample T-test was applied to compare the abundance in capsid particles with the abundance in tissues for each variant. The resulting p-values ​​were adjusted for multiple hypothesis testing using the Benjamini-Hochberg procedure. A variant was considered a significant transducer of an organ if it had an FDR-adjusted p-value <0.05 and a Log2FC >1 in both replicates. When selecting variants for validation experiments, we prioritized variants with inserted peptides identified as hits in multiple capsid / loop contexts, and variants with similar inserted peptides identified that infect the same organ.

[0158] Visualization of tissue transduction from pooled screening Heatmaps for visualizing AAV transduction were generated using the "clustermap" function in the seaborn Python package. Rows and columns were ordered via the scipy "optimal_leaf_ordering" function to minimize the Euclidean distance between adjacent leaves of the dendrogram. UMAP projections to visualize AAV tissue specificity were generated by embedding tissue-level log2 fold changes into two dimensions via the "plot" function in the UMAP Python package. All default parameters were used to generate the embedding. Variants were colored by the organ in which they had the greatest log2 fold change.

[0159] Accuracy assessment of predicted AAV variant tropism For each individually validated variant, the accuracy of both positive and negative predictions of tissue infectivity was assessed. For variants predicted to target a specific organ, the prediction was considered accurate if the individual validation showed greater than 50% of wild-type AAV9 infectivity in that organ. For variants predicted not to target a specific organ, the prediction was considered accurate if the individual validation showed less than 50% of wild-type AAV9 activity.

[0160] Peptide distance projection To calculate the Levenshtein distance between inserted peptides, we used the "Levenshtein" function from the Python package "rapidfuzz" with default parameters (100). After constructing a pairwise distance matrix between all significantly enriched peptides, the matrix was projected to two dimensions via UMAP with metric = "precomputed", n neighbors = 1500, and min dist = 0.1. Clusters of peptides with similar functions were then manually annotated on the resulting plot.

[0161] Convolutional Neural Networks To train a convolutional neural network (CNN) to predict AAV tissue specificity, the AA sequences of inserted peptides were converted to one-hot encodings via the "get_dummies" function from the pandas Python package. Among significantly enriched variants, a variant was considered a transducer for a given organ if its log2FC relative to the capsid in both replicates was greater than 0. The data were then randomly split into training (2 / 3) and validation (1 / 3) datasets. For each variant, the one-hot encoding was reshaped into a 20 x 20 matrix with rows indicating residue positions and columns indicating the presence or absence of specific amino acids. The model architecture was instantiated via the Keras sequential model 102. Briefly, a convolutional layer (Conv1D) with 32 filters, a kernel size of 3, and a "relu" activation was fed into a max pooling layer (MaxPool1D) with a pool size of 2. These layers were followed by another set of convolutional and max-pooling layers, this time with 64 filters within the convolutional layer. These layers were followed by a densely connected layer with unit = 20. Finally, a dropout layer was added with a dropout rate = 0.5. A flattening layer and a final densely connected layer (with sigmoid activation) were then used to output the resulting class probabilities. Separate, independent models were trained for each organ. When training the models, classes (infectious variants vs. non-infectious variants) were weighted proportionally to the inverse of the number of class examples. When calling "model.fit()" to train the CNN, a dictionary describing the class weights was passed via the "class_weight" parameter. Model performance was evaluated via accuracy, area under the receiver operating characteristic curve (AUROC), F1 score, and Matthew correlation coefficient (MCC). Evaluation functions were calculated via built-in Keras functions and plotted via matplotlib. To assess how model accuracy varies as a function of edit distance from the training data, the Levenshtein distance between peptides in the test and training datasets was calculated using the "rapidfuzz" Python package as described above.

[0162] Transmission electron microscopy (TEM) To obtain transmission electron microscopy images of selected AAV variants, 20 μL of AAV solution was applied to Formvar / carbon-coated EM grids. After three washes with water, the sample-containing grids were negatively stained with 2% aqueous uranyl acetate for 1 minute and blotted dry. The EM grids bearing each AAV variant were then imaged at 68,000x magnification using an FEI Tecnai Spirit G2 BioTWIN transmission electron microscope operated at 80 keV.

[0163] Mechanism knockout experiments sgRNA sequences targeting LRP6 or a non-targeting control were identified using CRISPick and cloned into the lentiCRISPR v2 plasmid backbone. Lentivirus was then produced as described (33). Briefly, HEK293T cells were seeded at ~40% confluency the day before transfection. On the day of transfection, Optimem serum-reduced medium was mixed with Lipofectamine 2000 (Thermo Fisher), 3 μg of pMD2.G plasmid, 12 μg of pCMV deltaR8.2 plasmid, and 9 μg of each lentiCRISPR v2 plasmid, and added dropwise onto HEK293T cells after a 30-minute incubation period. 48 hours after transfection, the medium was collected and replaced with fresh DMEM containing 10% FBS. Seventy-two hours after transfection, the supernatant containing viral particles was again collected, pooled, and concentrated to 1 mL using an Amicon-15 centrifugal filter (EMD Millipore) with a 100,000 nmWL cutoff. For lentiviral transduction, HEK293T cells were seeded into 12-well plates at ~20% confluency the day before transduction. On the day of transduction, lentivirus-containing DMEM containing 8 μg / mL polybrene was added to the cells. The medium was then replaced 24 hours later, and then changed to DMEM containing puromycin (2 μg / mL) 28 hours after transfection. After selection, once the cells reached confluency, they were passaged into 24-well plates at ~40% confluency. 24 hours later, AAV9 (4 × 10 9 viral genome) or AAV9.DKK1 (1 × 10 9 DMEM containing either the IgG or IgG virus genome was added to the cells. Cells were then harvested 24 hours later and flow cytometry was performed to quantify mCherry transgene expression.

[0164] Identification of immuno-orthogonal AAV capsids Potential AAV orthologs were first screened by identifying sequences that exhibited similarity to the AAV2 cap gene using the National Center for Biotechnology Information (NCBI) local alignment search tool (BLAST). Incomplete, highly truncated, and highly homologous sequences were then excluded from the selection criteria. Viruses within the human AAV clade and viruses derived from non-mammalian hosts were then excluded. Finally, sequences with high similarity to previously identified human serotypes were removed, resulting in 23 potential AAV orthologs being evaluated.

[0165] Immuno-orthogonal AAV production To clone computationally identified AAV orthologs, capsid sequences were codon-optimized and cloned downstream of the AAV2 rep gene using Gibson Assembly. Immune orthogonal AAV capsids were produced as described for AAV variant validation capsids, and AAV production titers were measured via qPCR using primers that bind to the AAV ITR region. Any ortholog with a production titer within a power of 10 of AAV5 was considered successfully packaged.

[0166] Immuno-orthogonal AAV in vivo transduction Immune orthogonal AAV capsids with sufficient packaging titers were then purified using a 1 x 10 12 C57BL / 6 mice were injected retroorbitally with a dose of 10 viral genomes / mouse. Livers were harvested 3 weeks post-injection, and total RNA was isolated as described above. cDNA was then generated, and transgene expression was quantified via qPCR using iTaq Universal SYBR green supermix and primers that bind to the mCherry transcript. mCherry transgene expression was then normalized to GAPDH, and relative expression was compared to AAV5.

[0167] Assessment of immune cross-reactivity Assays to assess the cross-reactivity of identified immuno-orthogonal antibodies were performed as previously described. Prior to injection, serum was collected via a tail snip procedure, and mice were then injected with 1 x 10 12 Three weeks later, serum was collected from each mouse and antibody cross-reactivity ELISA was performed. 1 x 10 viral genomes / mouse of AAV8, AAV MM2, AAV MG1, AAV MG2, or AAV CH1 were injected in triplicate. 9 Viral genomes were diluted in 1x coating buffer and incubated overnight in each well of a 96-well Nunc MaxiSorp plate. The plate was washed three times for 5 minutes with 1x wash buffer (Bethyl) and blocked with 1x BSA blocking buffer (Bethyl) for 2 hours at room temperature. The wells were then washed again, and serum samples were added at a 1:40 dilution. The plate was incubated for 5 hours at 4°C with shaking. The wells were washed three times, and 100 μL of HRP-labeled goat anti-mouse IgG1 (Bethyl; diluted 1:100,000 in 1% BSA) was added to each well. The secondary antibody was incubated for 1 hour at room temperature, the wells were washed three times, and 100 μL of TMB substrate was added to each well. The optical density at 450 nm was measured using a microplate absorbance reader (BioRad iMark).

[0168] Engineering immuno-orthogonal AAV variants To engineer the immunoorthogonal "AAV MG2" capsid, the AAV9 capsid amino acid sequence was aligned to the MG2 sequence using Clustal Omega (108). The PDGFC peptide coding sequence was then ligated into the "AAV MG2" vector at the appropriate loop 1 position using the same protocol as used to clone the peptide into AAV9. The engineered "AAV MG2.PDGFC" vector capsid was then assayed in vivo similarly to the validation experiments described above using AAV9.

[0169] Ethics and Compliance All experiments involving live vertebrate animals performed at UCSD were conducted in accordance with ethical regulations approved by the UCSD IACUC Committee.

[0170] statistics Unless otherwise noted, differences in means were calculated by unpaired t-test. Where necessary, p-values ​​were adjusted for multiple hypothesis testing via the Benjamini-Hochberg procedure (99). For all figures, unless otherwise noted, * p<0.05, ** p<0.01, *** p<0.001, **** p<0.0001.

[0171] A systematic library of AAV variants displaying fragmented proteins To generate the library, AAV5 and AAV9 were selected as the starting serotypes. This was due to their established clinical utility and two important features: first, AAV5 is evolutionarily more distant from other AAV serotypes in clinical use and has previously been shown to be immuno-orthogonal (relative to other common AAV serotypes, thereby enabling their continuous re-dosing); second, AAV9 has been used extensively for clinical trials and has been shown to cross the blood-brain barrier better than other AAV serotypes in most tissues. To generate a diverse library of AAV variants, we generated a DNA oligonucleotide pool of 275,298 gene fragments (Figures 1a-b). Each gene fragment encoded a 20-amino acid peptide derived from the coding sequence of a ligand with a known extracellular receptor or a gene predicted to have cell-penetrating or internalization properties (Figures 1a-b). Protein ligands were sourced from the Guide to Pharmacology database, a professionally curated list of pharmacological targets and their associated ligands, and cell permeability / internalization functions were inferred through text mining of UniProt entries. Examples of protein classes identified as having potential internalization functions include toxins, histones, granzymes, viral receptor binding domains, and nuclear localization signal domains (NLSs). Mouse and human genomes share 80% of their protein-coding genes and have 85% amino acid sequence identity between orthologs. These genomic similarities extend to the receptors for the ligands in the library (e.g., 86% of mouse GPCR proteins have human orthologs), implying that many human ligands are similarly functional across human and mouse contexts. Pools of single-stranded oligonucleotides encoding these gene fragments were synthesized, amplified into double-stranded DNA via PCR, digested, and ligated to two distinct locations on the AAV5 and AAV9 cap genes. Seamless cloning is made possible by type IIS restriction enzymes that cut outside their recognition sequences.PaqCI sites (encoding a glycine-serine two-amino acid linker) engineered to generate compatible overhangs were inserted into the ends of the peptide-encoding DNA library and AAV5 / AAV9 cap plasmid DNA at sites encoding two distinct surface loops, herein referred to as loop 1 and loop 2. Surface loop 1 (AA 443 and 456 on AAV5 and AAV9, respectively) and loop 2 (AA 576 and 587 on AAV5 and AAV9, respectively) were selected as peptide insertion sites due to their distance from the viral particle core, potentially promoting receptor binding (Figure 1b). While many previous AAV engineering attempts have focused on inserting peptides into surface loop 2, the loop 1 site has been shown to accommodate large insertions, even including full-length fluorescent proteins. Collectively, we generated four libraries of variants spanning two AAV capsids, each with two loop insertion sites. In addition to the protein-coding gene fragments, 444 stop codon-containing gene fragments were included as negative controls. This defined library synthesis methodology was employed to enable quantitative inference of variant packaging and transduction efficiency. The starting plasmid library was sequenced to establish initial relative variant abundances, and packaging efficiency was quantified through comparison to this initial baseline.

[0172] Biophysical factors in AAV encapsidation To quantify how well different AAV cap variants package into functional capsids, we generated recombinant AAV particles using engineered AAV5 and AAV9 cap plasmid libraries via transient triple transfection of HEK293T cells (Figure 2a). These viral particles were treated with benzonase to degrade residual plasmid DNA and then subjected to next-generation sequencing (NGS) to quantify relative variant abundance. Packaging efficiency was quantified by ranking AAV variants by the log2 fold change (log2FC) of their relative capsid abundance compared to their count in the plasmid pool (Figure 2b). Using this method, we identified over 250,000 AAV variants that were efficiently packaged into functional AAV capsids with a log2FC > 0. To validate this packaging evaluation function quantified from the screening data, we produced 25 AAV capsids, including 23 identified as successful packagers (log2FC>0) and 2 identified as non-packagers (log2FC<0). In these individual validations, the two non-packagers yielded titers >10-fold lower than those identified as packagers, thus providing confidence in the AAV packaging evaluation function. Consistent with their disruption of AAV capsid structure, there was also a depletion of non-functional stop codon-controlled AAV variants in the capsid pool. Importantly, this confirmed the lack of library cross-packaging during AAV production (Figure 2C).

[0173] To better understand which characteristics led to successful encapsidation, we analyzed the biophysical characteristics of the inserted peptides that yielded successfully packaged AAV variants. Peptide charge, flexibility, alpha-helical content, and hydrophobicity were all found to be significantly different between packaged and unpackaged AAV variants (Figure 2d). The set of successfully packaged variants had a narrower charge distribution than unpackaged variants, suggesting that peptides with extreme charge density negatively impact encapsidation. Successfully packaged AAV variants also had inserted peptides with higher flexibility, lower alpha-helical content, and lower hydrophobicity than unpackaged variants. The observed depletion of hydrophobic peptide-displaying variants is consistent with the solvent-exposed nature of AAV surface loops.

[0174] To build an integrated model predicting whether an AAV variant would be packaged based on the biophysical features of the inserted peptide, we trained a support vector machine classifier using the charge, flexibility, alpha-helical content, and hydrophobicity of the peptides in the dataset (Figure 2e). While all of these biophysical features were significantly different when comparing packaged and nonpackaged AAV variants, the magnitude of this difference was relatively moderate for each individual feature (Figure 2d). Collectively, however, these features were sufficient to train a model capable of distinguishing between packaged and nonpackaged AAV variants (area under the receiver operating characteristic curve = 0.89). Visualization of this class separation was possible by embedding the charge, flexibility, alpha-helical content, and hydrophobicity of the inserted peptide for each AAV variant into two dimensions using uniform manifold approximation and projection (UMAP) (Figure 2f). The resulting embedding placed the AAV variants into distinct clusters, indicating that while each underlying biophysical feature is continuous, there are separable groups of AAV variants with similar biophysical features. Packaging AAV variants tended to cluster with other packaging variants in this unsupervised embedding, further supporting the predictive power of these four biophysical features. This thorough quantification and analytical framework for assessing AAV packaging, enabled by the diversity and length of the inserted peptide library, is of critical translational importance for the identified AAV variants, as AAV production costs are directly related to their packaging capacity.

[0175] High-throughput mapping of engineered AAV tissue tropism After generating a library of recombinant AAV particles packaging their own cap genes, the following four virus pools were injected in duplicate into C57BL / 6 mice (Figure 3a). Two weeks later, the mice were sacrificed, and the peptide-containing region of the AAV cap gene was amplified from DNA isolated from mouse liver, kidney, spleen, brain, lung, heart, skeletal muscle, intestine, and pancreas. The relative abundance of each AAV variant was quantified via NGS, and log enrichment was calculated for each variant relative to its abundance in the capsid pool. Good correlation between replicates was observed (Figure 3b). Using this data, over 15,000 variants that efficiently infected at least one mouse tissue were identified (Figure 3c). The spleen and liver were the most frequent tissue targets of infectious variants, consistent with the established wild-type (WT) tropism of AAV5 and AAV9 for the liver and more recent studies showing that AAV5 and AAV9 readily transduce the spleen. Consistent with the high therapeutic AAV doses required to achieve clinical efficacy in muscle-targeted gene therapy and the challenges of delivery across the blood-brain barrier, the fewest AAV variants targeted to skeletal muscle and brain were identified. A significant number of AAV variants were composed of the same peptide inserted across different AAV capsids and insertion sites, lending credence to the hypothesis that tropism reprogramming is at least partially peptide-specific (Figure 3c).

[0176] Toward mapping and understanding tissue transduction patterns mediated by inserted peptides, we analyzed the feasibility of training a predictive model relating inserted peptide sequences to tissue tropism. Inspired by contemporary studies using convolutional neural networks (CNNs) to predict antibody specificity, we trained a CNN multi-label classifier to predict AAV tissue tropism using one-hot encoded inserted peptide sequences as input features (Figure 3d). To evaluate performance, we trained the model using a random selection of two-thirds of significantly enriched AAV variants and evaluated it on one-third of the test dataset. This CNN model architecture performed well across all organs, with a minimum accuracy of 72% in the kidney (Figure 3d). The highest F1 scores and Matthew correlation coefficients (MCCs) were observed for the liver and spleen, likely due to the large number of liver- and spleen-targeting variants identified in the pooled screen (Figure 3c). Additionally, for the liver and spleen, the CNN was observed to have only a slight reduction in F1 score when evaluated against peptides that were highly divergent (more than 10 edits away) from any of the training examples, implying that the model had learned generalizable features relevant to AAV transduction. For other organs, the model was much less accurate when predicting tissue tropism for peptides highly divergent from the training data, likely due to a reduced amount of positive training examples. Taken together, the ability to predict AAV tropism supports the conclusion that inserted peptides mediate tissue tropism retargeting and that a learnable relationship exists between peptide sequence and tissue tropism.

[0177] Engineered AAV variants with clinically relevant tissue tropism Similar to the WT scaffold from which they were derived, transduction of multiple organs was a nearly ubiquitous phenotype among the identified infectious AAV variants (Figure 4a). Significant transduction of the liver and spleen was observed for the majority of infectious variants, regardless of which other organs were co-transduced. This was also true for variants with insertions in surface loops known to be involved in WT capsid receptor binding. While liver and spleen targeting was nearly ubiquitous, we were able to identify variants that specifically targeted one other organ in addition to the liver / spleen, as well as variants that transduced all tissues at high levels (Figure 4a). When variants were hierarchically clustered based on their tissue detection levels, variants derived from the same sublibrary tended to cluster together, suggesting that the tissue specificity of the wild-type scaffold was, at least in part, a determinant of the engineered variant tropism. Hierarchical clustering of organ samples resulted in replicate clustering together, lending further confidence to the reliability of the screening results. To visualize the overall screening results, the tissue detection levels for each variant were embedded in two dimensions using UMAP, and variants were colored by the organs they most readily transduce (Figure 4b). In this reduced dimensional space, organ-specific clusters can be readily identified, with liver- and spleen-targeted variants being particularly prominent.

[0178] To confirm the tissue tropism of the novel AAV variants identified through the pooled screening, 21 variants were individually produced and validated by quantifying their ability to package and deliver the mCherry transgene in vivo (Figure 4c). All 21 variants were significantly enriched in at least one organ, and we prioritized variants for validation that were internally concordant within the screening data: concordant AAV capsids were defined as hits identified in which other variants with similar inserted peptides enriched in the same organ. Variants were first characterized by in vivo quantification of tissue tropism at the mRNA level (via RT-qPCR quantification of mCherry transgene expression) (Figure 4c). The tissue tropism of the variants largely recapitulated the screening predictions (Figure 4c), with 74.3% of tissue tropism predictions concordant. AAV variants were identified that specifically target organs difficult to infect, such as muscle, lung, and brain, while simultaneously detargeting away from the liver (Figure 4c). Notably, protein-level quantification of mCherry delivery to the liver, quantified via fluorescence microscopy, confirmed excellent agreement between mRNA and protein measurements of tissue transduction (R 2 = 0.97) (Fig. 4d). In summary, across the variants tested individually, 9 / 21 variants were found to exceed AAV9 infectivity in at least one organ. Variants were identified that exceeded AAV9 infectivity in all organs except the liver and pancreas, which had maximum relative transduction of 98.8% and 82.2% of WT AAV9, respectively. In addition, 18 / 21 variants had less than half the liver transduction of WT AAV9, and three variants had AAV9 liver transduction levels less than 5% (Fig. 4d), indicating clinically significant liver detargeting. To increase confidence in the fidelity of the variants individually validated in C67BL / 6 mice, three of the 21 variants were individually validated in BALB / c mice. Notably, their relative tropism was consistent across the two mouse strains.

[0179] Mechanistic insights into AAV reprogramming Having confirmed the efficacy of the AAV variants, we conducted experiments to explore the mechanisms underlying their reprogrammed tropism, particularly to investigate the hypothesis that AAV variants displaying peptides derived from distinct ligand domains could actually drive their transduction patterns. First, we performed experiments to examine how peptides derived from specific ligand protein regions alter AAV tropism. This was motivated by the observation that tiled peptide enrichment patterns in the screen can provide insight into the functional domains of the proteins from which they are derived. We then analyzed how peptides were derived. To do this, we performed experiments to examine the tissue-specific AAV variants identified in the screen, focusing initially on a lung-specific variant (termed AAV9.DKK1) induced via display of a peptide fragment from the DKK gene on the AAV9 loop2 scaffold. Specifically, as shown in Figure 5a (top left), we observed significantly enhanced lung transduction from peptides derived from specific domains in the N-terminus of the protein. After confirming by electron microscopy that this variant was packaged into functional capsids (Figure 5a, top right), the next experiment quantified its transduction profile in vivo, confirming that AAV9.DKK1 is indeed an efficient lung transduction variant, with expression >2-3-fold higher than AAV9 when quantified at both the RNA and protein levels, and consistent detargeting across all other organs compared to AAV9 (Figure 5a, bottom left and bottom right).

[0180] Next, we performed experiments on DKK1-displayed peptides to determine insights into the mechanism of transduction. Interestingly, the N-terminal region of the lung-enriched peptide (Figure 5a, top left) is centered on an evolutionarily conserved linear peptide motif that mediates binding to low-density lipoprotein receptor-related proteins 5 and 6 (LRP5 / 6). Because HEK293T cells strongly express LRP6, we utilized lentivirus-mediated CRISPR-Cas9 to disrupt LRP6 expression in these cells and examine the relative transduction potential of AAV9.DKK1 variants in this system (Figure 5b). Specifically, cells transduced with either a non-targeting control (NTC) guide or an LRP6-targeting guide were then transduced with either AAV9.DKK1 or wild-type AAV9. Using flow cytometry to quantify AAV delivery of the mCherry transgene, we observed a significant reduction in infectivity of AAV9.DKK1 in the LRP6 knockout population, without a corresponding reduction in infectivity of wild-type AAV9 (Figure 5b). The LRP6-dependent in vitro infectivity of AAV9.DKK1, combined with the known interaction between LRP6 and the N-terminus of DKK1 (crystal structure, Figure 5b) and its known robust expression in alveolar cells, provides strong support for the hypothesis that LRP6 mediates, at least in part, the reprogrammed tropism of AAV9.DKK1.

[0181] Based on these observations, and focusing on the brain as an example, we conducted experiments to more broadly examine the extent to which inserted peptides mediate retargeting. First, we generated a distance matrix quantifying the similarity between all peptide hits identified as significantly enriched in at least one organ (Figure 6a). Projecting this distance matrix in two dimensions allowed visualization of distinct clusters of homologous inserted peptides that infect the brain. We found that families of similar inserted peptides resulted in similar transduction rates (Figure 6a). This effect was not restricted to a specific capsid or inserted site, insofar as inserted peptides were often functional across all capsids and inserted sites tested (Figure 6a). The observation that inserted peptides resulted in consistent phenotypes across multiple capsids suggests that retargeting is directly attributable to a peptide-mediated mechanism. This result also highlights the power of peptide tiling library design, in that multiple overlapping peptides with similar sequences can serve as internal controls, conferring confidence in the identification of specific AAV variant hits.

[0182] We conducted experiments to specifically examine two brain-targeting variants (AAV5.APOA1 and AAV9.APOA1) that contained the same inserted apolipoprotein-A 1 (APOA1)-derived peptide. Electron microscopy confirmed that both variants were packaged into functional capsids (Figure 6b). Across the APOA1 protein, functional APOA1-derived peptides were primarily derived from the tandem repeat region at the C-terminus of the protein (Figure 6b). Highlighting the peptide-specific effects of engineered AAV tropism, we observed robust detargeting away from the liver at AAV9 levels below 2% for both variants, while brain transduction was unperturbed compared to wild-type AAV9 (Figure 6c). This was particularly striking for AAV5.APOA1, as the engineered variant was able to overcome the limited transduction of its parent AAV serotype across the blood-brain barrier. The improved specificity of AAV5.APOA1 was subsequently shown to be strain-independent, as it retained both its brain-targeting and liver-detargeting functions in BALB / c mice (Figure 6d).

[0183] Engineering retargeted immuno-orthogonal AAV vectors To generate additional variants with clinically relevant characteristics, we expanded our AAV engineering approach to include novel capsids beyond AAV5 and AAV9. In particular, we explored the possibility of engineering the tropism of highly diverged AAV orthologs. This was motivated by the fact that one of the greatest challenges associated with AAV therapy is the associated immune response: currently, a significant proportion of the human population either has pre-existing immunity that precludes eligibility for any AAV therapy, or induced immunity upon AAV injection that prevents subsequent AAV rechallenge. This allows for only a single opportunity to treat patients and therefore often requires very high titers for a single therapy, which can then lead to dangerous levels of toxicity, making it extremely difficult to create an effective therapy. Recently, the concept of immune orthogonality has been implemented to address this challenge. Specifically, it has been suggested that orthologs with sufficient sequence divergence do not cross-react with the immune response generated by exposure to other orthologs, thereby enabling rechallenge that avoids pre-existing neutralizing antibodies or clearance of treated cells by activated cytotoxic T cells. While the study focused on commonly used AAV orthologs, experiments were also conducted to explore highly divergent natural orthologs from the entire spectrum of mammalian species, including those not yet thoroughly tested for use as in vivo vectors. Toward this end, a computational experimental pipeline was established to identify novel immuno-orthogonal AAV serotypes. First, we used the BLAST (Blast Local Alignment Search Tool) to identify 687 capsids with sequence homology to the AAV2 cap gene. This initial list was then filtered to exclude truncated genomes, redundant samples, human and non-mammalian serotypes, and close orthologs, yielding 23 AAV capsid sequences for experimental investigation (Figure 7a). Analysis of the evolutionary distance of these potential AAV orthologs confirmed that they exhibited high dissimilarity from all wild-type AAV capsids (except for AAV5, previously shown to be immuno-orthogonal), with many exhibiting less than 60% sequence similarity to AAV2 (Figure 7b).To determine whether these computer-predicted AAV orthologs were functional, we examined their ability to package into capsids. Comparing their titers to AAV5, 11 of the 23 AAV capsids were identified with packaging titers within 10-fold of AAV5 (Figure 7c). Next, to assess their in vivo transduction ability, these 11 AAV capsids carrying the mCherry transgene were individually administered to C57BL / 6 mice. Given that most AAV capsids have at least some liver transduction ability, the livers of these mice were harvested 3 weeks post-infection. Four of the 11 analyzed AAV capsids had detectable levels of mCherry in the liver, two of which produced mCherry expression three-fold higher than AAV5 (Figure 7c).

[0184] Four functional AAV capsids were identified and their immunoorthogonality characterized by injecting each of them into C57BL / 6 mice and collecting serum at days 0 and 21 (Fig. 7d). AAV8 was chosen for comparison in this study due to its propensity to transduce the liver and its similar serological profile to other wild-type serotypes. ELISA at the 3-week time point confirmed that antibodies raised against the four studied immunoorthogonal AAVs did not indeed exhibit immunological cross-reactivity with AAV8 (Fig. 7d). This result highlights the utility of combining computational and experimental methods to identify and characterize novel AAV serotypes that can be repurposed to enable gene therapy re-administration.

[0185] The next experiment was performed to evaluate whether the tropism of the mined AAV immunoorthogonal peptides could be reengineered through the insertion of one of the hit peptides identified by the screening. Muscle-targeting AAV variants have broader tissue tropism than brain- and lung-targeting variants, insofar as they also readily infect the heart, lung, intestine, and spleen at levels comparable to or even exceeding those of AAV9. We hypothesized that grafting a peptide from one of these variants onto an AAV ortholog would result in enhanced muscle targeting with minimal effect on parental AAV infectivity. To evaluate this, we selected AAV9.PDGFC as a candidate variant and confirmed its ability to form functional capsids (Figure 6d) and enable efficient targeting of muscle tissue (>10-fold expression compared to AAV9, as quantified by RT-qPCR) (Figure 6d). Enhanced skeletal muscle and cardiac (heart) transduction compared to AAV9 was also confirmed by protein level visualization (Figure 6d). We then grafted a 20-mer PDGFC-derived peptide onto one of the surface loops of the immunoorthogonal AAV MGS (Figure 7e), and observed similar enhanced muscle transduction of the engineered variant over its parent AAV (Figure 7e). Collectively, these results highlight the fidelity of the ligand tiling display approach and lend credence to the exciting possibility that displayed peptides identified in the initial screening approach could be utilized to re-engineer diverse clades of AAV serotypes via peptide transfer.

[0186] It will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A method for improving the tropism of a virus or other delivery agent, comprising identifying ligand protein sequences derived from all known receptor-interacting ligands, systematically tiling the ligand peptides into 5-50 or 10-20 or 20 amino acid peptides that are inserted into surface-exposed loops of AAV capsids, and evaluating the engineered capsids for their packaging capacity, in vivo tropism, and enhanced protein interactions.

2. 2. The method of claim 1, wherein the virus is an adeno-associated virus (AAV).

3. 3. The method of claim 2, wherein the AAV is selected from the group consisting of AAV1, AAV2, AAV5, AAV6, AAV7, AAV8, and AAV9.

4. The method of claim 3, wherein the AAV is AAV5 or AAV9.

5. 4. The method of claim 3, wherein peptide sequences were generated via pooled oligonucleotide synthesis and inserted into four distinct loop regions: AAV5-loop1 (N443), AAV5-loop2 (S576), AAV9-loop1 (Q456), and AAV9-loop2 (A587), generating over one million AAV variants.

6. 4. The method of claim 3, wherein the 20'-mer ligand peptide is inserted into one or both of the two surface-exposed loops of AAV5 (SEQ ID NO: 2) or AAV9 (SEQ ID NO: 4).

7. 1. A recombinant vector comprising a capsid protein, said capsid protein comprising one or two peptides inserted into one or both of two surface-exposed loops in said capsid protein, said vector having a desired tropism or immuno-orthogonality, wherein said peptides are independently selected from the group consisting of SEQ ID NOs: 5-820 and 870-911.

8. The recombinant vector of claim 7, wherein the vector is an adeno-associated virus (AAV).

9. The recombinant vector of claim 8 , wherein the AAV vector is an AAV5 serotype.

10. 10. The recombinant vector of claim 9, wherein the AAV5 comprises a capsid protein having the sequence shown in SEQ ID NO:

2.

11. The recombinant vector of claim 8 , wherein the AAV vector is an AAV9 serotype.

12. The recombinant vector of claim 11, wherein the AAV9 comprises a capsid protein having the sequence set forth in SEQ ID NO:

4.

13. 8. The recombinant vector of claim 7, wherein the vector has tropism for the pancreas, heart, brain, lung, liver, kidney, muscle, spleen, or intestine.

14. 13. The recombinant vector of claim 10 or 12, wherein a peptide of SEQ ID NO: 530-820, 870-910, or 911 is inserted into loop 1 and / or loop 2 of an AAV5 capsid or an AAV9 capsid.

15. The recombinant vector of claim 7 , wherein the vector is immunoorthogonal.

16. 16. The recombinant vector of claim 15, wherein the vector is an adeno-associated virus (AAV).

17. 17. The recombinant vector of claim 16, wherein the AAV comprises a capsid protein of any one of SEQ ID NOs: 822, 824, 826, 828, 830, 832, 834, 836, 838, 840, 842, 844, 846, 848, 850, 852, 854, 856, 858, 860, 862, 864, or 866, or a sequence which is at least 85% to 99% identical to any of the foregoing sequences.

18. A method for producing a delivery vehicle or vector with a desired tropism, comprising: selecting a peptide sequence from any one of SEQ ID NOs: 5-820 or 870-911; and (i) cloning a nucleic acid sequence encoding the peptide into a coding sequence of a capsid protein at an exposed loop site to obtain a recombinant capsid coding sequence, and using the recombinant capsid coding sequence to produce the vector; or (ii) inserting the peptide into an exposed surface of the delivery vehicle.

19. 19. The method of claim 18, wherein the vector is an adeno-associated virus (AAV).

20. 20. The method of claim 19, wherein the AAV is selected from the group consisting of AAV1, AAV2, AAV5, AAV6, AAV7, AAV8, and AAV9.

21. 21. The method of claim 20, wherein the AAV is AAV5 or AAV9.

22. 22. The method of claim 21, wherein the AAV5 capsid coding sequence comprises SEQ ID NO:

1.

23. 22. The method of claim 21 , wherein the AAV9 capsid coding sequence comprises SEQ ID NO:

2.

24. 19. An AAV vector, wherein the capsid protein has been modified according to the method of claim 18.

25. An AAV vector comprising a capsid protein, wherein the capsid protein expresses a peptide of any one of SEQ ID NOs: 5-820 or 870-911 in surface-exposed loop 1 and / or loop 2 of the capsid protein.

26. 26. The AAV vector of claim 25, wherein the wild-type capsid protein sequence comprises SEQ ID NO: 2 or 4.

27. 26. The AAV vector of claim 25, wherein the AAV vector has an AAV5 or AAV9 serotype.

28. 26. The AAV vector of claim 25, wherein the vector has a desired tropism.

29. 29. The AAV vector of claim 28, wherein the AAV vector has tropism for the pancreas, heart, brain, lung, liver, kidney, muscle, spleen, or intestine.

30. A viral vector having a capsid protein, wherein the capsid protein comprises a heterologous targeting peptide ranging from 10 to 30 amino acids in length inserted into a surface-exposed portion of the capsid protein, wherein the targeting peptide is set forth in any one of SEQ ID NOs: 5-820 or 870-911.

31. 31. The viral vector of claim 30, wherein the heterologous targeting peptide is about 15 to 25 amino acids in length.

32. 32. The viral vector of claim 31, wherein the heterologous targeting peptide is about 20 amino acids in length.

33. The viral vector of claim 30, wherein the viral vector is an adeno-associated virus (AAV).

34. The viral vector of claim 30, wherein the viral vector is a lentiviral vector.

35. 31. The viral vector of claim 30, wherein the capsid protein is a VP1 capsid protein.

36. 31. The viral vector of claim 30, wherein the capsid protein is a VP2 capsid protein.

37. 31. The viral vector of claim 30, wherein the capsid protein is a VP3 capsid protein.

38. 31. The viral vector of claim 30, wherein the heterologous targeting peptide is inserted into the AAV capsid protein at loop 1 and / or loop 2.

39. The viral vector of claim 30, wherein the viral vector is AAV5.

40. The viral vector of claim 30, wherein the viral vector is AAV9.

41. 40. The viral vector of claim 39, wherein the heterologous targeting peptide is flanked by linker peptides at the N-terminus and C-terminus of the heterologous targeting peptide.

42. 41. The viral vector of claim 40, wherein the heterologous targeting peptide is flanked by linker peptides at the N-terminus and C-terminus of the heterologous targeting peptide.

43. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to hepatocytes or liver tissue.

44. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to neural cells or brain tissue.

45. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to pancreatic cells or tissue.

46. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to cardiac cells or tissue.

47. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to lung tissue.

48. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to intestinal tissue.

49. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to spleen tissue.

50. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to kidney cells or tissue.

51. 31. The viral vector of claim 30, wherein the heterologous targeting peptide targets the viral vector to a muscle cell or tissue.

52. 1. An adeno-associated virus (AAV) capsid protein comprising a heterologous targeting peptide cloned into loop 1 and / or loop 2 of said capsid protein, said heterologous targeting peptide being about 10-30 amino acids in length and being comprised in or containing any one of the peptides of SEQ ID NOs: 5-820 or 870-911.

53. 53. The AAV capsid protein of claim 52, wherein the capsid protein is a VP1 capsid protein.

54. 53. The AAV capsid protein of claim 52, wherein the capsid protein is a VP2 capsid protein.

55. 53. The AAV capsid protein of claim 52, wherein the capsid protein is a VP3 capsid protein.

56. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide is about 15 to 25 amino acids in length.

57. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide is about 20 amino acids in length.

58. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide is flanked by linker peptides at the N-terminus and C-terminus of the heterologous targeting peptide.

59. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets hepatocytes or liver tissue.

60. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets neural cells or brain tissue.

61. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets pancreatic cells or pancreatic tissue.

62. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets cardiac cells or tissue.

63. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets lung tissue.

64. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets intestinal tissue.

65. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets spleen tissue.

66. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets kidney cells or kidney tissue.

67. 53. The AAV capsid protein of claim 52, wherein the heterologous targeting peptide targets a muscle cell or tissue.

68. A recombinant AAV (rAAV) comprising a capsid protein according to any one of claims 52 to 67.

69. A recombinant AAV (rAAV) comprising a capsid protein having a targeting peptide in loop 1 and / or loop 2, wherein the targeting peptide is independently selected from SEQ ID NOs: 5-820 or 870-911.

70. The recombinant AAV of claim 69, wherein the recombinant AAV further comprises a heterologous polynucleotide for gene delivery.

71. 71. The recombinant AAV of claim 70, wherein the heterologous polynucleotide is a therapeutic gene.

72. A composition comprising the recombinant rAAV of any one of claims 69 to 71.

73. 73. The composition of claim 72, further comprising a pharmaceutically acceptable carrier.

73. 1. A method for delivering a transgene to a subject, comprising administering to the subject a recombinant AAV (rAAV), wherein the rAAV: (i) a capsid protein according to any one of claims 52 to 67; and (ii) at least one transgene, wherein the rAAV infects cells of a target tissue of the subject.

74. 74. The method of claim 73, wherein the at least one transgene encodes a protein.

75. 75. The method of claim 74, wherein the protein is an immunoglobulin heavy or light chain or a fragment thereof.

76. 74. The method of Claim 73, wherein said at least one transgene encodes a small interfering nucleic acid.

77. 77. The method of claim 76, wherein the small interfering nucleic acid is a miRNA.

78. 77. The method of claim 76, wherein the small interfering nucleic acid is a miRNA sponge or TuD RNA that inhibits the activity of at least one miRNA in the subject or animal.

79. 78. The method of claim 77, wherein the miRNA is expressed in cells of the target tissue.

80. 74. The method of claim 73, wherein the target tissue is skeletal muscle, heart, liver, pancreas, brain, or lung.

81. 74. The method of Claim 73, wherein the transgene expresses a transcript comprising at least one binding site for a miRNA, and the miRNA inhibits activity of the transgene in tissues other than the target tissue by hybridizing to the binding site.

82. 74. The method of Claim 73, wherein the at least one transgene encodes a gene product that mediates genome editing.

83. 74. The method of claim 73, wherein the transgene comprises a tissue-specific or inducible promoter.

84. 84. The method of claim 83, wherein the tissue-specific promoter is a liver-specific thyroxine-binding globulin (TBG) promoter, an insulin promoter, a glucagon promoter, a somatostatin promoter, a pancreatic polypeptide (PPY) promoter, a synapsin-1 (Syn) promoter, a creatine kinase (MCK) promoter, a mammalian desmin (DES) promoter, an a-myosin heavy chain (a-MHC) promoter, or a cardiac troponin T (cTnT) promoter.

85. 74. The method of claim 73, wherein the rAAV is administered intravenously, intravascularly, transdermally, intraocularly, intrathecally, orally, intramuscularly, subcutaneously, intranasally, or by inhalation.

86. 74. The method of claim 73, wherein the subject is selected from a mouse, rat, rabbit, dog, cat, sheep, pig, and non-human primate.

87. 74. The method of claim 73, wherein the subject is a human.

88. An isolated nucleic acid encoding an AAV capsid protein comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 5-820 and 870-911.

89. A delivery vehicle for delivery of a small molecule drug or biological agent having a desired tropism, comprising a peptide or peptide fragment of at least 10-20 amino acids of any one of SEQ ID NOs: 5-820 or 870-911.

90. 90. The delivery vehicle of claim 89, wherein the delivery vehicle is selected from the group consisting of a liposome, a nanoparticle, a bacterium, a bacteriophage, a virus-like particle (VLP), an erythrocyte ghost, and an exosome.

91. 90. The delivery vehicle of claim 89, wherein the biological agent comprises an siRNA, an antisense molecule, a protein or polypeptide, insulin, a vaccine, or an antibody.

92. 90. The delivery vehicle of claim 89, wherein the small molecule drug comprises a chemotherapeutic agent, an anti-inflammatory agent, a steroid, and an antibiotic.

93. A biological agent having a desired tropism linked to a peptide or peptide fragment of at least 10-20 amino acids of any one of SEQ ID NOs: 5-820 or 870-911.

94. 94. The biological agent of claim 93, wherein the biological agent is a nucleic acid, a protein, a polypeptide, a peptide, an antibody, an antibody fragment, a non-immunoglobulin binding agent, or an enzyme.