Methods and compositions for transducing hematopoietic cells

Hematopoietic cell-specific targeting moieties for AAV improve transduction efficiency and safety by reducing liver toxicity and enhancing accuracy in non-human primate models.

JP2025526766APending Publication Date: 2025-08-15PRESIDENT & FELLOWS OF HARVARD COLLEGE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025507585
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-10
Filing Date
2023-08-10
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Conventional recombinant adeno-associated virus (rAAV) delivery vehicles exhibit limited cell tropism, particularly for tissues and organs other than the liver, requiring high doses that can cause liver toxicity and are inefficient in non-human primate models, leading to inaccurate preclinical studies.

Method used

Development of hematopoietic cell-specific targeting moieties linked to AAV motifs, such as engineered AAV capsid polypeptides with mutations at specific positions to enhance targeting efficiency and reduce uptake in non-hematopoietic cells.

Benefits of technology

Enhances transduction efficiency in hematopoietic cells, reducing the need for high viral doses and minimizing liver toxicity, while improving accuracy in non-human primate models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025526766000001_ABST
    Figure 2025526766000001_ABST
Patent Text Reader

Abstract

Described herein are compositions comprising a hematopoietic cell-specific targeting moiety and a hematopoietic cell-specific targeting motif. Also described herein are uses of compositions comprising a hematopoietic cell-specific targeting motif and a hematopoietic cell-specific targeting moiety. In some embodiments, the hematopoietic cell-specific targeting moiety and compositions comprising the hematopoietic cell-specific targeting moiety can be used to direct delivery of cargo to hematopoietic cells. Embodiments disclosed herein provide hematopoietic cell-specific targeting moieties that can be linked to or otherwise associated with cargo.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 396,922, filed August 10, 2022. The entire teachings of the above application are incorporated herein by reference.

[0002] government support This invention was made with government support under AG063419 awarded by the National Institutes of Health (NIH). The U.S. Government has certain rights in this invention. [Background technology]

[0003] Recombinant AAV (rAAV) is the most commonly used delivery vehicle for gene therapy and gene editing. However, rAAV containing natural capsid variants has limited cell tropism. In fact, currently used rAAV mainly infects the liver after systemic delivery. Furthermore, the transduction efficiency of conventional rAAV in other cell types, tissues, and organs is limited by these conventional rAAVs with natural capsid variants. Therefore, AAV-mediated polynucleotide delivery for diseases affecting cells, tissues, and organs other than the liver (e.g., nervous system, skeletal muscle, and cardiac muscle) typically requires high doses of virus (usually about 1 × 10), which often results in liver toxicity. 14 The high doses required for conventional rAAVs make it extremely difficult to produce sufficient quantities of therapeutic rAAV for administration to adult patients. Additionally, due to differences in gene expression and physiology, mouse and primate models respond differently to viral capsids. The transduction efficiency of different viral particles varies between species, and as a result, preclinical studies in mice often do not accurately reflect results in primates, including humans. Thus, there is a need for improved rAAVs for use in the treatment of various genetic disorders. Summary of the Invention [Means for solving the problem]

[0004] Embodiments disclosed herein provide hematopoietic cell-specific targeting moieties that may be linked to or otherwise associated with cargo, which may be modified hematoAAV motifs and may include one or more n-mer motifs that can confer hematopoietic cell specificity to the targeting moiety.

[0005] Described herein are compositions comprising a targeting moiety effective in targeting hematopoietic cells, the targeting moiety comprising one or more n-mer motifs, wherein at least one n-mer motif comprises or consists of: a) any one of SEQ ID NOs: 1-1000; b) any one of SEQ ID NOs: 2001-3000; c) any one of SEQ ID NOs: 4001-5000; d) any one of SEQ ID NOs: 6001-7000; e) any one of SEQ ID NOs: 8001-9000; f) any one of SEQ ID NOs: 10001-11000; or g) any combination thereof.

[0006] In some embodiments, the hematopoietic cells are differentiated hematopoietic cells or progenitor cells.

[0007] In some embodiments, the targeting moiety comprises a polypeptide (e.g., a viral polypeptide or a viral capsid polypeptide), a polynucleotide, a lipid, a polymer, a sugar, or any combination thereof. In one embodiment, the targeting moiety comprises an adeno-associated virus (AAV) polypeptide or an adeno-associated virus (AAV) capsid polypeptide.

[0008] In some embodiments, the n-mer motif is inserted between any two amino acids of the viral polypeptide, viral capsid polypeptide, AAV polypeptide, or AAV capsid polypeptide. In some aspects, the n-mer motif is inserted into the viral polypeptide, viral capsid polypeptide, AAV polypeptide, or AAV capsid polypeptide such that one, two, or more amino acids at the N-terminus and / or C-terminus of the n-mer motif replace one, two, or more amino acids of the viral polypeptide, viral capsid polypeptide, AAV polypeptide, or AAV capsid polypeptide. In some embodiments, the n-mer motif is inserted between any two consecutive amino acids between amino acids 262-269, 327-332, 382-386, 452-460, 488-505, 527-539, 545-558, 581-593, 598-599, 704-714, or any combination thereof, of the capsid polypeptide of AAV9, or at an analogous position in the capsid polypeptide of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10.

[0009] In some embodiments, the AAV capsid polypeptide is an engineered AAV capsid polypeptide that has reduced or eliminated uptake into non-hematopoietic cells (e.g., liver cells) compared to a corresponding wild-type AAV capsid polypeptide (e.g., AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10 capsid polypeptide). In some embodiments, the engineered AAV capsid polypeptide comprises one or more mutations that result in reduced or eliminated uptake in non-hematopoietic cells. The one or more mutations may be at a) position 267, b) position 269, c) position 504, d) position 505, e) position 590, or f) any combination thereof in the AAV9 capsid protein (SEQ ID NO: 12001), or one or more corresponding positions in a non-AAV9 capsid polypeptide (e.g., an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10 capsid polypeptide). In one embodiment, the mutation at position 267 of the AAV9 capsid protein (SEQ ID NO: 12001) or a corresponding position in a non-AAV9 capsid polypeptide is a G or X to A mutation, where X is any amino acid. In one embodiment, the mutation at position 269 of the AAV9 capsid protein (SEQ ID NO: 12001), or a corresponding position in a non-AAV9 capsid polypeptide, is an S or X to T mutation, where X is any amino acid. In one embodiment, the mutation at position 504 of the AAV9 capsid protein (SEQ ID NO: 12001), or a corresponding position in a non-AAV9 capsid polypeptide, is a G or X to A mutation, where X is any amino acid. In one embodiment, the mutation at position 505 of the AAV9 capsid protein (SEQ ID NO: 12001), or a corresponding position in a non-AAV9 capsid polypeptide, is a P or X to A mutation, where X is any amino acid.In one embodiment, the mutation at position 590 of the AAV9 capsid protein (SEQ ID NO: 12001), or the corresponding position in a non-AAV9 capsid polypeptide, is a Q or X to A mutation, where X is any amino acid. In one embodiment, the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 267, 269, or both, of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 267 is a G to A mutation and the mutation at position 269 is an S to T mutation. In one embodiment, the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 590 of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 509 is a Q to A mutation. In one embodiment, the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 504, 505, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 504 is a G to A mutation and the mutation at position 505 is a P to A mutation.

[0010] In some embodiments, the composition is an engineered viral particle, optionally an engineered AAV particle (e.g., an engineered AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10 viral particle).

[0011] Also described herein is a composition comprising a targeting moiety effective in targeting hematopoietic cells, the targeting moiety comprising one or more n-mer motifs, at least one of the n-mer motifs being VKX. n contains or consists of X n are each selected from any amino acid, and n is 5. Further described herein are compositions comprising a targeting moiety effective in targeting hematopoietic cells, the targeting moiety comprising one or more n-mer motifs, at least one n-mer motif being selected from the group consisting of VKX, ... nYGAL, containing or consisting of X n are each selected from any amino acid, and n is 1.

[0012] The compositions described herein may further comprise a cargo linked to or otherwise associated with the targeting moiety. In some embodiments, the cargo is a) effective in treating or preventing a blood disease or disorder, b) effective in treating or preventing a non-blood disease or disorder, c) a vaccine, or d) any combination thereof.

[0013] Also described herein is a vector system comprising a vector comprising one or more polynucleotides, at least one of the one or more additional polynucleotides encoding all or a portion of a targeting moiety effective in targeting a hematopoietic cell, the targeting moiety comprising one or more n-mer motifs, wherein at least one n-mer motif comprises or consists of: a) any one of SEQ ID NOs: 1-1000, b) any one of SEQ ID NOs: 2001-3000, c) any one of SEQ ID NOs: 4001-5000, d) any one of SEQ ID NOs: 6001-7000, e) any one of SEQ ID NOs: 8001-9000, f) any one of SEQ ID NOs: 10001-11000, or g) any combination thereof. Optionally, the vector system comprises a regulatory element operably linked to one or more of the one or more polynucleotides.

[0014] In some embodiments, the vector system further comprises a cargo polynucleotide, optionally operably linked to at least one polynucleotide encoding all or a portion of the targeting moiety. In some embodiments, the vector system is capable of producing a polypeptide (e.g., a viral polypeptide, e.g., a viral capsid polypeptide) comprising or consisting of the targeting moiety. In some embodiments, the vector system is capable of producing an adeno-associated virus (AAV) polypeptide, optionally an AAV capsid polypeptide. In some embodiments, the vector system is capable of producing a viral particle, optionally an AAV particle, which optionally comprises cargo. In one embodiment, the AAV polypeptide, AAV capsid polypeptide, and / or AAV particle is an engineered AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10 viral particle or polypeptide.

[0015] In some embodiments, the polypeptide comprises one or more n-mer motifs inserted between two amino acids of the polypeptide, and optionally, the one or more n-mer motifs are inserted so that they are on the outside of the capsid of a virus produced by the vector system. In one embodiment, the AAV capsid polypeptide comprises one or more n-mer motifs inserted between any two consecutive amino acids (e.g., 262-269, 327-332, 382-386, 452-460, 488-505, 527-539, 545-558, 581-593, 598-599, 704-714, or any combination thereof) of the capsid polypeptide of AAV9, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10. In some embodiments, the AAV capsid polypeptide is an engineered AAV capsid polypeptide that has reduced or eliminated uptake into non-hematopoietic cells (e.g., liver cells) compared to a corresponding wild-type AAV capsid polypeptide (e.g., AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10 capsid polypeptide).

[0016] In some embodiments, the engineered AAV capsid polypeptide comprises one or more mutations that result in reduced or eliminated uptake into non-hematopoietic cells, and optionally the one or more mutations are at a) position 267, b) position 269, c) position 504, d) position 505, e) position 590, or f) any combination thereof, of the AAV9 capsid protein (SEQ ID NO: 12001), or one or more corresponding positions in a non-AAV9 capsid polypeptide (e.g., an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10 capsid polypeptide).In one embodiment, the mutation at position 267 of the AAV9 capsid protein (SEQ ID NO: 12001), or a corresponding position in a non-AAV9 capsid polypeptide, is a G or X to A mutation, where X is any amino acid; b) the mutation at position 269 of the AAV9 capsid protein (SEQ ID NO: 12001), or a corresponding position in a non-AAV9 capsid polypeptide, is a S or X to T mutation, where X is any amino acid; or c) the mutation at position 269 of the AAV9 capsid protein (SEQ ID NO: 12001), or a corresponding position in a non-AAV9 capsid polypeptide, is a S or X to T mutation, where X is any amino acid. a) a mutation at position 504 of the AAV9 capsid protein (SEQ ID NO: 12001) or a corresponding position in a non-AAV9 capsid polypeptide is a G or X to A mutation, where X is any amino acid; b) a mutation at position 505 of the AAV9 capsid protein (SEQ ID NO: 12001) or a corresponding position in a non-AAV9 capsid polypeptide is a P or X to A mutation, where X is any amino acid; c) a mutation at position 590 of the AAV9 capsid protein (SEQ ID NO: 12001) or a corresponding position in a non-AAV9 capsid polypeptide is a P or X to A mutation, where X is any amino acid; the mutation at the corresponding position in SEQ ID NO: 12001 is a mutation of Q or X to A, where X is any amino acid; f) the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 267, 269, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 267 is a mutation from G to A and the mutation at position 269 is a mutation from S to T; g) the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 267, 269, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 267 is a mutation from G to A and the mutation at position 269 is a mutation from S to T; h) the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 504, 505, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 504 is a G to A mutation and the mutation at position 505 is a P to A mutation; or i) any permissible combination thereof.

[0017] In some embodiments, the vector system further comprises a polynucleotide encoding a viral rep protein, optionally an AAV rep protein, hi one embodiment, the polynucleotide encoding the viral rep protein is on the same vector as the one or more polynucleotides or on a different vector, and is optionally operably linked to a regulatory element.

[0018] In some embodiments, the vector system is capable of producing a composition or portion thereof described herein.

[0019] Also described herein are polynucleotides encoding all or part of the compositions described herein. Also described herein are polypeptides encoded by the vector systems described herein, or encoded by the polynucleotides described herein, or both. In some embodiments, the polypeptides are linked to or otherwise associated with a cargo.

[0020] Further described herein are particles, optionally viral particles, produced by the vector systems and / or polynucleotides described herein, which optionally comprise a polypeptide described herein.

[0021] In some embodiments, the particle is an AAV particle and optionally comprises a cargo, hi some embodiments, the cargo or cargo polynucleotide is a) effective in treating or preventing a hematological disease or disorder, b) effective in treating or preventing a non-hematological disease or disorder, c) a vaccine, or d) any combination thereof.

[0022] Also disclosed herein are cells comprising the compositions, vector systems, polypeptides, or particles described herein, or any combination thereof. In some embodiments, the cells are hematopoietic cells (e.g., prokaryotic or eukaryotic cells).

[0023] Disclosed herein are pharmaceutical formulations comprising a composition, vector system, polypeptide, particle, or cell described herein, or any combination thereof, and a pharmaceutically acceptable carrier.

[0024] Also disclosed herein is a method of treating or preventing a disease, optionally a hematological disease, or a symptom thereof, in a subject in need thereof, the method comprising administering to a subject in need thereof a composition, vector system, polypeptide, particle, cell, or pharmaceutical formulation, or any combination thereof, as described herein.

[0025] In some embodiments, the blood disease or disorder is HIV / AIDs, blood cancer (e.g., leukemia, lymphoma, myeloma, monoclonal gammopathy of undetermined significance (MGUS)), bleeding disorder (e.g., acquired platelet dysfunction, congenital platelet dysfunction, disseminated intravascular coagulation (DIC), prothrombin deficiency, factor V deficiency, factor VII deficiency, factor X deficiency, factor XI deficiency (hemophilia C), Glanzmann's disease, hemophilia A, hemophilia B, idiopathic thrombocytopenic purpura (ITP), von Willebrand's disease (types I, II, and / or III)). , hemoglobinopathies (e.g., sickle cell disease (HbS), sickle cell trait (HbAS), sickle cell hemoglobin C (HbSC), sickle cell thalassemia (HbS and HbA), thalassemia (alpha thalassemia and beta thalassemia), hemoglobin C disease (HbCC), hemoglobin C trait (HbAC), primary immunodeficiencies (e.g., autoimmune lymphoproliferative syndrome (ALPS), APS-1 (APECED), BENTA disease, caspase 8 deficiency (CEDS), CARD9 deficiency and other candidiasis susceptibility syndromes), chronic granulomatous disease (CGD), unclassifiable Type 2 immunodeficiency (CVID), congenital neutropenic syndrome, CTLA4 deficiency, DOCK8 deficiency, GATA2 deficiency, glycosylation disorders associated with immunodeficiency, hyperimmunoglobulin E syndrome (HIES), hyperimmunoglobulin M syndrome, interferon gamma deficiency, interleukin-12 deficiency, and interleukin-23 deficiency, leukocyte adhesion deficiency (LAD), LRBA deficiency, PI3 kinase disease, PLCG2-associated antibody deficiency and immune dysregulation (PLAID), severe combined immunodeficiency (SCID), STAT3 dominant-negative disease, and STAT3 gain of function type diseases, Ward, hypogammaglobulinemia, infections, myeloid cellular pool (WHIM) syndrome, Wiskott-Aldrich syndrome (WAS), x-linked agammaglobulinemia (XLA), x-linked lymphoproliferative disorders (XLP), XMEN diseases), cytopenias (anemia, leukopenia, thrombocytopenia, pancytopenia, autoimmune cytopenia, refractory cytopenia), and storage and metabolic disorders (e.g., diabetes mellitus, familial hypercholesterolemia, Hunter syndrome, Krabbe disease, maple syrup urine disease, metachromatic leukodystrophy, Niemann-Pick disease, Gaucher disease,Hemochromatosis, phenylketonuria (PKU), mitochondrial disorders, porphyria, Tay-Sachs disease, and Wilson's disease. [Brief explanation of the drawings]

[0026] [Figure 1] 1 shows the mechanism of adeno-associated virus (AAV) transduction leading to the production of mRNA from a transgene.

[0027] [Figure 2] 1 shows a graph demonstrating that mRNA-based selection of AAV variants can be more stringent than DNA-based selection. The viral library was expressed under the control of a CMV promoter.

[0028] [Figure 3] AB show graphs demonstrating the correlation between the viral library and vector genomic DNA (A) and mRNA (B) in the liver.

[0029] [Figure 4] A-F show graphs demonstrating the capsid variants present at the DNA level and expressed at the mRNA level identified in various tissues. In this experiment, the viral library was expressed under the control of the CMV promoter.

[0030] [Figure 5A] Graphs are shown demonstrating capsid mRNA expression in different tissues under the control of cell-type specific promoters (shown on the x-axis). CMV is included as an exemplary constitutive promoter. CK8 is a muscle-specific promoter. MHCK7 is a muscle-specific promoter. hSyn is a neuron-specific promoter. Expression levels from these cell-type specific promoters are normalized based on the expression levels from the constitutive CMV promoter in each tissue. [Figure 5B]Graphs are shown demonstrating capsid mRNA expression in different tissues under the control of cell-type specific promoters (shown on the x-axis). CMV is included as an exemplary constitutive promoter. CK8 is a muscle-specific promoter. MHCK7 is a muscle-specific promoter. hSyn is a neuron-specific promoter. Expression levels from these cell-type specific promoters are normalized based on the expression levels from the constitutive CMV promoter in each tissue. [Figure 5C] Graphs are shown demonstrating capsid mRNA expression in different tissues under the control of cell-type specific promoters (shown on the x-axis). CMV is included as an exemplary constitutive promoter. CK8 is a muscle-specific promoter. MHCK7 is a muscle-specific promoter. hSyn is a neuron-specific promoter. Expression levels from these cell-type specific promoters are normalized based on the expression levels from the constitutive CMV promoter in each tissue.

[0031] [Figure 6] (A) Schematic showing an embodiment of a method for producing and selecting capsid variants for tissue-specific gene delivery across species, and (B) a schematic showing evaluation of the top selected capsids.

[0032] [Figure 7] 1 shows a schematic diagram illustrating an embodiment for generating an AAV capsid variant library, specifically the insertion of random n-mers (n=3-15 amino acids) into wild-type AAV, e.g., AAV9.

[0033] [Figure 8] 1 shows a schematic diagram illustrating an embodiment of generating an AAV capsid variant library, specifically variant AAV particle production, where each capsid variant encapsulates its own coding sequence as a vector genome.

[0034] [Figure 9]A schematic vector map of a representative AAV capsid plasmid library vector that can be used in an AAV vector system to generate an AAV capsid variant library is shown (see, e.g., Figure 8).

[0035] [Figure 10] A graph demonstrating the viral titers (calculated as AAV9 vector genomes / 15 cm dish) produced by constructs containing different constitutive and cell type-specific mammalian promoters is shown.

[0036] [Figure 11A] Figure 1 shows the superior functionality of HematoAAV variants. Figure 2 shows a schematic diagram for testing HematoAAV variants in vitro. [Figure 11B] The superior functionality of HematoAAV variants is shown. The results of testing HematoAAV variants are shown.

[0037] [Figure 12-1] Figure 1 shows AAV9 variants sorted based on their ability to target hematopoietic cells or progenitor cells. A pooled library of AAV9 capsid variants that differ within the inserted heptamer region was generated and screened to identify variants that specifically target human and mouse hematopoietic cells in vivo. [Figure 12-2] Same as above. [Figure 12-3] Same as above. [Figure 12-4] Same as above. [Figure 12-5] Same as above. [Figure 12-6] Same as above. [Figure 12-7] Same as above. [Figure 12-8] Same as above. [Figure 12-9] Same as above. [Figure 12-10] Same as above. [Figure 12-11] Same as above. [Figure 12-12] Same as above. [Figure 12-13] Same as above. [Figure 12-14] Same as above. [Figure 12-15] Same as above. [Figure 12-16] Same as above. [Figure 12-17] Same as above. [Figure 12-18] Same as above. [Figure 12-19] Same as above. [Figure 12-20] Same as above. [Figure 12-21] Same as above. [Figure 12-22] Same as above. [Figure 12-23] Same as above. [Figure 12-24] Same as above. [Figure 12-25] Same as above. [Figure 12-26] Same as above. [Figure 12-27] Same as above. [Figure 12-28] Same as above. [Figure 12-29] Same as above. [Figure 12-30] Same as above. [Figure 12-31] Same as above. [Figure 12-32] Same as above. [Figure 12-33] Same as above. [Figure 12-34] Same as above. [Figure 12-35] Same as above. [Figure 12-36] Same as above. [Figure 12-37] Same as above. [Figure 12-38] Same as above. [Figure 12-39] Same as above. [Figure 12-40] Same as above. [Figure 12-41] Same as above. [Figure 12-42] Same as above. [Figure 12-43] Same as above. [Figure 12-44] Same as above. [Figure 12-45] Same as above. [Figure 12-46] Same as above. [Figure 12-47] Same as above. [Figure 12-48] Same as above. [Figure 12-49] Same as above. [Figure 12-50] Same as above. [Figure 12-51] Same as above. [Figure 12-52] Same as above. [Figure 12-53] Same as above. [Figure 12-54] Same as above. [Figure 12-55] Same as above. [Figure 12-56] Same as above. [Figure 12-57] Same as above. [Figure 12-58] Same as above. [Figure 12-59] Same as above. [Figure 12-60] Same as above. [Figure 12-61] Same as above. [Figure 12-62] Same as above. [Figure 12-63] Same as above. [Figure 12-64] Same as above. [Figure 12-65] Same as above. [Figure 12-66] Same as above. [Figure 12-67] Same as above. [Figure 12-68] Same as above. [Figure 12-69] Same as above. [Figure 12-70] Same as above. [Figure 12-71] Same as above. [Figure 12-72] Same as above. [Figure 12-73] Same as above. [Figure 12-74] Same as above. [Figure 12-75] Same as above. [Figure 12-76] Same as above. [Figure 12-77] Same as above. [Figure 12-78] Same as above. [Figure 12-79] Same as above. [Figure 12-80] Same as above. [Figure 12-81] Same as above. [Figure 12-82] Same as above. [Figure 12-83] Same as above. [Figure 12-84] Same as above. [Figure 12-85] Same as above. [Figure 12-86] Same as above. [Figure 12-87] Same as above. [Figure 12-88] Same as above. [Figure 12-89] Same as above. [Figure 12-90] Same as above. [Figure 12-91] Same as above. [Figure 12-92] Same as above. [Figure 12-93] Same as above. [Figure 12-94] Same as above. [Figure 12-95] Same as above. [Figure 12-96] Same as above. [Figure 12-97] Same as above. [Figure 12-98] Same as above. [Figure 12-99] Same as above. [Figure 12-100] Same as above. [Figure 12-101] Same as above. [Figure 12-102] Same as above. [Figure 12-103] Same as above. [Figure 12-104] Same as above. [Figure 12-105] Same as above. [Figure 12-106] Same as above. [Figure 12-107] Same as above. [Figure 12-108] Same as above.

[0038] [Figure 13] 1 shows the self-complementary Cbh-Gfp transfer vector used for AAV production.

[0039] [Figure 14] A table summarizing the production yields of variant AAVs quantified using qPCR detection of BGh vector elements is shown.

[0040] [Figure 15] Panels A-B show data from a bioactivity assay to assess the functionality of HematoAAV variants. K562 cells were transduced at 1 x 105 vg / cell with the indicated HematoAAV capsids carrying the Cbh-Gfp expression cassette and compared with the parental vector (AAV9, produced in two independent batches). Panel A shows data representing the % of live cells expressing GFP 2 days after AAV exposure. Panel B shows data representing the median GFP fluorescence intensity in cells gated as positive for GFP expression.

[0041] [Figure 16] Figures A-B show in vitro transduction of human peripheral blood mononuclear cells (PBMCs) with S1-MF1-B, S2-MF1-E, or AAV9. PBMCs from a single human donor were exposed in culture to the indicated AAV variants carrying the Cbh-Gfp gene transfer vector. Two different vector doses were tested: 1e5 vg / cell (A) or 5e5 vg / cell (B). Forty-one hours after AAV exposure, cells were harvested for flow cytometry. Plots were gated to exclude dead cells and debris and then selected for the indicated cell surface markers to identify specific cell populations, which were scored and analyzed for % GFP+.

[0042] [Figure 17A] Motifs of interest enriched in human and mouse bone marrow cells are identified. Motifs of interest identified from initial screening of human and mouse bone marrow cells are shown. Motifs of interest identified in both human and mouse bone marrow cells are also identified. [Figure 17B] Motifs of interest enriched in human and mouse bone marrow cells are identified. Motifs of interest identified from a second screen of human and mouse bone marrow cells are provided. Motifs of interest identified in both human and mouse bone marrow cells are also identified. [Figure 17C]Identifying motifs of interest enriched in human and mouse bone marrow cells. A list of motifs of interest identified through primary and secondary screening is shown. DETAILED DESCRIPTION OF THE INVENTION

[0043] general definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Definitions of common terms and techniques in molecular biology can be found in Molecular Cloning: A Laboratory Manual, 2004. nd edition (1989) (Sambrook, Fritsch, and Maniatis), Molecular Cloning: A Laboratory Manual, 4 th edition (2012) (Green and Sambrook), Current Protocols in Molecular Biology (1987) (FMAusubel et al. eds.), the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (1995) (MJMacPherson, BD Hames, and GRTaylor eds.): Antibodies, A Laboratory Manual (1988) (Harlow and Lane, eds.): Antibodies A Laboratory Manual, 2 ndedition 2013 (EAGreenfield ed.), Animal Cell Culture (1987) (RIFreshney, ed.), Benjamin Lewin, Genes IX, published by Jones and Bartlet, 2008 (ISBN 0763752223), Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN 0632021829), Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 9780471185710), Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY1994), March, Advanced Organic Chemistry Reactions, Mechanisms and Structure 4th ed., John Wiley & Sons (New York, NY1992), and Marten H. Hofker and Jan van Deursen, Transgenic Mouse Methods and Protocols, 2 nd edition (2011).

[0044] As used herein, the singular forms "a," "an," and "the" include both singular and plural referents unless the context clearly indicates otherwise.

[0045] The term "optional" or "optionally" means that the subsequently described event, circumstance, or substitution may or may not occur, and that the description includes instances in which the event or circumstance occurs and instances in which it does not occur.

[0046] The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within each range, as well as the recited endpoint. It is further understood that each endpoint of the range is significant both in relation to the other endpoint, and independently of the other endpoint. It is also understood that there are a number of values disclosed herein, and that each value is also herein disclosed as "about" that particular value in addition to the value itself. For example, if the value "10" is disclosed, then "about 10" is also disclosed. Ranges can be expressed herein as from "about" one particular value and / or to "about" another particular value. Similarly, when values are expressed as approximations by use of the antecedent "about," it will be understood that the particular value forms a further embodiment. For example, if the value "about 10" is disclosed, then "10" is also disclosed.

[0047] It should be understood that such range formats are used for convenience and brevity and should thus be interpreted in a flexible manner to include not only the numerical values expressly recited as range limits, but also all individual numerical values or subranges subsumed within that range, as if each numerical value and subrange were expressly recited. To illustrate, a numerical range of "about 0.1% to 5%" should be interpreted to include not only the explicitly recited value of about 0.1% to about 5%, but also individual values within the stated range (e.g., about 1%, about 2%, about 3%, and about 4%) and subranges (e.g., about 0.5% to about 1.1%, about 5% to about 2.4%, about 0.5% to about 3.2%, and about 0.5% to about 4.4%, as well as other possible subranges). When a range is expressed, a further embodiment includes from the one particular value and / or to the other particular value.

[0048] When a range of values is provided, it is understood that each intervening value between the upper and lower limit of that range (to one-tenth of the unit of the lower limit, unless the context clearly dictates otherwise) and any other stated or intervening value in that stated range is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included in the disclosure. For example, when a stated range includes one or both of the limits, ranges excluding either or both of the included limits are also included in the disclosure; for example, the expression "from x to y" includes ranges from "x" to "y" as well as ranges from greater than "x" to less than "y." Ranges can also be expressed as upper limits, e.g., "about x, y, z, or less," and should be interpreted to include the specific ranges of "about x," "about y," and "about z," as well as the ranges "less than x," "less than y," and "less than z." Similarly, the phrase "about x, y, z, or more" should be interpreted to include the specific ranges of "about x," "about y," and "about z," as well as the ranges "greater than x," "greater than y," and "greater than z." Also, the phrase "about 'x' to 'y'" (where 'x' and 'y' are numbers) includes "about 'x' to about 'y'."

[0049] As used herein, the term "about" or "approximately," when referring to a measurable value, e.g., a parameter, amount, time duration, etc., is meant to encompass variations at and from the particular value, e.g., variations of + / - 10% or less, + / - 5% or less, + / - 1% or less, and + / - 0.1% or less, to the extent that such variations are appropriate for practicing the disclosed invention. It should be understood that the value to which the modifier "about" or "approximately" refers is itself specifically and preferably disclosed. As used herein, the terms "about," "approximate," "at or about," and "substantially" can mean that the amount or value in question may be the exact value or a value that provides results or effects equivalent to those described in the claims or taught herein, and are understood to reflect other known factors. That is, it is understood that amounts, sizes, formulations, parameters, and other quantitative values and characteristics are not, and need not be, exact, but may be approximate and / or larger or smaller, and, where appropriate, reflect tolerances, conversion factors, rounding, measurement error, and the like, as well as other factors known to those skilled in the art to produce equivalent results or effects. In some situations, a value that will produce equivalent results or effects cannot be reasonably determined. Typically, amounts, sizes, formulations, parameters, or other quantities or characteristics are "about," "approximately," or "at or about" (whether or not expressly stated as such). When "about," "approximately," or "at or about" is used before a quantitative value, it is understood that the parameter also includes the particular quantitative value itself, unless specifically stated otherwise.

[0050] As used herein, a "biological sample" may include whole cells and / or viable cells and / or cell debris. A biological sample may include (or be derived from) a "body fluid." The present invention encompasses embodiments in which the body fluid is selected from amniotic fluid, aqueous humor, vitreous humor, bile, serum, breast milk, cerebrospinal fluid, cerumen (earwax), chyle, chyme, endolymph, perilymph, exudate, feces, female vaginal fluid, gastric acid, gastric juice, lymph, mucus (including nasal secretions and phlegm), pericardial fluid, peritoneal fluid, pleural fluid, pus, secretions, saliva, sebum (skin oil), semen, sputum, synovial fluid, sweat, tears, urine, vaginal secretions, vomit, and mixtures of one or more thereof. Biological samples include cell cultures, body fluids, and cell cultures from body fluids. Body fluids may be obtained from a mammalian organism, for example, by paracentesis or other collection or sampling procedures.

[0051] The terms "subject," "individual," and "patient" are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Also encompassed are tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro.

[0052] Various embodiments are described below. It should be noted that a particular embodiment is not intended as an exhaustive description or as a limitation to the broader embodiments discussed herein. An embodiment described in conjunction with a particular embodiment is not necessarily limited to that embodiment and may be practiced with any other embodiment(s). References throughout this specification to "one embodiment," "an embodiment," or "an exemplary embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment," "in an embodiment," or "an exemplary embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, although they may. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments, as would be apparent to one of ordinary skill in the art from this disclosure. Furthermore, while some embodiments described herein may include some features but not other features included in other embodiments, it is intended that combinations of features from different embodiments are within the scope of the present invention. For example, in the appended claims, any of the claimed embodiments may be used in any combination.

[0053] All publications, published patent documents, and patent applications cited in this specification are herein incorporated by reference to the same extent as if each individual publication, published patent document, or patent application was specifically and individually indicated to be incorporated by reference.

[0054] overview

[0010] Embodiments disclosed herein provide hematopoietic cell-specific targeting moieties that may be linked to or otherwise associated with cargo.

[0011] Embodiments disclosed herein provide polypeptides and particles that may incorporate one or more hematopoietic cell-specific targeting moieties. The polypeptides and / or particles may be linked to, bound to, encapsulate, or otherwise incorporate cargo, thereby associating the cargo with the targeting moiety(s).

[0055] The embodiments disclosed herein provide hematopoietic cell-specific targeting moieties that can include one or more n-mer motifs as further described herein. In some embodiments, the n-mer motif is an improved hematoAAV motif. In some embodiments, the n-mer motif can confer hematopoietic cell specificity to the targeting moiety.

[0056] Embodiments disclosed herein provide engineered adeno-associated virus (AAV) capsids that can be engineered to confer cell-type-specific and / or species-specific tropism to the engineered AAV particles.

[0057] Embodiments disclosed herein also provide methods for generating rAAVs with engineered capsids, which may include systematically directing the generation of a diverse library of variants with modified surface structures, such as variant capsid proteins. Embodiments of the methods for generating rAAVs with engineered capsids may also include stringent selection of capsid variants capable of targeting specific cells, tissues, and / or organ types. Embodiments of the methods for generating rAAVs with engineered capsids may include stringent selection of capsid variants capable of efficient, selective, and / or uniform transduction in at least one or more species.

[0058] Embodiments disclosed herein provide vectors and systems capable of producing the engineered AAVs described herein.

[0059] Embodiments disclosed herein provide cells that may be capable of producing the engineered AAV particles described herein, hi some embodiments, the cells comprise one or more of the vectors or systems described herein.

[0060] The embodiments disclosed herein provide engineered AAVs that can comprise the engineered capsids described herein. In some embodiments, the engineered AAVs can comprise a cargo polynucleotide that is delivered to a cell. In some embodiments, the cargo polynucleotide is a recombinant polynucleotide.

[0061] Embodiments disclosed herein provide formulations, which may include an engineered AAV vector or system thereof, an engineered AAV capsid, an engineered AAV particle comprising an engineered AAV capsid described herein, and / or an engineered cell described herein comprising an engineered AAV capsid, and / or an engineered AAV vector or system thereof. In some embodiments, the formulation may also include a pharmaceutically acceptable carrier. The formulations described herein may be delivered to a subject in need thereof or to a cell.

[0062]

[0013] Embodiments disclosed herein also provide kits that include one or more of the polypeptides, polynucleotides, vectors, engineered AAV capsids, engineered AAV particles, cells, or other components described herein, and combinations thereof, as well as one or more of the pharmaceutical formulations described herein. In embodiments, one or more of the polypeptides, polynucleotides, vectors, engineered AAV capsids, engineered AAV particles, cells, and combinations thereof described herein may be provided as a combination kit.

[0063] Embodiments disclosed herein provide methods of using engineered AAVs with cell-specific tropism described herein, for example, to deliver therapeutic polynucleotides to cells. In this manner, the engineered AAVs described herein can be used to treat and / or prevent disease in subjects in need thereof. Embodiments disclosed herein also provide methods of delivering the engineered AAV capsids, engineered AAV viral particles, engineered AAV vectors or systems thereof, and / or formulations thereof to cells. Also provided herein are methods of treating a subject in need thereof by delivering engineered AAV particles, engineered AAV capsids, engineered AAV capsid vectors or systems thereof, engineered cells, and / or formulations thereof to the subject. Embodiments disclosed herein provide methods of using engineered AAVs with cell-specific tropism described herein, for example, to deliver polynucleotides designed to alter cell signaling / function to cells. In this manner, the engineered AAVs described herein can be used to discover disease mechanisms.

[0064] Additional features and advantages of embodiments of the engineered AAVs and methods of making and using the engineered AAVs are further described herein.

[0065] Hematopoietic cell-specific targeting moieties and compositions thereof Described herein are targeting moieties that may be capable of specifically targeting, binding to, associating with, or otherwise interacting specifically with cells (e.g., hematopoietic cells or cells found within hematopoietic organs, e.g., stromal and / or endothelial cells). In some embodiments, the targeting moiety may be or include an n-mer motif.

[0066] In some embodiments, the targeting moiety may comprise multiple n-mer motifs. In some embodiments, the targeting moiety may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more n-mer motifs. In some embodiments, all of the n-mer motifs comprised in the targeting moiety can be the same. In some embodiments where multiple n-mer motifs are comprised, at least two of the n-mer motifs are different from each other. In some embodiments where multiple n-mer motifs are comprised, all of the n-mer motifs are different from each other. In some embodiments, each n-mer motif comprised in the targeting moiety can be any one of those depicted in FIG. 12 or elsewhere herein. In some embodiments, each n-mer motif comprised in the targeting moiety is a VKX motif. n It may contain or consist of X n are each selected from any amino acid, and n is 1, 2, 3, 4, 5, 6, or 7. In some embodiments, each n-mer motif included in the targeting moiety is n Contains or consists of YGAL and X n are each selected from any amino acid, and n is 1, 2, or 3.

[0067] In some embodiments, the first 1, 2, 3, or 4 amino acids of an n-mer motif can replace 1, 2, 3, or 4 amino acids of the polypeptide into which it is inserted and preceding the insertion position. For example, in one or more of the 7-mer inserts shown in FIG. 12, the first three amino acids shown can replace 1 to 3 amino acids in the polypeptide into which they are inserted. Using AAV as another non-limiting example, one or more of the n-mer motifs can be inserted, for example, between amino acids 587 and 588 and between amino acids 588 and 589 of the capsid polypeptide of AAV9, where the insertion can replace amino acids 586, 587, and 588, with the amino acid immediately preceding the n-mer motif after insertion being residue 585. It will be understood that this principle can be applied to any other insertion situation and is not necessarily limited to insertions between residues 587 and 588 or between residues 588 and 589 of the capsid of AAV9, or to insertions at the equivalent positions in the capsid of another AAV. It will further be understood that in some embodiments the n-mer motif does not replace any amino acids in the polypeptide into which it is inserted.

[0068] The hematopoietic cell-specific targeting moiety may be linked to or otherwise associated with a cargo. In some embodiments, one or more hematopoietic cell-specific targeting moieties described herein are directly linked to a cargo. In some embodiments, one or more hematopoietic cell-specific targeting moieties described herein are indirectly linked to a cargo, for example, via a linker molecule. In some embodiments, one or more hematopoietic cell-specific targeting moieties described herein are linked to or associated with a polypeptide or other particle that is linked to, bound to, encapsulates, and / or comprises a cargo.

[0069] Exemplary particles include, but are not limited to, viral particles (e.g., viral capsids, including bacteriophage capsids), polysomes, liposomes, nanoparticles, microparticles, exosomes, micelles, and the like. As used herein, the term "nanoparticle" includes nanoscale precipitates of homogeneous or heterogeneous materials. Nanoparticles may be regular or irregular in shape and may be formed from multiple co-precipitated particles forming composite nanoscale particles. Nanoparticles may be generally spherical in shape or may have a composite shape formed from multiple co-precipitated, generally spherical particles. Exemplary shapes for nanoparticles include, but are not limited to, spherical, rod-shaped, ellipsoidal, cylindrical, discoidal, and the like. In some embodiments, the nanoparticles have a substantially spherical shape.

[0070] As used herein, the term "specific," when used in reference to describing an interaction between two moieties, refers to a non-covalent physical association of a first moiety and a second moiety, wherein the association between the first moiety and the second moiety is at least 2-fold stronger, at least 5-fold stronger, at least 10-fold stronger, at least 50-fold stronger, at least 100-fold stronger, or more than the association between either moiety and most or all other moieties present in the environment in which the binding occurs. The binding of two or more entities is such that the equilibrium dissociation constant, Kd, is greater than or equal to 10 under the conditions used, e.g., within a cell or within a vascular bed or lymphatic vessel, or within a body tissue or organ, or under physiological conditions consistent with cell survival. -3 M or less, 10 -4 M or less, 10 -5 M or less, 10 -6 M or less, 10 -7 M or less, 10 -8 M or less, 10 -9 M or less, 10 -10 M or less, 10 -11 M or less, or 10 -12 In some embodiments, specific binding can be considered specific if the binding is greater than or equal to 10 M. ... -3In some embodiments, specific binding, which may be referred to as "molecular recognition," is a saturable binding interaction between two entities that relies on the complementary arrangement of functional groups on each entity. Examples of specific interactions include primer-polynucleotide interactions, aptamer-aptamer target interactions, antibody-antigen interactions, avidin-biotin interactions, ligand-receptor interactions, metal-chelate interactions, hybridization between complementary nucleic acids, and the like. In some embodiments, in addition to the n-mer motif(s), the targeting moiety may comprise a polypeptide, a polynucleotide, a lipid, a polymer, a sugar, or a combination thereof.

[0071] In some embodiments, the targeting moiety is incorporated into a viral protein, such as a capsid protein, including but not limited to lentivirus, adenovirus, AAV, bacteriophage, or retrovirus protein. In some embodiments, the n-mer motif is positioned between two amino acids of the viral protein, so that the n-mer motif is on the outside of the viral capsid (i.e., displayed on its surface).

[0072] In some embodiments, compositions comprising one or more hematopoietic cell-specific targeting moieties described herein have increased hematopoietic cell potency, hematopoietic cell specificity, reduced immunogenicity, or any combination thereof. As used herein, the terms "hematopoietic cell-specific," "hematopoietic cell specificity," "hematopoietic cell potency," and the like refer to the increased specificity, selectivity, or potency of hematopoietic cell-specific targeting moieties of the invention and compositions incorporating such hematopoietic cell-specific targeting moieties for hematopoietic cells compared to non-hematopoietic cells. Furthermore, in some aspects, the hematopoietic cell-specific targeting moiety can distinguish one hematopoietic cell type from a second hematopoietic cell type. In some embodiments, the cell specificity, selectivity, or potency, or a combination thereof, of a hematopoietic cell-specific targeting moiety described herein or a composition comprising a hematopoietic cell-specific targeting moiety is at least 2-fold to at least 500-fold more specific, selective, and / or potent for / in hematopoietic cells compared to non-hematopoietic cells. In some embodiments, the specificity, or selectivity, or potency of / in a hematopoietic cell-specific targeting moiety described herein is at least 2-fold to / or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107 , 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143,144、145、146、147、148、149、150、151、152、153、154、155、156、157、158、159、160、161、162、163、164、165、166、167、168、169、170、171、172、173、174、175、176、177、178、179、180、181、182、183、184、185、186、187、188、189、190、191、192、193、194、195、196、197、198、199、200、201、202、203、204、205、206、207、208、209、210、211、212、213、214、215、216、217、218、219、220、221、222、223、224、225、226、227、228、229、230、231、232、233、234、235、236、237、238、239、240、241、242、243、244、245、246、247、248、249、250、251、252、253、254、255、256、257、258、259、260、261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500-fold specific or selective.

[0073] In some embodiments, the hematopoietic cell-specific targeting moieties described herein and / or compositions comprising one or more of said hematopoietic cell-specific targeting moieties have reduced non-hematopoietic cell potency, non-hematopoietic cell specificity, reduced immunogenicity, or any combination thereof. In some embodiments, the hematopoietic cell-specific targeting moieties described herein and / or compositions comprising one or more of said hematopoietic cell-specific targeting moieties have at least 2-fold to at least 500-fold reduced specificity, selectivity, and / or potency for / in non-hematopoietic cells compared to hematopoietic cells. In some embodiments, the specificity, selectivity, or potency of / in the hematopoietic cell-specific targeting moieties described herein is at least 2-fold to / or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 1 8, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 10 0, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 1 47, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193,194、195、196、197、198、199、200、201、202、203、204、205、206、207、208、209、210、211、212、213、214、215、216、217、218、219、220、221、222、223、224、225、226、227、228、229、230、231、232、233、234、235、236、237、238、239、240、241、242、243、244、245、246、247、248、249、250、251、252、253、254、255、256、257、258、259、260、261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500-fold less specific or selective.

[0074] The immunogenicity of a composition incorporating a hematopoietic cell-specific targeting moiety can be reduced, for example, by 1 to 100-fold or more. In some embodiments, immunogenicity is reduced by 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 1 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or a 100-fold or greater decrease.

[0075] Cargo can include any molecule capable of linking to or associating with a muscle-specific targeting moiety described herein. Cargo can include, but is not limited to, nucleotides, oligonucleotides, polynucleotides, amino acids, peptides, polypeptides, riboproteins, lipids, sugars, pharmaceutically active agents (e.g., drugs, imaging agents, and other diagnostic agents), chemical compounds, and combinations thereof. In some embodiments, the cargo is DNA, RNA, amino acids, peptides, polypeptides, antibodies, aptamers, ribozymes, guide sequences for ribozymes that inhibit the translation or transcription of essential tumor proteins and genes, hormones, immunomodulators, antipyretics, anxiolytics, antipsychotics, analgesics, anticonvulsants, anti-inflammatory drugs, antihistamines, anti-infectives, radiosensitizers, chemotherapeutics, radioactive compounds, imaging agents, and combinations thereof.

[0076] In some embodiments, the cargo is capable of treating or preventing a blood disease or disorder. Non-limiting examples of blood diseases or disorders include HIV / AIDs, blood cancers (e.g., leukemia, lymphoma, myeloma, monoclonal gammopathy of undetermined significance (MGUS)), bleeding disorders (e.g., acquired platelet dysfunction, congenital platelet dysfunction, disseminated intravascular coagulation (DIC), prothrombin deficiency, factor V deficiency, factor VII deficiency, factor X deficiency, factor XI deficiency (hemophilia C), Glanzmann's disease, hemophilia A, hemophilia B, idiopathic thrombocytopenic purpura (ITP), von Willebrand's disease (types I, II, and / or III), and / or IV. or III), hemoglobinopathies (e.g., sickle cell disease (HbS), sickle cell trait (HbAS), sickle cell hemoglobin C (HbSC), sickle cell thalassemia (HbS and HbA), thalassemia (alpha thalassemia and beta thalassemia), hemoglobin C disease (HbCC), hemoglobin C trait (HbAC), primary immunodeficiencies (e.g., autoimmune lymphoproliferative syndrome (ALPS), APS-1 (APECED), BENTA disease, caspase 8 deficiency (CEDS), CARD9 deficiency, and other candidiasis susceptibility syndromes) group, chronic granulomatous disease (CGD), common variable immunodeficiency (CVID), congenital neutropenic syndrome, CTLA4 deficiency, DOCK8 deficiency, GATA2 deficiency, glycosylation disorders with immunodeficiency, hyperimmunoglobulin E syndrome (HIES), hyperimmunoglobulin M syndrome, interferon gamma deficiency, interleukin-12 deficiency, and interleukin-23 deficiency, leukocyte adhesion deficiency (LAD), LRBA deficiency, PI3 kinase disease, PLCG2-associated antibody deficiency and immune dysregulation (PLAID), severe combined immunodeficiency (SCID), STAT3 dominant-negative diseases, STAT3 gain-of-function diseases, WARFARIN syndrome, hypogammaglobulinemia, infections, myeloid cellular pool (WHIM) syndrome, Wiskott-Aldrich syndrome (WAS), x-linked agammaglobulinemia (XLA), x-linked lymphoproliferative disorders (XLP), XMEN diseases), cytopenias (anemia, leukopenia, thrombocytopenia, pancytopenia, autoimmune cytopenia, refractory cytopenia), and / or storage and metabolic disorders (e.g., diabetes mellitus, familial hypercholesterolemia, Hunter syndrome, Krabbe disease,These include maple syrup urine disease, metachromatic leukodystrophy, Niemann-Pick disease, Gaucher disease, hemochromatosis, phenylketonuria (PKU), mitochondrial disorders, porphyria, Tay-Sachs disease, and Wilson's disease.

[0077] In some embodiments, the cargo is a morpholino, a peptide-linked morpholino, an antisense oligonucleotide, a PMO, a therapeutic transgene, a polynucleotide encoding a therapeutic polypeptide or peptide, a PPMO, one or more peptides, one or more polynucleotides encoding a CRISPR-Cas protein, a guide RNA, or both, a ribonucleoprotein comprising a CRISPR-Cas system molecule, a therapeutic transgene RNA, or other recombinant or therapeutic RNA and / or protein, or any combination thereof.

[0078] Engineered viral capsids and encoding polynucleotides Described herein are various embodiments of engineered or variant viral capsids, e.g., adeno-associated viral (AAV) capsids, that can be engineered to confer cell-specific tropism, e.g., hematopoietic cell-specific tropism, to the engineered viral particle. The engineered viral capsid can be a lentivirus, retrovirus, adenovirus, or AAV capsid. The engineered capsid can be included in an engineered viral particle (e.g., an engineered lentivirus, retrovirus, adenovirus, or AAV viral particle) and can confer cell-specific tropism, low immunogenicity, or both to the engineered viral particle. The engineered or variant viral capsids described herein can include one or more engineered or variant viral capsid proteins described herein. The engineered or variant viral capsids described herein can include a hematopoietic cell-specific targeting moiety that includes or is composed of an n-mer motif described elsewhere herein.

[0079] The engineered or variant viral capsid and / or capsid proteins may be encoded by one or more engineered or variant viral capsid polynucleotides. In some embodiments, the engineered viral capsid polynucleotide is an engineered AAV capsid polynucleotide, an engineered lentiviral capsid polynucleotide, an engineered retroviral capsid polynucleotide, or an engineered adenoviral capsid polynucleotide. In some embodiments, the engineered viral capsid polynucleotide (e.g., an engineered AAV capsid polynucleotide, an engineered lentiviral capsid polynucleotide, an engineered retroviral capsid polynucleotide, or an engineered adenoviral capsid polynucleotide) may include a 3' polyadenylation signal. The polyadenylation signal may be an SV40 polyadenylation signal.

[0080] The capsid of the engineered or variant virus can be a variant of the capsid of a wild-type virus. For example, in some embodiments, the capsid of the engineered AAV can be a variant of the capsid of a wild-type AAV. In some embodiments, the capsid of the wild-type AAV can be composed of VP1, VP2, VP3 capsid proteins, or a combination thereof. In other words, the capsid of the engineered AAV can include one or more variants of the wild-type VP1, wild-type VP2, and / or wild-type VP3 capsid proteins. In some embodiments, the serotype of the capsid of the reference wild-type AAV can be AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-8, AAV-9, or any combination thereof. In some embodiments, the serotype of the capsid of the wild-type AAV can be AAV-9. The capsid of the engineered AAV may have a different tropism than that of the capsid of the reference wild-type AAV.

[0081] The engineered or variant viral capsid may comprise between 1 and 60 engineered capsid proteins, in some embodiments, the engineered viral capsid may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 engineered capsid proteins. In some embodiments, the engineered viral capsid can comprise 0 to 59 wild-type viral capsid proteins. In some embodiments, the engineered viral capsid can comprise 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, or 59 wild-type viral capsid proteins.

[0082] In some embodiments, the engineered or variant AAV capsid may comprise between 1 and 60 engineered capsid proteins. In some embodiments, the engineered AAV capsid may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 engineered capsid proteins. In some embodiments, the engineered AAV capsid may comprise 0 to 59 wild-type AAV capsid proteins. In some embodiments, the engineered AAV capsid may comprise 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, or 59 wild-type AAV capsid proteins.

[0083] In some embodiments, the engineered or variant viral capsid protein can have an n-mer amino acid motif, where n can be at least 3 amino acids. In some embodiments, n can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids. In some embodiments, the engineered AAV capsid can have a hexameric or heptameric amino acid motif. In some embodiments, the n-mer amino acid motif can be inserted between two amino acids of a wild-type viral protein (VP) (or capsid protein). In some embodiments, the n-mer motif can be inserted between two amino acids within a variable amino acid region of a viral capsid protein.

[0084] In some embodiments, the n-mer motif can be inserted between two amino acids in the variable amino acid region of the AAV capsid protein. The core of each wild-type AAV viral protein contains an eight-stranded beta-barrel motif (beta B to beta I) and an alpha-helix (alpha A), which are conserved in the capsids of autonomously replicating parvoviruses (see, e.g., DiMattia et al. 2012. J. Virol. 86(12):6947-6958). Structural variable regions (VRs) occur in the surface loops connecting the beta strands, which cluster together to generate local variations on the capsid surface. AAV has 12 variable regions (also called hypervariable regions) (see, e.g., Weitzman and Linden. 2011. "Adeno-Associated Virus Biology." In Snyder, RO, Moullier, P. (eds.) Totowa, NJ: Humana Press). In some embodiments, one or more n-mer motifs may be inserted between two amino acids in one or more of the 12 variable regions in the wild-type AVV capsid protein. In some embodiments, the one or more n-mer motifs may each be inserted between two amino acids in VR-I, VR-II, VR-III, VR-IV, VR-V, VR-VI, VR-VII, VR-III, VR-IX, VR-X, VR-XI, VR-XII, or a combination thereof. In some embodiments, the n-mer may be inserted between two amino acids in VR-III of the capsid protein. In some embodiments, the n-mer may be inserted between two amino acids in VR-VIII of the capsid protein.In some embodiments, the engineered capsid may have an n-mer inserted between any two consecutive amino acids between amino acids 262 and 269 of the AAV9 viral protein, between any two consecutive amino acids between amino acids 327 and 332, between any two consecutive amino acids between amino acids 382 and 386, between any two consecutive amino acids between amino acids 452 and 460, between any two consecutive amino acids between amino acids 488 and 505, between any two consecutive amino acids between amino acids 545 and 558, between any two consecutive amino acids between amino acids 581 and 593, or between any two consecutive amino acids between amino acids 704 and 714. In some embodiments, the engineered capsid may have an n-mer inserted between amino acids 588 and 589 of the AAV9 viral protein. In some embodiments, the engineered capsid may have an n-mer inserted between amino acids 587 and 588 of the AAV9 viral protein. In some embodiments, the engineered capsid may have a heptamer motif inserted between amino acids 588 and 589 of the AAV9 viral protein. SEQ ID NO: 12001 is a reference AAV9 capsid sequence for referencing at least the insertion site discussed above. It will be understood that n-mers may be inserted at analogous positions in AAV viral proteins of other serotypes. In some embodiments, as discussed above, the n-mer(s) may be inserted between any two consecutive amino acids within the AAV viral protein, and in some embodiments, the insertion is made in the variable region.

[0085] In some embodiments, the first 1, 2, 3, or 4 amino acids of an n-mer motif can replace 1, 2, 3, or 4 amino acids of the polypeptide into which it is inserted and preceding the insertion position. For example, in one or more of the 7-mer inserts shown in FIG. 12, the first three amino acids shown can replace 1 to 3 amino acids in the polypeptide into which they are inserted. Using AAV as another non-limiting example, one or more of the n-mer motifs can be inserted, for example, between amino acids 587 and 588 or between amino acids 588 and 589 of the capsid polypeptide of AAV9, and the insertion can replace amino acids 586, 587, and 588, with the amino acid immediately preceding the n-mer motif after insertion being residue 585. It will be understood that this principle can be applied to any other insertion situation and is not necessarily limited to insertions between residues 587 and 588 or between residues 588 and 589 of the capsid of AAV9, or to insertions at the equivalent position in the capsid of another AAV. It will further be understood that in some embodiments the n-mer motif does not replace any amino acids in the polypeptide into which it is inserted.

[0086] SEQ ID NO: 12001 AAV9 capsid reference sequence.

[0087] MAADGYLPDWLEDNLSEGIREWWALKPGAPQPKANQQHQDNARGLVLPGYKYLGPGNGLDKGEPVNAADAAALEHDKAYDQQLKAGDNPYLKYNHADAEFQERLKEDTSFGGNLGRAVFQAKKRLLEPLGLVEEAAKTAPGKKRPVEQSPQEPDSSAGIGKSGAQPAKKRLNFGQTGDTESVPD PQPIGEPPAAPSGVGSLTMASGGGAPVADNNEGADGVGSSSGNWHCDSQWLGDRVITTSTRTWALPTYNNHLYKQISNSTSGGSSNDNAYFGYSTPWGYFDFNRFHCHFSPRDWQRLINNNWGFRPKRLNFKLFNIQVKEVTDNNGVKTIANNLTSTVQVFTDSDYQLPYVLGSAHEGCLPPFP ADVFMIPQYGYLTLNDGSQAVGRSSFYCLEYFPSQMLRTGNNFQFSYEFENVPFHSSYAHSQSLDRLMNPLIDQYLYYLSKTINGSGQNQQTLKFSVAGPSNMAVQGRNYIPGPSYRQQRVSTTVTQNNNSEFAWPGASSWALNGRNNSLMNPGPAMASHKEGEDRFFPLSGSLIFGKQGTGRDN VDADKVMITNEEEIKTTNPVATESYGQVATNHQSAQAQAQTGWVQNQGILPGMVWQDRDVYLQGPIWAKIPHTDGNFHPSPLMGGFGMKHPPPQILIKNTPVPADPPTAFNKDKLNSFITQYSTGQVSVEIEWELQKENSKRWNPEIQYTSNYYKSNNVEFAVNTEGVYSEPRPIGTRYLTRNL

[0088] In some embodiments, the n-mer can be any amino acid motif shown in Figure 12 or encoded by a nucleic acid shown in Figure 12. In some embodiments, the n-mer is VKX n Each X n is selected from any amino acid, and n is 1, 2, 3, 4, 5, 6, or 7, more particularly, n is 5. In some embodiments, the n-mer is VKX n Contains YGAL, each Xn is selected from any amino acid, and n is 1, 2, or 3, more particularly, n is 1. In some embodiments, insertion of the n-mer into an AAV or other viral capsid can result in cells, tissues, organs, specific engineered AAV or other viral capsids, or other compositions comprising the n-mer motif or capsid protein of the invention. In some embodiments, engineered or variant capsids or other compositions comprising the n-mer motif have specificity for bone tissue and / or cells, lung tissue and / or cells, liver tissue and / or cells, bladder tissue and / or cells, kidney tissue and / or cells, heart tissue and / or cells, skeletal muscle tissue and / or cells, smooth muscle and / or cells, neuronal tissue and / or cells, intestinal tissue and / or cells, pancreatic tissue and / or cells, adrenal tissue and / or cells, brain tissue and / or cells, tendon tissue or cells, skin tissue and / or cells, spleen tissue and / or cells, eye tissue and / or cells, blood or hematopoietic cells and / or hematopoietic organs, synovial cells, immune cells (including specificity for particular types of immune cells), and combinations thereof. In some embodiments, engineered or variant capsids or other compositions comprising the n-mer motif have specificity for hematopoietic cells.

[0089] In some embodiments, the AAV capsid or other viral capsid or composition may be hematopoietic cell-specific. In some embodiments, the cell specificity of the engineered AAV capsid or other viral capsid or composition is conferred by a hematopoietic cell-specific n-mer motif incorporated into the engineered or variant AAV or other viral capsid or other composition described herein. Without intending to be bound by theory, it is believed that the n-mer motif confers 3D structure to or within a domain or region of the engineered AAV capsid or other viral capsid or other composition, such that interaction of a viral particle or other composition comprising the engineered AAV capsid or other viral capsid or other composition described herein increases or improves (e.g., increases affinity for) interaction with cell surface receptors and / or other surface molecules of hematopoietic cells. In some embodiments, the cell surface receptor is an AAV receptor (AAVR). In some embodiments, the cell surface receptor is a hematopoietic cell-specific AAV receptor. In some embodiments, the cell surface receptor or other molecule is a cell surface receptor or other molecule that is selectively expressed on the surface of a hematopoietic cell.

[0090] In some embodiments, hematopoietic cell-specific engineered or variant viral particles or other compositions described herein that comprise a hematopoietic cell-specific capsid, n-mer motif, or hematopoietic cell-specific targeting moiety described herein may have improved uptake, delivery rate, transduction rate, efficiency, quantity, or a combination thereof in hematopoietic cells compared to other cell types and / or other viral particles (including, but not limited to, AAV) and other compositions that do not comprise a hematopoietic cell-specific n-mer motif of the invention.

[0091] Also described herein are polynucleotides described herein that encode the engineered hematopoietic cell-specific targeting moieties and other compositions described herein, including, but not limited to, engineered or variant AAV capsids.

[0092] In some embodiments, the engineered or variant polynucleotide may be included in a polynucleotide configured to be a viral genome donor of a viral vector system that may be used to generate engineered or variant viral particles as described elsewhere herein.

[0093] In some embodiments, the polynucleotide encoding the engineered or variant AAV capsid can be included in a polynucleotide configured to be an AAV genome donor in an AAV vector system that can be used to generate the engineered AAV particles described elsewhere herein. In some embodiments, the polynucleotide encoding the engineered AAV capsid can be operably linked to a polyadenylation tail. In some embodiments, the polyadenylation tail can be the SV40 polyadenylation tail. In some embodiments, the polynucleotide encoding the AAV capsid can be operably linked to a promoter. In some embodiments, the promoter can be a tissue-specific promoter. In some embodiments, the tissue-specific promoter is specific for muscle (e.g., cardiac, skeletal, and / or smooth muscle), neurons and support cells (e.g., astrocytes, glial cells, Schwann cells, etc.), adipose, spleen, liver, kidney, immune cells, cerebrospinal fluid cells, synovial cells, skin cells, cartilage, tendon, connective tissue, bone, pancreas, adrenal gland, blood cells, bone marrow cells, thymocytes, lymph node cells, placenta, endothelial cells, and combinations thereof. In some embodiments, the promoter may be a constitutive promoter. Suitable promoters are discussed elsewhere herein, are generally known in the art, and may be commercially available.

[0094] Suitable endothelial cell-specific promoters include, but are not limited to, the Fit-1 promoter and the ICAM-2 promoter.

[0095] Suitable neuronal tissue / cell-specific promoters include, but are not limited to, the GFAP promoter (astrocytes), the SYN1 promoter (neurons), and NSE / RU5' (mature neurons).

[0096] Suitable kidney-specific promoters include, but are not limited to, the NphsI promoter (podocyte).

[0097] Suitable bone-specific promoters include, but are not limited to, the OG-2 promoter (osteoblasts, odontoblasts).

[0098] Suitable lung-specific promoters include, but are not limited to, the SP-B promoter (lung).

[0099] Suitable liver-specific promoters include, but are not limited to, the SV40 / Alb promoter.

[0100] Suitable cardiac-specific promoters include, but are not limited to, alpha-MHC.

[0101] Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1 alpha, CAG, and beta-actin.

[0102] AAV with reduced non-hematopoietic cell specificity In some embodiments, the n-mer motif(s) described herein are inserted into an AAV protein (e.g., an AAV capsid protein) that has reduced specificity (or no detectable, measurable, or clinically significant interaction) with one or more non-hematopoietic cell types. Exemplary non-hematopoietic cell types include, but are not limited to, liver, kidney, lung, heart, spleen, central or peripheral nervous system cells, bone, immune, stomach, intestinal, eye, skin cells, etc. In some embodiments, the non-hematopoietic cell is a liver cell.

[0103] In certain exemplary embodiments, the AAV capsid protein is an engineered AAV capsid protein that has reduced or eliminated uptake in non-hematopoietic cells compared to the corresponding wild-type AAV capsid polypeptide.

[0104] In certain exemplary embodiments, the non-hematopoietic cells are liver cells.

[0105] In certain exemplary embodiments, the wild-type capsid polypeptide is a capsid polypeptide of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10.

[0106] In certain exemplary embodiments, the capsid protein of the engineered AAV comprises one or more mutations that reduce or eliminate uptake in non-hematopoietic cells.

[0107] In certain exemplary embodiments, the one or more mutations are in the capsid protein of AAV9 (SEQ ID NO: 12001). a.267th place, b.269th place, c.504th place, d.505th place, e.590th place, f. or any combination thereof. or at one or more corresponding positions in the non-AAV9 capsid polypeptide.

[0108] In certain exemplary embodiments, the non-AAV9 capsid protein is an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10 capsid polypeptide.

[0109] In certain exemplary embodiments, the mutation at position 267 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a G or X to A mutation, where X is any amino acid.

[0110] In certain exemplary embodiments, the mutation at position 269 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is an S or X to T mutation, where X is any amino acid.

[0111] In certain exemplary embodiments, the mutation at position 504 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a G or X to A mutation, where X is any amino acid.

[0112] In certain exemplary embodiments, the mutation at position 505 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a P or X to A mutation, where X is any amino acid.

[0113] In certain exemplary embodiments, the mutation at position 590 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a Q or X to A mutation, where X is any amino acid.

[0114] In certain exemplary embodiments, the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 267, 269, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 267 is a G to A mutation and the mutation at position 269 is an S to T mutation.

[0115] In certain exemplary embodiments, the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 590 of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 509 is a Q to A mutation.

[0116] In certain exemplary embodiments, the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 504, 505, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 504 is a G to A mutation and the mutation at position 505 is a P to A mutation.

[0117] In some embodiments, the AAV capsid protein into which the n-mer motif(s) can be inserted can be 80 to 100 (e.g., 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, to / or 100) percent identical to SEQ ID NO:4 or SEQ ID NO:5 set forth in International Patent Application Publication No. WO2019 / 217911, which is incorporated herein in its entirety. These sequences are also incorporated herein as SEQ ID NOs:12002 and 12003, respectively. When considering variants of these AAV9 capsid proteins with reduced liver specificity, it will be understood that residues 267 and / or 269 must contain the relevant mutation or equivalent.

[0118] SEQ ID NO: 12002 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser Glu Gly Ile Arg Glu Trp Trp Ala Leu Lys Pro Gly Ala Pro Gln Pro Lys Ala Asn Gln Gln His Gln Asp Asn Ala Arg Gly Leu Val Leu Pro Gly Tyr Lys Val Leu Gly Pro Gly Asn Gly Leu Asp Lys Gly Glu Pro Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp Gln Gln Leu Lys Ala Gly Asp Asn Pro Tyr Leu Lys Tyr Asn His Ala Asp Ala Glu Phe Gln Glu Arg Leu Lys Glu Asp Thr Ser Phe Gly Gly Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Leu Leu Glu Pro Leu Gly Leu Val Glu Glu Ala Ala Lys Thr Ala Pro Gly Lys Lys Arg Pro Val Glu Gln Ser Pro Gln Glu Pro Asp Ser Ser Ala Gly Ile Gly Lys Ser Gly Ala Gln Pro Ala Lys Lys Arg Leu Asn Phe Gly Gln Thr Gly Asp Thr Glu Ser Val Pro Asp Pro Gln Pro Ile Gly Glu Pro Pro Ala Ala Pro Ser Gly Val Gly Ser Leu Thr Met Ala Ser Gly Gly Gly Ala Pro Val Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Ser Ser Ser Gly Asn Trp His Cys Asp Ser Gln Trp Leu Gly Asp Arg Val Ile Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu Tyr Lys Gln Ile Ser Asn Ser Thr Ser Gly Ala Ser Ser Asn Asp Asn Ala Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn Asn Trp Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile Gln Val Lys Glu Val Thr Asp Asn Asn Gly Val Lys Thr Ile Ala Asn Asn Leu Thr Ser Thr Val Gln Val Phe Thr Asp Ser Asp Tyr Gln Leu Pro Tyr Val Leu Gly Ser Ala His Glu Gly Cys Leu Pro Pro Phe Pro Ala Asp Val Phe Met Ile Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asp Gly Ser Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Gln Phe Ser Tyr Glu Phe Glu Asn Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu Ser Lys Thr Ile Asn Gly Ser Gly Gln Asn Gln Gln Thr Leu Lys Phe Ser Val Ala Gly Pro Ser Asn Met Ala Val Gln Gly Arg Asn Tyr Ile Pro Gly Pro Ser Tyr Arg Gln Gln Arg Val Ser Thr Thr Val Thr Gln Asn Asn Asn Ser Glu Phe Ala Trp Pro Gly Ala Ser Ser Trp Ala Leu Asn Gly Arg Asn Ser Leu Met Asn Pro Gly Pro Ala Met Ala Ser His Lys Glu Gly Glu Asp Arg Phe Phe Pro Leu Ser Gly Ser Leu Ile Phe Gly Lys Gln Gly Thr Gly Arg Asp Asn Val Asp Ala Asp Lys Val Met Ile Thr Asn Glu Glu Glu Ile Lys Thr Thr Asn Pro Val Ala Thr Glu Ser Tyr Gly Gln Val Ala Thr Asn His Gln Ser Ala Gln Ala Gln Ala Gln Thr Gly Trp Val Gln Asn Gln Gly Ile Leu Pro Gly Met Val Trp Gln Asp Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Met Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala Asp Pro Pro Thr Ala Phe Asn Lys Asp Lys Leu Asn Ser Phe Ile Thr Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn Tyr Tyr Lys Ser Asn Asn Val Glu Phe Ala Val Asn Thr Glu Gly Val Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg Asn Leu sequence number 12003 Met Ala Ala Asp Gly Tyr Leu Pro Asp Trp Leu Glu Asp Asn Leu Ser Glu Gly Ile Arg Glu Trp Trp Ala Leu Lys Pro Gly Ala Pro Gln Pro Lys Ala Asn Gln Gln His Gln Asp Asn Ala Arg Gly Leu Val Leu Pro Gly Tyr Lys Val Leu Gly Pro Gly Asn Gly Leu Asp Lys Gly Glu Pro Val Asn Ala Ala Asp Ala Ala Ala Leu Glu His Asp Lys Ala Tyr Asp Gln Gln Leu Lys Ala Gly Asp Asn Pro Tyr Leu Lys Tyr Asn His Ala Asp Ala Glu Phe Gln Glu Arg Leu Lys Glu Asp Thr Ser Phe Gly Gly Asn Leu Gly Arg Ala Val Phe Gln Ala Lys Lys Arg Leu Leu Glu Pro Leu Gly Leu Val Glu Glu Ala Ala Lys Thr Ala Pro Gly Lys Lys Arg Pro Val Glu Gln Ser Pro Gln Glu Pro Asp Ser Ser Ala Gly Ile Gly Lys Ser Gly Ala Gln Pro Ala Lys Lys Arg Leu Asn Phe Gly Gln Thr Gly Asp Thr Glu Ser Val Pro Asp Pro Gln Pro Ile Gly Glu Pro Pro Ala Ala Pro Ser Gly Val Gly Ser Leu Thr Met Ala Ser Gly Gly Gly Ala Pro Val Ala Asp Asn Asn Glu Gly Ala Asp Gly Val Gly Ser Ser Ser Gly Asn Trp His Cys Asp Ser Gln Trp Leu Gly Asp Arg Val Ile Thr Thr Ser Thr Arg Thr Trp Ala Leu Pro Thr Tyr Asn Asn His Leu Tyr Lys Gln Ile Ser Asn Ser Thr Ser Gly Ala Ser Thr Asn Asp Asn Ala Tyr Phe Gly Tyr Ser Thr Pro Trp Gly Tyr Phe Asp Phe Asn Arg Phe His Cys His Phe Ser Pro Arg Asp Trp Gln Arg Leu Ile Asn Asn Asn Trp Gly Phe Arg Pro Lys Arg Leu Asn Phe Lys Leu Phe Asn Ile Gln Val Lys Glu Val Thr Asp Asn Asn Gly Val Lys Thr Ile Ala Asn Asn Leu Thr Ser Thr Val Gln Val Phe Thr Asp Ser Asp Tyr Gln Leu Pro Tyr Val Leu Gly Ser Ala His Glu Gly Cys Leu Pro Pro Phe Pro Ala Asp Val Phe Met Ile Pro Gln Tyr Gly Tyr Leu Thr Leu Asn Asp Gly Ser Gln Ala Val Gly Arg Ser Ser Phe Tyr Cys Leu Glu Tyr Phe Pro Ser Gln Met Leu Arg Thr Gly Asn Asn Phe Gln Phe Ser Tyr Glu Phe Glu Asn Val Pro Phe His Ser Ser Tyr Ala His Ser Gln Ser Leu Asp Arg Leu Met Asn Pro Leu Ile Asp Gln Tyr Leu Tyr Tyr Leu Ser Lys Thr Ile Asn Gly Ser Gly Gln Asn Gln Gln Thr Leu Lys Phe Ser Val Ala Gly Pro Ser Asn Met Ala Val Gln Gly Arg Asn Tyr Ile Pro Gly Pro Ser Tyr Arg Gln Gln Arg Val Ser Thr Thr Val Thr Gln Asn Asn Asn Ser Glu Phe Ala Trp Pro Gly Ala Ser Ser Trp Ala Leu Asn Gly Arg Asn Ser Leu Met Asn Pro Gly Pro Ala Met Ala Ser His Lys Glu Gly Glu Asp Arg Phe Phe Pro Leu Ser Gly Ser Leu Ile Phe Gly Lys Gln Gly Thr Gly Arg Asp Asn Val Asp Ala Asp Lys Val Met Ile Thr Asn Glu Glu Glu Ile Lys Thr Thr Asn Pro Val Ala Thr Glu Ser Tyr Gly Gln Val Ala Thr Asn His Gln Ser Ala Gln Ala Gln Ala Gln Thr Gly Trp Val Gln Asn Gln Gly Ile Leu Pro Gly Met Val Trp Gln Asp Arg Asp Val Tyr Leu Gln Gly Pro Ile Trp Ala Lys Ile Pro His Thr Asp Gly Asn Phe His Pro Ser Pro Leu Met Gly Gly Phe Gly Met Lys His Pro Pro Pro Gln Ile Leu Ile Lys Asn Thr Pro Val Pro Ala Asp Pro Pro Thr Ala Phe Asn Lys Asp Lys Leu Asn Ser Phe Ile Thr Gln Tyr Ser Thr Gly Gln Val Ser Val Glu Ile Glu Trp Glu Leu Gln Lys Glu Asn Ser Lys Arg Trp Asn Pro Glu Ile Gln Tyr Thr Ser Asn Tyr Tyr Lys Ser Asn Asn Val Glu Phe Ala Val Asn Thr Glu Gly Val Tyr Ser Glu Pro Arg Pro Ile Gly Thr Arg Tyr Leu Thr Arg Asn Leu

[0119] In some embodiments, the capsid protein of the AAV into which the n-mer motif(s) can be inserted can be 80-100 (e.g., 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, to / or 100) percent identical to any of those described in Adachi et al., (Nat.Comm. 2014.5:3075, DOI:10.1038 / ncomms4075), which has reduced specificity for non-CNS cells, particularly liver cells. Adachi et al., (Nat.Comm. 2014.5:3075, DOI:10.1038 / ncomms4075) is incorporated herein by reference in its entirety.

[0120] In some embodiments, the modified AAV has a specificity for non-hematopoietic cells that is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, In some embodiments, the modified AAV may have no measurable or detectable uptake and / or expression in one or more non-hematopoietic cells.

[0121] Methods for generating engineered AAV capsids Also provided herein are methods for generating engineered AAV capsids. The engineered AAV capsid variants can be variants of wild-type AAV capsids. Figures 6-8 illustrate various embodiments of methods capable of generating engineered or variant AAV capsids having the variant motifs described herein. Generally, an AAV capsid library can be generated by expressing engineered capsid vectors, each containing a previously described engineered AAV capsid polynucleotide, in a suitable AAV producer cell line. See, for example, Figure 8. While Figure 8 illustrates a method for AAV particle production using a helper, it will be understood that this can also be achieved via a helper-free method. This allows for the generation of a library of AAV capsids that can contain one or more desired cell-specific engineered AAV capsid variants. As shown in Figure 6, the AAV capsid library can be administered to various non-human animals, or non-human animals transplanted with human hematopoietic cells, for a first round of mRNA-based selection. As shown in Figure 1, the transduction process with AAV and related vectors can result in the production of mRNA molecules that reflect the genome of the virus that transduced the cells. mRNA-based selection can be more specific and effective in determining viral particles that are capable of functionally transducing cells because it is based on the functional product produced, as opposed to simply detecting the presence of viral particles in the cells by measuring the presence of viral DNA.

[0122] After the first administration, one or more engineered AAV viral particles having the desired capsid variant can then be used to form a filtered AAV capsid library. Desired AAV viral particles can be identified by measuring the mRNA expression of the capsid variant and determining which variant is highly expressed in the desired cell type(s) compared to undesired cell type(s). Those highly expressed in the desired cell, tissue, and / or organ type are the desired AAV capsid variant particles. In some embodiments, the polynucleotide encoding the AAV capsid variant is under the control of a tissue-specific promoter that has selective activity in the desired cell, tissue, or organ.

[0123] The engineered AAV capsid variant particles identified in the first round can then be administered to various non-human animals or non-human animals transplanted with human hematopoietic cells. In some embodiments, the animals used in the second round of selection and identification are not the same as those used in the first round of selection and identification. As in the first round, after administration, the top-expressing variants in the desired cell, tissue, and / or organ type(s) can be identified by measuring viral mRNA expression in the cells. The top variants identified after the second round can then optionally be barcoded and optionally pooled. In some embodiments, the top variants from the second round can then be administered to non-human primates to identify the top cell-specific variant(s), especially if the ultimate use of the top variants is in humans. Administration in each round can be systemic.

[0124] In some embodiments, the method for producing the AAV capsid variant may include the following steps: (a) expressing a vector system described herein comprising an engineered AAV capsid polynucleotide in a cell to produce an engineered AAV viral particle capsid variant; (b) harvesting the engineered AAV viral particle capsid variant produced in step (a); (c) administering the engineered AAV viral particle capsid variant to one or more first subjects, wherein the engineered AAV viral particle capsid variant is produced by expressing an engineered AAV capsid variant vector or system thereof in a cell and harvesting the engineered AAV viral particle capsid variant produced by the cell; and (d) identifying one or more engineered AAV capsid variants produced at significantly higher levels by one or more specific cells or specific cell types or specific tissues / organs in the one or more first subjects. In this context, "significantly higher" means approximately 2 × 10 per 15 cm dish. 11 ~about 6×10 12 It can refer to the titer that can be achieved by the vector genome.

[0125] The method may further include the following steps: (e) administering some or all of the engineered AAV viral particle capsid variants identified in step (d) to one or more second subjects, and (f) identifying one or more engineered AAV viral particle capsid variants produced at significantly higher levels in one or more specific cells or specific cell types in the one or more second subjects. The cells in step (a) may be prokaryotic or eukaryotic. In some embodiments, the administration in step (c), step (e), or both, is systemic. In some embodiments, the one or more first subjects, the one or more second subjects, or both, are non-human mammals or non-human mammals transplanted with human cells. In some embodiments, the one or more first subjects, the one or more second subjects, or both, are each independently selected from the group consisting of wild-type non-human mammals, humanized non-human mammals, disease-specific non-human mammal models, and non-human primates.

[0126] In some embodiments, further optimization of the variant motif may be performed.

[0127] The polynucleotide and vector systems described herein can also be used to generate viral particles and other compositions that can be generated to contain cargo molecules that can be delivered to cells.

[0128] Engineered Vectors and Vector Systems Also provided herein are vectors and vector systems that can include one or more of the engineered polynucleotides described herein that can encode one or more n-mer motifs of the invention, including, but not limited to, engineered viral polynucleotides (e.g., engineered AAV polynucleotides). In some embodiments, the polynucleotide(s) that can encode an n-mer motif of the invention can be any of those described in FIG. 12 and / or described elsewhere herein. In some embodiments, the polynucleotides can encode any n-mer motif described in FIG. 12 and / or described elsewhere herein. As used in this context, an engineered viral capsid polynucleotide refers to any one or more of the polynucleotides described herein that can encode an engineered viral capsid described elsewhere herein and / or a polynucleotide(s) that can encode one or more engineered viral capsid proteins described elsewhere herein. Furthermore, when the vector comprises an engineered viral capsid polynucleotide described herein, the vector can also be referred to and considered an engineered vector or system thereof, even if not specifically referred to as such. In some embodiments, the vectors can include one or more polynucleotides encoding one or more elements of the engineered viral capsids described herein. The vectors and systems can be useful for producing bacteria, fungi, yeast, plant cells, animal cells, and transgenic animals that can express one or more components of the engineered viral capsids, particles, or other compositions described herein. Included within the scope of the present disclosure are vectors that include one or more of the polynucleotide sequences described herein. One or more of the polynucleotides that are part of the engineered viral capsids and systems described herein can be included in a vector or vector system.

[0129] In some embodiments, the vector can comprise an engineered viral (e.g., AAV) capsid polynucleotide having a 3' polyadenylation signal. In some embodiments, the 3' polyadenylation signal is an SV40 polyadenylation signal. In some embodiments, the vector does not have a splice control element. In some embodiments, the vector comprises one or more minimal splice control elements. In some embodiments, the vector can further comprise a modified splice control element, wherein the modification inactivates the splice control element. In some embodiments, the modified splice control element is a polynucleotide sequence sufficient to direct splicing between a rep protein polynucleotide and the engineered viral (e.g., AAV) capsid protein variant polynucleotide. In some embodiments, the polynucleotide sequence that may be sufficient to direct splicing is a splice acceptor or a splice donor. In some embodiments, the viral (e.g., AAV) capsid polynucleotide is an engineered viral (e.g., AAV) capsid polynucleotide described elsewhere herein. In some embodiments, the vector does not comprise one or more minimal splice control elements, modified splice regulators, splice acceptors, and / or splice donors.

[0130] The vectors and / or vector systems can be used, for example, to express one or more of the engineered viral (e.g., AAV) capsids and / or other polynucleotides in a cell, such as a producer cell, to produce engineered viral (e.g., AAV) particles and / or other compositions (e.g., polypeptides, particles, etc.) comprising engineered viral (e.g., AAV) capsids or other compositions comprising the n-mer motifs of the invention described elsewhere herein. Other uses of the vectors and vector systems described herein are also within the scope of this disclosure. In some contexts, as will be understood by those skilled in the art, a "vector" can be a term of the art to refer to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. A vector can be a replicon, such as a plasmid, phage, or cosmid, into which another DNA segment can be inserted to effect replication of the inserted segment. Typically, a vector is capable of replication when associated with appropriate control elements.

[0131] Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other types of polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus (AAV)). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Certain vectors are capable of autonomous replication in host cells into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, thereby replicating along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.

[0132] A recombinant expression vector can be comprised of a nucleic acid (e.g., a polynucleotide) of the invention in a form suitable for expression of the nucleic acid in a host cell. This means that the recombinant expression vector contains one or more regulatory elements operably linked to the nucleic acid sequence to be expressed, which can be selected based on the host cell to be used for expression. Within a recombinant expression vector, "operably linked" and "operably linked" are used interchangeably herein and are further defined elsewhere herein. In the context of a vector, the term "operably linked" is intended to mean that the nucleotide sequence of interest is linked to regulatory element(s) in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). Advantageous vectors include adeno-associated viruses, and such vector types can also be selected to target specific cell types, for example, engineered viral (e.g., AAV) vectors containing engineered viral (e.g., AAV) capsid polynucleotides with desired cell-specific tropism. These and other embodiments of the vectors and vector systems are described elsewhere herein.

[0133] In some embodiments, the vector may be a bicistronic vector. In some embodiments, a bicistronic vector may be used with one or more engineered viral (e.g., AAV) capsid system elements described herein. In some embodiments, expression of the engineered viral (e.g., AAV) capsid system elements described herein may be driven by an appropriate constitutive or tissue-specific promoter. If the engineered viral (e.g., AAV) capsid system element is RNA, its expression may be driven by a Pol III promoter, such as the U6 promoter. In some embodiments, the two are combined.

[0134] Cell-based vector amplification and expression Vectors can be designed for expression of one or more elements (e.g., nucleic acid transcripts, proteins, enzymes, and combinations thereof) of an engineered viral (e.g., AAV) capsid system or other composition comprising the n-mer motif of the invention described herein in a suitable host cell. In some embodiments, the suitable host cell is a prokaryotic cell. Suitable host cells include, but are not limited to, bacterial cells, yeast cells, insect cells, and mammalian cells. The vectors can be viral or non-viral based. In some embodiments, the suitable host cell is a eukaryotic cell. In some embodiments, the suitable host cell is a suitable bacterial cell. Suitable bacterial cells include, but are not limited to, bacterial cells from the Escherichia coli species of bacteria. Many suitable strains of E. coli are known in the art for expression of vectors. These include, but are not limited to, Pir1, Stbl2, Stbl3, Stbl4, TOP10, XL1 Blue, and XL10 Gold. In some embodiments, the host cell is a suitable insect cell. Suitable insect cells include those from Spodoptera frugiperda. Suitable strains of S. frugiperda cells include, but are not limited to, Sf9 and Sf21. In some embodiments, the host cell is a suitable yeast cell. In some embodiments, the yeast cell may be from Saccharomyces cerevisiae. In some embodiments, the host cell is a suitable mammalian cell. Many mammalian cell types have been developed for expressing vectors. Suitable mammalian cells include, but are not limited to, HEK293, Chinese hamster ovary cells (CHO), mouse myeloma cells, HeLa, U2OS, A549, HT1080, CAD, P19, NIH 3T3, L929, N2a, MCF-7, Y79, SO-Rb50, HepG G2, DIKX-X11, J558L, baby hamster kidney cells (BHK), and chicken embryo fibroblasts (CEF).Suitable host cells are further discussed in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990).

[0135] In some embodiments, the vector may be a yeast expression vector. Examples of vectors for expression in the yeast Saccharomyces cerevisiae include pYepSec1 (Baldari, et al., 1987, EMBO J. 6:229-234), pMFa (Kuijan and Herskowitz, 1982, Cell 30:933-943), pJRY88 (Schultz et al., 1987, Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.). As used herein, "yeast expression vector" refers to a nucleic acid that contains one or more sequences encoding RNA and / or polypeptides and may further contain any desired elements that control expression of the nucleic acid(s), as well as any elements that allow replication and maintenance of the expression vector within the yeast cell. Many suitable yeast expression vectors and their characteristics are known in the art. For example, various vectors and techniques are described in "Yeast Protocols, 2nd edition," Xiao, W., ed. (Humana Press, New York, 2007) and Buckholz, RG, and Gleeson, MA (1991) Biotechnology (NY) 9(11):1067-72. Yeast vectors may include, but are not limited to, a centromeric (CEN) sequence, an autonomously replicating sequence (ARS), a promoter such as an RNA polymerase III promoter operably linked to a sequence or gene of interest, a terminator such as an RNA polymerase III terminator, an origin of replication, and a marker gene (e.g., auxotrophic, antibiotic, or other selectable marker). Examples of expression vectors for use in yeast include plasmids, yeast artificial chromosomes, 2μ plasmids, yeast integrating plasmids, yeast replicating plasmids, shuttle vectors, and episomal plasmids.

[0136] In some embodiments, the vector may be a baculovirus vector or expression vector suitable for expressing polynucleotides and / or proteins in insect cells. Baculovirus vectors available for expressing proteins in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39). rAAV (recombinant adeno-associated virus) vectors are preferably produced in insect cells, such as Spodoptera frugiperda Sf9 insect cells, and grown in serum-free suspension culture. Serum-free insect cells can be purchased from commercial vendors, such as Sigma-Aldrich (EX-CELL 405).

[0137] In some embodiments, the vector is a mammalian expression vector. In some embodiments, the mammalian expression vector is capable of expressing one or more polynucleotides and / or polypeptides in mammalian cells. Examples of mammalian expression vectors include, but are not limited to, pCDM8 (Seed, 1987. Nature 329:840) and pMT2PC (Kaufman, et al., 1987. EMBO J. 6:187-195). The mammalian expression vector may contain one or more appropriate regulatory elements capable of controlling the expression of one or more polynucleotides and / or proteins in mammalian cells. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. Further details regarding suitable regulatory elements are provided elsewhere herein.

[0138] For other suitable expression vectors and vector systems for both prokaryotic and eukaryotic cells, see, e.g., chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0139] In some embodiments, the recombinant mammalian expression vector is capable of directing expression of the nucleic acid preferentially in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific, Pinkert, et al., 1987. Genes Dev. 1:268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43:235-275), promoters of T-cell receptors in particular (Winoto and Baltimore, 1989. EMBO J. 8:729-733) and immunoglobulins (Baneiji, et al., 1983. Cell 33:729-740; Queen and Baltimore, 1983. Cell 33:741-748), neuron-specific promoters (e.g., the neurofilament promoter, Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86:5473-5477), pancreatic-specific promoters (Edlund, et al., 1989. Proc. Natl. Acad. Sci. USA 86:5473-5477), and the like. al., 1985, Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoters, U.S. Pat. No. 4,873,316 and European Patent Publication No. 264,166). Developmentally regulated promoters, such as the mouse hox promoters (Kessel and Gruss, 1990, Science 249:374-379) and the alpha-fetoprotein promoter (Campes and Tilghman, 1989, Genes Dev. 3:537-546), are also encompassed. With regard to these prokaryotic and eukaryotic vectors, reference is made to U.S. Pat. No. 6,750,059, the contents of which are incorporated herein by reference in their entirety. Other embodiments may utilize viral vectors, as referenced in U.S. Patent Application No. 13 / 092,085, the contents of which are incorporated herein by reference in their entirety.Tissue-specific regulatory elements are known in the art, and in this regard reference is made to U.S. Patent No. 7,776,321, the contents of which are incorporated herein by reference in their entirety. In some embodiments, a regulatory element is operably linked to and capable of driving expression of one or more elements of the engineered AAV capsid system described herein.

[0140] Vectors may be introduced and propagated in prokaryotes or prokaryotic cells. In some embodiments, prokaryotes are used to amplify copies of vectors that are introduced into eukaryotic cells or as intermediate vectors in the production of vectors that are introduced into eukaryotic cells (e.g., amplifying plasmids as part of a packaging system for viral vectors). In some embodiments, prokaryotes are used to amplify copies of vectors and express one or more nucleic acids to provide a source of one or more proteins for delivery to a host cell or host organism.

[0141] In some embodiments, the vector may be a fusion vector or a fusion expression vector. In some embodiments, a fusion vector adds multiple amino acids to a protein encoded therein, for example, to the amino terminus, carboxy terminus, or both of the recombinant protein. Such fusion vectors may serve one or more purposes, for example, (i) increasing the expression of the recombinant protein, (ii) increasing the solubility of the recombinant protein, and (iii) aiding in the purification of the recombinant protein by acting as a ligand in affinity purification. In some embodiments, expression of polynucleotides (such as non-coding polynucleotides) and proteins in prokaryotes may be carried out in Escherichia coli using vectors containing constitutive or inducible promoters directing the expression of either fusion or non-fusion polynucleotides and / or proteins. In some embodiments, the fusion expression vector may contain a proteolytic cleavage site that may be introduced at the junction of the fusion vector backbone or other fusion moiety and the recombinant polynucleotide or protein, thereby allowing for separation of the recombinant polynucleotide or protein from the fusion vector backbone or other fusion moiety following purification of the fusion polynucleotide or protein. Such enzymes, and their cognate recognition sequences, include factor Xa, thrombin, and enterokinase. Exemplary fusion expression vectors include pGEX (Pharmacia Biotech Inc., Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, Mass.), and pRIT5 (Pharmacia, Piscataway, NJ), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A, respectively, to the target recombinant protein.Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET 11d (Studier et al., GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990) 60-89).

[0142] In some embodiments, one or more vectors driving the expression of one or more elements of an engineered or variant viral (e.g., AAV) capsid system or other composition comprising an n-mer motif described herein are introduced into a host cell, such that expression of the elements of the engineered delivery system described herein directs the production of an engineered viral (e.g., AAV) capsid system or other composition comprising an n-mer motif described herein (including, but not limited to, engineered gene transfer agent particles described in more detail elsewhere herein). For example, different elements of an engineered viral (e.g., AAV) capsid system or other composition comprising an n-mer motif described herein can each be operably linked to separate control elements on separate vectors. The RNA(s) of the different elements of the engineered delivery systems described herein can be delivered to an animal or mammal or cells thereof that constitutively, inducibly, or conditionally express the different elements of an engineered viral (e.g., AAV) capsid system or other composition comprising an n-mer motif as described herein, to produce an animal or mammal or cells thereof that contain one or more cells that incorporate one or more elements of an engineered viral (e.g., AAV) capsid system or other composition comprising an n-mer motif as described herein, or that incorporate and / or express one or more elements of an engineered viral (e.g., AAV) capsid system or other composition comprising an n-mer motif as described herein.

[0143] In some embodiments, two or more of the elements expressed from the same or different regulatory element(s) may be combined in a single vector, with one or more additional vectors providing any components of the system not included in the first vector. Engineered polynucleotides of the invention combined in a single vector may be positioned in any suitable orientation, e.g., one element may be located 5' ("upstream") or 3' ("downstream") relative to the second element. The coding sequence of one element may be located on the same or opposite strand as the coding sequence of the second element, and may be oriented in the same or opposite direction. In some embodiments, a single promoter drives expression of transcripts encoding one or more (e.g., each contained in a different intron, two or more contained in at least one intron, or all contained in a single intron) capsid proteins or other compositions of an engineered virus (e.g., AAV) comprising the n-mer motif described herein embedded in one or more intron sequences. In some embodiments, engineered polynucleotides of the invention (including but not limited to engineered viral polynucleotides) may be operably linked to and expressed from the same promoter.

[0144] Vector characteristics The vector may include additional features that may confer one or more functions to the vector, the polynucleotide delivered, the viral particle produced therefrom, or the polypeptide expressed therefrom. Such features include, but are not limited to, regulatory elements, selectable markers, molecular identifiers (e.g., molecular barcodes), stabilizing elements, etc. Those skilled in the art will understand that the design of the expression vector and the additional features included may depend on factors such as the choice of host cell to be transformed, the desired expression level, etc.

[0145] Control Elements In some embodiments, the polynucleotides and / or vectors thereof described herein (including, but not limited to, engineered AAV capsid polynucleotides of the invention) may comprise one or more regulatory elements that may be operably linked to the polynucleotide. The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, e.g., polyadenylation signals, and poly-U sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many host cell types and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). A tissue-specific promoter may direct expression primarily in a desired tissue of interest, e.g., muscle, neurons, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). Regulatory elements may also direct expression in a time-dependent manner, e.g., cell cycle-dependent or developmental stage-dependent, which may or may not be tissue- or cell-type-specific. In some embodiments, the vector comprises one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or a combination thereof. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters.Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) (see, e.g., Boshart et al., Cell, 41:521-530 (1985)), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. Also encompassed within the term "regulatory element" are enhancer elements, such as the WPRE, the CMV enhancer, the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988), the SV40 enhancer, and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).

[0146] In some embodiments, the regulatory sequences may be those described in U.S. Patent No. 7,776,321, U.S. Patent Publication No. 2011 / 0027239, and PCT Publication No. WO2011 / 028929, the contents of which are incorporated herein by reference in their entireties. In some embodiments, the vector may include a minimal promoter. In some embodiments, the minimal promoter is a Mecp2 promoter, a tRNA promoter, or a U6 promoter. In further embodiments, the minimal promoter is tissue-specific. In some embodiments, the length of the minimal promoter and polynucleotide sequence of the vector polynucleotide is less than 4.4 Kb.

[0147] To express a polynucleotide, the vector may include one or more transcriptional and / or translational initiation control sequences, such as promoters, that direct the transcription of the gene and / or the translation of the encoded protein in a cell. In some embodiments, a constitutive promoter may be used. Constitutive promoters suitable for mammalian cells are generally known in the art and include, but are not limited to, SV40, CAG, CMV, EF-1α, β-actin, RSV, and PGK. Constitutive promoters suitable for bacterial, yeast, and fungal cells are generally known in the art, such as the T-7 promoter for bacterial expression and the alcohol dehydrogenase promoter for yeast expression.

[0148] In some embodiments, the control element may be a regulated promoter. A "regulated promoter" refers to a promoter that directs gene expression in a temporally and / or spatially regulated manner, rather than constitutively, and includes tissue-specific, tissue-preferred, and inducible promoters. In some embodiments, the regulated promoter is a tissue-specific promoter as previously discussed elsewhere herein. Regulated promoters include conditional promoters and inducible promoters. In some embodiments, a conditional promoter may be used to direct expression of a polynucleotide in a particular cell type under certain environmental conditions and / or during a particular state of development. Suitable tissue-specific promoters include liver-specific promoters (e.g., APOA2, SERPIN A1 (hAAT), CYP3A4, and MIR122), pancreatic cell promoters (e.g., INS, IRS2, Pdx1, Alx3, Ppy), heart-specific promoters (e.g., Myh6 (alpha MHC), MYL2 (MLC-2v), TNI3 (cTnl), NPPA (ANF), Slc8a1 (Ncx1)), central nervous system cell promoters (SYN1, GFAP, INA, NES, MOBP, MBP, TH, FOXA2 (HNF3 beta)), skin cell-specific promoters (e.g., FLG, K14, TGM3), immune cell-specific promoters Examples of suitable promoters include, but are not limited to, promoters specific to genitourinary tract cells (e.g., ITGAM, CD43 promoter, CD14 promoter, CD45 promoter, CD68 promoter), urogenital cell-specific promoters (e.g., Pbsn, Upk2, Sbp, Fer114), endothelial cell-specific promoters (e.g., ENG), pluripotent and embryonic germ layer cell-specific promoters (e.g., Oct4, NANOG, synthetic Oct4, T-Brachyury, NES, SOX17, FOXA2, MIR122), and muscle cell-specific promoters (e.g., desmin). Other tissue- and / or cell-specific promoters are discussed elsewhere herein, may be commonly known in the art, and are within the scope of the present disclosure.

[0149] Inducible / conditional promoters can be positive inducible / conditional promoters (e.g., promoters that activate transcription of a polynucleotide upon appropriate interaction with an activated activator or inducer (e.g., a compound, environmental condition, or other stimulus)), or negative / conditionally inducible promoters (e.g., promoters that are repressed (e.g., bound by a repressor) until the promoter's repressor condition is removed (e.g., an inducer binds to a repressor bound to the promoter, resulting in the release of the promoter by the repressor or a chemical repressor from the promoter environment). The inducer may be a compound, an environmental condition, or other stimulus. Thus, an inducible / conditional promoter may be responsive to any suitable stimulus, such as a chemical, biological, or other molecular substance, temperature, light, and / or pH. Suitable inducible / conditional promoters include, but are not limited to, Tet-On, Tet-Off, Lac promoter, pBad, AlcA, LexA, Hsp70 promoter, Hsp90 promoter, pDawn, XVE / OlexA, GVG, and pOp / LhGR.

[0150] When desired to express in plant cells, the components of the engineered AAV capsid system described herein are typically placed under the control of a plant promoter, i.e., a promoter that can operate in plant cells.The use of different types of promoters is envisioned.In some embodiments, the inclusion of engineered virus (e.g., AAV) capsid system vector in plants can be for the purpose of producing viral vectors.

[0151] A constitutive plant promoter is a promoter that can cause the open reading frame (ORF) it controls to be expressed (referred to as "constitutive expression") in all or nearly all plant tissues during all or nearly all developmental stages of the plant. One non-limiting example of a constitutive promoter is the cauliflower mosaic virus 35S promoter. Different promoters can direct the expression of genes in different tissues or cell types, or at different developmental stages, or in response to different environmental conditions. In certain embodiments, one or more components of the engineered AAV capsid system are expressed under the control of a constitutive promoter, such as the cauliflower mosaic virus 35S promoter, and a tissue-preferential promoter can be used to target enhanced expression in certain cell types within specific plant tissues, such as vascular cells in leaves or roots, or specific cells in seeds. Examples of specific promoters for use in the engineered AAV capsid systems and other compositions of the invention are found in Kawamata et al., (1997) Plant Cell Physiol 38:792-803, Yamamoto et al., (1997) Plant J 12:255-65, Hire et al., (1992) Plant Mol Biol 20:207-18, Kuster et al., (1995) Plant Mol Biol 29:759-72, and Capana et al., (1994) Plant Mol Biol 25:681-91.

[0152] Examples of promoters that are inducible and can enable gene editing or spatiotemporal control of gene expression include energy forms. Energy forms may include, but are not limited to, sound energy, electromagnetic radiation, chemical energy, and / or thermal energy. Examples of inducible systems include tetracycline-inducible promoters (Tet-On or Tet-Off), small molecule two-hybrid transcription activation systems (FKBP, ABA, etc.), or light-inducible systems (phytochrome, LOV domain, or cryptochrome), such as light-inducible transcription effectors (LITEs) that direct changes in transcription activity in a sequence-specific manner. Components of the light-inducible system may include one or more elements of the engineered AAV capsid system or other compositions of the present invention described herein, a light-responsive cytochrome heterodimer (e.g., from Arabidopsis thaliana), and a transcription activation / repression domain. In some embodiments, the vector can include one or more of the inducible DNA binding proteins provided in PCT Publication No. WO2014 / 018423 and U.S. Publication Nos. 2015 / 0291966, 2017 / 0166903, 2019 / 0203212, which describe, for example, embodiments of inducible DNA binding proteins and methods of use that can be adapted for use in the present invention.

[0153] In some embodiments, transient or inducible expression can be achieved, for example, by including a chemically regulated promoter, i.e., application of an exogenous chemical induces gene expression. Regulation of gene expression can also be achieved by including a chemically repressible promoter, in which case application of a chemical represses gene expression. Chemically inducible promoters include, but are not limited to, the maize ln2-2 promoter, which is activated by benzenesulfonamide herbicide safeners (De Veylder et al., (1997) Plant Cell Physiol 38:568-77), the maize GST promoter (GST-11-27, WO93 / 01294), which is activated by hydrophobic electrophilic compounds used as pre-emergence herbicides, and the tobacco PR-1a promoter, which is activated by salicylic acid (Ono et al., (2004) Biosci Biotechnol Biochem 68:803-7). Antibiotic-regulated promoters, such as tetracycline-inducible and tetracycline-repressible promoters (Gatz et al., (1991) Mol Gen Genet 227:229-37; U.S. Patent Nos. 5,814,618 and 5,789,156), can also be used herein.

[0154] In some embodiments, the vector or system may comprise one or more elements capable of translocating and / or expressing an engineered polynucleotide of the invention (e.g., a capsid polynucleotide of an engineered or variant virus (e.g., AAV)) to / in a particular cellular component or organelle, including, but not limited to, the nucleus, ribosomes, endoplasmic reticulum, Golgi apparatus, chloroplasts, mitochondria, vacuoles, lysosomes, cytoskeleton, cell membrane, cell wall, peroxisomes, centrioles, etc.

[0155] Selectable markers and tags One or more of the engineered polynucleotides of the invention (e.g., engineered or variant virus (e.g., AAV) capsid polynucleotides) can be operably linked to, fused to, or otherwise modified to include a polynucleotide encoding or being a selectable marker or tag, which can be a polynucleotide or a polypeptide. In some embodiments, a polypeptide encoding the polypeptide selectable marker can be incorporated into an engineered polynucleotide of the invention (e.g., engineered or variant virus (e.g., AAV) capsid polynucleotide) such that the selectable marker polypeptide, when translated, is inserted between two amino acids between the N-terminus and C-terminus of the engineered polypeptide (e.g., engineered AAV capsid polypeptide) or at the N-terminus and / or C-terminus of the engineered polypeptide (e.g., engineered AAV capsid polypeptide). In some embodiments, the selectable marker or tag is a polynucleotide barcode or unique molecular identifier (UMI).

[0156] It will be understood that polynucleotides encoding such selectable markers or tags can be incorporated into polynucleotides encoding one or more components of the engineered AAV capsid system described herein in an appropriate manner to allow for expression of the selectable marker or tag. Such techniques and methods are described elsewhere herein and will be readily apparent to those of skill in the art in view of this disclosure. Many such selectable markers and tags are generally known in the art and are intended to be within the scope of this disclosure.

[0157] Suitable selection markers and tags include affinity tags, e.g., chitin-binding protein (CBP), maltose-binding protein (MBP), glutathione-S-transferase (GST), poly(His) tags, solubilization tags, e.g., thioredoxin (TRX) and poly(NANP), MBP, and GST, chromatography tags, e.g., those composed of polyanionic amino acids, e.g., FLAG tags, epitope tags, e.g., V5 tags, Myc tags, HA tags, and NE tags, protein tags that may allow for specific enzymatic modification (e.g., biotinylation with biotin ligase) or chemical modification (e.g., reaction with FlAsH-EDT2 for fluorescent imaging), DNA and / or RNA segments containing restriction enzyme or other enzyme cleavage sites, and tags encoding products that provide resistance to other toxic compounds, including antibiotics such as spectinomycin, ampicillin, kanamycin, tetracycline, basta, neomycin phosphotransferase II (NEO), hygromycin phosphotransferase (HPT), etc. These include, but are not limited to, DNA segments, DNA and / or RNA segments encoding products otherwise lacking in the recipient cell (e.g., tRNA genes, auxotrophic markers), DNA and / or RNA segments encoding products that can be easily identified (e.g., phenotypic markers, e.g., β-galactosidase, GUS, fluorescent proteins, e.g., green fluorescent protein (GFP), cyan (CFP), yellow (YFP), red (RFP), luciferase, and cell surface proteins), polynucleotides capable of generating one or more new primer sites for PCR (e.g., juxtaposition of two DNA sequences that were not previously juxtaposed), DNA sequences that are insensitive or responsive to restriction endonucleases or other DNA-modifying enzymes, chemicals, etc., epitope tags (e.g., GFP, FLAG, and His tags), and DNA sequences that generate molecular barcodes or unique molecular identifiers (UMIs), DNA sequences required for specific modifications (e.g., methylation) that enable their identification. Other suitable markers will be apparent to those of skill in the art.

[0158] Selectable markers and tags may be operably linked to one or more components of the engineered AAV capsid system or other compositions and / or systems described herein via a suitable linker, such as a glycine or glycine-serine linker as short as GS or GG, up to (GGGGG)3 (SEQ ID NO: 12004) or (GGGGS)3 (SEQ ID NO: 12005). Other suitable linkers are described elsewhere herein.

[0159] The vector or vector system may comprise one or more polynucleotides encoding one or more targeting moieties. In some embodiments, polynucleotides encoding the targeting moieties may be included in the vector or vector system, e.g., a viral vector system, such that they are expressed within and / or on the produced viral particle(s) so that the viral particles can be targeted to specific cells, tissues, organs, etc. In some embodiments, polynucleotides encoding the targeting moieties may be included in the vector or vector system such that the engineered polynucleotide(s) of the invention (e.g., the capsid polynucleotide(s) of an engineered virus (e.g., AAV)) and / or products expressed therefrom comprise the targeting moiety and can be targeted to specific cells, tissues, organs, etc. In some embodiments, e.g., in non-viral carriers, the targeting moiety may be attached to the carrier (e.g., a polymer, lipid, inorganic particle, etc.), allowing targeting of the carrier and any attached or associated engineered polynucleotide(s) of the invention, engineered polypeptide(s) of the invention described herein, or other compositions to specific cells, tissues, organs, etc. In some embodiments, the particular cell is a hematopoietic cell.

[0160] Cell-free vectors and polynucleotide expression In some embodiments, polynucleotide(s) encoding the n-mer motif of the present invention can be expressed from a vector or suitable polynucleotide in a cell-free in vitro system. In some embodiments, polynucleotides encoding one or more features of the engineered AAV capsid system can be expressed from a vector or suitable polynucleotide in a cell-free in vitro system. In other words, the polynucleotide can be transcribed and optionally translated in vitro. In vitro transcription / translation systems and suitable vectors are generally known in the art and commercially available. Typically, in vitro transcription and in vitro translation systems reproduce the processes of RNA and protein synthesis, respectively, outside a cellular environment. Vectors and suitable polynucleotides for in vitro transcription can include promoter regulatory sequences that can be recognized and acted upon by T7, SP6, T3, or an appropriate polymerase to transcribe the polynucleotide or vector.

[0161] In vitro translation can be stand-alone (e.g., translation of purified polyribonucleotides) or linked / coupled to transcription. In some embodiments, the cell-free (or in vitro) translation system can include extracts from rabbit reticulocytes, wheat germ, and / or E. coli. The extracts can contain various macromolecular components required for translation of exogenous RNA (e.g., 70S or 80S ribosomes, tRNAs, aminoacyl-tRNAs, synthases, initiation, elongation, and termination factors, etc.). Other components can be included or added to the translation reaction, including, but not limited to, amino acids, energy sources (ATP, GTP), energy regeneration systems (creatine phosphate and creatine phosphokinase (eukaryotic systems)) (pyruvate phosphoenol and pyruvate kinase for bacterial systems), and other cofactors (Mg2+, K+, etc.). As previously mentioned, in vitro translation can be based on RNA or DNA starting materials. Some translation systems can utilize an RNA template as the starting material (e.g., reticulocyte lysate and wheat germ extract). Some translation systems can utilize a DNA template as the starting material (e.g., E. coli-based systems). In these systems, transcription and translation are coupled, with DNA first being transcribed into RNA, which is then translated. Suitable standard coupled cell-free translation systems are generally known in the art and are commercially available.

[0162] Codon optimization of vector polynucleotides As described elsewhere herein, polynucleotides encoding the n-mer motifs and / or other polynucleotides of the invention described herein may be codon-optimized. In some embodiments, polynucleotides of the engineered AAV capsid systems described herein may be codon-optimized. In some embodiments, in addition to any codon-optimized polynucleotides encoding n-mer motifs, one or more polynucleotides contained in the vectors described herein ("vector polynucleotides"), including but not limited to the engineered AAV capsid system embodiments described herein, may be codon-optimized. Generally, codon optimization refers to the process of modifying a nucleic acid sequence for improved expression in a host cell of interest by replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of the native sequence with a codon more frequently or most frequently used in the genes of that host cell, while maintaining the native amino acid sequence. Various species exhibit particular biases for certain codons for particular amino acids. Codon bias (differences in codon usage among organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, among other things, on the properties of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell typically reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database," available at kazusa.orjp / codon / , and these tables can be adapted in numerous ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000).Computer algorithms are also available for codon-optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen, Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in the coding sequence of a DNA / RNA-targeted Cas protein correspond to the most frequently used codon for a particular amino acid. For codon usage in yeast, see the online yeast genome database available at yeastgenome.org / community / codon_usage.shtml, or see "Codon selection in yeast," Bennetzen and Hall, J Biol Chem. 1982 Mar 25;257(6):3026-31. For information on codon usage in plants, including algae, see "Codon usage in higher plants, green algae, and cyanobacteria," Campbell and Gowri, Plant Physiol. 1990 Jan;92(1):1-11, and "Codon usage in plant genes," Murray et al., Nucleic Acids Res. 1989 Jan 25;17(2):477-98, or "Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages," Morton BR, J Mol Evol. 1998 Apr;46(4):449-59.

[0163] The vector polynucleotide may be codon-optimized for expression in a particular cell type, tissue type, organ type, and / or subject type. In some embodiments, the codon-optimized sequence is a sequence optimized for expression in a eukaryote, e.g., a human (i.e., optimized for expression in a human or human cells), or another eukaryote, such as another animal (e.g., a mammal or bird) described elsewhere herein. Such codon-optimized sequences are within the realm of one of skill in the art in view of the description herein. In some embodiments, the polynucleotide is codon-optimized for a particular cell type. Such cell types can include, but are not limited to, epithelial cells (including skin cells, cells lining the gastrointestinal tract, cells lining other hollow organs), neural cells (nerves, brain cells, spinal column cells, neural support cells (e.g., astrocytes, glial cells, Schwann cells, etc.), muscle cells (e.g., cardiac muscle cells, smooth muscle cells, and skeletal muscle cells), connective tissue cells (adipose and other soft tissue fill cells, bone cells, tendon cells, chondrocytes), blood cells, stromal cells (e.g., bone marrow stromal cells), stem and other progenitor cells, immune system cells, embryonic cells, and combinations thereof. Such codon-optimized sequences are within the purview of one of skill in the art in view of the description herein. In some embodiments, the polynucleotides In some embodiments, the polynucleotide is codon optimized for a particular tissue type. Such tissue types may include, but are not limited to, muscle tissue, connective tissue, nervous tissue, and epithelial tissue. Such codon-optimized sequences are within the realm of one of ordinary skill in the art in light of the description herein. In some embodiments, the polynucleotide is codon optimized for a particular organ. Such organs include, but are not limited to, muscle, skin, intestine, liver, spleen, brain, lung, stomach, heart, kidney, gallbladder, pancreas, bladder, thyroid, bone, blood vessels, blood, and combinations thereof. Such codon-optimized sequences are within the realm of one of ordinary skill in the art in light of the description herein.

[0164] In some embodiments, the vector polynucleotide is codon-optimized for expression in a particular cell, such as a prokaryotic cell or a eukaryotic cell, which may be of or derived from a particular organism, e.g., a plant or mammal, including, but not limited to, a human or a non-human eukaryote or animal or mammal discussed herein, e.g., a mouse, rat, rabbit, dog, livestock, or a non-human mammal or primate.

[0165] Non-viral vectors and carriers In some embodiments, the vector is a non-viral vector or carrier. In some embodiments, non-viral vectors may have the advantage(s) of reduced toxicity and / or immunogenicity and / or increased biosafety compared to viral vectors. The term "non-viral vector and carrier," as used herein in this context, refers to molecules and / or compositions that are not based on a virus or one or more components of a viral genome (excluding any nucleotides delivered and / or expressed by a non-viral vector) and that may be capable of binding to, incorporating into, linking to, and / or otherwise interacting with the engineered capsid polynucleotides of the invention described herein (e.g., engineered AAV capsid polynucleotides) or other compositions, and may be capable of delivering the polynucleotides to cells and / or expressing the polynucleotides. It will be understood that this does not exclude the inclusion of viral-based polynucleotides being delivered. For example, if the gRNA being delivered is directed against a viral component and is inserted into or otherwise linked to another non-viral vector or carrier, such a vector would not be considered a "viral vector." Non-viral vectors and carriers include naked polynucleotides, chemical-based carriers, polynucleotide (non-viral)-based vectors, and particle-based carriers. It will be understood that the term "vector," when used in the context of non-viral vectors and carriers, refers to a polynucleotide vector, and that "carrier," as used in this context, refers to a non-nucleic acid or polynucleotide molecule or composition that is bound to or otherwise interacts with the polynucleotide to be delivered, e.g., the capsid polynucleotide of an engineered or variant AAV of the invention.

[0166] Naked polynucleotides In some embodiments, one or more engineered AAV capsid polynucleotides or other polynucleotides of the invention described elsewhere herein can be included in a naked polynucleotide. As used herein, the term "naked polynucleotide" refers to a polynucleotide that is not associated with other molecules (e.g., proteins, lipids, and / or other molecules) that may often help protect it from environmental factors and / or degradation. As used herein, associated with includes, but is not limited to, linked to, attached to, adsorbed to, encapsulated in, encapsulated in or within, mixed with, etc. Naked polynucleotides comprising one or more engineered AAV capsid polynucleotides or other polynucleotides of the invention described herein can be delivered directly to a host cell and, optionally, expressed therein. The naked polynucleotide can have any suitable two-dimensional and three-dimensional structure. By way of non-limiting example, a naked polynucleotide can be a single-stranded molecule, a double-stranded molecule, a circular molecule (e.g., plasmids and artificial chromosomes), a molecule comprising a single-stranded portion and a double-stranded portion (e.g., ribozyme), etc. In some embodiments, the naked polynucleotide comprises only the engineered AAV capsid polynucleotide(s) or other polynucleotides of the invention. In some embodiments, the naked polynucleotide can comprise other nucleic acids and / or polynucleotides in addition to the engineered or variant AAV capsid polynucleotide(s) or other polynucleotides of the invention described elsewhere herein. The naked polynucleotide can comprise elements of one or more transposon systems. Transposons and their systems are described in more detail elsewhere herein.

[0167] Non-viral polynucleotide vectors In some embodiments, one or more of the engineered AAV capsid polynucleotides or other polynucleotides of the present invention can be contained in a non-viral polynucleotide vector. Suitable non-viral polynucleotide vectors include, but are not limited to, transposon vectors and vector systems, plasmids, bacterial artificial chromosomes, yeast artificial chromosomes, AR (antibiotic resistance)-free plasmids and mini-plasmids, circular covalently closed vectors (e.g., minicircles, minivectors, miniknots), linear covalently closed vectors ("dumbbell-shaped"), MIDGE (minimalistic immunologically defined gene expression) vectors, MiLV (micro-linear vector) vectors, ministrings, mini-intron plasmids, PSK systems (post-segregational killing systems), ORT (operator repressor titration) plasmids, and the like. See, e.g., Hardee et al. 2017. Genes. 8(2):65.

[0168] In some embodiments, the non-viral polynucleotide vector may have a conditional origin of replication. In some embodiments, the non-viral polynucleotide vector may be an ORT plasmid. In some embodiments, the non-viral polynucleotide vector may have minimal immunologically defined gene expression. In some embodiments, the non-viral polynucleotide vector may have one or more post-segregational killing system genes. In some embodiments, the non-viral polynucleotide vector is AR-free. In some embodiments, the non-viral polynucleotide vector is a minivector. In some embodiments, the non-viral polynucleotide vector comprises a nuclear localization signal. In some embodiments, the non-viral polynucleotide vector may comprise one or more CpG motifs. In some embodiments, the non-viral polynucleotide vector may comprise one or more scaffold / matrix attachment regions (S / MARs). See, e.g., Mirkovitch et al. 1984. Cell. 39:223-232, Wong et al. 2015. Adv. Genet. 89:113-152. This technology and vectors can be adapted for use in the present invention. S / MARs are AT-rich sequences that play a role in the spatial organization of chromosomes through the attachment of DNA loop bases to the nuclear matrix. S / MARs are often found near regulatory elements such as promoters, enhancers, and DNA replication origins. The inclusion of an S / MAR can facilitate replication once per cell cycle, maintaining the non-viral polynucleotide vector as an episome in daughter cells. In some embodiments, the S / MAR sequence is located downstream of an actively transcribed polynucleotide (e.g., one or more engineered AAV capsid polynucleotides of the present invention or other polynucleotides or molecules) contained in the non-viral polynucleotide vector. In some embodiments, the S / MAR can be an S / MAR from the beta-interferon gene cluster.See, e.g., Verghese et al. 2014. Nucleic Acid Res. 42:e53; Xu et al. 2016. Sci. China Life Sci. 59:1024-1033; Jin et al. 2016. 8:702-711; Koirala et al. 2014. Adv. Exp. Med. Biol. 801:703-709; and Nehlsen et al. 2006. Gene Ther. Mol. Biol. 10:233-244. This technology and vectors can be adapted for use in the present invention.

[0169] In some embodiments, the non-viral vector is a transposon vector or system thereof. As used herein, "transposon" (also referred to as a transposable element) refers to a polynucleotide sequence that can move from one location in a genome to another. There are several classes of transposons. Transposons include retrotransposons and DNA transposons. Retrotransposons require transcription of the moving (or transposing) polynucleotide in order for the polynucleotide to transpose into a new genome or polynucleotide. DNA transposons do not require reverse transcription of the moving (or transposing) polynucleotide in order for the polynucleotide to transpose into a new genome or polynucleotide. In some embodiments, the non-viral polynucleotide vector can be a retrotransposon vector. In some embodiments, the retrotransposon vector comprises long terminal repeats. In some embodiments, the retrotransposon vector does not comprise long terminal repeats. In some embodiments, the non-viral polynucleotide vector can be a DNA transposon vector. A DNA transposon vector can comprise a polynucleotide sequence encoding a transposase. In some embodiments, the transposon-based vector is configured as a non-autonomous transposon-based vector, meaning that transposition does not occur spontaneously by itself. In some of these embodiments, the transposon-based vector lacks one or more polynucleotide sequences encoding proteins required for transposition. In some embodiments, the non-autonomous transposon-based vector lacks one or more Ac elements.

[0170] In some embodiments, a non-viral polynucleotide transposon vector system can include a first polynucleotide vector comprising an engineered AAV capsid polynucleotide(s) of the invention or other polynucleotides or molecules described herein flanked at the 5' and 3' ends by transposon inverted terminal repeats (TIRs), and a second polynucleotide vector comprising a polynucleotide capable of encoding a transposase linked to a promoter for driving expression of the transposase. When both are expressed in the same cell, the transposase can be expressed from the second vector and can transpose material (e.g., an engineered AAV capsid polynucleotide(s) of the invention or other polynucleotides or molecules) between the TIRs on the first vector and integrate it into one or more locations in the genome of the host cell. In some embodiments, the transposon vector or system can be configured as a gene trap. In some embodiments, the TIR can be configured to flank a strong splice acceptor site followed by a reporter and / or other gene (e.g., one or more of the capsid polynucleotide(s) or other polynucleotides or molecules of an engineered or variant AAV of the invention) and a strong polyA tail. When transposition occurs using this vector or system, the transposon can insert into an intron of a gene, and the inserted reporter or other gene can induce a mis-splicing process, resulting in inactivation of the trapped gene.

[0171] Any suitable transposon system can be used. Suitable transposons and their systems can include the Sleeping Beauty transposon system (Tc1 / mariner superfamily) (see, for example, Ivics et al. 1997. Cell. 91(4):501-510), piggyBac (piggyBac superfamily) (see, for example, Li et al. 2013 110(25):E2279-E2287 and Yusa et al. 2011. PNAS. 108(4):1531-1536), Tol2 (hAT superfamily), Frog Prince (Tc1 / mariner superfamily) (see, for example, Miskey et al. 2003 Nucleic Acid Res. 31(23):6873-6881) and variants thereof.

[0172] Chemical carriers In some embodiments, the engineered AAV capsid polynucleotide(s) of the present invention or other polynucleotides or other molecules described herein can be linked to a chemical carrier. Chemical carriers that may be suitable for polynucleotide delivery can be broadly classified into the following classes: (i) inorganic particles, (ii) lipid-based, (iii) polymer-based, and (iv) peptide-based. They can be classified as: (1) capable of forming a condensation complex with a polynucleotide (such as an engineered AAV capsid polynucleotide(s) of the present invention), (2) capable of targeting specific cells, (3) capable of increasing delivery of a polynucleotide or other molecule (such as an engineered AAV capsid polynucleotide(s)) of the present invention to the nucleus or cytosol of a host cell, (4) capable of degrading DNA / RNA in the cytosol of a host cell, and (5) capable of sustained or controlled release. It will be understood that any given chemical carrier may include features from more than one class. As used herein, the term "particle" refers to a particle of any suitable size for delivery of the compositions of the invention described herein (including the particles, polypeptides, polynucleotides, and other compositions described herein). Suitable sizes include macro-, micro-, and nano-sized particles.

[0173] In some embodiments, the non-viral carrier can be an inorganic particle. In some embodiments, the inorganic particle can be a nanoparticle. The inorganic particle can be configured and optimized with various sizes, shapes, and / or porosities. In some embodiments, the inorganic particle is optimized to escape the reticuloendothelial system. In some embodiments, the inorganic particle can be optimized to protect encapsulated molecules from degradation. Suitable inorganic particles that can be used as non-viral carriers in this context include, but are not limited to, calcium phosphate, silica, metals (e.g., gold, platinum, silver, palladium, rhodium, osmium, iridium, ruthenium, mercury, copper, rhenium, titanium, niobium, tantalum, and combinations thereof), magnetic compounds, particles, and materials (e.g., supermagnetic iron oxide and magnetite), quantum dots, fullerenes (e.g., carbon nanoparticles, nanotubes, nanostrings, etc.), and combinations thereof. Other suitable inorganic non-viral carriers are discussed elsewhere herein.

[0174] In some embodiments, the non-viral carrier can be lipid-based. Suitable lipid-based carriers are also described in more detail herein. In some embodiments, the lipid-based carrier includes a cationic lipid or an amphipathic lipid capable of binding to or otherwise interacting with the negative charge of the polynucleotide to be delivered (e.g., the engineered AAV capsid polynucleotide(s) of the present invention). In some embodiments, the non-viral chemical carrier system can include a polynucleotide (e.g., the engineered AAV capsid polynucleotide(s) of the present invention or other compositions or other molecules) and a lipid (e.g., a cationic lipid). These are also referred to in the art as lipoplexes. Other embodiments of lipoplexes are described elsewhere herein. In some embodiments, the non-viral lipid-based carrier can be a lipid nanoemulsion. A lipid nanoemulsion can be formed by a stabilized emulsifier dispersion of immiscible liquids, having particles of approximately 200 nm, which consist of the lipid, water, and surfactant, and which can contain the polynucleotide to be delivered (e.g., the engineered AAV capsid polynucleotide(s) of the present invention). In some embodiments, the lipid-based non-viral carrier may be a solid lipid particle or nanoparticle.

[0175] In some embodiments, the non-viral carrier can be peptide-based. In some embodiments, the peptide-based non-viral carrier can include one or more cationic amino acids. In some embodiments, 35-40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 99, or 100% of the amino acids are cationic. In some embodiments, peptide carriers can be used in conjunction with other types of carriers (e.g., polymer-based carriers and lipid-based carriers that functionalize these carriers). In some embodiments, the functionalization is for host cell targeting. Suitable polymers that may be included in the polymer-based non-viral carriers include, but are not limited to, polyethyleneimine (PEI), chitosan, poly(DL-lactide) (PLA), poly(DL-lactide-co-glycolide) (PLGA), dendrimers (see, e.g., U.S. Patent Publication No. 2017 / 0079916, the technology and compositions of which can be adapted for use with the engineered AAV capsid polynucleotides of the invention), polymethacrylates, and combinations thereof.

[0176] In some embodiments, the non-viral carrier may be configured to release the polynucleotide of the engineered delivery system associated with or bound to the non-viral carrier in response to an external stimulus, such as pH, temperature, osmotic pressure, the concentration of a particular molecule or composition (e.g., calcium, NaCl, etc.), pressure, etc. In some embodiments, the non-viral carrier may be a particle configured to contain one or more of the engineered or variant AAV capsid polynucleotides or other compositions of the invention described herein and a response element for an environmental stimulus, and optionally the stimulus. In some embodiments, the particle may comprise a polymer selected from the group of polymethacrylate and polyacrylate. In some embodiments, the non-viral particle may comprise one or more embodiments of the microparticles of the compositions described in U.S. Patent Publication Nos. 2015 / 0232883 and 2005 / 0123596. These technologies and compositions may be adapted for use in the present invention.

[0177] In some embodiments, the non-viral carrier can be a polymer-based carrier. In some embodiments, the polymer is cationic or predominantly cationic and can interact with the negatively charged polynucleotide to be delivered (e.g., the capsid polynucleotide(s) of the engineered AAV of the present invention) in a charge-dependent manner. Polymer-based systems are described in more detail elsewhere herein.

[0178] viral vectors In some embodiments, the vector is a viral vector. As used herein in this context, the term "viral vector" refers to a polynucleotide-based vector that includes or is based on one or more elements of a virus that may be capable of expressing and packaging a polynucleotide, e.g., an engineered AAV capsid polynucleotide, cargo, or other composition or molecule, into a viral particle and producing such a viral particle when used alone or with one or more other viral vectors (e.g., in a viral vector system). Viral vectors and systems thereof may be used to produce viral particles for delivery and / or expression and / or production of one or more compositions of the invention described herein (including, but not limited to, any viral particles and associated cargo). The viral vector may be part of a viral vector system comprising multiple vectors. In some embodiments, systems incorporating multiple viral vectors can enhance the safety of these systems. Suitable viral vectors may include adenovirus-based vectors, adeno-associated vectors, helper-dependent adenovirus (HdAd) vectors, hybrid adenovirus vectors, and the like. Other embodiments of viral vectors and viral particles produced therefrom are described elsewhere herein, hi some embodiments, the viral vectors are configured to produce replication-incompetent viral particles to improve the safety of these systems.

[0179] Adenovirus vectors, helper-dependent adenovirus vectors, and hybrid adenovirus vectors In some embodiments, the vector may be an adenoviral vector. In some embodiments, the adenoviral vector may contain elements such that viral particles produced using the vector or system may be serotype 2, 5, or 9. In some embodiments, the polynucleotide delivered via the adenoviral particle may be up to about 8 kb. Thus, in some embodiments, the adenoviral vector may contain a delivered DNA polynucleotide that may range in size from about 0.001 kb to about 8 kb. Adenoviral vectors have been used successfully in several situations (see, e.g., Teramato et al. 2000. Lancet. 355:1911-1912; Lai et al. 2002. DNA Cell. Biol. 21:895-913; Flotte et al., 1996. Hum. Gene. Ther. 7:1145-1159; and Kay et al. 2000. Nat. Genet. 24:257-261). The engineered AAV capsid can be included in an adenoviral vector to produce adenoviral particles containing such engineered AAV capsid.

[0180] In some embodiments, the vector can be a helper-dependent adenoviral vector or system thereof. These are also referred to in the art as "gutless" or "gutted" vectors and are a modified generation of adenoviral vectors (see, e.g., Thrasher et al. 2006. Nature. 443:E5-7). In some embodiments of a helper-dependent adenoviral vector system, one vector (the helper) can contain all of the viral genes required for replication but contain a conditional gene deletion in the packaging domain. The second vector in the system can contain only the ends of the viral genome, one or more engineered AAV capsid polynucleotides, and the native packaging recognition signal, which can enable selective packaging release from cells (see, e.g., Cideciyan et al. 2009. N Engl J Med. 361:725-727). Helper-dependent adenoviral vector systems have been successful for gene delivery in some situations (see, e.g., Simonelli et al. 2010. J Am Soc Gene Ther. 18:643-650; Cideciyan et al. 2009. N Engl J Med. 361:725-727; Crane et al. 2012. Gene Ther. 19(4):443-452; Alba et al. 2005. Gene Ther. 12:18-S27; Croyle et al. 2005. Gene Ther. 12:579-587; Amalfitano et al. 1998. J. Virol. 72:926-933; and Morral et al. 1999. PNAS. 96:12816-12821). The techniques and vectors described in these publications can be adapted for the inclusion and delivery of the engineered AAV capsid polynucleotides described herein. In some embodiments, the polynucleotides delivered via helper-dependent adenoviral vectors or viral particles produced therefrom can be up to about 38 kb.Thus, in some embodiments, the adenoviral vector can comprise a delivered DNA polynucleotide that can range in size from about 0.001 kb to about 37 kb (see, e.g., Rosewell et al. 2011. J. Genet. Syndr. Gene Ther. Suppl. 5:001).

[0181] In some embodiments, the vector is a hybrid adenoviral vector or system thereof. Hybrid adenoviral vectors combine the high transduction efficiency of gene-deleted adenoviral vectors with the long-term genomic integration potential of adeno-associated, retroviral, lentiviral, and transposon-based gene transfer. In some embodiments, such hybrid vector systems can provide stable transduction and limited integration sites. See, e.g., Balague et al. 2000. Blood. 95:820-828; Morral et al. 1998. Hum. Gene Ther. 9:2709-2716; Kubo and Mitani. 2003. J. Virol. 77(5):2964-2971; Zhang et al. 2013. PloS One. 8(10)e76771; and Cooney et al. 2015. Mol. Ther. 23(4):667-674. The techniques and vectors described therein can be modified and adapted for use in the engineered AAV capsid systems of the invention. In some embodiments, the hybrid adenoviral vector can comprise one or more features of retrovirus and / or adeno-associated virus. In some embodiments, the hybrid adenoviral vector can comprise one or more features of spumaretrovirus or foamy virus (FV). See, e.g., Ehrhardt et al. 2007. Mol. Ther. 15:146-156 and Liu et al. 2007. Mol. Ther. 15:1834-1841. The techniques and vectors described therein can be modified and adapted for use in the engineered AAV capsid systems of the invention. Advantages of using one or more features from FV in the hybrid adenoviral vector or system include the ability of viral particles produced therefrom to infect a broad range of cells, a large packaging capacity compared to other retroviruses, and the ability to persist in dormant (non-dividing) cells.See also, for example, Ehrhardt et al. 2007. Mol. Ther. 156:146-156 and Shuji et al. 2011. Mol. Ther. 19:76-82. The techniques and vectors described therein can be modified and adapted for use in the engineered AAV capsid systems of the present invention.

[0182] Adeno-associated vector In one embodiment, the engineered vector or system thereof may be an adeno-associated vector (AAV). See, e.g., West et al., Virology 160:38-47 (1987), U.S. Patent No. 4,797,368, WO 93 / 24641, Kotin, Human Gene Therapy 5:793-801 (1994), and Muzyczka, J. Clin. Invest. 94:1351 (1994). AAVs are similar to adenovirus vectors in some of their characteristics, but have some defects in their replication and / or pathogenicity, and therefore may be safer than adenovirus vectors. In some embodiments, the AAV can integrate into a specific site on chromosome 19 of human cells without observable side effects. In some embodiments, the volume of the AAV vector, its system, and / or AAV particle may be up to about 4.7 kb. The AAV vector or system thereof may comprise one or more engineered capsid polynucleotides described herein.

[0183] The AAV vector or system thereof may comprise one or more regulatory molecules. In some embodiments, the regulatory molecules may be promoters, enhancers, repressors, etc., which are described in more detail elsewhere herein. In some embodiments, the AAV vector or system thereof may comprise one or more polynucleotides capable of encoding one or more regulatory proteins. In some embodiments, the one or more regulatory proteins may be selected from Rep78, Rep68, Rep52, Rep40, variants thereof, and combinations thereof. In some embodiments, the promoter may be a tissue-specific promoter as discussed above. In some embodiments, the tissue-specific promoter may drive expression of the capsid polynucleotide of the engineered capsid AAV described herein.

[0184] The AAV vector or system thereof can include one or more polynucleotides capable of encoding one or more capsid proteins, such as the engineered AAV capsid proteins described elsewhere herein. The engineered capsid proteins can be capable of assembling into the protein shell (engineered capsid) of an AAV viral particle. The engineered capsid can have cell-, tissue-, and / or organ-specific tropism.

[0185] In some embodiments, the AAV vector or system thereof may include one or more adenoviral helper factors or polynucleotides capable of encoding one or more adenoviral helper factors, including, but not limited to, E1A, E1B, E2A, E4ORF6, and VA RNA. In some embodiments, the producer host cell line expresses one or more of the adenoviral helper factors.

[0186] The AAV vector or system thereof can be configured to produce AAV particles having a specific serotype. In some embodiments, the serotype can be AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-8, AAV-9, or any combination thereof. In some embodiments, the AAV can be AAV1, AAV-2, AAV-5, AAV-9, or any combination thereof. AAV can be selected from AAVs related to the cells to be targeted. For example, AAV serotypes 1, 2, 5, 9, or hybrid capsids AAV-1, AAV-2, AAV-5, AAV-9, or any combination thereof can be selected for targeting brain and / or neural cells; AAV-4 can be selected for targeting cardiac tissue; and AAV-8 can be selected for delivery to the liver. Thus, in some embodiments, an AAV vector or system thereof capable of producing AAV particles capable of targeting the brain and / or neural cells can be configured to generate AAV particles having serotypes 1, 2, 5, or hybrid capsids AAV-1, AAV-2, AAV-5, or any combination thereof. In some embodiments, an AAV vector or system thereof capable of producing AAV particles capable of targeting cardiac tissue can be configured to generate AAV particles having the AAV-4 serotype. In some embodiments, an AAV vector or system thereof capable of producing AAV particles capable of targeting the liver can be configured to generate AAV particles having the AAV-8 serotype. See also Srivastava. 2017. Curr. Opin. Virol. 21:75-80.

[0187] It will be understood that, while different serotypes may provide some level of cell, tissue, and / or organ specificity, each serotype remains multitropic and may therefore result in tissue toxicity when used to target tissues that the serotype is less efficient at transducing. Thus, in addition to achieving some tissue targeting capability by selecting an AAV of a particular serotype, it will be understood that the tropism of that AAV serotype can be modified by the engineered AAV capsids described herein. As described elsewhere herein, wild-type AAV variants of any serotype can be generated via the methods described herein and determined to have a particular cell-specific tropism that may be the same as or different from that of a reference wild-type AAV serotype. In some embodiments, the cell, tissue, and / or specificity of the wild-type serotype can be enhanced (e.g., becoming more selective or specific for a particular cell type for which the serotype is already biased). For example, wild-type AAV-9 is biased toward muscle and brain in humans (see, e.g., Srivastava. 2017. Curr. Opin. Virol. 21:75-80). The inclusion of an engineered AAV capsid and / or capsid protein variant of wild-type AAV-9 described herein can, for example, reduce or eliminate the brain bias and / or increase muscle specificity such that brain specificity appears reduced in comparison, thereby improving muscle specificity compared to wild-type AAV-9. As previously discussed, the inclusion of an engineered capsid and / or capsid protein variant of a wild-type AAV serotype can have a different tropism than the wild-type reference AAV serotype. For example, an engineered AAV capsid and / or capsid protein variant of AAV-9 can have specificity for tissues other than muscle or brain in humans.

[0188] In some embodiments, the AAV vector is a hybrid AAV vector or a system thereof. Hybrid AAV is an AAV containing a genome with elements from one serotype packaged in a capsid derived from at least one different serotype. For example, if rAAV2 / 5 is produced, and the production method is based on the helper-free transient transfection method discussed above, the first and third plasmids (adeno-helper plasmids) will be the same as those described for rAAV2 production. However, the second plasmid, pRepCap, is different. In this plasmid, called pRep2 / Cap5, the Rep gene is still derived from AAV2, while the Cap gene can be derived from AAV5. The production scheme is the same as the approach described above for AAV2 production. The resulting rAAV is called rAAV2 / 5, and its genome is based on recombinant AAV2, while the capsid is based on AAV5. It is expected that the cell or tissue tropism exhibited by this AAV2 / 5 hybrid virus will be the same as that of AAV5. It will be appreciated that wild-type hybrid AAV particles suffer from the same specificity issues as the non-hybrid wild-type serotypes discussed above.

[0189] The advantages achieved by the wild-type-based hybrid AAV system can be combined with the improved and customized cell specificity that can be achieved with engineered AAV capsids by generating hybrid AAVs that can contain engineered AAV capsids described elsewhere herein.It will be understood that hybrid AAVs can contain engineered AAV capsids that contain genomes with elements from serotypes different from the reference wild-type serotype of which the engineered AAV capsid is a variant.For example, hybrid AAVs can be produced that contain engineered AAV capsids that are variants of the AAV-9 serotype, which are used to package genomes that contain components (e.g., rep elements) from the AAV-2 serotype.As with the wild-type-based hybrid AAVs discussed above, the tropism of the resulting AAV particles will be that of the engineered AAV capsid.

[0190] A tabulation of certain wild-type AAV serotypes for these cells can be found in Grimm, D. et al, J. Virol. 82:5887-5911 (2008), reproduced below as Table 1. Further tropism details can be found in Srivastava. 2017. Curr. Opin. Virol. 21:75-80, as previously discussed. [Table 1]

[0191] In some embodiments, the AAV vector or strain thereof is AAV rh.74 or AAV rh.10.

[0192] In some embodiments, the AAV vector or system thereof is configured as a "gutless" vector, similar to those described for retroviral vectors. In some embodiments, a "gutless" AAV vector or system thereof may have cis-acting viral DNA elements involved in genome amplification and packaging in conjunction with a heterologous sequence of interest (e.g., an engineered AAV capsid polynucleotide(s)).

[0193] Vector construction The vectors described herein can be constructed using any suitable process or technique. In some embodiments, one or more suitable recombination and / or cloning methods or techniques can be used to form the vector(s) described herein. Suitable recombination and / or cloning techniques and / or methods can include, but are not limited to, those described in U.S. Application Publication No. US2004 / 0171156 A1. Other suitable methods and techniques are described elsewhere herein.

[0194] The construction of recombinant AAV vectors has been described in numerous publications, including U.S. Patent No. 5,173,414, Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985), Tratschin et al., Mol. Cell. Biol. 4:2072-2081 (1984), Hermonat & Muzyczka, PNAS 81:6466-6470 (1984), and Samulski et al., J. Virol. 63:03822-3828 (1989). Any of these techniques and / or methods can be used and / or adapted to construct the AAV or other vectors described herein. AAV vectors are discussed elsewhere herein.

[0195] In some embodiments, the vectors can have one or more insertion sites, e.g., restriction endonuclease recognition sequences (also referred to as "cloning sites"). In some embodiments, one or more insertion sites (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors.

[0196] The delivery vehicles, vectors, particles, nanoparticles, formulations and components thereof for expression of one or more elements of the engineered AAV capsid system described herein are as used in the aforementioned documents, e.g., International Patent Application Publication No. WO2014 / 093622 (PCT / US2013 / 074667), and are discussed in more detail herein.

[0197] Production of viral particles from viral vectors Production of AAV particles There are two major strategies for producing AAV particles from AAV vectors and systems such as those described herein, depending on how adenoviral helper factors are provided (helper versus helper-free). In some embodiments, methods for producing AAV particles from AAV vectors and systems can involve adenoviral infection of a cell line stably carrying a polynucleotide encoding AAV replication and capsid, along with an AAV vector containing a polynucleotide to be packaged and delivered by the resulting AAV particles (e.g., engineered AAV capsid polynucleotide(s)). In some embodiments, methods for producing AAV particles from AAV vectors and systems can be a "helper-free" method, which involves co-transfecting a suitable producer cell line with three vectors (e.g., plasmid vectors): (1) an AAV vector containing a polynucleotide of interest (e.g., engineered AAV capsid polynucleotide(s)) between two ITRs, (2) a vector carrying a polynucleotide encoding AAV Rep-Cap, and (3) a helper polynucleotide. Those skilled in the art will appreciate the various methods and variations thereof, both helper and helper-free, and the different advantages of each system.

[0198] The engineered AAV vectors and systems described herein can be produced by any of these methods.

[0199] Delivery of vectors and viral particles The vectors described herein (including non-viral carriers) can be introduced into host cells to produce transcripts, proteins, or peptides (including fusion proteins or peptides encoded by the nucleic acids described herein) (e.g., engineered AAV capsid system transcripts, proteins, enzymes, mutant forms thereof, fusion proteins thereof, etc.), and viral particles (e.g., from viral vectors and systems thereof).

[0200] One or more engineered AAV capsid polynucleotides can be delivered using previously described adeno-associated viruses (AAV), adenoviruses, or other plasmid or viral vector types, particularly formulations and dosages from, for example, U.S. Patent Nos. 8,454,972 (formulations, dosages for adenoviruses), 8,404,658 (formulations, dosages for AAVs), and 5,846,946 (formulations, dosages for DNA plasmids), as well as clinical trials involving lentiviruses, AAVs, and adenoviruses and publications related to such clinical trials. For example, in the case of AAV, the route of administration, formulation, and dosage can be as described in U.S. Patent No. 8,454,972 and in clinical trials involving AAVs. In the case of adenoviruses, the route of administration, formulation, and dosage can be as described in U.S. Patent No. 8,404,658 and in clinical trials involving adenoviruses.

[0201] For plasmid delivery, the route of administration, formulation, and dosage may be as described in U.S. Patent No. 5,846,946 and in clinical studies involving plasmids. In some embodiments, dosages may be based on or extrapolated to an average 70 kg individual (e.g., a male adult human) and may be adjusted for patients, subjects, or mammals of different weights and species. The frequency of administration is within the purview of a medical or veterinary practitioner (e.g., a physician or veterinarian), depending on usual factors including the patient's or subject's age, sex, general health, other conditions, and the specific condition or symptom being addressed. The viral vector may be injected or otherwise delivered to the tissue or cell of interest.

[0202] From the perspective of in vivo delivery, AAV is advantageous over other viral vectors for several reasons, including its low toxicity (which may be due to a purification method that does not require ultracentrifugation of cellular particles that may activate an immune response) and its low potential to cause insertional mutagenesis due to its lack of integration into the host genome.

[0203] The vector(s) and viral particles described herein can be delivered to host cells in vitro, in vivo, and / or ex vivo. Delivery can be achieved by any suitable method, including, but not limited to, physical, chemical, and biological methods. Physical delivery methods use physical forces to weaken the membrane barrier of cells to facilitate intracellular delivery of the vector. Suitable physical methods include, but are not limited to, needles (e.g., injections), ballistic polynucleotides (e.g., particle bombardment, microprojectile gene transfer, and gene guns), electroporation, sonoporation, photoporation, magnetofection, hydroporation, and mechanical massage. Chemical methods use chemicals to induce changes in cell membrane permeability or other property(ies) to facilitate vector entry into cells. For example, environmental pH can be altered, which can cause a change in cell membrane permeability. Biological methods rely on and utilize the biological processes or properties of host cells to facilitate the transport of vectors (with or without a carrier) into cells. For example, the vector and / or its carrier may stimulate endocytosis or a similar process in the cell to facilitate uptake of the vector into the cell.

[0204] Delivery of engineered AAV capsid system components (e.g., engineered AAV capsids and / or polynucleotides encoding capsid proteins) to cells via particles. As used herein, the term "particle" refers to a particle of any appropriate size for delivery of the engineered AAV capsid system components described herein. Suitable sizes include macro-, micro-, and nano-sized particles. In some embodiments, any of the engineered AAV capsid system components (e.g., polypeptides, polynucleotides, vectors, and combinations thereof described herein) can be bound, linked, incorporated, or otherwise associated with one or more particles or components thereof described herein. The particles described herein can then be administered to a cell or organism by an appropriate route and / or technique. In some embodiments, particle delivery can be advantageous for delivery of the polynucleotide or vector components selected for delivery. It will be understood that in embodiments, particle delivery can also be advantageous for other engineered capsid system molecules and formulations described elsewhere herein.

[0205] Engineered or variant viral particles containing the capsid of an engineered or variant virus (e.g., AAV) Also described herein are engineered or variant virus particles (also referred to herein and elsewhere herein as "engineered viral particles" or "variant viral particles"), as described in detail elsewhere herein, which may comprise an engineered or variant viral capsid (e.g., an AAV capsid, referred to as an "engineered AAV particle" or "variant AAV particle"). It will be understood that the engineered AAV particle may be an adenovirus-based particle, a helper adenovirus-based particle, an AAV-based particle, or a hybrid adenovirus-based particle comprising at least one engineered AAV capsid protein, as previously described. An engineered AAV capsid is one that comprises one or more engineered AAV capsid proteins described elsewhere herein. In some embodiments, the engineered AAV particle may comprise between 1 and 60 of the engineered AAV capsid proteins described herein. In some embodiments, the engineered AAV particles may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 engineered capsid proteins. In some embodiments, the engineered AAV particles may comprise 0 to 59 wild-type AAV capsid proteins.In some embodiments, the engineered AAV particles may comprise 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, or 59 wild-type AAV capsid proteins. The engineered AAV particles may therefore comprise one or more of the n-mer motifs described above.

[0206] The engineered AAV particles can contain one or more cargo polynucleotides. Cargo polynucleotides are discussed in more detail elsewhere herein. Methods for producing the engineered AAV particles from viral and non-viral vectors are described elsewhere herein. Formulations containing the engineered viral particles are described elsewhere herein.

[0207] Cargo Polynucleotide Cargos are also described elsewhere herein. In some embodiments, the cargo is a cargo polynucleotide, which can be packaged into an engineered or variant viral particle and then delivered to a cell. In some embodiments, delivery is hematopoietic cell-specific. The engineered or variant viral (e.g., AAV) capsid polynucleotide, other viral (e.g., AAV) polynucleotide(s), and / or vector polynucleotide can comprise one or more cargo polynucleotides. In some embodiments, the one or more cargo polynucleotides can be operably linked to the engineered viral (e.g., AAV) capsid polynucleotide(s) and can become part of the genome of the engineered viral (e.g., AAV) of the viral (e.g., AAV) system of the present invention. The cargo polynucleotide can be packaged into an engineered viral (e.g., AAV) particle, which can be delivered, for example, to a cell. In some embodiments, the cargo polynucleotide can be capable of modifying a polynucleotide (e.g., a gene or transcript) of the cell to which it is delivered. As used herein, a "gene" can refer to a hereditary unit that occupies a specific location on a chromosome and corresponds to a sequence of DNA that contains the genetic instructions for a characteristic(s) or trait(s) of an organism. The term gene can refer to translated and / or untranslated regions of a genome. A "gene" can refer to a specific DNA sequence, which may be transcribed into an RNA transcript, which may be translated into a polypeptide, or a catalytic RNA molecule, including, but not limited to, tRNA, siRNA, piRNA, miRNA, long non-coding RNA, and shRNA. Modifications of polynucleotides, genes, transcripts, etc., include all genetic engineering techniques, including, but not limited to, gene editing and conventional recombinant genetic engineering techniques (e.g., total or partial gene insertion, deletion, and mutagenesis (e.g., total or partial gene insertion, deletion, and mutagenesis (e.g., insertion and deletion mutagenesis) techniques).

[0208] In some embodiments, the cargo molecule is a vaccine or a polynucleotide that can encode it. In some embodiments, the vaccine can stimulate an immune response against cancer.

[0209] Genetically modified cargo polynucleotides In some embodiments, the cargo molecule may be a polynucleotide or polypeptide that, when delivered alone or as part of a system, can act to modify the genome, epigenome, and / or transcriptome of the cell to which it is delivered, regardless of whether it is delivered with other components of the system. Such systems include, but are not limited to, CRISPR-Cas systems. Other genetic engineering systems, such as TALEN, zinc finger nucleases, Cre-Lox, morpholinos, etc., are other non-limiting examples of genetic engineering systems in which one or more components can be delivered by the engineered viral (e.g., AAV) particles described herein.

[0210] In some embodiments, the cargo molecule is a gene editing system or its components. In some embodiments, the cargo molecule is a CRISPR-Cas system molecule or its components. In some embodiments, the cargo molecule is a polynucleotide that encodes one or more components of a genetic modification system (for example, a CRISPR-Cas system). In some embodiments, the cargo molecule is a gRNA.

[0211] In some embodiments, the cargo molecule, when delivered alone or as part of a system, can act to modify the genome, epigenome, and / or transcriptome of a cell to which it is delivered so as to treat or prevent a hematological disease or disorder, or a symptom thereof, whether delivered with other components of the system. In some embodiments, the cargo molecule (e.g., beta-globin, secreted therapeutic proteins, Cas9, gRNA, etc.), whether delivered with other components of the system, can be used to treat or prevent a hematological disease or disorder, such as HIV / AIDs, hematological cancers (e.g., leukemia, lymphoma, myeloma, monoclonal gammopathy of undetermined significance (MGUS)), bleeding disorders (e.g., acquired platelet dysfunction, congenital platelet dysfunction, disseminated intravascular coagulation (DIC), prothrombin deficiency, factor V), or other hematological disorders (e.g., HIV / AIDS ... Hemoglobin deficiency, factor VII deficiency, factor X deficiency, factor XI deficiency (hemophilia C), Glanzmann's disease, hemophilia A, hemophilia B, idiopathic thrombocytopenic purpura (ITP), von Willebrand's disease (types I, II, and / or III)), hemoglobinopathies (e.g., sickle cell disease (HbS), sickle cell trait (HbAS), sickle cell hemoglobin C (HbSC), sickle cell thalassemia (HbS and HbA), thalassemia (alpha thalassemia and beta-thalassemia), hemoglobin C disease (HbCC), hemoglobin C trait (HbAC), primary immunodeficiencies (e.g., autoimmune lymphoproliferative syndrome (ALPS), APS-1 (APECED), BENTA disease, caspase 8 deficiency (CEDS), CARD9 deficiency and other candidiasis susceptibility syndromes, chronic granulomatous disease (CGD), common variable immunodeficiency (CVID), congenital neutropenic syndromes, CTLA4 deficiency, DOCK8 deficiency, GATA2 deficiency, immunodeficiency Glycosylation disorders with immune deficiency, hyperimmunoglobulin E syndrome (HIES), hyperimmunoglobulin M syndrome, interferon gamma deficiency, interleukin-12 deficiency, and interleukin-23 deficiency, leukocyte adhesion deficiency (LAD), LRBA deficiency, PI3 kinase disease, PLCG2-associated antibody deficiency and immune dysregulation (PLAID), severe combined immunodeficiency (SCID), STAT3 dominant-negative disease, STAT3 gain-of-function disease, Ward,and / or act to modify the genome, epigenome, and / or transcriptome of the cell to which it is delivered to treat or prevent blood diseases or disorders, including hypogammaglobulinemia, infections, myeloid cellular pool (WHIM) syndrome, Wiskott-Aldrich syndrome (WAS), x-linked agammaglobulinemia (XLA), x-linked lymphoproliferative disorders (XLP), XMEN diseases), cytopenias (anemia, leukopenia, thrombocytopenia, pancytopenia, autoimmune cytopenia, refractory cytopenia), and / or storage and metabolic disorders (e.g., diabetes, familial hypercholesterolemia, Hunter syndrome, Krabbe disease, maple syrup urine disease, metachromatic leukodystrophy, Niemann-Pick disease, Gaucher disease, hemochromatosis, phenylketonuria (PKU), mitochondrial disorders, porphyria, Tay-Sachs disease, Wilson disease).

[0212] Exon skipping In some embodiments, the nucleotide sequence may encode a nucleic acid capable of inducing exon skipping. Such an encoded nucleic acid may be an antisense oligonucleotide or an antisense nucleotide system. As used herein, the term "exon skipping" refers to the modification of pre-mRNA splicing by targeting splice donor and / or acceptor sites within the pre-mRNA with one or more complementary antisense oligonucleotides (AONs). By blocking spliceosome access to one or more splice donor or acceptor sites, AONs can prevent the splicing reaction, thereby causing the deletion of one or more exons from the fully processed mRNA. Exon skipping can be achieved in the nucleus during the pre-mRNA maturation process. In some examples, exon skipping can involve masking key sequences involved in the splicing of targeted exons by using antisense oligonucleotides (AONs) complementary to splice donor sequences within the pre-mRNA.

[0213] CRISPR-Cas-based cargo molecules In some embodiments, the engineered viruses (e.g., AAV) or other particles described herein can include one or more CRISPR-Cas system molecules, which can be polynucleotides or polypeptides. In some embodiments, the polynucleotides can encode one or more CRISPR-Cas system molecules. In some embodiments, the polynucleotides encode Cas proteins, CRISPR cascade proteins, gRNAs, or combinations thereof. Other CRISPR-Cas system molecules are discussed elsewhere herein and can be delivered either as polypeptides or polynucleotides.

[0214] Generally, CRISPR-Cas or CRISPR system, as used herein and in the literature, such as International Patent Application Publication No. WO2014 / 093622 (PCT / US2013 / 074667), refers collectively to the transcripts and other elements involved in the expression of or directing the activity of CRISPR-associated ("Cas") genes, including sequences encoding the Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (including "direct repeats" and tracrRNA-processed partial direct repeats in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), or "RNA(s)" as that term is used herein (e.g., RNA(s) to guide a Cas, such as Cas9, e.g., CRISPR CRISPR systems include RNA and transactivating (tracr) RNA or single guide RNA (sgRNA) (chimeric RNA)) or other sequences and transcripts from the CRISPR locus. Typically, CRISPR systems are characterized by elements that promote the formation of CRISPR complexes at the site of the target sequence (also referred to as a protospacer in the context of endogenous CRISPR systems). See, e.g., Shmakov et al. (2015) "Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems," Molecular Cell, DOI: dx.doi.org / 10.1016 / j.molcel.2015.10.008.

[0215] In certain embodiments, a protospacer adjacent motif (PAM) or PAM-like motif directs the binding of the effector protein complex disclosed herein to a target locus of interest. In some embodiments, the PAM may be a 5' PAM (i.e., located upstream of the 5' end of the protospacer). In other embodiments, the PAM may be a 3' PAM (i.e., located downstream of the 5' end of the protospacer). The term "PAM" may be used interchangeably with the terms "PFS" or "protospacer flanking site" or "protospacer flanking sequence."

[0216] In a preferred embodiment, the CRISPR effector protein can recognize a 3' PAM. In certain embodiments, the CRISPR effector protein can recognize a 3' PAM that is 5'H (H is A, C, or U).

[0217] In the context of CRISPR complex formation, "target sequence" refers to a sequence to which a guide sequence is designed to have complementarity, and hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. The target sequence may comprise an RNA polynucleotide. The term "target RNA" refers to an RNA polynucleotide that is or contains a target sequence. In other words, the target RNA may be a gRNA, i.e., an RNA polynucleotide or a portion of an RNA polynucleotide to which a portion of the guide sequence is designed to have complementarity and which directs the effector function mediated by a complex comprising a CRISPR effector protein and a gRNA. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell.

[0218] In certain exemplary embodiments, the CRISPR effector protein can be delivered using a nucleic acid molecule encoding the CRISPR effector protein.The nucleic acid molecule encoding the CRISPR effector protein can advantageously be a codon-optimized CRISPR effector protein.An example of a codon-optimized sequence is a sequence optimized for expression in eukaryotes, for example, humans (i.e., optimized for expression in humans), or a sequence optimized for another eukaryote, animal, or mammal discussed herein.See, for example, the SaCas9 human codon-optimized sequence in International Patent Application Publication No. WO2014 / 093622 (PCT / US2013 / 074667).Although this is preferred, it is understood that other examples are possible, and codon optimization for host species other than humans, or codon optimization for specific organs is known.In some embodiments, the enzyme-coding sequence encoding the CRISPR effector protein is codon-optimized for expression in specific cells, for example, eukaryotic cells. The eukaryotic cell may be of or derived from a specific organism, such as a plant or mammal, including, but not limited to, a human or a non-human eukaryote, animal, or mammal discussed herein, such as a mouse, rat, rabbit, dog, livestock, or non-human mammal or primate. In some embodiments, processes that modify the genetic identity of a human germline and / or processes that modify the genetic identity of an animal that may cause suffering to humans or animals without providing substantial medical benefit, as well as the animals resulting from such processes, may also be excluded. Generally, codon optimization refers to the process of modifying a nucleic acid sequence to improve expression in a host cell of interest by replacing at least one codon (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of the native sequence with a codon that is more frequently or most frequently used in the genes of that host cell, while maintaining the native amino acid sequence. Various species exhibit specific biases for certain codons of specific amino acids.Codon bias (differences in codon usage among organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, among other things, on the properties of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell typically reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in the "Codon Usage Database," available at kazusa.orjp / codon / , and these tables can be adapted in numerous ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, for example, in Gene Forge (Aptagen, Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in the coding sequence of a Cas correspond to the most frequently used codon for a particular amino acid.

[0219] In certain embodiments, the methods described herein may include providing a Cas transgenic cell in which one or more nucleic acids encoding one or more guide RNAs are provided or introduced into the cell in operably linked relationship to a regulatory element, including a promoter of one or more genes of interest. As used herein, the term "Cas transgenic cell" refers to a cell, e.g., a eukaryotic cell, into which a Cas gene has been genomically integrated. The nature, type, or origin of the cell is not particularly limited according to the present invention. Furthermore, the method by which the Cas transgene is introduced into a cell may vary and may be any method known in the art. In certain embodiments, the Cas transgenic cell is obtained by introducing the Cas transgene into an isolated cell. In certain other embodiments, the Cas transgenic cell is obtained by isolating cells from a Cas transgenic organism. By way of example and without limitation, the Cas transgenic cells referred to herein may be obtained from a Cas transgenic eukaryotic organism, e.g., a Cas knock-in eukaryotic organism. Reference is made to International Patent Application Publication No. WO2014 / 093622 (PCT / US13 / 74667), which is incorporated herein by reference. U.S. Patent Publication Nos. 20120017290 and 20110265198, assigned to Sangamo BioSciences, Inc., for methods of targeting the Rosa locus, can be modified to utilize the CRISPR-Cas system of the present invention. U.S. Patent Publication No. 20130236946, assigned to Cellectis, for methods of targeting the Rosa locus, can also be modified to utilize the CRISPR-Cas system of the present invention. By way of further example, reference is made to Platt et al. (Cell;159(2):440-455(2014)), which is incorporated herein by reference, describing Cas9 knock-in mice. The Cas transgene may further comprise a Lox-Stop-PolyA-Lox (LSL) cassette, making Cas expression inducible by Cre recombinase. Alternatively, the Cas transgenic cells may be obtained by introducing the Cas transgene into isolated cells.Delivery systems for transgenes are well known in the art. For example, the Cas transgene can be delivered to eukaryotic cells using vectors (e.g., AAV, adenovirus, lentivirus) and / or particle and / or nanoparticle delivery, as described elsewhere herein. Lentivirus and retrovirus systems, as well as non-viral systems, for delivering CRISPR-Cas system components are generally known in the art. AAV and adenovirus-based systems for CRISPR-Cas system components are generally known in the art and are described herein (e.g., the engineered AAV of the present invention).

[0220] Those skilled in the art will understand that the cells, e.g., Cas transgenic cells, as referred to herein, may further comprise genomic alterations in addition to having an integration of a Cas gene or in addition to mutations resulting from the sequence-specific action of Cas when complexed with an RNA capable of guiding Cas to a target locus.

[0221] In certain embodiments, the present invention includes vectors for, for example, delivering or introducing Cas and / or RNA capable of guiding Cas to a target locus (i.e., guide RNA) into a cell, but also for propagating these components (e.g., in a prokaryotic cell). This can be in addition to delivery of one or more CRISPR-Cas components or other genetic recombination system components not already delivered by the engineered AAV particles described herein. As used herein, a "vector" is a tool that allows or facilitates the movement of an entity from one environment to another. It is a replicon, such as a plasmid, phage, or cosmid, into which another DNA segment can be inserted to result in replication of the inserted segment. Typically, a vector is capable of replication when associated with appropriate control elements. In general, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, or no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other types of polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus (AAV)). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Certain vectors are capable of autonomous replication in host cells into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of a host cell upon introduction into the host cell, thereby replicating along with the host genome.Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." Common expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.

[0222] A recombinant expression vector can contain the nucleic acid of the present invention in a form suitable for the expression of the nucleic acid in a host cell. This means that the recombinant expression vector contains one or more control elements operably linked to the nucleic acid sequence to be expressed, which can be selected based on the host cell used for expression. Within the recombinant expression vector, "operably linked" is intended to mean that the nucleotide sequence of interest is linked to a control element(s) in a manner that allows the expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). For recombination and cloning methods, see U.S. Patent Application Publication No. 2004 / 0171156, the contents of which are incorporated herein by reference in their entirety. Therefore, the embodiments disclosed herein can also include transgenic cells containing the CRISPR effector system. In certain exemplary embodiments, the transgenic cells can function in separate, individual amounts. In other words, a sample containing a masking construct may be delivered to a cell, for example, in an appropriate delivery vesicle, and if the target is present in the delivery vesicle, the CRISPR effector is activated and a detectable signal is generated.

[0223] The vector(s) can include regulatory element(s), e.g., promoter(s). The vector(s) can include a coding sequence for Cas and / or can also include coding sequences for a single, and optionally at least 3, or 8, or 16, or 32, or 48, or 50 guide RNAs (e.g., sgRNAs), e.g., 1 to 2, 1 to 3, 1 to 4, 1 to 5, 3 to 6, 3 to 7, 3 to 8, 3 to 9, 3 to 10, 3 to 8, 3 to 16, 3 to 30, 3 to 32, 3 to 48, or 3 to 50 RNAs (e.g., sgRNAs). A single vector can contain a promoter for each RNA (e.g., sgRNA), advantageously up to about 16 RNAs, and if a single vector provides more than 16 RNAs, one or more promoters can drive the expression of the multiple RNAs; for example, if 32 RNAs are present, each promoter can drive the expression of two RNAs, and if 48 RNAs are present, each promoter can drive the expression of three RNAs. With simple calculations, known cloning protocols, and the teachings of this disclosure, one skilled in the art can easily implement the present invention for an appropriate exemplary vector, such as AAV, and an appropriate promoter, such as the U6 promoter, for the RNA(s). For example, the packaging limit for AAV is about 4.7 kb. The length of a single U6-gRNA (plus a restriction enzyme recognition site for cloning) is 361 bp. Therefore, one skilled in the art can easily fit about 12 to 16, e.g., 13, U6-gRNA cassettes into a single vector. This can be assembled using any suitable method, for example, the Golden Gate method used for TALE assembly (genome-engineering.org / taleffectors / ). Those skilled in the art can also use the tandem guide method to increase the number of U6-gRNAs by about 1.5-fold, for example, from 12 to 16, e.g., 13, to about 18 to 24, e.g., about 19 U6-gRNAs. Therefore, those skilled in the art can easily incorporate about 18 to 24, e.g., about 19 promoter-RNAs, e.g., U6-gRNAs, into a single vector, e.g., an AAV vector.Another way to increase the number of promoters and RNAs in a vector is to use a single promoter (e.g., U6) to express an array of RNAs separated by cleavable sequences. Yet another way to increase the number of promoter-RNAs in a vector is to express an array of promoter-RNAs separated by cleavable sequences contained in the coding sequence or intron of a gene; in this example, it is advantageous to use a polymerase II promoter, which can show improved expression and can also enable tissue-specific transcription of long RNAs (see, for example, nar.oxfordjournals.org / content / 34 / 7 / e53.short and nature.com / mt / journal / v16 / n9 / abs / mt2008144a.html). In an advantageous embodiment, AAV may package U6 tandem gRNAs targeting up to about 50 genes. Thus, from knowledge in the art and from the teachings in this disclosure, one of ordinary skill in the art will be readily able to make and use vector(s), e.g., a single vector, that expresses multiple RNAs or guides under the control of, or operably or functionally linked to, one or more promoters without undue experimentation, particularly with respect to the number of RNAs or guides discussed herein.

[0224] The coding sequence of the guide RNA(s) and / or the coding sequence of the Cas can be functionally or operably linked to a regulatory element(s), such that the regulatory element(s) drive expression. The promoter(s) can be constitutive promoter(s) and / or conditional promoter(s) and / or inducible promoter(s) and / or tissue-specific promoter(s). The promoter can be selected from the group consisting of RNA polymerase, pol I, pol II, pol III, T7, U6, H1, retroviral Rous sarcoma virus (RSV) LTR promoter, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, β-actin promoter, phosphoglycerol kinase (PGK) promoter, and EF1α promoter. A preferred promoter is U6.

[0225] Additional effectors for use in accordance with the invention can be identified by their proximity to the Cas1 gene, for example, but not limited to, within a region 20 kb from the start of the Cas1 gene and 20 kb from the end of the Cas1 gene. In certain embodiments, the effector protein comprises at least one HEPN domain and at least 500 amino acids, and the C2c2 effector protein is naturally present within 20 kb upstream or downstream of the Cas gene or CRISPR array in the prokaryotic genome. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas12a, Cas13a, Cas13b, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csa6, Csa7, Csa8, Csa9 (also known as Csn1 and Csx12), Csa9, Csa10, Csa12a, Csa13a, Csa13b, Csa1, Csa2, Csa3, Csa1, Csa2, Csa5, Csa6, Csa7, Csa8, Csa9 (also known as Csn1 and Csx12), ... Examples of C2c2 effector proteins include sn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified forms thereof. In certain exemplary embodiments, the C2c2 effector protein is naturally present within 20 kb upstream or downstream of the Cas1 gene in a prokaryotic genome. The terms "orthologue" (also referred to herein as "ortholog") and "homologue" (also referred to herein as "homolog") are well known in the art. With further guidance, as used herein, a "homolog" of a protein is a protein of the same species that performs the same or similar function as the protein to which it is a homolog. Homologous proteins may, but are not necessarily, structurally related, or are only partially structurally related. As used herein, an "ortholog" of a protein is a protein of a different species that performs the same or similar function as the protein to which it is an ortholog.Orthologous proteins may, but are not necessarily, structurally related, or are only partially structurally related.

[0226] In some embodiments, one or more elements of the nucleic acid targeting system are derived from a specific organism that contains an endogenous CRISPR RNA targeting system. In certain embodiments, the CRISPR RNA targeting system is found in Eubacterium and Ruminococcus. In certain embodiments, the effector protein contains targeted collateral ssRNA cleavage activity. In certain embodiments, the effector protein contains a dual HEPN domain. In certain embodiments, the effector protein lacks the corresponding Helical-1 domain of Cas13a. In certain embodiments, the effector protein is smaller than previously characterized Class 2 CRISPR effectors, with a median size of 928 aa. This median size is 190 aa (17%) smaller than that of Cas13c, more than 200 aa (18%) smaller than that of Cas13b, and more than 300 aa (26%) smaller than that of Cas13a. In certain embodiments, the effector protein does not require flanking sequences (e.g., PFS, PAM).

[0227] In certain embodiments, the locus structure of the effector protein comprises a WYL domain-containing accessory protein (designated after the three amino acids conserved in the first group of these domains identified; see, e.g., WYL domain IPR026881). In certain embodiments, the WYL domain accessory protein comprises at least one helix-turn-helix (HTH) or ribbon-helix-helix (RHH) DNA-binding protein. In certain embodiments, the WYL domain-containing accessory protein increases both the targeted and collateral ssRNA cleavage activity of the RNA-targeting effector protein. In certain embodiments, the WYL domain-containing accessory protein comprises an N-terminal RHH domain and a pattern of primarily hydrophobic conserved residues, including an invariant tyrosine-leucine doublet corresponding to the original WYL motif. In certain embodiments, the WYL domain-containing accessory protein is WYL1. WYL1 is a single WYL domain protein primarily associated with Ruminococcus.

[0228] In another exemplary embodiment, the type VI RNA-targeting Cas enzyme is Cas13d. In certain embodiments, Cas13d is Eubacterium siraeum DSM15702 (EsCas13d) or Ruminococcus sp. N15. MGS-57 (RspCas13d) (see, e.g., Yan et al., "Cas13d Is a Compact RNA-Targeting Type VI CRISPR Effector Positively Modulated by a WYL-Domain-Containing Accessory Protein," Molecular Cell (2018), doi.org / 10.1016 / j.molcel.2018.02.028). RspCas13d and EsCas13d do not require flanking sequences (e.g., PFS, PAM).

[0229] The methods, systems, and tools provided herein can be designed for use with Class 1 CRISPR proteins, which can be Type I, Type III, or Type IV Cas proteins, as described in Makarova et al., The CRISPR Journal, v. 1, n., 5 (2018); DOI: 10.1089 / crispr.2018.0033, particularly Figure 1 on p. 326, which is incorporated herein by reference in its entirety. Such Class 1 systems typically use multiprotein effector complexes, which in some embodiments can include auxiliary proteins, e.g., one or more proteins in a complex referred to as a CRISPR-associated complex for antiviral defense (cascade), one or more adaptation proteins (e.g., Cas1, Cas2, RNA nuclease), and / or one or more accessory proteins (e.g., Cas4, DNA nuclease), a CRISPR-associated Rossmann fold (CARF) domain-containing protein, and / or an RNA transcriptase. Although Class 1 systems share limited sequence similarity, proteins in Class 1 systems can be distinguished by their similar structure, including one or more subunits of the repeat-associated mysterious protein (RAMP) family, such as Cas5, Cas6, and Cas7. RAMP proteins are characterized by having one or more RNA recognition motif domains. The large subunit (e.g., Cas8 or Cas10) and small subunit (e.g., Cas11) are also unique to Class 1 systems. See, e.g., Figures 1 and 2. Koonin EV, Makarova KS. 2019 Origins and evolution of CRISPR-Cas systems. Phil. Trans. R. Soc. B 374:20180087, DOI: 10.1098 / rstb.2018.0087. In one embodiment, Class 1 systems are characterized by the signature protein Cas3. A cascade in a particular Class 1 protein may comprise a dedicated complex of multiple Cas proteins that bind to the pre-crRNA and recruit additional Cas proteins, e.g., Cas6 or Cas5, which are nucleases directly involved in processing the pre-crRNA.In one embodiment, type I CRISPR proteins comprise an effector complex comprising one or more Cas5 subunits and two or more Cas7 subunits. Class 1 subtypes include types IA, IB, IC, IU, ID, IE, and IF, types IV-A and IV-B, and types III-A, III-D, III-C, and III-B. Class 1 systems also include CRISPR-Cas variants, including type IA, IB, IE, IF, and IU variants, which can include transposon- and plasmid-borne variants, including versions of subtype IF encoded by the large family of Tn7-like transposons and a smaller group of Tn7-like transposons encoding similarly resolved subtype IB systems. See also Peters et al., PNAS 114(35)(2017); DOI:10.1073 / pnas.1709035114; Makarova et al, the CRISPR Journal, v.1, n5, Figure 5.

[0230] Cas molecule In some embodiments, the cargo molecule can be or include a Cas polypeptide and / or a polynucleotide capable of encoding a Cas polypeptide or a fragment thereof. Any Cas molecule can be a cargo molecule. In some embodiments, the cargo molecule is a Cas polypeptide of a Class I CRISPR-Cas system. In some embodiments, the cargo molecule is a Cas polypeptide of a Class II CRISPR-Cas system. In some embodiments, the Cas polypeptide is a Type I Cas polypeptide. In some embodiments, the Cas polypeptide is a Type II Cas polypeptide. In some embodiments, the Cas polypeptide is a Type III Cas polypeptide. In some embodiments, the Cas polypeptide is a Type IV Cas polypeptide. In some embodiments, the Cas polypeptide is a Type V Cas polypeptide. In some embodiments, the Cas polypeptide is a Type VI Cas polypeptide. In some embodiments, the Cas polypeptide is a Type VII Cas polypeptide. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Cas12, Cas12a, Cas13a, Cas13b, Cas13c, Cas13d, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, and Csc2. , Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified forms thereof.

[0231] Guide Sequence As used herein, the terms "guide sequence" and "guide molecule" in the context of the CRISPR-Cas system include any polynucleotide sequence that has sufficient complementarity with a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid targeting complex to the target nucleic acid sequence. Guide sequences generated using the methods disclosed herein can be full-length guide sequences, truncated guide sequences, full-length sgRNA sequences, truncated sgRNA sequences, or E+F sgRNA sequences. Each gRNA can be designed to contain multiple binding recognition sites (e.g., aptamers) specific for the same or different adapter proteins. Each gRNA can be designed to bind to the promoter region -1000 to +1 nucleic acid, preferably -200 nucleic acid, upstream of the transcription start site (i.e., TSS). This positioning allows for improved functional domains that affect gene activation (e.g., transcriptional activators) or gene inhibition (e.g., transcriptional repressors). The modified gRNA can be one or more modified gRNAs (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs) that target one or more target loci included in the composition. Such multiple gRNA sequences can be arranged in tandem, preferably separated by direct repeats.

[0232] In some embodiments, the degree of complementarity of the guide sequence to a given target sequence is about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more when optimally aligned using a suitable alignment algorithm. In certain exemplary embodiments, the guide molecule comprises a guide sequence that can be designed to have at least one mismatch with the target sequence so that an RNA duplex is formed between the guide sequence and the target sequence. Thus, the degree of complementarity is preferably less than 99%. For example, if the guide sequence consists of 24 nucleotides, the degree of complementarity is more specifically about 96% or less. In certain embodiments, the guide sequence is designed to have a stretch of two or more adjacent mismatched nucleotides, so that the degree of complementarity across the entire guide sequence is further reduced. For example, if the guide sequence consists of 24 nucleotides, the degree of complementarity is more particularly about 96% or less, more particularly about 92% or less, more particularly about 88% or less, more particularly about 84% or less, more particularly about 80% or less, more particularly about 76% or less, more particularly about 72% or less, depending on whether the stretch of two or more mismatched nucleotides includes 2, 3, 4, 5, 6, or 7 nucleotides, etc. In some embodiments, in addition to the stretch of one or more mismatched nucleotides, the degree of complementarity is about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more, or more, when optimally aligned using a suitable alignment algorithm.Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, available at novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of a guide sequence (within a nucleic acid-targeting guide RNA) to direct sequence-specific binding of a nucleic acid-targeting complex to a target nucleic acid sequence can be assessed by any suitable assay. For example, sufficient components of a nucleic acid-targeting CRISPR system to form a nucleic acid-targeting complex, including a test guide sequence, can be provided to a host cell having a corresponding target nucleic acid sequence, such as by transfection with a vector encoding the components of the nucleic acid-targeting complex, followed by evaluation of preferential targeting (e.g., cleavage) within the target nucleic acid sequence, such as by a Surveyor assay as described herein. Similarly, cleavage of a target nucleic acid sequence (or sequences before and after it) can be evaluated in a test tube by providing components of a nucleic acid-targeting complex, including the target nucleic acid sequence, the test guide sequence, and a control guide sequence different from the test guide sequence, and comparing the binding or cleavage rate at or before and after the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art. Guide sequences, and thus nucleic acid-targeting guide RNAs, can be selected to target any target nucleic acid sequence.

[0233] As used herein, the term "crRNA" or "guide RNA" or "single guide RNA" or "sgRNA" or "one or more nucleic acid components" of a Type V or Type VI CRISPR-Cas locus effector protein includes any polynucleotide sequence that has sufficient complementarity with a target nucleic acid sequence to hybridize with the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid targeting complex to the target nucleic acid sequence. In some embodiments, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more, or greater, when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, available at novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). The ability of a guide sequence (within a nucleic acid-targeting guide RNA) to direct sequence-specific binding of a nucleic acid-targeting complex to a target nucleic acid sequence can be assessed by any suitable assay. For example, components of a nucleic acid-targeted CRISPR system sufficient to form a nucleic acid-targeted complex, including the guide sequence to be tested, can be provided to a host cell having a corresponding target nucleic acid sequence, such as by transfection with a vector encoding the components of the nucleic acid-targeted complex, followed by assessment of preferential targeting (e.g., cleavage) within the target nucleic acid sequence, such as by a Surveyor assay described herein.Similarly, cleavage of a target nucleic acid sequence can be assessed in a test tube by providing components of a nucleic acid targeting complex, including the target nucleic acid sequence, a test guide sequence, and a control guide sequence different from the test guide sequence, and comparing the rate of binding or cleavage at the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art. The guide sequence, and thus the nucleic acid targeting guide, can be selected to target any target nucleic acid sequence. The target sequence can be DNA. The target sequence can be any RNA sequence. In some embodiments, the target sequence can be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double-stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (lncRNA), and small cytoplasmic RNA (scRNA). In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of mRNA, pre-mRNA, and rRNA. In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of ncRNA and lncRNA. In some more preferred embodiments, the target sequence may be a sequence within an mRNA molecule or a pre-mRNA molecule.

[0234] In some embodiments, nucleic acid targeting guides are selected to reduce the degree of secondary structure within the nucleic acid targeting guide. In some embodiments, about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1% or less of the nucleotides of the nucleic acid targeting guide are involved in self-complementary base pairing when optimally folded. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of minimum Gibbs free energy. An example of such an algorithm is mFold, described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold developed at the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., A.R. Gruber et al., 2008, Cell 106(1):23-24, and P.A. Carr and G.M. Church, 2009, Nature Biotechnology 27(12):1151-62). In certain embodiments, a guide RNA or crRNA can comprise, consist essentially of, or consist of a direct repeat (DR) sequence and a guide sequence or spacer sequence. In certain embodiments, a guide RNA or crRNA can comprise, consist essentially of, or consist of a direct repeat sequence fused or linked to a guide sequence or spacer sequence. In certain embodiments, the direct repeat sequence can be located upstream (i.e., 5') from the guide sequence or spacer sequence. In other embodiments, the direct repeat sequence can be located downstream (i.e., 3') from the guide sequence or spacer sequence.

[0235] In certain embodiments, the crRNA comprises a stem-loop, preferably a single stem-loop. In certain embodiments, the direct repeat sequence forms a stem-loop, preferably a single stem-loop.

[0236] In certain embodiments, the length of the spacer of the guide RNA is 15 to 35 nt. In certain embodiments, the length of the spacer of the guide RNA is at least 15 nucleotides. In certain embodiments, the length of the spacer is 15 to 17 nt, for example, 15, 16, or 17 nt, 17 to 20 nt, for example, 17, 18, 19, or 20 nt, 20 to 24 nt, for example, 20, 21, 22, 23, or 24 nt, 23 to 25 nt, for example, 23, 24, or 25 nt, 24 to 27 nt, for example, 24, 25, 26, or 27 nt, 27 to 30 nt, for example, 27, 28, 29, or 30 nt, 30 to 35 nt, for example, 30, 31, 32, 33, 34, or 35 nt, or 35 nt or more.

[0237] The term "tracrRNA" or similar term includes any polynucleotide sequence that has sufficient complementarity with the crRNA sequence to which it hybridizes. In some embodiments, the degree of complementarity between the tracrRNA sequence and the crRNA sequence along the shorter length of the two, when optimally aligned, is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more. In some embodiments, the tracr sequence is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and the crRNA sequence are contained within a single transcript such that hybridization between the two produces a transcript with secondary structure, such as a hairpin. In an embodiment of the invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins. In a preferred embodiment, the transcript has two, three, four, or five hairpins. In a further embodiment of the invention, the transcript has up to five hairpins. In the hairpin structure, the portion of the sequence 5' of the last "N" and upstream of the loop corresponds to the tracr mate sequence, and the portion of the sequence 3' of the loop corresponds to the tracr sequence.

[0238] Generally, the degree of complementarity is measured with respect to optimal alignment of the sca and tracr sequences along the shorter length of the two sequences. Optimal alignment can be determined by any suitable alignment algorithm, which may further take into account secondary structure, such as self-complementarity, in either the sca or tracr sequences. In some embodiments, the degree of complementarity between the tracr and sca sequences along the shorter length of the two sequences when optimally aligned is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99% or more.

[0239] Generally, the CRISPR-Cas, CRISPR-Cas9, or CRISPR system may be as used in the aforementioned documents, e.g., International Patent Application Publication No. WO2014 / 093622 (PCT / US2013 / 074667), and collectively refer to transcripts and other elements involved in expressing or directing the activity of CRISPR-associated ("Cas") genes, including sequences encoding a Cas gene, particularly the Cas9 gene in the case of CRISPR-Cas9, a tracr (trans-activating CRISPR) sequence (e.g., a tracrRNA or an active partial tracrRNA), a tracr mate sequence (including "direct repeats" and tracrRNA-processed partial direct repeats in the context of an endogenous CRISPR system), a guide sequence (also referred to as a "spacer" in the context of an endogenous CRISPR system), or "RNA(s)" as that term is used herein (e.g., an RNA(s) that guides Cas9, e.g., a CRISPR These include RNA and transactivating (tracr) RNA or single guide RNA (sgRNA) (chimeric RNA), or other sequences and transcripts from the CRISPR locus. Typically, CRISPR systems feature elements that promote the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of endogenous CRISPR systems). In the context of CRISPR complex formation, a "target sequence" refers to a sequence to which a guide sequence is designed to have complementarity, and hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. The section of a guide sequence whose complementarity to the target sequence is important for cleavage activity is referred to herein as a seed sequence. A target sequence can include any polynucleotide, e.g., a DNA or RNA polynucleotide. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell, and can include nucleic acids within or derived from mitochondria, organelles, vesicles, liposomes, or particles present within the cell. In some embodiments, NLSs are not preferred, particularly for non-nuclear use. In some embodiments, the CRISPR system includes one or more nuclear export signals (NESs).In some embodiments, the CRISPR system comprises one or more NLSs and one or more NESs. In some embodiments, direct repeats can be identified in silico by searching for repetitive motifs that meet any or all of the following criteria: 1. found in a 2 Kb window of genomic sequence flanking the Type II CRISPR locus; 2. spanning 20-50 bp; and 3. spaced 20-50 bp apart. In some embodiments, two of these criteria can be used, e.g., 1 and 2, 2 and 3, or 1 and 3. In some embodiments, all three criteria can be used.

[0240] In embodiments of the present invention, the terms guide sequence and guide RNA, i.e., RNA capable of guiding Cas to a target genomic locus, are used interchangeably as described in the aforementioned documents, for example, International Patent Application Publication No. WO2014 / 093622 (PCT / US2013 / 074667). Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between the guide sequence and its corresponding target sequence is about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99% or more, or more, when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler transformation (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, available at www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, the guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. Preferably, the guide sequence is 10-30 nucleotides in length.The ability of a guide sequence to direct the sequence-specific binding of a CRISPR complex to a target sequence can be evaluated by any suitable assay. For example, the components of a CRISPR system sufficient to form a CRISPR complex, including the guide sequence to be tested, can be provided to a host cell having the corresponding target sequence, such as by transfection with a vector encoding the components of the CRISPR sequence, and then the preferential cleavage within the target sequence can be evaluated by, for example, the Surveyor assay described herein. Similarly, the cleavage of a target polynucleotide sequence can be evaluated in a test tube by providing the components of a CRISPR complex, including the target sequence, the guide sequence to be tested, and a control guide sequence that is different from the test guide sequence, and comparing the binding or cleavage rate at the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art.

[0241] In some embodiments of a CRISPR-Cas system, the degree of complementarity between the guide sequence and its corresponding target sequence can be about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or 100% or greater, and the guide or RNA or sgRNA can be about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110 , 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides, or the guide or RNA or sgRNA may be less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length, advantageously the tracrRNA is 30 or 50 nucleotides in length. However, embodiments of the invention are directed to reducing off-target interactions, for example, reducing guide interactions with target sequences that have low complementarity. Indeed, the examples demonstrate that the invention includes mutations that result in a CRISPR-Cas system that can distinguish between target and off-target sequences having greater than 80% to about 95% complementarity, e.g., 83% to 84%, or 88 to 89%, or 94 to 95% complementarity (e.g., distinguishing between a target having 18 nucleotides and an off-target having 1, 2, or 3 mismatches). Thus, in the context of the present invention, the degree of complementarity between a guide sequence and its corresponding target sequence is greater than 94.5%, 95%, 95.5%, 96%, 96.5%, 97%, 97.5%, 98%, 98.5%, 99%, 99.5%, or 99.9%, or 100%.An off-target is less than 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% or 94% or 93% or 92% or 91% or 90% or 89% or 88% or 87% or 86% or 85% or 84% or 83% or 82% or 81% or 80% complementarity between the sequence and the guide, and advantageously an off-target is less than 100% or 99.9% or 99.5% or 99% or 99% or 98.5% or 98% or 97.5% or 97% or 96.5% or 96% or 95.5% or 95% or 94.5% complementarity between the sequence and the guide.

[0242] In a particularly preferred embodiment according to the present invention, the guide RNA (capable of guiding Cas to a target locus) can include (1) a guide sequence capable of hybridizing to a genomic target locus in a eukaryotic cell, (2) a tracr sequence, and (3) a tracr mate sequence. All of (1) through (3) can be present in a single RNA, i.e., the sgRNA (oriented 5' to 3'), or the tracrRNA can be a separate RNA from the RNA containing the guide and tracr sequences. The tracr hybridizes to the tracr mate sequence and directs the CRISPR / Cas complex to the target sequence. When the tracrRNA is present in a separate RNA from the RNA containing the guide and tracr sequences, the length of each RNA can be optimized to be shorter than its native length, and each can be independently chemically modified to protect against degradation by cellular RNases or otherwise increase stability.

[0243] The methods of the present invention described herein include inducing one or more mutations in a eukaryotic cell (in vitro, i.e., in an isolated eukaryotic cell) as discussed herein, which comprises delivering a vector as discussed herein to the cell. The mutation(s) can include the introduction, deletion, or substitution of one or more nucleotides at each target sequence in the cell(s) via guide(s) RNA(s) or sgRNA(s). The mutation(s) can include the introduction, deletion, or substitution of 1 to 75 nucleotides at each target sequence in such cell(s) via guide(s) RNA(s) or sgRNA(s). The mutations can include the introduction, deletion, or substitution of 1, 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides at each target sequence in such cell(s) via the guide(s) RNA(s) or sgRNA(s). The mutations can include the introduction, deletion, or substitution of 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides at each target sequence in such cell(s) via the guide(s) RNA(s) or sgRNA(s). The mutations include the introduction, deletion, or substitution of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides at each target sequence of the cell(s) via the guide(s) RNA(s) or sgRNA(s). The mutations can include the introduction, deletion, or substitution of 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides at each target sequence of the cell(s) via the guide(s) RNA(s) or sgRNA(s).The mutations can include the introduction, deletion, or substitution of 40, 45, 50, 75, 100, 200, 300, 400, or 500 nucleotides at each target sequence in such cell(s) via guide(s) RNA(s) or sgRNA(s).

[0244] To minimize toxicity and off-target effects, it may be important to control the concentration of delivered Cas mRNA and guide RNA. The optimal concentration of Cas mRNA and guide RNA can be determined by testing different concentrations in cells or non-human eukaryotic animal models and analyzing the degree of modification at potential off-target genomic loci using deep sequencing. Alternatively, to minimize the level of toxicity and off-target effects, Cas nickase mRNA (e.g., S. pyogenes Cas9 with a D10A mutation) can be delivered together with a pair of guide RNAs targeting the desired site. Guide sequences and strategies that minimize toxicity and off-target effects can be those in WO2014 / 093622 (PCT / US2013 / 074667) or via mutations described herein.

[0245] Typically, in the context of an endogenous CRISPR system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands at or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs therefrom). Without wishing to be bound by theory, a tracr sequence that may comprise or consist of all or a portion of a wild-type tracr sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracr sequence) may also form part of a CRISPR complex, such as by hybridization along at least a portion of the tracr sequence to all or a portion of a tracr mate sequence operably linked to a guide sequence.

[0246] In certain embodiments, a guide of the invention comprises a non-naturally occurring nucleic acid and / or a non-naturally occurring nucleotide and / or a nucleotide analog, and / or a chemical modification. Non-naturally occurring nucleic acids can include, for example, a mixture of naturally occurring and non-naturally occurring nucleotides. The non-naturally occurring nucleotide and / or nucleotide analog can be modified at the ribose, phosphate, and / or base moiety. In embodiments of the invention, a guide nucleic acid comprises ribonucleotides and non-ribonucleotides. In one such embodiment, a guide comprises one or more ribonucleotides and one or more deoxyribonucleotides. In embodiments of the invention, the guide comprises one or more non-naturally occurring nucleotides or nucleotide analogs, such as a nucleotide having a phosphorothioate linkage, a boranophosphate linkage, a locked nucleic acid (LNA) nucleotide containing a methylene bridge between the 2' and 4' carbons of the ribose ring, a peptide nucleic acid (PNA), or a bridged nucleic acid (BNA). Other examples of modified nucleotides include 2'-O-methyl analogs, 2'-deoxy analogs, 2-thiouridine analogs, N6-methyladenosine analogs, or 2'-fluoro analogs. Further examples of modified nucleotides include the attachment of a chemical moiety at the 2' position, including, but not limited to, peptides, nuclear localization sequences (NLS), peptide nucleic acids (PNAs), polyethylene glycol (PEG), triethylene glycol, or tetraethylene glycol (TEG). Further examples of modified bases include 2-aminopurine, 5-bromo-uridine, pseudouridine (Ψ), N 1 -Methylpseudouridine (me 1Examples of chemical modifications of guide RNAs include, but are not limited to, 2'-O-methyl (M), 2'-O-methyl-3'-phosphorothioate (MS), phosphorothioate (PS), S-constrained ethyl (cEt), 2'-O-methyl-3'-thioPACE (MSP), or 2'-O-methyl-3'-phosphonoacetate (MP) at one or more terminal nucleotides. Such chemically modified guides may include improved stability and activity compared to unmodified guides, but on-target versus off-target specificity is not predicted (Hendel, 2015, Nat Biotechnol. 33(9):985-9, doi:10.1038 / nbt.3290, published online June 29, 2015; Ragdarm et al., 0215, PNAS, E7110-E7111; Allerson et al., J. Med. Chem. 2005, 48:901-904; Bramsen et al., Front. Genet., 2012, 3:154; Deng et al., PNAS, 2015, 112:11870-11875; Sharma et al., MedChemComm., 2014, 5:1454-1471; Hendel et al. (See, e.g., Li et al., Nat. Biotechnol. (2015) 33(9):985-989; Li et al., Nature Biomedical Engineering, 2017, 1,0066 DOI:10.1038 / s41551-017-0066; Ryan et al., Nucleic Acids Res. (2018) 46(2):792-803). In some embodiments, the 5' and / or 3' ends of the guide RNA are modified with various functional groups, including fluorescent dyes, polyethylene glycol, cholesterol, proteins, or detection tags (see, e.g., Kelly et al., 2016, J. Biotech. 233:74-83). In certain embodiments, the guide comprises ribonucleotides in the region that binds to the target DNA and one or more deoxyribonucleotides and / or nucleotide analogs in the region that binds to Cas9, Cpf1, or C2c1.In embodiments of the present invention, deoxyribonucleotides and / or nucleotide analogs are incorporated into engineered guide structures, including, but not limited to, the 5' and / or 3' ends, stem-loop regions, and seed regions. In certain embodiments, the modifications are absent from the 5' handle of the stem-loop region. Chemical modifications at the 5' handle of the stem-loop region of a guide can abolish its function (see Li, et al., Nature Biomedical Engineering, 2017, 1:0066). In certain embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, or 75 nucleotides of the guide are chemically modified. In some embodiments, 3 to 5 nucleotides are chemically modified at either the 3' or 5' end of the guide. In some embodiments, only minor modifications, such as 2'-F modifications, are introduced in the seed region. In some embodiments, 2'-F modifications are introduced at the 3' end of the guide. In certain embodiments, 3 to 5 nucleotides at the 5' and / or 3' end of the guide are chemically modified with 2'-O-methyl (M), 2'-O-methyl-3'-phosphorothioate (MS), S-constrained ethyl (cEt), 2'-O-methyl-3'-thioPACE (MSP), or 2'-O-methyl-3'-phosphonoacetate (MP). Such modifications can improve genome editing efficiency (see Hendel et al., Nat. Biotechnol. (2015) 33(9):985-989; Ryan et al., Nucleic Acids Res. (2018) 46(2):792-803). In certain embodiments, all of the phosphodiester bonds of the guide are replaced with phosphorothioate (PS) to enhance the level of gene disruption, and in certain embodiments, more than five nucleotides at the 5' and / or 3' ends of the guide are chemically modified with 2'-O-Me, 2'-F, or S-constrained ethyl (cEt).Such chemically modified guides can mediate high levels of gene disruption (see Ragdarm et al., 2015, PNAS, E7110-E7111). In embodiments of the present invention, guides are modified to contain chemical moieties at their 3' and / or 5' ends. Such moieties include, but are not limited to, amines, azides, alkynes, thiols, dibenzocyclooctynes (DBCO), rhodamines, peptides, nuclear localization sequences (NLSs), peptide nucleic acids (PNAs), polyethylene glycols (PEGs), triethylene glycols, or tetraethylene glycols (TEGs). In certain embodiments, the chemical moieties are conjugated to the guides via linkers, e.g., alkyl chains. In certain embodiments, the chemical moieties of the modified guides can be used to attach the guides to other molecules, e.g., DNA, RNA, proteins, or nanoparticles. Such chemically modified guides can be used to identify or enrich cells commonly edited by CRISPR systems (see Lee et al., eLife, 2017, 6:e25312, DOI:10.7554). In some embodiments, three nucleotides at each of the 3' and 5' ends are chemically modified. In particular embodiments, the modifications include 2'-O-methyl or phosphorothioate analogs. In particular embodiments, 12 nucleotides of the tetraloop and 16 nucleotides of the stem-loop region are substituted with 2'-O-methyl analogs. Such chemical modifications improve in vivo editing and stability (see Finn et al., Cell Reports (2018), 22:2227-2235). In some embodiments, more than 60 or more than 70 nucleotides of the guide are chemically modified. In some embodiments, the modifications include substitution of nucleotides with 2'-O-methyl or 2'-fluoro nucleotide analogs or phosphorothioate (PS) modifications of phosphodiester linkages.In some embodiments, the chemical modification comprises 2'-O-methyl or 2'-fluoro modification of the guide nucleotides that extend outside the nuclease protein when the CRISPR complex is formed, or PS modification of 20 to 30 or more nucleotides at the 3' end of the guide. In certain embodiments, the chemical modification further comprises 2'-O-methyl analogs at the 5' end of the guide or 2'-fluoro analogs in the seed and tail regions. Such chemical modifications improve stability against nuclease degradation and maintain or enhance genome editing activity or efficiency, but modifying all nucleotides may abolish the function of the guide (see Yin et al., Nat. Biotech. (2018), 35(12):1179-1187). Such chemical modifications can be guided by knowledge of the structure of the CRISPR complex, including knowledge of the interactions of a limited number of nucleases with the 2'-OH of RNA (see Yin et al., Nat. Biotech. (2018), 35(12):1179-1187). In some embodiments, one or more guide RNA nucleotides may be replaced with DNA nucleotides. In some embodiments, up to 2, 4, 6, 8, 10, or 12 RNA nucleotides in the 5'-terminal tail / seed guide region are replaced with DNA nucleotides. In certain embodiments, the majority of the 3'-terminal guide RNA nucleotides are replaced with DNA nucleotides. In certain embodiments, the 3'-terminal 16 guide RNA nucleotides are replaced with DNA nucleotides. In certain embodiments, 8 guide RNA nucleotides and the 3'-terminal 16 RNA nucleotides in the 5'-terminal tail / seed region are replaced with DNA nucleotides. In certain embodiments, the guide RNA nucleotides that extend outside the nuclease protein when the CRISPR complex is formed are replaced with DNA nucleotides. Such substitution of multiple RNA nucleotides with DNA nucleotides results in reduced off-target activity but similar on-target activity compared to the unmodified guide.However, substitution of all 3'-terminal RNA nucleotides can abolish the function of the guide (see Yin et al., Nat. Chem. Biol. (2018) 14, 311-316). Such modifications can be guided by knowledge of the structure of the CRISPR complex, including knowledge of the interactions of a limited number of nucleases with the 2'-OH of RNA (see Yin et al., Nat. Chem. Biol. (2018) 14, 311-316).

[0247] In one embodiment of the invention, the guide comprises a Cpf1 modified crRNA having a 5' handle and the guide segment further comprising a seed region and a 3' end. In some embodiments, the modified guide is selected from the group consisting of Acidaminococcus sp. BV3L6 Cpf1 (AsCpf1), Francisella tularensis subsp. Novicida U112 Cpf1 (FnCpf1), L. bacterium MC2017 Cpf1 (Lb3Cpf1), Butyrivibrio proteoclasticus Cpf1 (BpCpf1), Parcubacteria bacterium GWC2011_GWC2_44_17 Cpf1 (PbCpf1), Peregrinibacteria bacterium GW2011_GWA_33_10 Cpf1 (PeCpf1), Leptospira inadai Cpf1 (LiCpf1), Smithella sp. SC_K08D17 Cpf1 (SsCpf1), L. bacterium It may be used with any one of the following Cpf1s: MA2020 Cpf1 (Lb2Cpf1), Porphyromonas crevioricanis Cpf1 (PcCpf1), Porphyromonas macacae Cpf1 (PmCpf1), Candidatus Methanoplasma termitum Cpf1 (CMtCpf1), Eubacterium eligens Cpf1 (EeCpf1), Moraxella bovoculi 237 Cpf1 (MbCpf1), Prevotella disiens Cpf1 (PdCpf1), or L. bacterium ND2006 Cpf1 (LbCpf1).

[0248] In some embodiments, the modification to the guide is a chemical modification, an insertion, a deletion, or a split. In some embodiments, the chemical modification includes a 2'-O-methyl (M) analog, a 2'-deoxy analog, a 2-thiouridine analog, an N6-methyladenosine analog, a 2'-fluoro analog, a 2-aminopurine, a 5-bromo-uridine, a pseudouridine (Ψ), a N 1 -Methylpseudouridine (me1These include, but are not limited to, incorporation of Ψ), 5-methoxyuridine (5moU), inosine, 7-methylguanosine, 2'-O-methyl-3'-phosphorothioate (MS), S-constrained ethyl (cEt), phosphorothioate (PS), 2'-O-methyl-3'-thioPACE (MSP), or 2'-O-methyl-3'-phosphonoacetate (MP). In some embodiments, the guide comprises one or more phosphorothioate modifications. In certain embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 25 nucleotides of the guide are chemically modified. In some embodiments, all nucleotides are chemically modified. In certain embodiments, one or more nucleotides of the seed region are chemically modified. In certain embodiments, one or more nucleotides at the 3' end are chemically modified. In certain embodiments, none of the nucleotides in the 5' handle are chemically modified. In some embodiments, the chemical modification of the seed region is minor, such as the incorporation of a 2'-fluoro analog. In certain embodiments, one nucleotide in the seed region is substituted with a 2'-fluoro analog. In some embodiments, five or ten nucleotides at the 3' end are chemically modified. Such chemical modifications at the 3' end of the Cpf1 CrRNA improve gene cleavage efficiency (see Li, et al., Nature Biomedical Engineering, 2017, 1:0066). In certain embodiments, five nucleotides at the 3' end are substituted with a 2'-fluoro analog. In certain embodiments, ten nucleotides at the 3' end are substituted with a 2'-fluoro analog. In certain embodiments, five nucleotides at the 3' end are substituted with a 2'-O-methyl (M) analog. In some embodiments, three nucleotides at each of the 3' and 5' ends are chemically modified. In certain embodiments, the modifications include 2'-O-methyl or phosphorothioate analogs, hi certain embodiments, 12 nucleotides of the tetraloop and 16 nucleotides of the stem-loop region are substituted with 2'-O-methyl analogs.Such chemical modifications improve in vivo editing and stability (see Finn et al., Cell Reports (2018), 22:2227-2235).

[0249] In some embodiments, the loop of the 5' handle of the guide is modified. In some embodiments, the loop of the 5' handle of the guide is modified to have a deletion, insertion, split, or chemical modification. In certain embodiments, the loop comprises 3, 4, or 5 nucleotides. In certain embodiments, the loop comprises the sequence UCUU, UUUU, UAUU, or UGUU. In some embodiments, the guide molecule forms a stem loop with a separate, non-covalently linked sequence, which may be DNA or RNA.

[0250] Synthetically linked guides In one embodiment, the guide comprises a tracr sequence and a tracr mate sequence, which are chemically linked or conjugated via a non-phosphodiester bond. In another embodiment, the guide comprises a tracr sequence and a tracr mate sequence, which are chemically linked or conjugated via a non-nucleotide loop. In some embodiments, the tracr and tracr mate sequences are linked via a non-phosphodiester covalent linker. Examples of covalent linkers include, but are not limited to, chemical moieties selected from the group consisting of carbamates, ethers, esters, amides, imines, amidines, aminotriazines, hydrozones, disulfides, thioethers, thioesters, phosphorothioates, phosphorodithioates, sulfonamides, sulfonates, sulfones, sulfoxides, ureas, thioureas, hydrazides, oximes, triazoles, photolabile bonds, C—C bonds forming groups such as Diels-Alder cycloaddition pairs or ring-closing metathesis pairs, and Michael reaction pairs.

[0251] In some embodiments, the tracr and tracr mate sequences are first synthesized using standard phosphoramidite synthesis protocols (Herdewijn, P., ed., Methods in Molecular Biology Col. 288, Oligonucleotide Synthesis: Methods and Applications, Humana Press, New Jersey (2012)). In some embodiments, the tracr or tracr mate sequence can be functionalized to include appropriate functional groups for ligation reactions using standard protocols known in the art (Hermanson, GT, Bioconjugate Techniques, Academic Press (2013)). Examples of functional groups include, but are not limited to, hydroxyl, amine, carboxylic acid, carboxylic acid halide, activated ester of carboxylic acid, aldehyde, carbonyl, chlorocarbonyl, imidazolylcarbonyl, hydrozide, semicarbazide, thiosemicarbazide, thiol, maleimide, haloalkyl, sulfonyl, allyl, propargyl, diene, alkyne, and azide. After the tracr and tracr mate sequences are functionalized, a covalent chemical bond or linkage can be formed between the two oligonucleotides, including, but not limited to, carbamates, ethers, esters, amides, imines, amidines, aminotriazines, hydrozones, disulfides, thioethers, thioesters, phosphorothioates, phosphorodithioates, sulfonamides, sulfonates, sulphonates, sulfoxides, ureas, thioureas, hydrazides, oximes, triazoles, photolabile bonds, C—C bond forming groups such as Diels-Alder cycloaddition pairs or ring-closing metathesis pairs, and Michael reaction pairs.

[0252] In some embodiments, the tracr and tracr mate sequences can be chemically synthesized using automated solid-phase oligonucleotide synthesis machines using 2'-acetoxyethyl orthoester (2'-ACE) (Scaringe et al., J. Am. Chem. Soc. (1998) 120:11820-11821; Scaringe, Methods Enzymol. (2000) 317:3-18) or 2'-thionocarbamate (2'-TC) chemistry (Dellinger et al., J. Am. Chem. Soc. (2011) 133:11540-11546; Hendel et al., Nat. Biotechnol. (2015) 33:985-989).

[0253] In some embodiments, the tracr and tracr mate sequences can be covalently linked via sugar, internucleotide phosphodiester bond, or modification of purine and pyrimidine residues using various bioconjugation reactions, loops, bridges, and non-nucleotide linkages (Sletten et al., Angew. Chem. Int. Ed. (2009) 48:6974-6998; Manoharan, M. Curr. Opin. Chem. Biol. (2004) 8:570-9; Behlke et al., Oligonucleotides (2008) 18:305-19; Watts, et al., Drug. Discov. Today (2008) 13:842-55; Shukla, et al., ChemMedChem (2010) 5:328-49).

[0254] In some embodiments, the tracr and tracr mate sequences can be covalently linked using click chemistry. In some embodiments, the tracr and tracr mate sequences can be covalently linked using a triazole linker. In some embodiments, the tracr and tracr mate sequences can be covalently linked using a Huisgen 1,3-dipolar cycloaddition reaction involving an alkyne and an azide to yield a highly stable triazole linker (He et al., ChemBioChem (2015) 17:1809-1812, WO2016 / 186745). In some embodiments, the tracr and tracr mate sequences are covalently linked by combining a 5'-hexyne tracrRNA and a 3'-azide crRNA. In some embodiments, either or both of the 5'-hexyne tracrRNA and 3'-azido crRNA can be protected with 2'-acetoxyethyl orthoester (2'-ACE) groups, which can then be removed using Dharmacon protocols (Scaringe et al., J. Am. Chem. Soc. (1998) 120:11820-11821; Scaringe, Methods Enzymol. (2000) 317:3-18).

[0255] In some embodiments, the tracr and tracr mate sequences can be covalently linked via a linker (e.g., a non-nucleotide loop) containing moieties such as spacers, attachments, bioconjugates, chromophores, reporter groups, dye-labeled RNA, and non-naturally occurring nucleotide analogs. More specifically, suitable spacers for the purposes of the present invention include, but are not limited to, polyethers (e.g., polyethylene glycol, polyalcohols, polypropylene glycol, or mixtures of ethylene glycol and propylene glycol), polyamine groups (e.g., spermine, spermidine, and their polymeric derivatives), polyesters (e.g., poly(ethyl acrylate)), polyphosphodiesters, alkylenes, and combinations thereof. Suitable attachments include any moiety that can be added to the linker to impart additional properties to the linker, including, but not limited to, fluorescent labels. Suitable bioconjugates include, but are not limited to, peptides, glycosides, lipids, cholesterol, phospholipids, diacylglycerols and dialkylglycerols, fatty acids, hydrocarbons, enzyme substrates, steroids, biotin, digoxigenin, carbohydrates, and polysaccharides. Suitable chromophores, reporter groups, and dye-labeled RNAs include, but are not limited to, fluorescent dyes such as fluorescein and rhodamine, chemiluminescent, electrochemiluminescent, and bioluminescent marker compounds. Exemplary linker designs for conjugating two RNA components are also described in International Patent Application Publication No. WO2004 / 015075.

[0256] The linker (e.g., non-nucleotide loop) can be of any length. In some embodiments, the linker has a length equivalent to about 0-16 nucleotides. In some embodiments, the linker has a length equivalent to about 0-8 nucleotides. In some embodiments, the linker has a length equivalent to about 0-4 nucleotides. In some embodiments, the linker has a length equivalent to about 2 nucleotides. Exemplary linker designs are also described in International Patent Application Publication No. WO 2011 / 008730.

[0257] A typical Type II Cas9 sgRNA comprises (from 5' to 3'): a guide sequence, a poly-U tract, a first complementary stretch (the "repeat"), a loop (tetraloop), a second complementary stretch (the "anti-repeat" complementary to the repeat), a stem, and an additional stem-loop and stem and poly-A (often poly-U in RNA) tail (terminator). In preferred embodiments, certain embodiments of the guide structure are retained, and certain embodiments of the guide structure may be modified, for example, by adding, subtracting, or substituting features, while certain other embodiments of the guide structure are maintained. Preferred locations for engineered sgRNA modifications, including, but not limited to, insertions, deletions, and substitutions, include the guide ends and regions of the sgRNA that are exposed upon complexing with a CRISPR protein and / or target, e.g., the tetraloop and / or loop 2.

[0258] In certain embodiments, the guide of the present invention comprises a specific binding site (e.g., an aptamer) for an adaptor protein, which may comprise one or more functional domains (e.g., via a fusion protein). When such a guide forms a CRISPR complex (i.e., a CRISPR enzyme that binds to the guide and the target), the adaptor protein binds and the functional domain associated with the adaptor protein is positioned in a spatial orientation favorable for the attributed function to be effective. For example, if the functional domain is a transcriptional activator (e.g., VP64 or p65), the transcriptional activator is positioned in a spatial orientation that affects the transcription of the target. Similarly, a transcriptional repressor is positioned favorably to affect the transcription of the target, and a nuclease (e.g., Fok1) is positioned favorably to cleave or partially cleave the target.

[0259] Those skilled in the art will understand that modifications to a guide that allow for binding of an adaptor plus functional domain but do not allow for proper positioning of the adaptor plus functional domain (e.g., due to steric hindrance within the three-dimensional structure of the CRISPR complex) are unintended modifications. One or more modified guides can be modified at the tetraloop, stem-loop 1, stem-loop 2, or stem-loop 3, preferably at either the tetraloop or stem-loop 2, and most preferably at both the tetraloop and stem-loop 2, as described herein.

[0260] The repeat:anti-repeat duplex will be evident from the secondary structure of the sgRNA. It will typically consist of a first complementary stretch after the poly-U tract (5' to 3' direction) and before the tetraloop, and a second complementary stretch after the tetraloop (5' to 3' direction) and before the poly-A tract. The first complementary stretch (the "repeat") is complementary to the second complementary stretch (the "anti-repeat"). Thus, they undergo Watson-Crick base pairing and form the dsRNA duplex when folded back together. The anti-repeat sequence is therefore complementary to the repeat in terms of AU or CG base pairing, but also in terms of the anti-repeat being in the opposite direction due to the tetraloop.

[0261] In embodiments of the invention, modifications to the guide structure include substitutions of bases in stem-loop 2. For example, in some embodiments, the "actt" ("acuu" in RNA) and "aagt" ("aagu" in RNA) bases in stem-loop 2 are substituted with "cgcc" and "gcgg". In some embodiments, the "actt" and "aagt" bases in stem-loop 2 are substituted with a four-nucleotide complementary GC-rich region. In some embodiments, the four-nucleotide complementary GC-rich region is "cgcc" and "gcgg" (both in the 5' to 3' direction). In some embodiments, the four-nucleotide complementary GC-rich region is "gcgg" and "cgcc" (both in the 5' to 3' direction). Other C and G combinations of the four-nucleotide complementary GC-rich region will be apparent, including CCCC and GGGG.

[0262] In one embodiment, stem loop 2, e.g., "ACTTgtttAAGT" (SEQ ID NO: 12006), can be replaced with any "XXXXgtttYYYY" (SEQ ID NO: 12007), e.g., where XXXX and YYYY represent any complementary set of nucleotides that base pair with each other to create the stem.

[0263] In one embodiment, the stem comprises at least about 4 bp containing complementary X and Y sequences, although stems of more, e.g., 5, 6, 7, 8, 9, 10, 11, or 12 base pairs, or fewer, e.g., 3 or 2 base pairs, are also contemplated. Thus, for example, X2-12 and Y2-12 (X and Y represent any complementary set of nucleotides) may be contemplated. In one embodiment, a stem made of X and Y nucleotides, together with "gttt," forms a perfect hairpin in the overall secondary structure, which may be advantageous, and the amount of base pairs can be any amount that forms a perfect hairpin. In one embodiment, any complementary X:Y base-pairing sequence (e.g., in terms of length) is permissible as long as the secondary structure of the entire sgRNA is maintained. In one embodiment, the stem can be in the form of X:Y base-pairing that does not disrupt the secondary structure of the entire sgRNA, in that it has a DR:tracr duplex and three stem-loops. In one embodiment, the "gttt" tetraloop connecting ACTT and AAGT (or any alternative stem made of X:Y base pairs) can be any sequence of the same length (e.g., 4 base pairs) or longer that does not disrupt the overall secondary structure of the sgRNA. In one embodiment, the stem loop can further extend stem loop 2, e.g., an MS2 aptamer. In one embodiment, stem loop 3 "GGCACCGagtCGGTGC" (SEQ ID NO: 12008) can similarly take the form "XXXXXXXagtYYYYYYY" (SEQ ID NO: 12009), e.g., where X7 and Y7 represent any complementary set of nucleotides that base pair with each other to create the stem. In one embodiment, the stem comprises approximately 7 bp comprising complementary X and Y sequences, although stems of more or fewer base pairs are also contemplated. In one embodiment, the stem made of X and Y nucleotides, together with "agt", form a perfect hairpin in the overall secondary structure. In one embodiment, any complementary X:Y base pairing sequence is tolerated as long as the secondary structure of the entire sgRNA is maintained.In one embodiment, the stem can be in the form of X:Y base pairing that does not disrupt the secondary structure of the entire sgRNA, in that it has a DR:tracr duplex and three stem-loops. In one embodiment, the "agt" sequence of stem-loop 3 can be extended or replaced by an aptamer, e.g., an MS2 aptamer, or a sequence that otherwise generally maintains the structure of stem-loop 3. In an alternative embodiment of stem-loop 2 and / or 3, each X and Y pair can refer to any base pair. In one embodiment, non-Watson-Crick base pairings are contemplated, where such pairings otherwise normally maintain the structure of that position of the stem-loop.

[0264] In one embodiment, the DR:tracrRNA duplex can be substituted with a sequence of the form: gYYYYag(N)NNNNxxxxNNNN(AAN)uuRRRRu (SEQ ID NO: 12010) (using standard IUPAC nomenclature for nucleotides), where (N) and (AAN) represent a portion of the bulge in the duplex, and "xxxx" represents a linker sequence. The NNNN in the direct repeat can be any sequence as long as it base-pairs with the corresponding NNNN portion of the tracrRNA. In one embodiment, the DR:tracrRNA duplex can be connected with a linker of any length (xxxx...) and base composition, as long as it does not change the overall structure.

[0265] In one embodiment, the structural requirements for the sgRNA are that it has a double strand and three stem loops. In most embodiments, the actual structural requirements for many specific base requirements are relaxed in that the structure of the DR:tracrRNA duplex should be maintained, but the sequences that create the structure, i.e., stems, loops, bulges, etc., may be altered.

[0266] Aptamers One guide with a first aptamer / RNA binding protein pair can be linked or fused to an activator, while a second guide with a second aptamer / RNA binding protein pair can be linked or fused to a repressor. The guides are directed to different targets (gene loci), so one gene is activated and one is repressed. For example, the following outline illustrates such an approach: Guide1-MS2 aptamer ------- MS2 RNA-binding protein --- VP64 activator, and Guide2-PP7 aptamer-------PP7 RNA binding protein-------SID4x repressor.

[0267] The present invention also relates to orthogonal PP7 / MS2 gene targeting. In this example, sgRNAs targeting different loci are modified with different RNA loops to recruit MS2-VP64 or PP7-SID4X, which activate and repress the target loci, respectively. PP7 is an RNA-binding coat protein of the bacteriophage Pseudomonas. Like MS2, it binds to specific RNA sequences and secondary structures. The PP7-RNA recognition motif differs from that of MS2. Consequently, PP7 and MS2 can be multiplexed to mediate different effects simultaneously at different genomic loci. For example, an sgRNA targeting locus A can be modified with an MS2 loop to recruit the MS2-VP64 activator, while another sgRNA targeting locus B can be modified with a PP7 loop to recruit the PP7-SID4X repressor domain. Thus, dCas9 can mediate orthogonal locus-specific modifications in the same cell. This principle can be extended to incorporate other orthogonal RNA-binding proteins, such as Q-beta.

[0268] An alternative option for orthogonal suppression is to incorporate a non-coding RNA loop with an interactive inhibitory function into the guide (either at the same position as the MS2 / PP7 loop incorporated into the guide or at the 3' end of the guide). For example, a guide has been designed with a non-coding (but known to be inhibitory) RNA loop (e.g., using an Alu repressor (in RNA) that inhibits RNA polymerase II in mammalian cells). The Alu RNA sequence, as used herein, is placed in place of the MS2 RNA sequence (e.g., at the tetraloop and / or stem-loop 2) and / or at the 3' end of the guide. This allows for possible combinations of MS2, PP7, or Alu at the tetraloop and / or stem-loop 2 positions, and optionally adds an Alu at the 3' end of the guide (with or without a linker).

[0269] The use of two different aptamers (different RNAs) allows an activator-adapter protein fusion and a repressor-adapter protein fusion to be used with different guides, activating expression of one gene while suppressing expression of another. These, along with their different guides, can be administered together or substantially together in a multiplexed approach. While many such modified guides, e.g., 10, 20, or 30, can all be used simultaneously, a relatively small number of Cas9s can be used with many modified guides, allowing only one (or at least a minimal number) Cas9 to be delivered. The adapter protein can be associated (preferably linked or fused) with one or more activators or one or more repressors. For example, the adapter protein can be associated with a first activator and a second activator. The first and second activators can be the same, but are preferably different activators. For example, one can be VP64 and the other can be p65, although these are merely examples and other transcriptional activators are contemplated. Three or more, or even four or more, activators (or repressors) may be used, although package size may limit the number of different functional domains to more than five. A linker is preferably used over direct fusion to the adaptor protein, with two or more functional domains associated with the adaptor protein. Suitable linkers may include GlySer linkers.

[0270] It is also envisioned that the enzyme-guide complex as a whole may be associated with two or more functional domains, for example, two or more functional domains may be associated with the enzyme, two or more functional domains may be associated with the guide (through one or more adaptor proteins), or one or more functional domains may be associated with the enzyme and one or more functional domains may be associated with the guide (through one or more adaptor proteins).

[0271] The fusion between the adaptor protein and the activator or repressor may include a linker. For example, the GlySer linker GGGS can be used. They can be used in triplicate ((GGGGS)3 (SEQ ID NO: 12011)) or in repeats of 6, 9, 12, or more, as needed, to obtain the appropriate length. A linker can be used between the RNA binding protein and the functional domain (activator or repressor), or between the CRISPR enzyme (Cas9) and the functional domain (activator or repressor). The linker allows the user to manipulate the appropriate amount of "mechanical flexibility."

[0272] Dead guide In one embodiment, the present invention provides guide sequences that are modified in a manner that allows successful CRISPR complex formation and target binding while preventing nuclease activity (i.e., no nuclease activity / no indel activity). For illustrative purposes, such modified guide sequences are referred to as "inactive guides" or "inactive guide sequences." These inactive guides or inactive guide sequences may be considered catalytically inactive or conformationally inactive with respect to nuclease activity. Nuclease activity can be measured using surveyor analysis or deep sequencing, preferably surveyor analysis, commonly used in the art. Similarly, inactive guide sequences may not sufficiently participate in productive base pairing with respect to their ability to promote catalytic activity or distinguish between on-target and off-target binding activity. Briefly, the surveyor assay involves purifying and amplifying a CRISPR target portion relative to a gene and forming a heteroduplex with a primer that amplifies the CRISPR target site. After reannealing, the products are treated with SURVEYOR nuclease and SURVEYOR enhancer S (Transgenoics) according to the manufacturer's recommended protocol, analyzed on a gel, and quantified based on the intensity of the reference band.

[0273] Thus, in related embodiments, the invention provides a non-naturally occurring or engineered composition of a Cas9 CRISPR-Cas system, comprising a functional Cas9 as described herein and a guide RNA (gRNA), wherein the gRNA comprises an inactive guide sequence such that the gRNA is capable of hybridizing to a target sequence such that the Cas9 CRISPR-Cas system is directed to a genomic locus of interest in a cell without detectable indel activity due to the nuclease activity of the non-mutated Cas9 enzyme of the system, as detected by a SURVEYOR assay. For purposes of brevity, a gRNA comprising an inactive guide sequence capable of hybridizing to a target sequence such that the Cas9 CRISPR-Cas system is directed to a genomic locus of interest in a cell without detectable indel activity due to the nuclease activity of the non-mutated Cas9 enzyme of the system, as detected by a SURVEYOR assay, is referred to herein as an "inactive gRNA." It will be understood that any of the gRNAs of the invention described elsewhere herein can be used as an inactive gRNA / gRNA comprising an inactive guide sequence as described herein below. Any of the methods, products, compositions and uses described elsewhere herein are equally applicable with inactive gRNAs / gRNAs comprising an inactive guide sequence as described in more detail below. With further guidance, the following specific embodiments and embodiments are provided.

[0274] The ability of an inactive guide sequence to direct the sequence-specific binding of a CRISPR complex to a target sequence can be evaluated by any suitable assay. For example, the components of a CRISPR system sufficient to form a CRISPR complex, including the inactive guide sequence to be tested, can be provided to a host cell having the corresponding target sequence, such as by transfection with a vector encoding the components of the CRISPR sequence, and then the preferential cleavage within the target sequence can be evaluated by, for example, the Surveyor assay described herein. Similarly, the cleavage of a target polynucleotide sequence can be evaluated in a test tube by providing the components of a CRISPR complex, including the target sequence, the inactive guide sequence to be tested, and a control guide sequence different from the test inactive guide sequence, and comparing the binding or cleavage rate at the target sequence between the test and control guide sequence reactions. Other assays are possible and will occur to those skilled in the art. The inactive guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of a cell.

[0275] As further described herein, several structural parameters enable a suitable framework for arriving at such inactive guides. The inactive guide sequence is shorter than the respective guide sequence that results in active Cas9-specific indel formation. The inactive guide is 5%, 10%, 20%, 30%, 40%, or 50% shorter than the respective guide directed to the same Cas9 that results in active Cas9-specific indel formation.

[0276] As described below and known in the art, one embodiment of gRNA-Cas9 specificity is a direct repeat sequence that is appropriately linked to such a guide. In particular, this implies that the direct repeat sequence is designed according to the origin of Cas9. Therefore, available structural data for effective inactive guide sequences can be used to design a Cas9-specific equivalent. For example, the structural similarity between the orthologous nuclease domains RuvC of two or more Cas9 effector proteins can be used to transcribe an equivalent inactive guide design. Therefore, the inactive guide herein can be appropriately modified in length and sequence to reflect such a Cas9-specific equivalent that successfully forms a CRISPR complex and binds to the target, while not achieving nuclease activity.

[0277] The use of inactive guides in the context of this specification and the latest technology provides a surprising and unexpected platform for network biology and / or systems biology, which enables multiple gene targeting, especially bidirectional multiple gene targeting, in in vitro, ex vivo, and in vivo applications.Before the use of inactive guides, for example, it is difficult, and sometimes impossible, to address multiple targets of gene activity activation, suppression, and / or silencing.Using inactive guides, multiple targets, and therefore multiple activities, can be addressed, for example, in the same cell, the same animal, or the same patient.Such multiplexing can be performed simultaneously or can be adjusted to a desired time frame.

[0278] For example, the inactive guide allows gRNA to be used as a means of gene targeting without first causing nuclease activity, and at the same time provides a means for activation or repression. Guide RNAs comprising inactive guides can be further modified to include protein adaptors (e.g., aptamers) as described elsewhere herein, which allow for the functional placement of elements, particularly gene effectors (e.g., activators or repressors of gene activity), in a way that allows for the activation or repression of gene activity. One example is the incorporation of aptamers, which is described herein and is a current technology. By engineering gRNAs comprising inactive guides to incorporate aptamers that interact with proteins (Konermann et al., "Genome-scale transcription activation by an engineered CRISPR-Cas9 complex," doi:10.1038 / nature14136, incorporated herein by reference), synthetic transcription activation complexes consisting of multiple different effector domains can be assembled. This can be modeled after the natural transcription activation process. For example, an aptamer that selectively binds to an effector (e.g., an activator or repressor, dimerized MS2 bacteriophage coat protein as a fusion protein with an activator or repressor), or a protein that itself binds to an effector (e.g., an activator or repressor), can be added to the tetraloop and / or stem-loop 2 of an inactive gRNA. In the case of MS2, the fusion protein MS2-VP64 binds to the tetraloop and / or stem-loop 2 and then mediates transcriptional upregulation, for example, for Neurog2. Other transcriptional activators are, for example, VP64.P65, HSF1, and MyoD1. Simply by way of example of this concept, replacement of the MS2 stem-loop with a stem-loop that interacts with PP7 can be used to recruit a repression element.

[0279] Thus, one embodiment is a gRNA of the present invention comprising an inactive guide, wherein the gRNA further comprises a modification that results in gene activation or repression as described herein. The inactive gRNA may comprise one or more aptamers. The aptamer may be specific to a gene effector, gene activator, or gene repressor. Alternatively, the aptamer may be specific to a protein that is specific to and recruits / binds to a particular gene effector, gene activator, or gene repressor. When multiple sites exist for recruiting an activator or repressor, these sites are preferably specific to either an activator or a repressor. When multiple sites exist for binding an activator or a repressor, the sites may be specific to the same activator or the same repressor. The sites may also be specific to different activators or different repressors. The gene effector, gene activator, or gene repressor may exist in the form of a fusion protein.

[0280] In embodiments, the inactive gRNA or Cas9 CRISPR-Cas complex described herein comprises a non-naturally occurring or engineered composition comprising two or more adaptor proteins, each protein associated with one or more functional domains, and the adaptor proteins bind to different RNA sequence(s) inserted into at least one loop of the inactive gRNA. Accordingly, embodiments provide a non-naturally occurring or engineered composition comprising a guide RNA (gRNA) comprising an inactive guide sequence as defined herein capable of hybridizing to a target sequence at a genomic locus of interest in a cell, a Cas9 comprising at least one or more nuclear localization sequences and optionally comprising at least one mutation, at least one loop of the inactive gRNA modified by the insertion of different RNA sequence(s) that bind to one or more adaptor proteins, and the adaptor proteins are associated with one or more functional domains, or the inactive gRNA is modified to have at least one non-coding functional loop, and the composition comprises two or more adaptor proteins, each protein associated with one or more functional domains.

[0281] In certain embodiments, the adaptor protein is a fusion protein comprising the functional domain, wherein the fusion protein optionally comprises a linker between the adaptor protein and the functional domain, wherein the linker optionally comprises a GlySer linker.

[0282] In certain embodiments, at least one loop of said inactive gRNA is not modified by the insertion of different RNA sequence(s) that bind to said two or more adaptor proteins.

[0283] In certain embodiments, the one or more functional domains associated with the adaptor protein are transcription activation domains.

[0284] In certain embodiments, the one or more functional domains associated with the adaptor protein are transcription activation domains, including VP64, p65, MyoD1, HSF1, RTA, or SET7 / 9.

[0285] In certain embodiments, the one or more functional domains associated with the adaptor protein is a transcriptional repressor domain.

[0286] In certain embodiments, the transcriptional repressor domain is a KRAB domain.

[0287] In certain embodiments, the transcriptional repressor domain is a NuE domain, an NcoR domain, a SID domain, or a SID4X domain.

[0288] In certain embodiments, at least one of the one or more functional domains associated with the adaptor protein has one or more activities including methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional termination factor activity, histone modification activity, DNA incorporating activity, RNA cleavage activity, DNA cleavage activity, or nucleic acid binding activity.

[0289] In certain embodiments, the DNA cleavage activity is due to a Fok1 nuclease.

[0290] In certain embodiments, the inactive gRNA is modified such that after the inactive gRNA binds to the adaptor protein and further binds to the Cas9 and target, the functional domain is in a spatial orientation that allows it to function in its attributed function.

[0291] In certain embodiments, at least one loop of the inactive gRNA is a tetraloop and / or loop 2. In certain embodiments, the tetraloop and loop 2 of the inactive gRNA are modified by the insertion of the different RNA sequence(s).

[0292] In certain embodiments, the insertion of different RNA sequence(s) that bind to one or more adaptor proteins is an aptamer sequence.In certain embodiments, the aptamer sequence is two or more aptamer sequences that are specific to the same aptamer protein.In certain embodiments, the aptamer sequence is two or more aptamer sequences that are specific to different adaptor proteins.

[0293] In certain embodiments, the adaptor proteins include MS2, PP7, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, ΦCb5, ΦCb8r, ΦCb12r, ΦCb23r, 7s, PRR1.

[0294] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the eukaryotic cell is a mammalian cell, optionally a mouse cell. In certain embodiments, the mammalian cell is a human cell.

[0295] In certain embodiments, the first adaptor protein is associated with the p65 domain and the second adaptor protein is associated with the HSF1 domain.

[0296] In certain embodiments, the composition comprises a Cas9 CRISPR-Cas complex having at least three functional domains, at least one of which is associated with the Cas9 and at least two of which are associated with an inactive gRNA.

[0297] In certain embodiments, the composition further comprises a second gRNA, wherein the second gRNA is an active gRNA capable of hybridizing to a second target sequence such that the second Cas9 CRISPR-Cas system is directed to a second genomic locus of interest in the cell with detectable indel activity at the second genomic locus due to the nuclease activity of the Cas9 enzyme of the system.

[0298] In certain embodiments, the composition further comprises a plurality of inactive gRNAs and / or a plurality of active gRNAs.

[0299] One embodiment of the present invention utilizes the modularity and customizability of the gRNA scaffold to establish a series of gRNA scaffolds with different binding sites (particularly aptamers) to orthogonally recruit different types of effectors. Again, for illustrative purposes and to illustrate a broader concept, the replacement of the MS2 stem-loop with a stem-loop that interacts with PP7 may be used to bind / recruit repression elements, enabling multiplexed and bidirectional transcriptional regulation. Thus, in general, gRNAs containing inactive guides can be used to achieve multiplexed and favorable bidirectional transcriptional regulation, with this transcriptional regulation being the most favorable of the gene. For example, one or more gRNAs containing inactive guide(s) can be used to target activation of one or more target genes. At the same time, one or more gRNAs containing inactive guide(s) can be used to target repression of one or more target genes. Such sequences may be applied in a variety of different combinations, for example, the target gene is first repressed and then, at an appropriate time, other targets are activated, or selected genes are repressed while simultaneously activating selected genes, followed by further activation and / or repression, etc. As a result, multiple components of one or more biological systems can be advantageously addressed together.

[0300] In embodiments, the present invention provides nucleic acid molecule(s) encoding an inactive gRNA or Cas9 CRISPR-Cas complex or composition described herein.

[0301] In embodiments, the present invention provides a vector system comprising a nucleic acid molecule encoding an inactive guide RNA as defined herein. In certain embodiments, the vector system further comprises a nucleic acid molecule(s) encoding Cas9. In certain embodiments, the vector system further comprises a nucleic acid molecule(s) encoding an (active) gRNA. In certain embodiments, the nucleic acid molecule or vector further comprises a control element(s) operable in a eukaryotic cell operably linked to the nucleic acid molecule encoding the guide sequence (gRNA) and / or Cas9 and / or optionally the nucleic acid molecule encoding each localization sequence(s).

[0302] In another embodiment, structural analysis can be used to examine the interaction between the inactive guide and the active Cas9 nuclease, which allows DNA binding but does not cleave DNA.In this way, the amino acids that are important for the nuclease activity of Cas9 can be determined.The modification of these amino acids can improve the Cas9 enzyme used in gene editing.

[0303] A further embodiment is to combine the use of the inactive guides described herein with other applications of CRISPR known in the art, as described herein. For example, a gRNA containing an inactive guide(s) for targeted multiple gene activation or repression or targeted multiple bidirectional gene activation / repression may be combined with a gRNA containing a guide that maintains nuclease activity, as described herein. Such gRNAs containing guides that maintain nuclease activity may or may not further contain modifications that suppress gene activity (e.g., aptamers). Such gRNAs containing guides that maintain nuclease activity may or may not further contain modifications that activate gene activity (e.g., aptamers). In this way, additional means for multigene control are introduced (e.g., nuclease-free / indel-free multigene-targeted activation can be provided simultaneously or in combination with gene-targeted repression with nuclease activity).

[0304] For example, 1) using one or more gRNAs (e.g., 1-50, 1-40, 1-30, 1-20, preferably 1-10, more preferably 1-5) that target one or more genes and include inactive guide(s) modified with appropriate aptamers for recruitment of gene activators may be combined with 2) one or more gRNAs (e.g., 1-50, 1-40, 1-30, 1-20, preferably 1-10, more preferably 1-5) that target one or more genes and include inactive guide(s) modified with appropriate aptamers for recruitment of gene repressors. 1) and 2) may then be combined with 3) one or more gRNAs (e.g., 1-50, 1-40, 1-30, 1-20, preferably 1-10, more preferably 1-5) that target one or more genes. This combination may then be performed with 1) + 2) + 3) together with 4) one or more gRNAs (e.g., 1-50, 1-40, 1-30, 1-20, preferably 1-10, more preferably 1-5) that target one or more genes and are further modified with appropriate aptamers for recruitment of gene activators. This combination may then be performed with 1) + 2) + 3) + 4) together with 5) one or more gRNAs (e.g., 1-50, 1-40, 1-30, 1-20, preferably 1-10, more preferably 1-5) that target one or more genes and are further modified with appropriate aptamers for recruitment of gene repressors. Consequently, a variety of uses and combinations are encompassed by the present invention. For example, combination 1)+2), combination 1)+3), combination 2)+3), combination 1)+2)+3), combination 1)+2)+3), combination 1)+2)+3)+4), combination 1)+3)+4), combination 2)+3)+4), combination 1)+2)+4), combination 1)+2)+3)+4)+5), combination 1)+3)+4)+5), combination 2)+3)+4)+5), combination 1)+2)+4)+5), combination 1)+2)+3)+5), combination 1)+3)+5), combination 2)+3)+5), combination 1)+2)+5).

[0305] In embodiments, the present invention provides an algorithm for designing, evaluating, or selecting an inactive guide RNA targeting sequence (inactive guide sequence) for guiding a Cas9 CRISPR-Cas system to a target locus. In particular, it has been determined that the specificity of an inactive guide RNA is related to and can be optimized by varying i) the GC content and ii) the length of the targeting sequence. In embodiments, the present invention provides an algorithm for designing or evaluating an inactive guide RNA targeting sequence that minimizes off-target binding or interactions of the inactive guide RNA. In an embodiment of the present invention, the algorithm for selecting inactive guide RNA targeting sequence for directing CRISPR system to a locus in an organism comprises: a) locating one or more CRISPR motifs at the locus; and analyzing the downstream 20 nucleotide (nt) sequence of each CRISPR motif by i) measuring the GC content of the sequence, and ii) determining whether there is an off-target match of the downstream 15 nucleotides closest to the CRISPR motif in the genome of the organism; and c) if the GC content of the sequence is 70% or less and no off-target match is identified, select the sequence of 15 nucleotides for use in inactive guide RNA.In an embodiment, if the GC content of the sequence is 60% or less, the sequence is selected as targeting sequence.In some embodiments, if the GC content of the sequence is 55% or less, 50% or less, 45% or less, 40% or less, 35% or less or 30% or less, the sequence is selected as targeting sequence. In embodiments, two or more sequences of the locus are analyzed, and the sequence with the lowest GC content, or the next lowest GC content, or the next lowest GC content is selected.In embodiments, when no off-target match is identified in the genome of the organism, this sequence is selected as targeting sequence.In embodiments, when no off-target match is identified in the regulatory sequence of the genome, this targeting sequence is selected.

[0306] In an embodiment, the present invention provides a method for selecting an inactive guide RNA targeting sequence for directing a functionalized CRISPR system to a locus in an organism, the method comprising: a) placing one or more CRISPR motifs at the locus; b) analyzing a 20 nt sequence downstream of each CRISPR motif by i) measuring the GC content of the sequence, and ii) determining whether there is an off-target match of the first 15 nt of the sequence in the genome of the organism; c) selecting the sequence for use in the guide RNA if the GC content of the sequence is 70% or less and no off-target match is identified. In an embodiment, the sequence is selected if the GC content is 50% or less. In an embodiment, the sequence is selected if the GC content is 40% or less. In an embodiment, the sequence is selected if the GC content is 30% or less. In an embodiment, two or more sequences are analyzed, and the sequence with the lowest GC content is selected. In an embodiment, an off-target match is identified in a regulatory sequence of the organism. In an embodiment, the locus is a regulatory region. Embodiments provide an inactive guide RNA comprising a targeting sequence selected according to the methods described above.

[0307] In an embodiment, the present invention provides an inactive guide RNA for targeting a functionalized CRISPR system to a locus in an organism. In an embodiment of the present invention, the inactive guide RNA comprises a targeting sequence, and the CG content of the targeting sequence is 70% or less, and the first 15 nt of the targeting sequence does not match an off-target sequence downstream from a CRISPR motif in a regulatory sequence of another locus of the organism. In certain embodiments, the GC content of the targeting sequence is 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, or 30% or less. In certain embodiments, the GC content of the targeting sequence is 70% to 60%, 60% to 50%, 50% to 40%, or 40% to 30%. In an embodiment, the targeting sequence has the lowest CG content among the potential targeting sequences of the locus.

[0308] In an embodiment of the present invention, the first 15nt of the inactive guide matches the target sequence.In another embodiment, the first 14nt of the inactive guide matches the target sequence.In another embodiment, the first 13nt of the inactive guide matches the target sequence.In another embodiment, the first 12nt of the inactive guide matches the target sequence.In another embodiment, the first 11nt of the inactive guide matches the target sequence.In another embodiment, the first 10nt of the inactive guide matches the target sequence.In an embodiment of the present invention, the first 15nt of the inactive guide does not match the off-target sequence downstream from the CRISPR motif in the regulatory region of another gene locus.In other embodiments, the first 14nt, or the first 13nt of the inactive guide, or the first 12nt of the guide, or the first 11nt of the inactive guide, or the first 10nt of the inactive guide does not match the off-target sequence downstream from the CRISPR motif in the regulatory region of another gene locus. In other embodiments, the first 15 nt, or 14 nt, or 13 nt, or 12 nt, or 11 nt of the inactive guide do not match an off-target sequence downstream from the CRISPR motif in the genome.

[0309] In certain embodiments, the inactive guide RNA comprises additional nucleotides at 3' end that do not match the target sequence.Therefore, the inactive guide RNA that comprises the first 15nt, or 14nt, or 13nt, or 12nt, or 11nt downstream of CRISPR motif can be extended at 3' end by 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt or more.

[0310] The present invention provides methods for targeting a Cas9 CRISPR-Cas system, including but not limited to, an inactive Cas9 (dCas9) or a functionalized Cas9 system (which may include a functionalized Cas9 or a functionalized guide), to a locus. In embodiments, the present invention provides methods for selecting an inactive guide RNA targeting sequence and targeting a functionalized CRISPR system to a locus in an organism. In embodiments, the present invention provides methods for selecting an inactive guide RNA targeting sequence and effecting gene regulation of a target locus by a functionalized Cas9 CRISPR-Cas system. In certain embodiments, the method is used to effect target gene regulation while minimizing off-target effects. In embodiments, the present invention provides methods for selecting two or more inactive guide RNA targeting sequences and effecting gene regulation of two or more target loci by a functionalized Cas9 CRISPR-Cas system. In certain embodiments, the method is used to effect regulation of two or more target loci while minimizing off-target effects.

[0311] In embodiments, the present invention provides a method for selecting an inactive guide RNA targeting sequence for directing a functionalized Cas9 to a genetic locus in an organism, the method comprising: a) placing one or more CRISPR motifs at the genetic locus; b) analyzing the sequence downstream of each CRISPR motif by: i) selecting 10-15 nt adjacent to the CRISPR motif; ii) measuring the GC content of the sequence; and c) selecting the 10-15 nt sequence as a targeting sequence for use in the guide RNA if the GC content of the sequence is 40% or greater. In embodiments, the sequence is selected if the GC content is 50% or greater. In embodiments, the sequence is selected if the GC content is 60% or greater. In embodiments, the sequence is selected if the GC content is 70% or greater. In embodiments, two or more sequences are analyzed, and the sequence with the highest GC content is selected. In embodiments, the method further comprises adding nucleotides to the 3' end of the selected sequence that do not match the sequence downstream of the CRISPR motif. Embodiments provide an inactive guide RNA comprising a targeting sequence selected according to the methods described above.

[0312] In some embodiments, the present invention provides an inactive guide RNA for targeting a functionalized CRISPR system to a locus in an organism, wherein the targeting sequence of the inactive guide RNA consists of 10 to 15 nucleotides adjacent to a CRISPR motif at the locus, and the CG content of the targeting sequence is 50% or more. In certain embodiments, the inactive guide RNA further comprises additional nucleotides at the 3' end of the targeting sequence that do not match the sequence downstream of the CRISPR motif at the lo...

Claims

1. A composition comprising:

1. A targeting moiety effective in targeting hematopoietic cells, comprising one or more n-mer motifs, wherein at least one n-mer motif is a. Any one of SEQ ID NOs: 1-1000; b. any one of SEQ ID NOs: 2001-3000; c. any one of SEQ ID NOs: 4001-5000; d. any one of SEQ ID NOs: 6001-7000; e. any one of SEQ ID NOs: 8001-9000; f. any one of SEQ ID NOs: 10001-11000; g. or any combination thereof.

2. The composition of claim 1 , wherein the hematopoietic cells are differentiated hematopoietic cells or progenitor cells.

3. 10. The composition of any one of the preceding claims, wherein the targeting moiety comprises a polypeptide, a polynucleotide, a lipid, a polymer, a sugar, or any combination thereof.

4. 10. The composition of any one of the preceding claims, wherein the targeting moiety comprises a viral polypeptide.

5. 10. The composition of any one of the preceding claims, wherein the targeting moiety comprises a viral capsid polypeptide.

6. 10. The composition of any one of the preceding claims, wherein the targeting moiety comprises an adeno-associated virus (AAV) polypeptide.

7. 10. The composition of any one of the preceding claims, wherein the targeting moiety comprises an adeno-associated virus (AAV) capsid polypeptide.

8. The composition of any one of claims 4 to 7, wherein the n-mer motif is inserted between any two amino acids of the viral polypeptide, viral capsid polypeptide, AAV polypeptide, or AAV capsid polypeptide.

9. 8. The composition of any one of claims 4 to 7, wherein the n-mer motif is inserted into the viral polypeptide, viral capsid polypeptide, AAV polypeptide, or AAV capsid polypeptide such that one, two, or more amino acids at the N-terminus and / or C-terminus of the n-mer motif replace one, two, or more amino acids of the viral polypeptide, viral capsid polypeptide, AAV polypeptide, or AAV capsid polypeptide.

10. 10. The composition of any one of claims 7 to 9, wherein the n-mer motif is inserted between any two consecutive amino acids between amino acids 262-269, 327-332, 382-386, 452-460, 488-505, 527-539, 545-558, 581-593, 598-599, 704-714, or any combination thereof, of the capsid polypeptide of AAV9, or at an analogous position in the capsid polypeptide of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.

10.

11. The composition of any one of claims 7 to 10, wherein the AAV capsid polypeptide is an engineered AAV capsid polypeptide that has reduced or eliminated uptake in non-hematopoietic cells compared to a corresponding wild-type AAV capsid polypeptide.

12. The composition of claim 11 , wherein the non-hematopoietic cells are liver cells.

13. The composition of any one of claims 11 to 12, wherein the capsid polypeptide of the corresponding wild-type AAV is a capsid polypeptide of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.

10.

14. 14. The composition of any one of claims 11 to 13, wherein the engineered AAV capsid polypeptide comprises one or more mutations that result in reduced or eliminated uptake in non-hematopoietic cells.

15. The one or more mutations are in the capsid protein of AAV9 (SEQ ID NO: 12001) a. 267th place, b. 269th place, c. 504th place, d. 505th place, e. 590th place, f. or any combination thereof; or at one or more corresponding positions in a non-AAV9 capsid polypeptide.

16. 16. The composition of claim 15, wherein the non-AAV9 capsid polypeptide is an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10 capsid polypeptide.

17. 17. The composition of any one of claims 15-16, wherein the mutation at position 267 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a G or X to A mutation, wherein X is any amino acid.

18. 18. The composition of any one of claims 15 to 17, wherein the mutation at position 269 of the AAV9 capsid protein (SEQ ID NO: 12001) or a corresponding position in a non-AAV9 capsid polypeptide is an S or X to T mutation, wherein X is any amino acid.

19. 19. The composition of any one of claims 15 to 18, wherein the mutation at position 504 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a G or X to A mutation, wherein X is any amino acid.

20. 20. The composition of any one of claims 15 to 19, wherein the mutation at position 505 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a mutation of P or X to A, wherein X is any amino acid.

21. 21. The composition of any one of claims 15 to 20, wherein the mutation at position 590 of the AAV9 capsid protein (SEQ ID NO: 12001) or a corresponding position in a non-AAV9 capsid polypeptide is a mutation of Q or X to A, wherein X is any amino acid.

22. 17. The composition of any one of claims 15 to 16, wherein the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 267, 269, or both of wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 267 is a G to A mutation and the mutation at position 269 is an S to T mutation.

23. 17. The composition of any one of claims 15 to 16, wherein the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 590 of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), and the mutation at position 509 is a Q to A mutation.

24. 17. The composition of any one of claims 15 to 16, wherein the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 504, 505, or both, of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 504 is a G to A mutation and the mutation at position 505 is a P to A mutation.

25. 25. The composition of any one of claims 1 to 24, which is an engineered viral particle, optionally an engineered AAV particle.

26. 26. The composition of claim 25, wherein the viral particle is an engineered AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10 viral particle.

27. 27. The composition of any one of claims 1 to 26, further comprising a cargo, said cargo being linked to or otherwise associated with said targeting moiety.

28. The cargo is a. is effective in treating or preventing a blood disease or disorder; b. Is effective in treating or preventing a non-hematological disease or disorder; c. It is a vaccine; d. or any combination thereof.

29. A vector system comprising a vector comprising: one or more polynucleotides, at least one of said one or more polynucleotides encoding all or a portion of a targeting moiety effective in targeting a hematopoietic cell, said targeting moiety comprising one or more n-mer motifs, wherein at least one n-mer motif comprises: a. Any one of SEQ ID NOs: 1-1000; b. any one of SEQ ID NOs: 2001-3000; c. any one of SEQ ID NOs: 4001-5000; d. any one of SEQ ID NOs: 6001-7000; e. any one of SEQ ID NOs: 8001-9000; f. any one of SEQ ID NOs: 10001-11000; g. The one or more polynucleotides comprising or consisting of: Optionally, a regulatory element operably linked to one or more of said one or more polynucleotides.

30. 30. The vector system of claim 29, further comprising a cargo polynucleotide, optionally operably linked to at least one polynucleotide encoding all or part of the targeting moiety.

31. 10. A vector system according to any one of the preceding claims, capable of producing a polypeptide comprising or consisting of said targeting moiety.

32. 10. A vector system according to any one of the preceding claims, capable of producing viral polypeptides, optionally viral capsid polypeptides.

33. 10. A vector system according to any one of the preceding claims, capable of producing an adeno-associated virus (AAV) polypeptide, optionally a capsid polypeptide of AAV.

34. 10. A vector system according to any one of the preceding claims, capable of producing viral particles, optionally AAV particles, said viral particles optionally comprising cargo.

35. 35. The vector system of any one of claims 33-34, wherein the AAV polypeptide, AAV capsid polypeptide, and / or AAV particle is an engineered AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.10 viral particle or polypeptide.

36. 36. The vector system of any one of claims 32 to 35, wherein the polypeptide comprises one or more n-mer motifs inserted between two amino acids of the polypeptide, and optionally, the one or more n-mer motifs are inserted such that they are on the outside of the capsid of a virus produced by the vector system.

37. The vector system of claim 36, wherein the AAV capsid polypeptide comprises one or more n-mer motifs inserted between any two consecutive amino acids of the capsid polypeptide of AAV9 AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.

10.

38. 38. The vector system of claim 37, wherein the two consecutive amino acids are independently selected from amino acids 262-269, 327-332, 382-386, 452-460, 488-505, 527-539, 545-558, 581-593, 598-599, 704-714 of the capsid polypeptide of AAV9, or any combination thereof, or at analogous positions in the capsid polypeptide of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.

10.

39. 39. The vector system of any one of claims 33 to 38, wherein the AAV capsid polypeptide is an engineered AAV capsid polypeptide that has reduced or eliminated uptake in non-hematopoietic cells compared to a corresponding wild-type AAV capsid polypeptide.

40. 40. The vector system of claim 39, wherein the non-hematopoietic cells are liver cells.

41. The vector system of any one of claims 39 to 40, wherein the wild-type AAV capsid polypeptide is a capsid polypeptide of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV rh.74, or AAV rh.

10.

42. 42. The vector system of any one of claims 39 to 41, wherein the engineered AAV capsid polypeptide comprises one or more mutations that result in reduced or eliminated uptake in non-hematopoietic cells.

43. The one or more mutations are in the capsid protein of AAV9 (SEQ ID NO: 12001) a. 267th place, b. 269th place, c. 504th place, d. 505th place, e. 590th place, f. or any combination thereof; or at one or more corresponding positions in a non-AAV9 capsid polypeptide.

44. 44. The vector system of claim 43, wherein the non-AAV9 capsid polypeptide is an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV rh.74, or AAV rh.10 capsid polypeptide.

45. A vector system according to any one of claims 43 to 44, a. the mutation at position 267 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a mutation of G or X to A, where X is any amino acid; or b. The mutation at position 269 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a mutation of S or X to T, where X is any amino acid, or c. The mutation at position 504 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a mutation of G or X to A, where X is any amino acid, or d. The mutation at position 505 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a mutation of P or X to A, where X is any amino acid, or e. The mutation at position 590 of the AAV9 capsid protein (SEQ ID NO: 12001) or the corresponding position in a non-AAV9 capsid polypeptide is a mutation of Q or X to A, where X is any amino acid, or f. The engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 267, 269, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 267 is a G to A mutation and the mutation at position 269 is an S to T mutation; g. the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 590 of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 509 is a Q to A mutation; h. the engineered AAV capsid protein is an engineered AAV9 capsid polypeptide comprising a mutation at position 504, 505, or both of the wild-type AAV9 capsid protein (SEQ ID NO: 12001), wherein the mutation at position 504 is a G to A mutation and the mutation at position 505 is a P to A mutation; i. or any acceptable combination thereof.

46. 10. A vector system according to any one of the preceding claims, further comprising a polynucleotide encoding viral rep proteins, optionally AAV rep proteins.

47. 47. The vector system of claim 46, wherein the polynucleotide encoding the viral rep protein is on the same vector as the one or more polynucleotides or on a different vector, optionally operably linked to a regulatory element.

48. A vector system according to any one of the preceding claims, capable of producing a composition or part thereof according to any one of claims 1 to 28.

49. A polynucleotide encoding all or part of the composition of any one of claims 1 to 28.

50. A polypeptide encoded by the vector system of any one of claims 29 to 48, or a polynucleotide of claim 49, or both.

51. 51. The polypeptide of claim 50, wherein the polypeptide is linked to or otherwise associated with a cargo.

52. A particle, optionally a viral particle, produced by a vector system according to any one of claims 29 to 48 and / or a polynucleotide according to claim 49, said particle optionally comprising a polypeptide according to any one of claims 50 to 51.

53. 53. The particle of claim 52, which is an AAV particle.

54. 54. A particle according to any one of claims 52 to 53, comprising a cargo.

55. A particle according to claim 54 or a protein according to claim 51 and / or a vector system according to any one of claims 30 to 49, wherein the cargo or cargo polynucleotide is a. is effective in treating or preventing a blood disease or disorder; b. Is effective in treating or preventing a non-hematological disease or disorder; c. It is a vaccine; d. or any combination thereof.

56. A cell comprising: A composition according to any one of the preceding claims, a vector system according to any one of the preceding claims, a polypeptide according to any one of the preceding claims, a particle according to any one of the preceding claims, or any combination thereof.

57. 57. The cell of claim 56, which is a hematopoietic cell.

58. 57. The cell of claim 56, which is a prokaryotic or eukaryotic cell.

59. A pharmaceutical formulation comprising: A composition according to any one of the preceding claims, a vector system according to any one of the preceding claims, a polypeptide according to any one of the preceding claims, a particle according to any one of the preceding claims, a cell according to any one of the preceding claims, or any combination thereof, and A pharmaceutically acceptable carrier.

60. said method in a subject in need of treatment or prevention of a disease, optionally a blood disorder, or a symptom thereof, comprising: The method comprises administering to a subject in need thereof a composition described in any one of the preceding claims, a vector system described in any one of the preceding claims, a polypeptide described in any one of the preceding claims, a particle described in any one of the preceding claims, a cell described in any one of the preceding claims, a pharmaceutical formulation described in any one of the preceding claims, or any combination thereof.

61. A targeting moiety effective in targeting hematopoietic cells, comprising one or more n-mer motifs, wherein at least one n-mer motif is VKX. n and X n are each selected from any amino acid, and n is 5; and A composition optionally comprising a cargo, said cargo being linked to or otherwise associated with said targeting moiety.

62. A targeting moiety effective in targeting hematopoietic cells, comprising one or more n-mer motifs, wherein at least one n-mer motif is VKX. n YGAL, or consisting of X n are each selected from any amino acid, and n is 1; and A composition optionally comprising a cargo, said cargo being linked to or otherwise associated with said targeting moiety.