Cas9 fusion proteins with internal NLS: compositions and methods

Internal insertion of hairpin NLS sequences in Cas9 proteins addresses efficiency and recovery issues, enhancing genome editing precision and reducing off-target effects in CRISPR-Cas9 systems.

WO2026050560A1PCT designated stage Publication Date: 2026-03-05RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/044038
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-30
Filing Date
2025-08-28
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing CRISPR-Cas9 systems face challenges in achieving high editing efficiency with low, transient doses, particularly in therapeutic applications, due to limitations in protein recovery and nuclear localization efficiency when fused with multiple NLS motifs.

Method used

Incorporation of internal hairpin NLS sequences (hiNLS) at specific sites within the Cas9 protein backbone enhances editing efficiency and protein yield, improving the delivery and localization of CRISPR enzymes.

Benefits of technology

The internal insertion of hiNLS modules in Cas9 proteins results in improved editing efficiency and purity, facilitating precise genome editing with reduced off-target effects and minimizing immune response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000017_0001
    Figure IMGF000017_0001
  • Figure IMGF000018_0001
    Figure IMGF000018_0001
  • Figure IMGF000120_0001
    Figure IMGF000120_0001
Patent Text Reader

Abstract

Provided are Cas9 fusion polypeptides that include (i) a Cas9 protein; and (ii) an internal insertion of one or more linker-NLS1-linker-NLS2-linker (hiNLS) modules, as well as methods that employ such polypeptides. In some cases, the one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205, K468, I1022, or S1248, of SEQ ID NO: 1, or any combination thereof. In some cases, the one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to a residue selected from positions: 202-208, 465-471, 1019-1025, and 1245-1251 of the Cas9 protein set forth in SEQ ID NO: 1, or any combination thereof.
Need to check novelty before this filing date? Find Prior Art

Description

CAS9 FUSION PROTEINS WITH INTERNAL NLS: COMPOSITIONS AND METHODSCROSS-REFERENCE

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 689,450 filed August 30, 2024, which application is incorporated herein by reference in its entirety.STATEMENT REGARDING RESEARCH FUNDING

[0002] Research for this invention was supported by awards from the Cystic Fibrosis Foundation.INCORPORATION BY REFERENCE OF SEQUENCE LISTING PROVIDED AS AN XM L FILE

[0003] A Sequence Listing is provided herewith as a Sequence Listing XML, “BERK- 529WO_SEQLIST.xml” created on August 27, 2025 and having a size of 141 ,559 bytes. The contents of the Sequence Listing XML are incorporated by reference herein in their entirety.I. INTRODUCTION

[0004] Genome editing and modification technologies have experienced a transformative evolution with the emergence of CRISPR-mediated engineering, offering unprecedented precision in modifying genetic material. The capacity to manipulate the genome using these programmable enzymes has led to groundbreaking advances in various fields, including biology, medicine, and agriculture. However, challenges persist, particularly in the context of optimizing editing efficiency of enzymes such as the widely used nuclease Cas9 (e.g., S. pyogenes Cas9). This is especially true in therapeutic use cases, when it would be ideal to attain high rates of editing via a low, transient dose of the enzyme in the ribonucleoprotein (RNP) format used for multiple ex vivo clinical trials.

[0005] One of the first ways Cas9 was engineered for efficient eukaryotic genome editing was to fuse nuclear localization signal (NLS) sequences to one or both Cas9 termini. NLS motifs are short (7-16 amino acids) sequences found in many eukaryotic nuclear proteins as well as viral proteins, with the latter (e.g. SV40) typically having exceptional potency for nuclear import. Proteins containing NLS motifs are actively imported into the nucleus by endogenous proteins such as importin-a, and increased NLS density is typically associated with increased nuclear localization, both in termsof speed and partitioning. Because many CRISPR enzymes are of bacterial origin, fusion to NLS motifs can greatly enhance editing efficiency.

[0006] An NLS-rich Cas9 variant bearing six SV40 motifs at its termini - surgically administered in RNP format - exhibited the capacity for self-delivery into neurons of the mouse brain and has since demonstrated activity comparable to an adeno- associated (AAV)-Cas9 delivery system in the same murine context. The same 6 SV40 Cas9 construct was a top performer in peptide-enabled RNP delivery for CRISPR engineering (PERC) in human primary lymphocytes, as was a Cas9 construct bearing a trio of different NLS motifs, referred as “Cas9-triNLS”. A similar study of human and murine lymphocytes found that peptide-mediated delivery of Cas9 or Cas12a worked well with protein constructs featuring 6-8 terminal NLS motifs.

[0007] Electroporation of Cas9 RNP has been the basis for several clinical trials, as well as the first-ever approved CRISPR therapy, and even via this potent route of delivery there is evidence that NLS engineering can be used to attain optimal editing rates following electroporation. RNP delivery is appealing for therapeutic use because it is transient, which minimizes off-target editing and decreases risks of toxic, cell- mediated immune response to the microbial protein. The transient nature of RNP may explain why NLS density has a pronounced impact on efficiency in this context: the enzyme must quickly reach the nucleus so that it can induce editing before it is metabolized by the cell. The RNP format is also appealing for therapeutic use because it can be readily manufactured: solid-state synthesis of guide RNA (gRNA) is straightforward, as is recombinant production of CRISPR proteins. However, CRISPR protein yields can decrease - sometimes dramatically - if the construct bears too many NLSs.

[0008] There is a need for compositions and methods that provide Cas9 proteins with increased efficiency, and such is provided herein.II. SUMMARY

[0009] Fusion of nuclear localization signal (NLS) sequences at one or both termini of CRISPR enzymes is a widely adopted strategy to promote efficient genome editing. Engineered variants of CRISPR enzymes with diverse NLS sequences have demonstrated superior performance, expediting nuclear localization and enabling precise DNA editing. However, NLS fusion to the CRISPR protein’s termini can have limitations related to low protein recovery via recombinant expression.

[0010] The work described in the experimental examples below led to the surprising finding by the inventors that internal insertion of hairpin internal NLS sequences (hiNLS) fused at rationally selected sites within the backbone of the Cas9 protein enhances the protein’s editing efficiency. The inventors evaluated the performance of hiNLS Cas9 variants by editing genes in human primary T cells following delivery of ribonucleoproteins (RNPs) via electroporation or incubation with amphiphilic peptides. The work described below demonstrated that subject hiNLS Cas9 variants can improve editing efficiency and can be produced with high purity and yield. These findings represent a key advance in improving CRISPR-Cas effector protein design.

[0011] The present disclosure provides Cas9 fusion polypeptides that include (i) a Cas9 protein; and (ii) an internal insertion of one or more linker-NLS1-linker-NLS2-linker (hiNLS) modules, as well as methods that employ such polypeptides. In some cases, the one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248, of SEQ ID NO: 1 , or any combination thereof. In some cases, the one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to a residue selected from positions: 202-208, 465-471 , 1019-1025, and 1245-1251 of the Cas9 protein set forth in SEQ ID NO: 1 , or any combination thereof. The present disclosure provides nucleic acids encoding a subject fusion polypeptide, as well as cells comprising such polypeptides and / or nucleic acids. The present disclosure provides compositions that include a subject fusion polypeptide (and / or a nucleic acid encoding same) and a Cas9 guide RNA (and / or a nucleic acid encoding the guide RNA). The present disclosure provides methods of binding, modifying, and / or modulating transcription of a target nucleic acid (e.g., target DNA), involving use of a subject fusion polypeptide and a guide RNA.

[0012] Reagents, compositions, and kits / systems that find use in practicing the subject methods are provided.III. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The following detailed description of embodiments of the invention will be better understood when read in conjunction with the appended drawings. It should be understood that the invention is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings.

[0014] FIG. 1A-1C Amino acid and surface positions of NLS in hiNLS constructs used in primary human T cell editing. (FIG. 1a) Surface representation of the Cas9-sgRNA- dsDNA structure with designated hiNLS installation sites. hiNLS were installed at amino acid residues G205, K468, 11022, and S1248 (magenta) inside Cas9. The target DNA strand and displaced nontarget strand are colored dark blue and purple, respectively. The sgRNA sequence is highlighted in orange. (FIG. 1 b) Cartoon illustration of Cas9 terminal and internal positions of NLS. hiNLS were installed inside two Cas9 scaffolds with terminal NLS sequences: s-Cas9 bears a single SV40 motif at its C-terminus, and t-Cas9 bears three different NLS motifs fused to the termini. Numbered Cas9 amino acid residues with positions labeled 1-4 (circled) where hiNLS motifs were inserted. (FIG. 1c) Cas9 constructs used in T cell editing. Each scaffold bearing terminal NLS is denoted in lower case letters (s, SV40; t, triNLS). Type of hiNLS is denoted in upper case letters (S, SV40; M, c-Myc). The number following each capital letter (S or M) corresponds to hiNLS position in the Cas9 backbone.

[0015] FIG. 2A-2C hiNLS Cas9 RNP knockout efficiencies in primary human T cells. hiNLS Cas9 RNPs were delivered via electroporation (e-por) and peptides (PERC) targeting genomic loci (FIG. 2a) B2M, (FIG. 2b) TRAC, and (FIG. 2c) CD3. Knockout efficiency was assayed 4 days after delivery by flow cytometry to evaluate surface expression. 20 pmol and 50 pmol of RNP were delivered by e-por and PERC, respectively (each left bar is result for e-por and each right bar is result for PERC). Horizontal dotted line marks the editing efficiencies of electroporated s-Cas9. Lighter shaded bars represent terminally-fused constructs, s-Cas9 and t-Cas9. NT, nontreated.

[0016] FIG. 3A-3B hiNLS Cas9 RNP knockout (FIG. 3a). Primary human T cell viability assessed at day 2 and is reported in relative light units (RLU). Cas9 B2M-RNP was delivered using peptides or electroporation. Mock is electroporation only (no RNP). (FIG. 3b) Representative flow cytometry gating to assess knockout of B2M surface expression in CD4+ T cells. Data are related to Fig. 2. Dnr, donor.

[0017] FIG. 4 Primary human T cell viability assessed at day 2 and is reported in relative light units (RLU). Cas9 TRAC-RNP was delivered using peptides or electroporation. Mock is electroporation only (no RNP). Data are related to Fig. 4. NT, nontreated.

[0018] FIG. 5A-5C. Sequence and position of hiNLS motifs incorporated into CRISPR- Cas9. (a) Model of Cas9 ribonucleoprotein (protein, gray; RNA, orange) in complex with a DNA substrate (pink) showing the hiNLS installation sites (black spheres) as well as the N- and C-termini (white spheres). Each hiNLS module was installed at amino acid residues G205 (“1”), K468 (“2”), 11022 (“3”), and / or S1248 (“4”) in thebackbone of S. pyogenes Cas9. (b) Primary structure representation of two previously-reported Cas9 constructs that were used as the backbone for generation of hiNLS constructs. Colored bars represent NLS sequences: SV40 (S, red); c-Myc (M, blue); nucleoplasmin (gray), (c) Primary sequence of NLS motifs and hiNLS modules. Flexible linkers (italics) separate the tandem NLS motifs from each other and from the Cas9 backbone.

[0019] FIG. 6. Properties of hiNLS Cas9 constructs. Summary of NLS composition and properties of all the hiNLS constructs designed and tested, alongside s-Cas9, t-Cas9, and 6xNLS-Cas9. The last column shows an expected average yield of protein expressed and purified for each construct. Each scaffold bearing terminal NLS is denoted in lower case letters (s, SV40; t, triNLS). Type of hiNLS module is denoted in upper case letters (S, SV40; M, c-Myc). The number following each capital letter (S or M) corresponds to hiNLS position in the Cas9 backbone. “+” denotes RNPs that were soluble only in the trehalose / arginine buffer, whereas “++” indicates RNPs that remained soluble in the trehalose / arginine buffer and a standard buffer (lacking additional stabilizers).

[0020] FIG. 7A-7E. Relative editing efficiencies mediated by hiNLS Cas9 in primary human T cells. Cas9 B2 / W-RNPs and TRAC-RNPs were delivered via peptides (PERC) (a,c) or electroporation (EP) (b,d). Knockout efficiencies of B2 / W-RNPs (green bars) and TRAC-RNPs (orange bars) were assayed 4 days after delivery by flow cytometry to evaluate B2M surface expression. Knockout efficiencies are represented as fold change compared to the average editing efficiency in the experiment. The hiNLS constructs (colored bars) were compared to previously-reported constructs lacking hiNLS modules: s-Cas9 (white bars) and t-Cas9 (gray bars). 50 pmol and 20 pmol of RNP were delivered by PERC and EP, respectively. NT, non-treated. n=4 distinct T cell donors for B2M editing; n=2 distinct T cell donors for TRAC editing, (e) schematic of the constructs tested in FIG. 7A-7D.

[0021] FIG. 8A-8G. The impact of NLS amount and diversity on editing outcomes and recombinant protein yield, (a) A red-to-blue color gradient representing the ratio of c-Myc:SV40 NLS content within each Cas9 construct. The blue and red denote the maximum relative c-Myc and SV40 content, respectively, (b) Distribution of B2M knockout efficiencies via peptide-enabled delivery (PERC) as measured by fold change (compared to the average editing efficiency) as it relates to the total number of NLS motifs in each construct, or (c) the total number of positive charges contributed by NLS motifs, (d) Distribution of B2M knockout efficiencies via electroporation (EP)as measured by fold change (compared to the average editing efficiency) as it relates to the total number of NLS motifs in each construct, or (e) the total number of positive charges contributed by NLS motifs, (f) Comparison of protein yields from recombinantly expressed hiNLS and terminally-fused NLS Cas9 constructs (outlined shapes) corresponding to the total number of NLS motifs in each construct, or (g) the total number of positive charges contributed by NLS motifs. The nucleoplasmin NLS that is present in the t-Cas9 backbone is accounted for in the count of NLS motifs or positive charges in the Cas9 constructs, but it is not considered as part of the color gradient. Labels are applied to terminally-fused NLS Cas9 constructs and the seven most active hiNLS constructs of each editing dataset.

[0022] FIG. 9. Ranked editing performance of hiNLS Cas9 constructs via peptide- mediated delivery. The B2M knockout efficiency, with hiNLS constructs presented according to their performance. The Cas9 B2 / W-RNPs were delivered via peptides (PERC), with knockout efficiencies assayed 4 days after delivery by flow cytometry to evaluate surface expression. Here, knockout efficiencies are represented as fold change compared to the average editing efficiency in the experiment. 50 pmol of RNP were delivered. Single asterisk, P < 0.05; double asterisk, P < 0.01. NT, nontreated; n=4 distinct T cell donors. These values are derived from the same experiment underlying Fig. 7 and Figures 10 and 11.

[0023] FIG. 10A-10C. PERC-mediated B2M editing efficiencies of hiNLS constructs in primary human T cells. A single experiment wherein hiNLS Cas9 B2 / W-RNPs were delivered via peptides (PERC), with and knockout efficiencies (a) assessed by flow cytometry to evaluate surface expression and (b) represented as fold change compared to the average editing efficiency. 50 pmol of RNP were delivered. The data in panel (b) are also presented in Fig. 5a. (c) Primary human T cells were stained with live / dead stain and surface marker- targeting antibodies and sampled at defined volumes (60 pl per well) to quantify cell counts. Green bars represent novel hiNLS Cas9 constructs; white and gray bars represent previously published constructs s- Cas9 and t-Cas9, respectively. Rightmost bars: no treatment (NT).

[0024] FIG. 11A-11C. Electroporation-mediated B2M editing efficiencies of hiNLS constructs in primary human T cells. A single experiment wherein hiNLS Cas9 82 / W-RNPs were delivered via electroporation (EP), with knockout efficiencies (a) assessed by flow cytometry to evaluate surface expression and (b) represented as fold change compared to the average editing efficiency. 20 pmol of RNP were delivered. The data in panel (b) are also presented in Fig. 5b. (c) Primary human Tcells were stained with live / dead stain and surface marker-targeting antibodies and sampled at defined volumes (60 pl per well) to quantify cell counts. Green bars represent novel hiNLS Cas9 constructs; white and gray bars represent previously published constructs s-Cas9 and t-Cas9, respectively.

[0025] FIG. 12A-12C. PERC-mediated TRAC editing efficiencies of hiNLS constructs in primary human T cells. A single experiment wherein hiNLS Cas9 TRAC-RNPs were delivered via peptides (PERC), with knockout efficiencies (a) assessed by flow cytometry to evaluate surface expression and (b) represented as fold change compared to the average editing efficiency. 50 pmol of RNP were delivered. The data in panel (b) are also presented in Fig. 5c. (c) Primary human T cells were stained with live / dead stain and surface marker-targeting antibodies and sampled at defined volumes (60 pl per well) to quantify cell counts. Orange bars represent novel hiNLS Cas9 constructs; white and gray bars represent previously published constructs s- Cas9 and t-Cas9, respectively. Rightmost bars: no treatment (NT).

[0026] FIG. 13A-13C. Electroporation-mediated TRAC editing efficiencies of hiNLS constructs in primary human T cells. A single experiment wherein hiNLS Cas9 TRAC-RNPs were delivered via electroporation (EP), with knockout efficiencies (a) assessed by flow cytometry to evaluate surface expression and (b) represented as fold change compared to the average editing efficiency. 20 pmol of RNP were delivered. The data in panel (b) are also presented in Fig. 5d. (c) Primary human T cells were stained with live / dead stain and surface marker-targeting antibodies and sampled at defined volumes (60 pl per well) to quantify cell counts. Orange bars represent novel hiNLS Cas9 constructs; white and gray bars represent previously published constructs s-Cas9 and t-Cas9, respectively. Rightmost bars: no treatment.

[0027] FIG. 14A-14C. On- and off-target editing at the EMX1 locus. The addition of hiNLS modules affects both on-target editing efficiency and off-target activity, with specific hiNLS configurations enhancing overall editing performance and specificity, (a) On- target and (b) off-target editing rates measured as the indel frequency (% detected using NGS) at the EMX1 target site, (c) Off-target to on-target editing ratio, a metric of enzyme specificity. Error bars represent standard deviation. The X-axis label illustrates the hiNLS configurations of Cas9 constructs. The triNLS variant includes three NLS sequences, the 6xNLS variant has six repeats of NLS, while t-M4 and t- M1M4 variants incorporate specific hiNLS modules. Color coding denotes the NLS configurations (See Fig. 5). NT, nontreated.IV. DEFINITIONS

[0028] “Heterologous,” as used herein, means a nucleotide or polypeptide sequence that is not found in the native nucleic acid or protein, respectively. For example, a Cas9 fusion polypeptide of the present disclosure includes a Cas9 protein with one or more linker-NLS1-linker-NLS2-linker (hiNLS) modules inserted internally within the Cas9 protein. The hiNLS is considered a heterologous polypeptide - it is heterologous to the Cas9 protein (the hiNLS includes amino acid sequence from a protein other than the Cas9 protein). Likewise, in some cases a non-CRISPR-Cas effector protein (e.g., a histone deacetylase) can be fused to a subject Cas9 fusion polypeptide (e.g., the Cas9 protein can be internally fused to one or more hiNLS modules and also fused, e.g., at or near an N- or C- terminus, to a heterologous protein that provides an activity such as a DNA modifying activity, a protein modifying activity, transcription modulation activity, etc.). In such a case the non-CRISPR-Cas effector protein (e.g., a histone deacetylase) is considered heterologous to the Cas9 protein (i.e. , includes amino acid sequence from a protein other than the Cas9 protein). As another example, a guide sequence of a guide RNA can be heterologous to the proteinbinding sequence (a scaffold) of the guide RNA - as such, the guide sequence is not found in nature together with the protein-binding sequence (the scaffold).

[0029] The terms “polynucleotide” and “nucleic acid,” used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxynucleotides. Thus, this term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The terms “polynucleotide” and “nucleic acid” should be understood to include, as applicable to the embodiment being described, single-stranded (such as sense or antisense) and double-stranded polynucleotides.

[0030] The term “naturally-occurring” as used herein as applied to a protein, a nucleic acid, a cell, or an organism, refers to a protein, a nucleic acid, cell, or organism that is found in nature. For example, a polypeptide or polynucleotide sequence that is present in an organism (including viruses) that can be isolated from a source in nature is naturally occurring.

[0031] As used herein the term “isolated” is meant to describe a polynucleotide, a polypeptide, or a cell that is in an environment different from that in which the polynucleotide, the polypeptide, or the cell naturally occurs. An isolated geneticallymodified host cell may be present in a mixed population of genetically modified host cells.

[0032] As used herein, the term “exogenous nucleic acid” (sometimes referred to as a ‘heterologous nucleic acid’) refers to a nucleic acid that is not normally or naturally found in and / or produced by a given bacterium, organism, or cell in nature. As used herein, the term “endogenous nucleic acid” refers to a nucleic acid that is normally found in and / or produced by a given bacterium, organism, or cell in nature. An “endogenous nucleic acid” is also referred to as a “native nucleic acid” or a nucleic acid that is “native” to a given bacterium, organism, or cell.

[0033] “Recombinant,” as used herein, means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps resulting in a construct having a structural coding or non-coding sequence distinguishable from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding the structural coding sequence can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide a synthetic nucleic acid which is capable of being expressed from a recombinant transcription unit contained in a cell or in a cell-free transcription and translation system. Such sequences can be provided in the form of an open reading frame uninterrupted by internal non-translated sequences, or introns, which are typically present in eukaryotic genes. Genomic DNA comprising the relevant sequences can also be used in the formation of a recombinant gene or transcription unit. Sequences of non-translated DNA may be present 5’ or 3’ from the open reading frame, where such sequences do not interfere with manipulation or expression of the coding regions, and may indeed act to modulate production of a desired product by various mechanisms (see “DNA regulatory sequences”, below).

[0034] Thus, e.g., the term “recombinant” polynucleotide or “recombinant” nucleic acid refers to one which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of sequence through human intervention. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques. Such can be done to replace a codon with a redundant codon encoding the same or a conservative amino acid, while typically introducing or removing a sequence recognition site. Alternatively, it is performed to join together nucleic acid segments of desired functions to generate a desired combination of functions. This artificial combination is often accomplished by either chemicalsynthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques.

[0035] Similarly, the term “recombinant” polypeptide refers to a polypeptide which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of amino sequence through human intervention. Thus, e.g., a polypeptide that comprises a heterologous amino acid sequence is recombinant.

[0036] By “construct” is meant a recombinant nucleic acid, generally recombinant DNA, which has been generated for the purpose of the expression and / or propagation of a specific nucleotide sequence(s), or is to be used in the construction of other recombinant nucleotide sequences.

[0037] A "vector" or “expression vector” is a replicon, such as plasmid, phage, virus, or cosmid, to which another DNA segment, i.e. an “insert”, may be attached so as to bring about the replication of the attached segment in a cell.

[0038] The terms “recombinant expression vector,” or “DNA construct” are used interchangeably herein to refer to a DNA molecule comprising a vector and at least one insert. Recombinant expression vectors are usually generated for the purpose of expressing and / or propagating the insert(s), or for the construction of other recombinant nucleotide sequences. The insert(s) may or may not be operably linked to a promoter sequence and may or may not be operably linked to DNA regulatory sequences.

[0039] The terms “DNA regulatory sequences,” “control elements,” and “regulatory elements,” used interchangeably herein, refer to transcription and translational control sequences, such as promoters, enhancers, polyadenylation signals, terminators, protein degradation signals, and the like, that provide for and / or regulate expression of a coding sequence and / or production of an encoded polypeptide in a host cell.

[0040] An “expression cassette” comprises a DNA coding sequence operably linked to a promoter. "Operably linked" refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. For instance, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression (the coding sequence can also be said to be operably linked to the promoter).

[0041] The term “transformation” is used interchangeably herein with “genetic modification” and refers to a permanent or transient genetic change induced in a cell following introduction of new nucleic acid (i.e., DNA exogenous to the cell). Genetic change (“modification”) can be accomplished either by incorporation of the new DNA into thegenome of the host cell, or by transient or stable maintenance of the new DNA as an episomal element. Where the cell is a eukaryotic cell, a permanent genetic change is generally achieved by introduction of the DNA into the genome of the cell. In prokaryotic cells, permanent changes can be introduced into the chromosome or via extrachromosomal elements such as plasmids and expression vectors, which may contain one or more selectable markers to aid in their maintenance in the recombinant host cell. Suitable methods of genetic modification include viral infection, transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, lipid nanoparticle delivery, and the like. The choice of method is generally dependent on the type of cell being transformed and the circumstances under which the transformation is taking place (i.e. in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al, Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.

[0042] “Operably linked” refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. For instance, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression (The coding sequence can also be said to be operably linked to the promoter). As used herein, the terms “heterologous promoter” and “heterologous control regions” refer to promoters and other control regions that are not normally associated with a particular nucleic acid in nature. For example, a “transcription control region heterologous to a coding region” is a transcription control region that is not normally associated with the coding region in nature.

[0043] A “host cell,” as used herein, denotes an in vivo or in vitro eukaryotic cell, a prokaryotic cell, or a cell from a multicellular organism (e.g., a cell line) cultured as a unicellular entity, which eukaryotic or prokaryotic cells can be, or have been, used as recipients for a nucleic acid (e.g., an expression vector that comprises a nucleotide sequence encoding a subject Cas9 fusion polypeptide), and include the progeny of the original cell which has been genetically modified by the nucleic acid. It is understood that the progeny of a single cell may not necessarily be completely identical in morphology or in genomic or total DNA complement as the original parent, due to natural, accidental, or deliberate mutation. A “recombinant host cell” (also referred to as a “genetically modified host cell”) is a host cell into which has been introduced a heterologous nucleic acid, e.g., an expression vector. For example, a subject prokaryotic host cell is a genetically modified prokaryotic host cell (e.g., a bacterium), by virtue of introduction into a suitable prokaryotic host cell of aheterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to (not normally found in nature in) the prokaryotic host cell, or a recombinant nucleic acid that is not normally found in the prokaryotic host cell; and a subject eukaryotic host cell is a genetically modified eukaryotic host cell, by virtue of introduction into a suitable eukaryotic host cell of a heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the eukaryotic host cell, or a recombinant nucleic acid that is not normally found in the eukaryotic host cell.

[0044] The term “conservative amino acid substitution” refers to the interchangeability in proteins of amino acid residues having similar side chains. For example, a group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; a group of amino acids having aliphatic-hydroxyl side chains consists of serine and threonine; a group of amino acids having amide-containing side chains consists of asparagine and glutamine; a group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; a group of amino acids having basic side chains consists of lysine, arginine, and histidine; and a group of amino acids having sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are: valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.

[0045] Percent complementarity between particular stretches of nucleotide sequences within nucleic acids and / or amino acid sequences within proteins can be determined routinely, e.g., using convenient BLAST programs (basic local alignment search tools) and PowerBLAST programs known in the art (e.g., Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656) (e.g., BLASTP, BLASTN) (e.g., blast.ncbi.nlm.nih.gov / Blast.cgi).

[0046] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0047] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excludedlimit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0048] Certain ranges are presented herein with numerical values being preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.

[0049] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.

[0050] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.

[0051] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. As such, the articles “a” and “an” are used herein to refer to one or to more than one (i.e. , to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Thus, for example, reference to “a cell” includes a plurality of such cells and reference to “the polypeptide” includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.

[0052] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. For example, it is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0053] While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in any way by the construction of "means" or "steps" limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112.V. DETAILED DESCRIPTION

[0054] As noted above, the present disclosure provides Cas9 fusion polypeptides that include (i) a Cas9 protein; and (ii) an internal insertion of one or more linker-NLS1- linker-NLS2-linker (hiNLS) modules, as well as methods that employ such polypeptides. The one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248, of SEQ ID NO: 1 , or any combination thereof. The present disclosure provides nucleic acids encoding a subject fusion polypeptide, as well as cells comprising such polypeptides and / or nucleic acids. The present disclosure provides compositions that include a subject fusion polypeptide (and / or anucleic acid encoding same) and a Cas9 guide RNA (and / or a nucleic acid encoding the guide RNA). The present disclosure provides methods of binding, modifying, and / or modulating transcription of a target nucleic acid (e.g., target DNA), involving use of a subject fusion polypeptide and a guide RNA. hiNLS modules

[0055] A subject Cas9 fusion polypeptide includes one or more linker-NLS1-linker-NLS2- linker (hiNLS) modules inserted internally within a Cas9 protein.Nuclear localization signals (NLSs)

[0056] In some cases, the first nuclear localization signal (NLS), referred to herein as NLS1 , and the second NLS, referred to herein as NLS2, are the same sequence. In such cases, the hiNLS module can be referred to as linker-NLS-linker-NLS-linker. In some cases, the first and second NLSs (i.e., NLS1 and NLS2) are not the same - they are different from one another, i.e., they differ by amino acid sequence.

[0057] An NLS (NLS1 and / or NLS2) can be any convenient NLS. Examples of NLSs will be known to one of ordinary skill in the art and any convenient NLS can be used for NLS1 , just as any convenient NLS can be used for NLS2. Examples of NLSs include, but are not limited to, those listed in Table 1. In some embodiments, NLS1 and NLS2 are each independently selected from Table 1. Additional examples include, but are not limited to: RQRRNELKRSP (SEQ ID NO: 49); the hRNPAI M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 50); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 51) of the IBB domain from importin-alpha; the sequences VSRKRPRP (SEQ ID NO: 52) and PPKKARED (SEQ ID NO: 53) of the myoma T protein; the sequence PQPKKKPL (SEQ ID NO: 54) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 55) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO: 56) and PKQKKRK (SEQ ID NO: 57) of the influenza virus NS1 ; the sequence RKLKKKIKKL (SEQ ID NO: 58) of the Hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 59) of the mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 60) of the human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 61) of the steroid hormone receptors (human) glucocorticoid. For additional NLS sequences see, e.g., Lu et al., Cell Commun Signal. 2021 May 22; 19(1 ):60; and Kosugi et al., J Biol Chem. 2009 Jan 2;284(1):478-485.

[0058] As a non-limiting illustrative example, in some cases, NLS1 and NLS2 are different, NLS1 is PKKKRKV (SEQ ID NO: 41), and NLS2 is selected from Table 1. Likewise, in some cases, NLS1 and NLS2 are different, NLS2 is PKKKRKV (SEQ ID NO: 41), and NLS1 is selected from Table 1. In some cases, NLS1 and NLS2 are different, NLS1 is PAAKKKKLD (SEQ ID NO: 42), and NLS2 is selected from Table 1. Likewise, in some cases, NLS1 and NLS2 are different, NLS2 is PAAKKKKLD (SEQ ID NO: 42), and NLS1 is selected from Table 1.

[0059] In some embodiments, NLS1 and NLS2 are the same NLS, which is selected from Table 1. For example, in some embodiments, NLS1 and NLS2 are PKKKRKV (SEQ ID NO: 41). In some cases, NLS1 and NLS2 are PAAKKKKLD (SEQ ID NO: 42). In some embodiments, NLS1 and NLS2 are KRPAATKKAGQAKKKK (SEQ ID NO: 43). In some cases, NLS1 and NLS2 are PAAKRVKLD (SEQ ID NO: 44). In some embodiments, NLS1 and NLS2 are KRTADGSEFESPKKKRKVE (SEQ ID NO: 45). In some cases, NLS1 and NLS2 are KRTADGSEFESPKKARKVE (SEQ ID NO: 46). In some embodiments, NLS1 and NLS2 are KRTADGSEFESPKKKAKVE (SEQ ID NO: 47). In some cases, NLS1 and NLS2 are KR(X1)5-I5KK(X2)(X3)KV (SEQ ID NO: 48), where X1 can be any amino acid, the “5-15” denotes a string of from 5-15 X1 residues, X2 is lysine or alanine, and X3 is lysine, arginine, or alanine.

[0060] Table 1. Example NLS Amino acid sequences. Linker sequences are italicized.Notes:• BP-SV40C is a BP-SV40 bipartite NLS consensus sequence. X1 can be any amino acid, and the “5-15” denotes a string of from 5-15 X residues. In some cases, X2 is lysine or alanine. In some cases, X3 is lysine, arginine, or alanine.• Basic charges (R,H,K): SV40 - 5; c-Myc - 4; NP - 8• [ADE] is any amino acid except Asp or GluLinker(s) of an hiNLS module

[0061] Each of the three linkers of an hiNLS module (which can be referred to as linkerl , Iinker2, and linkers - for example: Iinker1-NLS1-Iinker2-NLS2-Iinker3) can be the same or different from one another. In some cases, all three of the linkers are the same. In some cases, two of the linkers are the same (where the different linker may be in position linkerl , Iinker2 or Iinker3). In some cases, the 3 linkers are different from one another. Flexible amino acid linkers will be known to one of ordinary skill in the art and any convenient linker can be used. Examples of linkers that can be used as part of an hiNLS module include, but are not necessarily limited to: glycine polymers (G)n, glycine-serine polymers (including, for example, (GS)n, GSGGSn (SEQ ID NO: 82), GGSGGSn (SEQ ID NO: 83), and GGGSn (SEQ ID NO: 84), where n is an integer of at least one), glycine-alanine polymers, and alanine-serine polymers. Example linkers can comprise amino acid sequences such as: GGSG (SEQ ID NO: 77), GGSGG (SEQ ID NO: 78), GSGSG (SEQ ID NO: 76), GSGGG (SEQ ID NO: 79), GGGSG (SEQ ID NO: 80), GSSSG (SEQ ID NO: 81), and the like. As such, in some cases, all three of the linkers are independently selected from GGSG (SEQ ID NO: 77), GGSGG (SEQ ID NO: 78), GSGSG (SEQ ID NO: 76), GSGGG (SEQ ID NO: 79), GGGSG (SEQ ID NO: 80), and GSSSG (SEQ ID NO: 81).

[0062] In some cases, at least one of the linkers of a given hiNLS module is GSGSG (SEQ ID NO: 76). In some cases, at least two of the linkers of a given hiNLS module are GSGSG (SEQ ID NO: 76). In some cases, all 3 of the linkers of a given hiNLS module are GSGSG (SEQ ID NO: 76). See, e.g., the example hiNLS module sequences of Table 1.

[0063] In some cases, at least one of the linkers of a given hiNLS module is GGSG (SEQ ID NO: 77). In some cases, at least two of the linkers of a given hiNLS module are GGSG (SEQ ID NO: 77). In some cases, all 3 of the linkers of a given hiNLS module are GGSG (SEQ ID NO: 77).

[0064] In some cases, at least one of the linkers of a given hiNLS module is GGSGG (SEQ ID NO: 78). In some cases, at least two of the linkers of a given hiNLS module are GGSGG (SEQ ID NO: 78). In some cases, all 3 of the linkers of a given hiNLS module are GGSGG (SEQ ID NO: 78).

[0065] In some cases, at least one of the linkers of a given hiNLS module is GSGGG (SEQ ID NO: 79). In some cases, at least two of the linkers of a given hiNLS module are GSGGG (SEQ ID NO: 79). In some cases, all 3 of the linkers of a given hiNLS module are GSGGG (SEQ ID NO: 79).

[0066] In some cases, at least one of the linkers of a given hiNLS module is GGGSG (SEQ ID NO: 80). In some cases, at least two of the linkers of a given hiNLS module are GGGSG (SEQ ID NO: 80). In some cases, all 3 of the linkers of a given hiNLS module are GGGSG (SEQ ID NO: 80).

[0067] In some cases, at least one of the linkers of a given hiNLS module is GSSSG (SEQ ID NO: 81). In some cases, at least two of the linkers of a given hiNLS module are GSSSG (SEQ ID NO: 81). In some cases, all 3 of the linkers of a given hiNLS module are GSSSG (SEQ ID NO: 81).Cas9 Fusion Polypeptides (internal insertions)

[0068] A subject Cas9 fusion polypeptide includes one or more hiNLS modules (in some cases one hiNLS) inserted within a Cas9 protein. For example, the inventors discovered several sites that tolerate insertions without negatively affecting the function of the Cas9 protein (e.g., improving the efficiency of cleaving of a target nucleic acid and / or the binding of a target nucleic acid).Cas9 protein

[0069] Cas9 proteins are known in the art and any convenient Cas9 protein can be used as part of a subject Cas9 fusion polypeptide. Because the Cas9 fusion polypeptide is the result of an internal insertion into a Cas9 protein, the Cas9 protein prior to insertion (i.e., in the absence of the inserted one or more hiNLS modules) is sometimes referred to herein as a 'parent’ protein, e.g., a parent Cas9 protein. In some cases, the Cas9 protein is a Streptococcus pyogenes (S. pyogenes) Cas9. In some cases, the Cas9 protein is a Staphylococcus aureus (S. aureus) Cas9. In some cases, the Cas9 protein is a Streptococcus thermophilus (S. thermophilus) Cas9. In some cases, the Cas9 protein is a Neisseria meningitidis (N. meningitidis) Cas9. In some cases, the Cas9 is a Geobacillus stearothermophilus (G. stearothermophilus) Cas9. Examples ofnaturally occurring Cas9 proteins, each of which is suitable to be a parent protein, include, but are not limited to, those provided as SEQ ID NOs: 1-33.

[0070] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e. , the Cas9 protein prior to insertion) comprises an amino acid sequence that is 80% or more (e.g. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-33. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1- 33.

[0071] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 80% or more (e.g. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-4. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1- 4.

[0072] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 80% or more (e.g. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 95% or more (e.g. 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 1.

[0073] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 80% or more (e.g. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S.pyogenes amino acid sequence set forth in SEQ ID NO: 2. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 95% or more (e.g. 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 2.

[0074] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 80% or more (e.g. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 3. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 95% or more (e.g. 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 3.

[0075] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 80% or more (e.g. 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 4. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., the Cas9 protein prior to insertion) comprises an amino acid sequence that is 95% or more (e.g. 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 4.

[0076] As would be known to one of ordinary skill in the art, a Cas9 protein includes an HNH nuclease domain and a RuvC nuclease domain. In a wild type Cas9 protein the RuvC domain cleaves the non-complementary strand (the strand that does not directly hybridized with the guide RNA) of a dsDNA target while the HNH domain cleaves the complementary strand (the strand that does directly hybridized with the guide RNA). Thus, for example a Cas9 protein lacking a catalytically active RuvC domain cannot cleave the non-complementary strand of a target dsDNA, but may be able to cleave the complementary strand (e.g., if the HNH domain is intact). Likewise, a Cas9 protein lacking a catalytically active HNH domain cannot cleave the complementary strand of a target dsDNA, but may be able to cleave the non-complementary strand (e.g., if the RuvC domain is intact). If a Cas9 protein lacks a catalytically active HNH domain andlacks a catalytically active RuvC domain, then the protein does not cleave target DNA, unless such an activity is provided by a heterologous protein (e.g., a fusion partner).

[0077] The RuvC domain of a wild type Cas9 protein is referred to herein as a split RuvC domain because the primary amino acid sequence of a wild type Cas9 protein includes 3 separate stretches (referred to herein as “RuvC subdomains”) of primary amino acid sequence that each make up a portion of the RuvC domain despite the fact that the three separate stretches are separated from one another by intervening primary amino acid sequence. The three subdomains fold to form a RuvC domain. In other words, a wild type Cas9 protein has 3 different regions (sometimes referred to as subdomains RuvC-l, RuvC-ll, and RucC-l 11), that are not contiguous with respect to the primary amino acid sequence of the Cas9 protein, but fold together to form a RuvC domain once the protein is produced and folds. As would be known to one of ordinary skill in the art, a wild type Cas9 protein includes a split RuvC domain and an HNH domain. A suitable Cas9 protein can include one or more amino acid mutations that render the protein to be a nickase (e.g., RuvC domain or HNH domain catalytically inactive) or a nuclease inactive protein (e.g., RuvC domain and HNH domain catalytically inactive).Sites for insertion

[0078] The present disclosure provides Cas9 fusion polypeptides that include one or more (in some cases one) linker-NLS1-linker-NLS2-linker (hiNLS) modules inserted internally within a Cas9 protein. The insertion can be at one or more specified internal insertion sites. When an internal insertion site is specified, an hiNLS module(s) is inserted immediately adjacent and C-terminal to the specified residue. As an illustrative nonlimiting example, if the specified amino acid residue is at position 205 of the Cas9 protein, then the hiNLS module(s) is inserted immediately adjacent and C-terminal to position 205 such that the first residue of the hiNLS module would be at position 206.

[0079] Suitable insertion sites include G205, K468, 11022, and S1248 with reference to the S. pyogenes Cas9 protein of SEQ ID NO: 1. In some cases, the insertion site is within 3 amino acids of these positions, e.g., positions 202-208, 465-471 , 1019-1025, and 1245-1251 of SEQ ID NO: 1 . If the Cas9 protein is a protein other than the protein set forth as SEQ ID NO: 1 (e.g., any of the proteins set forth in SEQ ID NOs: 2-38, or any other Cas9 protein, of which many will be known to one of ordinary skill in the art), sequence alignments and / or structural alignments can be used to determine which amino acid ‘corresponds’ to the G205, K468, 11022, and / or S1248 insertion residue(or to positions 202-208, 465-471 , 1019-1025, and 1245-1251 of SEQ ID NO:1). In other words, the amino acid positions listed herein (G205, K468, 11022, and S1248) use the protein set forth in SEQ ID NO: 1 as the reference sequence, but any Cas9 protein can serve as the reference sequence and the ‘corresponding’ amino acid position can be readily determined, e.g., using a sequence and / or structural alignment.

[0080] As such, an hiNLS module(s) can be inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to a residue selected from: G205, K468, 11022, and S1248 based on the numbering of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to a residue selected from: G205, K468, 11022, and S1248 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to a residue selected from positions: 202-208, 465-471, 1019-1025, and 1245-1251 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to a position at 202-208 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to K468 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to a position at 465-471 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to 11022 of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to a position at 1019-1025of the Cas9 protein set forth in SEQ ID NO: 1. In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to S1248 of the Cas9 protein set forth in SEQ ID NO: 1 . In some cases, an hiNLS module(s) is inserted at a position that is immediately adjacentand C-terminal to an amino acid residue corresponding to a position at 1245-1251 of the Cas9 protein set forth in SEQ ID NO: 1.

[0081] For example, in some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to SEQ ID NO: 1 , and the suitable insertion sites are G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to SEQ ID NO: 1 , and the suitable insertion sites are G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 99% or more (e.g., 99.5% or more, or 100%) identical to SEQ ID NO: 1 , and the suitable insertion sites are G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to SEQ ID NO: 1 , and one or more hiNLS modules (e.g., one hiNLS module) are inserted at G205 of SEQ ID NO: 1 . In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to SEQ ID NO: 1, and one or more hiNLS modules (e.g., one hiNLS module) are inserted at K468 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to SEQ ID NO: 1 , and one or more hiNLS modules (e.g., one hiNLS module) are inserted at 11022 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is80% or more (e g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to SEQ ID NO: 1, and one or more hiNLS modules (e.g., one hiNLS module) are inserted at S1248 of SEQ ID NO: 1.

[0082] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion sites correspond to G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion sites correspond to positions 202-208, 465-471, 1019-1025, and 1245-1251 of SEQ ID NO:1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion sites correspond to G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion site corresponds to G205 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion site corresponds to K468 of SEQ ID NO: 1 . In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion site corresponds to 11022 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-33, and the suitable insertion site corresponds to S1248 of SEQ ID NO: 1.

[0083] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion sites correspond to G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion sites correspond to positions 202-208, 465-471, 1019-1025, and 1245-1251 of SEQ ID NO: 1.

[0084] In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion sites correspond to G205, K468, 11022, and S1248 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion site corresponds to G205 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion site corresponds to K468 of SEQ ID NO: 1. In some cases, the Cas9 protein (whenconsidered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion site corresponds to 11022 of SEQ ID NO: 1. In some cases, the Cas9 protein (when considered in the absence of the inserted one or more hiNLS modules, i.e., when considered prior to insertion) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, 99.5% or more, or 100%) identical to any one of SEQ ID NOs: 1-4, and the suitable insertion site corresponds to S1248 of SEQ ID NO: 1.

[0085] In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted immediately adjacent and C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248 of SEQ ID NO: 1 , or any combination thereof. In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted immediately adjacent and C-terminal to an amino acid residue corresponding to positions 202-208, 465-471 , 1019-1025, or 1245-1251 of SEQ ID NO: 1 , or any combination thereof.

[0086] In some cases, at a given site, one module is inserted. In some cases, at a given site, more than one module (e.g., two modules, three modules) is inserted. In some cases, one module is inserted at each of more than one site (e.g., two sites, three sites, or all four sites). In some cases, more than one module (e.g., two modules, three modules) is inserted at each of more than one site (e.g., two sites, three sites, or all four sites).

[0087] In some cases, one hiNLS module is inserted C-terminal to an amino acid residue corresponding to G205 of SEQ ID NO: 1. In some cases, one hiNLS module is inserted C-terminal to an amino acid residue corresponding to K468 of SEQ ID NO: 1. In some cases, one hiNLS module is inserted C-terminal to an amino acid residue corresponding to 11022 of SEQ ID NO: 1. In some cases, one hiNLS module is inserted C-terminal to an amino acid residue corresponding to S1248 of SEQ ID NO: 1. In some cases, one hiNLS module is inserted C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248 of SEQ ID NO: 1. In some cases, one hiNLS module is inserted C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248 of SEQ ID NO: 1 , or any combination thereof. As such, in some cases, one hiNLS module is inserted at each of more than one site (e.g., two sites, three sites, or all four sites).

[0088] In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted C-terminal to an amino acid residue corresponding to G205 of SEQ ID NO: 1. In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted C-terminal to an amino acid residue corresponding to K468 of SEQ ID NO: 1. In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted C-terminal to an amino acid residue corresponding to 11022 of SEQ ID NO: 1. In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted C-terminal to an amino acid residue corresponding to S1248 of SEQ ID NO: 1 . In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248 of SEQ ID NO: 1 . In some cases, one or more hiNLS modules (e.g., two or more, three or more) are inserted C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248 of SEQ ID NO: 1 , or any combination thereof (i.e., one or more modules, such as two or more or three or more, can be inserted at one of the sties, or can be inserted at each of two sites, three sites, or all four sites).Protein variants

[0089] Any Cas9 protein can be used as a parent protein (as described above), i.e., as a Cas9 protein portion of a subject Cas9 fusion polypeptide. A Cas9 protein of a subject Cas9 fusion polypeptide can lack a catalytically active RuvC domain and / or can lack a catalytically active HNH domain. Thus, a Cas9 protein of a subject Cas9 fusion polypeptide can be nickase or can be a catalytically inactive protein.

[0090] In some cases, a Cas9 protein of a subject Cas9 fusion polypeptide has reduced catalytic activity (e.g., a Cas9 protein with nickase activity or a Cas9 protein that is catalytically inactive, e.g., dCas9). For example, when a Cas9 protein has a mutation at one or more amino acid positions corresponding to D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or a A987 of the Cas9 protein set forth in SEQ ID NO: 1 (e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), the variant Cas9 protein can still bind to target DNA in a site-specific manner (because it is still guided to a target DNA sequence by a guide RNA) as long as it retains the ability to interact with the guide RNA. In some cases, Cas9 protein of a subject Cas9 fusion polypeptide is a nickase (e.g., cleaves one strand of a double stranded target nucleic acid but not the other strand) (e.g., the Cas9 protein can be a nickase, e.g., can include one or more amino acid mutations that make it a nickase). For example, in some cases, Cas9 protein of a subject Cas9fusion polypeptide has a mutation in a catalytic domain (e.g., a mutation in a RuvC or HNH domain).

[0091] For example, in some cases, a Cas9 protein of a subject Cas9 fusion polypeptide can cleave the complementary strand of a target nucleic acid but has reduced ability to cleave the non-complementary strand of a target nucleic acid. For example, the Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some cases, a Cas9 protein has a mutation at residue D10 (e.g., D10A, aspartate to alanine) of SEQ ID NO: 1 (or the corresponding position of any Cas9 protein, e.g., any of the proteins set forth in SEQ ID NOs: 1-33) and can therefore cleave the complementary strand of a double stranded target nucleic acid but has reduced ability to cleave the non-complementary strand of a double stranded target nucleic acid (thus resulting in a single strand break (SSB) instead of a double strand break (DSB) when the variant Cas9 protein cleaves a double stranded target nucleic acid) (see, for example, Jinek et al., Science. 2012 Aug 17;337(6096):816-21). Examples of such amino acid positions in a RuvC domain can include: D10, G12, G17, E762, H982, H983, A984, D986, and / or A987 of the Cas9 protein set forth in SEQ ID NO: 1 (e.g., D10A, G12A, G17A, E762A, H982A, H983A, A984A, and / or D986A).

[0092] In some cases, Cas9 protein of a subject Cas9 fusion polypeptide can cleave the non- complementary strand of a target nucleic acid but has reduced ability to cleave the complementary strand of the target nucleic acid. For example, the Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain. Thus, the Cas9 protein can be a nickase that cleaves the non- complementary strand, but does not cleave the complementary strand (e.g., does not cleave a single stranded target nucleic acid). As a non-limiting example, in some embodiments, the Cas9 protein has a mutation at position H840 (e.g., an H840A mutation, histidine to alanine) of SEQ ID NO: 1 (or the corresponding position of any Cas9 protein, e.g., the Cas9 proteins set forth as SEQ ID NOs: 1-33 and can therefore cleave the non-complementary strand of the target nucleic acid but has reduced ability to cleave (e.g., does not cleave) the complementary strand of the target nucleic acid. Such a Cas9 protein has a reduced ability to cleave a target nucleic acid (e.g., a single stranded target nucleic acid). Examples of such amino acid positions in an HNH domain can include: H840, N854, and / or N863 of the Cas9 protein set forth in SEQ ID NO: 1 (e.g., H840A, N854A, and / or N863A).

[0093] In some cases, a Cas9 protein of a subject Cas9 fusion polypeptide has a reduced ability to cleave both the complementary and the non-complementary strands of a double stranded target nucleic acid. In some cases, the Cas9 protein is a dCas9 protein. As a non-limiting example, in some cases, the Cas9 protein harbors mutations at residues D10 and H840 (e.g., D10A and H840A) of SEQ ID NO: 1 (or the corresponding residues of another Cas9 protein, e.g., any of the proteins set forth as SEQ ID NOs: 1-33) such that the polypeptide has a reduced ability to cleave (e.g., does not cleave) both the complementary and the non-complementary strands of a target nucleic acid. Such a Cas9 protein has a reduced ability to cleave a target nucleic acid (e.g., a single stranded or double stranded target nucleic acid) but retains the ability to bind a target nucleic acid. For example, a Cas9 protein of a subject Cas9 fusion polypeptide can have a mutation in one or more of amino acid positions in (i) a RuvC domain: corresponding to D10, G12, G17, E762, H982, H983, A984, D986, and / or A987 of the Cas9 protein set forth in SEQ ID NO: 1 (e.g., D10A, G12A, G17A, E762A, H982A, H983A, A984A, and / or D986A); and one or more of amino acid positions in (ii) an HNH domain: corresponding to H840, N854, and / or N863 of the Cas9 protein set forth in SEQ ID NO: 1 (e.g., H840A, N854A, and / or N863A).

[0094] In some cases, the parent Cas9 protein of a subject Cas9 fusion polypeptide is a high fidelity (HF) Cas9 protein (also referred to as SpCas9-HF1 or HF1 - and also -HF2, - HF3, -HF4) (e.g., see Kleinstiver et al. (2016) Nature 529:490), and thus the generated subject Cas9 fusion polypeptide includes the same amino acid mutations that caused the parent protein to be high fidelity. For example, amino acids N497, R661, Q695, and Q926 of the amino acid sequence set forth as (SEQ ID NO:1) (or the corresponding position of another Cas9 protein, e.g., a protein having the amino acid sequence of any of the sequences set forth as SEQ ID NOs: 1-33) can be substituted, e.g., with alanine. In some cases, a suitable parent Cas9 protein exhibits altered PAM specificity. See, e.g., Kleinstiver et al. (2015) Nature 523:481. Additional examples of Cas9 variants that can be used, include, but are not limited to: HiFiCas9 (e.g., R691A), eSpCas9 (e.g., K810A, K1003A, R1060A), eSpCas9 (e.g., D1135E), HypaCas9 (e.g., N692A, M694A, Q695A, H698A), xCas9 (e.g., E108G, S217A, A262T, S409I, E480K, E543D, M694I, E1219V), Sniper-Cas9 (e.g., F539S, M763I, K890N), evoCas9 (e.g., M495V, Y515N, K526E, R661Q), SpartaCas (e.g., D23A, T67L, Y128V, D1251G), LZ3Cas9 (e.g., N690C, T769I, G915M, N980K), miCas9 (e.g., SV40 NLS linker fused with brex27 motif), SuperFi-Cas9 (e.g., Y1010D, Y1013D, Y1016D, V1018D, R1019D, Q1027D, K1031D) (see, e.g., Allemailem et al,Int J Mol Sci. 2023 Apr 11 ;24(8):7052). In some cases, the Cas9 is an iGeoCas9 (see, e.g., international patent publication WO2024112479, which is incorporated herein by reference for such disclosure).Heterologous polypeptides / Fusion partners

[0095] In some embodiments, a subject Cas9 fusion polypeptide, in addition to including an internal insertion, also includes one or more heterologous proteins, e.g., at or near the N-terminus and / or C-terminus (such a heterologous protein is also referred to herein as a “fusion partner”) - and the Cas9 fusion polypeptide will exhibit the activity of the protein to which it is fused. In some such cases, the Cas9 protein of the Cas9 fusion polypeptide has nickase activity (nCas9) or is catalytically deactivated (i.e. , is a ‘dead’ Cas9, i.e., dCas9). Nickase and catalytically inactive Cas9 proteins are described in more detail elsewhere herein.

[0096] In other words, a subject Cas9 fusion polypeptide recognizes and binds specific DNA sequences when guided by a complementary RNA molecule. By fusing a subject Cas9 fusion polypeptide to DNA-modifying enzymes, chimeric proteins can be generated for various purposes, including precise genome editing and / or recruitment of genomic modifying machinery. Fusion strategies such as SpyTag / SpyCatcher and CHARM (Chromoaffinity-based Rapid and Modular) systems facilitate the covalent or non-covalent coupling of a subject Cas9 fusion polypeptide with enzymes capable of diverse DNA modifications, including base editing, insertion, deletion, and replacement. CHARM utilizes a split protein tag system, where one part of the tag is fused to the CRISPR domain and the other part to the DNA-modifying domain. The two parts of the tag interact to bring the CRISPR domain and DNA-modifying domain into close proximity, facilitating their interaction. The BEaST protocol specifically employs SpyTag fusion to assemble deaminase and nickase domains into a base editor complex, facilitating efficient and precise nucleotide conversion at targeted genomic loci as directed by the subject Cas9 fusion polypeptide.

[0097] In some cases, a subject Cas9 fusion polypeptide includes a variant Cas9 protein (e.g., a nickase or ‘dead’ version), and is fused to a heterologous protein, e.g., one that has transcription repressor activity (e.g., includes a transcription repression domain) and thereby reduces transcription (and therefore expression) of a target gene. In some cases, a subject Cas9 fusion polypeptide includes a variant Cas9 protein (e.g., a nickase or ‘dead’ version), and is fused to a heterologous protein that has transcription activating activity (e.g., includes a transcription activation domain)and thereby increases transcription (and therefore expression) of a target gene. In some cases, at least one of the one or more fusion partners includes a deaminase, a reverse transcriptase, a transcription modulator, or an epigenetic modulator.

[0098] A variety of heterologous polypeptides are suitable for inclusion in a subject Cas9 fusion polypeptide of the present disclosure. In some embodiments, a heterologous polypeptide is fused at or near (within 50 amino acids of) the C-terminus of a subject Cas9 fusion polypeptide (e.g., fused to the C-terminus of a subject Cas9 fusion polypeptide). In some embodiments, a heterologous polypeptide is fused at or near (within 50 amino acids of) the N-terminus of a subject Cas9 fusion polypeptide (e.g., fused to the N-terminus of a subject Cas9 fusion polypeptide). In some embodiments, a heterologous polypeptide is fused at or near (within 50 amino acids of) the C- terminus of a subject Cas9 fusion polypeptide (e.g., fused to the C-terminus of a subject Cas9 fusion polypeptide); and a heterologous polypeptide (which can be the same or different relative to the C-terminal heterologous polypeptide) is fused at or near (within 50 amino acids of) the N-terminus of a subject Cas9 fusion polypeptide (e.g., fused to the N-terminus of a subject Cas9 fusion polypeptide).

[0099] A Cas9 fusion polypeptide of the present disclosure can contain any convenient number of heterologous polypeptides. Thus, the subject Cas9 fusion polypeptide can contain one or more, e.g., 2 or more, 3 or more, 4 or more, 5 or more, 7 or more, including 10 or more heterologous polypeptides, and can contain 15 or fewer, e.g., 12 or fewer, 10 or fewer, 8 or fewer, 6 or fewer, 4 or fewer, 3 or fewer, including 2 or fewer distinct heterologous polypeptides. In some cases, the subject Cas9 fusion polypeptide contains between 1 to 10, e.g., between 1 to 8, between 1 to 5, between 1 to 4, between 1 to 3, including between 1 to 2 heterologous polypeptides. Where the subject Cas9 fusion polypeptide contains more than one heterologous polypeptide, any two of the heterologous polypeptides may be the same or may be distinct heterologous polypeptides. For a Cas9 fusion polypeptide, a heterologous polypeptide (fusion partner) can be fused to the N-terminus, the C-terminus, or both.

[0100] A heterologous polypeptide to which a subject Cas9 fusion polypeptide is fused can also be referred to herein as a ‘fusion partner.’

[0101] It is well recognized that a Cas9 protein can be used to modify target nucleic acids (e.g., DNA and / or RNA) in a variety of ways without creating a double strand break (DSB) in the target DNA. For example, in some cases a double stranded target DNA is nicked (one strand is cleaved), and in some cases (e.g., in some cases where the Cas9 protein is devoid of nuclease activity, e.g., a Cas9 protein may harbor mutationsin the catalytic nuclease domains), the target nucleic acid is not cleaved at all. For example, in some cases a Cas9 protein with or without nuclease activity (e.g., nickase activity or dead), is fused to a heterologous protein domain. The heterologous protein domain can provide an activity to the fusion protein such as (i) a DNA-modifying activity (e.g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity or glycosylase activity), (ii) a transcription modulation activity (e.g., fusion to a transcription repressor or activator), or (iii) an activity that modifies a protein (e.g., a histone) that is associated with target DNA (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity or demyristoylation activity). As such, a gene editing system can be used in applications that modify a target nucleic acid in way that do not cleave the target nucleic acid, and can also be used in applications that modulate transcription from a target DNA.

[0102] As would be known to one of ordinary skill in the art, some heterologous proteins (their activity) can cause transcription activation (and this can be referred to as CRISPRa), while other heterologous proteins (their activity) can cause transcription repression (and this can be referred to as CRISPRi).

[0103] Examples of CRISPRi and CRISPRa fusion proteins (e.g., a CRISPR-Cas effector protein - in this case a subject Cas9 fusion polypeptide - fused to a heterologous protein) will be known to one of ordinary skill in the art and any convenient CRISPRi or CRISPRa fusion partner can be used. A CRISPRi fusion protein or CRISPRa fusion protein includes a Cas9 fusion polypeptide (e.g., one with a variant Cas9 protein) linked (covalently or non-covalently) to a transcription repressing or activating protein. In some cases, the Cas9 fusion polypeptide is fused to a transcription repressing protein (and is called a CRISPRi fusion protein or simply a CRISPRi protein). In some cases, the Cas9 fusion polypeptide is fused to a transcription activating protein (and is called a CRISPRa fusion protein or simply a CRISPRa protein). In some cases, the Cas9 fusion polypeptide is linked (covalently or non- covalently) directly to the transcription repressing protein (and is called a CRISPRifusion protein or simply a CRISPRi protein). In some cases, the Cas9 fusion polypeptide is linked (covalently or non-covalently) directly to the transcription activating protein (and is called a CRISPRa fusion protein or simply a CRISPRa protein). In some cases, the Cas9 fusion polypeptide is linked (covalently or non- covalently) indirectly to the transcription repressing or activating protein, e.g., by being linked a protein that recruits a transcription repressing or activating protein (see, e.g., Griffith et al., Cell Genom. 2023 Sep 1;3(9): 100387).

[0104] In some cases, the fusion partner can modulate transcription (e.g., inhibit transcription, increase transcription) of a target DNA. For example, in some cases the fusion partner is a protein (or a domain from a protein) that inhibits transcription (e.g., a transcription repressor, a protein that functions via recruitment of transcription inhibitor proteins, modification of target DNA such as methylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of a histone modifier such as those that modify acetylation and / or methylation of histones, and the like). In some cases, the fusion partner is a protein (or a domain from a protein) that increases transcription (e.g., a transcription activator, a protein that acts via recruitment of transcription activator proteins, modification of target DNA such as demethylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of a histone modifier such as those that modify acetylation and / or methylation of histones, and the like). In some cases, the fusion partner is a reverse transcriptase. In some cases, the fusion partner is a base editor. In some cases, the fusion partner is a deaminase.

[0105] In some cases, the one or more fusion partners includes a heterologous polypeptide that has enzymatic activity that modifies a target nucleic acid (e.g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity).

[0106] In some cases, the one or more fusion partners includes a heterologous polypeptide that has enzymatic activity that modifies a polypeptide (e.g., a histone) associated with a target nucleic acid (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylationactivity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity or demyristoylation activity).

[0107] Examples of proteins (or fragments thereof) that can be used in increase transcription include but are not limited to: transcription activators such as VP16, VP64, VP48, VP160, p65 subdomain (e.g., from NFkB), and activation domain of EDLL and / or TAL activation domain (e.g., for activity in plants); histone lysine methyltransferases such as SET1A, SET1 B, MLL1 to 5, ASH1 , SYMD2, NSD1 , and the like; histone lysine demethylases such as JHDM2a / b, UTX, JMJD3, and the like; histone acetyltransferases such as GCN5, PCAF, CBP, p300, TAF1 , TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1 , ACTR, P160, CLOCK, and the like; and DNA demethylases such as Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), TET1 , DME, DML1 , DML2, ROS1 , and the like. See, e.g., Chavez et al., Nat Methods. 2015 Apr; 12(4): 326-328. In some cases, the CRISPRa system is a SAM system, which includes 3 components that form the DNA-binding complex: (1) a CRISPRa fusion protein (e.g., dTnpB fused to VP64), (2) MS2 aptamer(s) added to the coRNA (forming a characteristic stem loop structure recognized by MS2), and (3) transcription activators P65 (Nuclear Factor NF-KB p65) and HSF1 (Heat Shock Factor 1) fused with an MS2-tag corresponding to the minimal aptamer-binding peptide of the MS2 coat protein. See, e.g., review articles such as Adli, Nat Commun. 2018 May 15;9(1):1911 ; Becirovic, Cell Mol Life Sci. 2022 Feb 12;79(2):130; and Nidhi S, et al., Int J Mol Sci. 2021 Mar 24;22(7):3327 for examples with fusions to other CRISPR-Cas effector proteins.

[0108] Examples of proteins (or fragments thereof) that can be used in decrease transcription include but are not limited to: transcription repressors such as the Kruppel associated box (KRAB or SKD); KOX1 repression domain; the Mad mSIN3 interaction domain (SID); the ERF repressor domain (ERD), the SRDX repression domain (e.g., for repression in plants), and the like; histone lysine methyltransferases such as Pr-SET7 / 8, SUV4-20H1, RIZ1 , and the like; histone lysine demethylases such as JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1 , JMJD2D, JARID1A / RBP2, JARID1 B / PLU-1 , JARID1C / SMCX, JARID1 D / SMCY, and the like; histone lysine deacetylases such as HDAC1 , HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1 , SIRT2, HDAC11 , and the like; DNA methylases such as Hhal DNA m5c-methyltransferase (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3(plants), ZMET2, CMT1 , CMT2 (plants), and the like; and periphery recruitment elements such as Lamin A, Lamin B, and the like.

[0109] In some cases, a fusion partner has enzymatic activity that modifies the target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activity that can be provided by the fusion partner include but are not limited to: nuclease activity (e.g., provided by Fokl nuclease), methyltransferase activity such as that provided by a methyltransferase (e.g., Hhal DNA m5c-methyltransferase (M.Hhal), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1 , CMT2 (plants), and the like); demethylase activity such as that provided by a demethylase (e.g., Ten-Eleven Translocation (TET) dioxygenase 1 (TET1CD), TET1 , DME, DML1, DML2, ROS1 , and the like), DNA repair activity, DNA damage activity, deamination activity such as that provided by a deaminase (e.g., a cytosine deaminase enzyme such as rat APOBEC1), dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity such as that provided by an integrase and / or resolvase (e.g., Gin invertase such as the hyperactive mutant of the Gin invertase, GinH106Y; human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase; and the like), transposase activity, recombinase activity such as that provided by a recombinase (e.g., a Ore recombinase; a Hin recombinase; a Tre recombinase; a FLP recombinase; catalytic domain of Gin recombinase, and the like), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity).

[0110] In some cases, the fusion partner has enzymatic activity that modifies a protein associated with the target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA) (e.g., a histone, an RNA binding protein, a DNA binding protein, and the like). Examples of enzymatic activity (that modifies a protein associated with a target nucleic acid) that can be provided by the fusion partner include but are not limited to: methyltransferase activity such as that provided by a histone methyltransferase (HMT) (e.g., suppressor of variegation 3-9 homolog 1 (SUV39H1 , also known as KMT1A), euchromatic histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1 , and the like, SET1A, SET1 B, MLL1 to 5, ASH1 , SYMD2, NSD1 , DOT1 L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1), demethylase activity such as that provided by a histone demethylase (e.g., Lysine Demethylase 1A (KDM1A also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1 , JMJD2D, JARID1A / RBP2, JARID1 B / PLU-1 , JARID1C / SMCX, JARID1D / SMCY, UTX,JMJD3, and the like), acetyltransferase activity such as that provided by a histone acetylase transferase (e.g., catalytic core / fragment of the human acetyltransferase p300, GCN5, PCAF, CBP, TAF1 , TIP60 / PLIP, M0Z / MYST3, M0RF / MYST4, HB01 / MYST2, HM0F / MYST1, SRC1, ACTR, P160, CLOCK, and the like), deacetylase activity such as that provided by a histone deacetylase (e.g., HDAC1 , HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1 , SIRT2, HDAC11 , and the like), kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, ribosylation activity, deribosylation activity, myristoylation activity, and demyristoylation activity.

[0111] Additional examples of a suitable fusion partners are dihydrofolate reductase (DHFR) destabilization domain (e.g., to generate a chemically controllable fusion polypeptide), and a chloroplast transit peptide.

[0112] In some cases, a fusion partner can be a chloroplast transit peptide. Thus, for example, a ribonucleoprotein (RNP) complex, comprising a Cas9 fusion polypeptide of the present disclosure and a guide RNA, can be targeted to the chloroplast. In some cases, this targeting may be achieved by the presence of an N-terminal extension, called a chloroplast transit peptide (CTP) or plastid transit peptide. Chromosomal transgenes from bacterial sources must have a sequence encoding a CTP sequence fused to a sequence encoding an expressed polypeptide if the expressed polypeptide is to be compartmentalized in the plant plastid (e.g. chloroplast). Accordingly, localization of an exogenous polypeptide to a chloroplast is often 1 accomplished by means of operably linking a polynucleotide sequence encoding a CTP sequence to the 5' region of a polynucleotide encoding the exogenous polypeptide. The CTP is removed in a processing step during translocation into the plastid. Processing efficiency may, however, be affected by the amino acid sequence of the CTP and nearby sequences at the amino terminus of the peptide. Other options for targeting to the chloroplast which have been described are the maize cab-m7 signal sequence (U.S. Pat. No. 7,022,896, WO 97 / 41228) a pea glutathione reductase signal sequence (WO 97 / 41228) and the CTP described in US2009029861 .

[0113] In some cases, a fusion partner can be an endosomal escape peptide. In some cases, an endosomal escape polypeptide comprises the amino acid sequence GLFXALLXLLXSLWXLLLXA (SEQ ID NO:24), wherein each X is independently selected from lysine, histidine, and arginine. In some cases, an endosomal escapepolypeptide comprises the amino acid sequence GLFHALLHLLHSLWHLLLHA (SEQ ID NO:25).

[0114] Additional suitable heterologous polypeptides (fusions partners) include, but are not limited to, a polypeptide that directly and / or indirectly provides for increased or decreased transcription and / or translation of a target nucleic acid (e.g., a transcription activator or a fragment thereof, a protein or fragment thereof that recruits a transcription activator, a small molecule / drug-responsive transcription and / or translation regulator, a translation-regulating protein, etc.). Non-limiting examples of heterologous polypeptides to accomplish increased or decreased transcription include transcription activator and transcription repressor domains. In some such cases, a CRISPR-Cas fusion polypeptide is targeted by the guide nucleic acid (guide RNA) to a specific location (i.e., sequence) in the target nucleic acid and exerts locus-specific regulation such as blocking RNA polymerase binding to a promoter (which selectively inhibits transcription activator function), and / or modifying the local chromatin status (e.g., when a fusion sequence is used that modifies the target nucleic acid or modifies a polypeptide associated with the target nucleic acid). In some cases, the changes are transient (e.g., transcription repression or activation). In some cases, the changes are inheritable (e.g., when epigenetic modifications are made to the target nucleic acid or to proteins associated with the target nucleic acid, e.g., nucleosomal histones).

[0115] Non-limiting examples of heterologous polypeptides for use when targeting ssRNA target nucleic acids include (but are not limited to): splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or release factors; e.g., elF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases, e.g., adenosine deaminase acting on RNA (ADAR), including A to I and / or C to U editing enzymes); helicases; RNA-binding proteins; and the like. It is understood that a heterologous polypeptide can include the entire protein or in some cases can include a fragment of the protein (e.g., a functional domain).

[0116] The heterologous polypeptide can be any domain capable of interacting with ssRNA (which, for the purposes of this disclosure, includes intramolecular and / or intermolecular secondary structures, e.g., double-stranded RNA duplexes such as hairpins, stem-loops, etc.), whether transiently or irreversibly, directly or indirectly, including but not limited to an effector domain selected from the group comprising; endonucleases (for example RNase III, the CRR22 DYW domain, Dicer, and PIN (PilT N-terminus) domains from proteins such as SMG5 and SMG6); proteins and protein domains responsible for stimulating RNA cleavage (for example CPSF, CstF, CFImand CFIIm); Exonucleases (for example XRN-1 or Exonuclease T); Deadenylases (for example HNT3); proteins and protein domains responsible for nonsense mediated RNA decay (for example UPF1 , UPF2, UPF3, UPF3b, RNP S1 , Y14, DEK, REF2, and SRm160); proteins and protein domains responsible for stabilizing RNA (for example PABP); proteins and protein domains responsible for repressing translation (for example Ago2 and Ago4); proteins and protein domains responsible for stimulating translation (for example Staufen); proteins and protein domains responsible for (e.g., capable of) modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, etc., e.g., elF4G); proteins and protein domains responsible for polyadenylation of RNA (for example PAP1 , GLD-2, and Star- PAP); proteins and protein domains responsible for polyuridinylation of RNA (for example Cl D1 and terminal uridylate transferase); proteins and protein domains responsible for RNA localization (for example from IMP1 , ZBP1 , She2p, She3p, and Bicaudal-D); proteins and protein domains responsible for nuclear retention of RNA (for example Rrp6); proteins and protein domains responsible for nuclear export of RNA (for example TAP, NXF1 , THO, TREX, REF, and Aly); proteins and protein domains responsible for repression of RNA splicing (for example PTB, Sam68, and hnRNP A1); proteins and protein domains responsible for stimulation of RNA splicing (for example Serine / Arginine-rich (SR) domains); proteins and protein domains responsible for reducing the efficiency of transcription (for example FUS (TLS)); and proteins and protein domains responsible for stimulating transcription (for example CDK7 and HIV Tat). Alternatively, the effector domain may be selected from the group comprising endonucleases; proteins and protein domains capable of stimulating RNA cleavage; exonucleases; deadenylases; proteins and protein domains having nonsense mediated RNA decay activity; proteins and protein domains capable of stabilizing RNA; proteins and protein domains capable of repressing translation; proteins and protein domains capable of stimulating translation; proteins and protein domains capable of modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, etc., e.g., elF4G); proteins and protein domains capable of polyadenylation of RNA; proteins and protein domains capable of polyuridinylation of RNA; proteins and protein domains having RNA localization activity; proteins and protein domains capable of nuclear retention of RNA; proteins and protein domains having RNA nuclear export activity; proteins and protein domains capable of repression of RNA splicing; proteins and protein domains capable of stimulation of RNA splicing; proteins and protein domains capable of reducing theefficiency of transcription; and proteins and protein domains capable of stimulating transcription. Another suitable heterologous polypeptide is a PUF RNA-binding domain, which is described in more detail in WO2012068627, which is hereby incorporated by reference in its entirety.

[0117] Some RNA splicing factors that can be used (in whole or as fragments thereof) as heterologous polypeptides for a fusion polypeptide of the present disclosure have modular organization, with separate sequence-specific RNA binding modules and splicing effector domains. For example, members of the Serine / Arginine-rich (SR) protein family contain N-terminal RNA recognition motifs (RRMs) that bind to exonic splicing enhancers (ESEs) in pre-mRNAs and C-terminal RS domains that promote exon inclusion. As another example, the hnRNP protein hnRNP Al binds to exonic splicing silencers (ESSs) through its RRM domains and inhibits exon inclusion through a C-terminal Glycine-rich domain. Some splicing factors can regulate alternative use of splice site (ss) by binding to regulatory sequences between the two alternative sites. For example, ASF / SF2 can recognize ESEs and promote the use of intron proximal sites, whereas hnRNP Al can bind to ESSs and shift splicing towards the use of intron distal sites. One application for such factors is to generate ESFs that modulate alternative splicing of endogenous genes, particularly disease associated genes. For example, Bcl-x pre-mRNA produces two splicing isoforms with two alternative 5' splice sites to encode proteins of opposite functions. The long splicing isoform Bcl-xL is a potent apoptosis inhibitor expressed in long-lived postmitotic cells and is up-regulated in many cancer cells, protecting cells against apoptotic signals. The short isoform Bcl-xS is a pro-apoptotic isoform and expressed at high levels in cells with a high turnover rate (e.g., developing lymphocytes). The ratio of the two Bcl- x splicing isoforms is regulated by multiple coo-elements that are located in either the core exon region or the exon extension region (i.e. , between the two alternative 5' splice sites). For more examples, see WO2010075303, which is hereby incorporated by reference in its entirety.

[0118] Further suitable fusion partners include, but are not limited to, proteins (or fragments thereof) that are boundary elements (e.g., CTCF), proteins and fragments thereof that provide periphery recruitment (e.g., Lamin A, Lamin B, etc.), protein docking elements (e.g., FKBP / FRB, Pil1 / Aby1 , etc.).Reverse transcriptases / prime editor

[0119] In some cases, a subject Cas9 fusion polypeptide is a prime editor (see, e.g. Anzalone et al (2019) Nature 576: 149-157). As such, in some cases, a Cas9 fusion polypeptide (e.g., one with a nickase version of the Cas9 protein, such as a Cas9 nickase that cuts the non-complementary strand of the DNA) is fused to a reverse transcriptase (RT) [e.g., M-MLV RT], As would be known to one of ordinary skill in the art, the guide RNA for prime editing can be fused to an RNA sequence that serves as a donor template for the RT.

[0120] As such, in some cases, a fusion partner is a reverse transcriptase polypeptide. Reverse transcriptases are known in the art; see, e.g., Cote and Roth (2008) Virus Res. 134:186. Suitable reverse transcriptases include, e.g., a murine leukemia virus reverse transcriptase; a Rous sarcoma virus reverse transcriptase; a human immunodeficiency virus type I reverse transcriptase; a Moloney murine leukemia virus reverse transcriptase; a transcription xenopolymerase (RTX); avian myeloblastosis virus reverse transcriptase (AMV-RT); a Eubacterium rectale maturase reverse transcriptase (Marathon®; and the like. The reverse transcriptase fusion partner can include one or more mutations. For example, in some cases, the reverse transcriptase is a M-MLV reverse transcriptase polypeptide that comprises one or more mutations selected from the group consisting of D200N, T306K, W313F, T330P and L603W. In some cases, the reverse transcriptase is a pentamutant of M-MLV RT (e.g., comprising the following substitutions: D200N / L603W / T330P / T306K / W313F) (where D200, L603, T330, T306, and W313 correspond to D199, L602, T329, T305, and W312 of the M-MLV RT amino acid sequence of SEQ ID NO:157).Base editors

[0121] In some cases, a subject Cas9 fusion polypeptide is a base editor. As such, in some cases, a Cas9 fusion polypeptide (e.g., one with a dead or nickase version of the Cas9 protein, i.e. , a dCas9 or nCas9) is fused to a base editing enzyme such as a deaminase enzyme such as APOBEC. Examples of base editing enzymes include, but are not necessarily limited to: cytosine base editors (CBEs), adenine base editors (ABEs), cytosine-to-guanine base editors (CGBEs), and simultaneous adenine and cytosine base editors (ACBEs). The deaminase fused to the Cas9 fusion polypeptide - be it a cytosine deaminase (e.g. APOBEC) or an adenosine deaminase (e.g. engineered TadA) - will dictate whether the base editor is a CBE or an ABE. CBEs contain cytosine deaminases, while ABEs typically contain adenine deaminases(though several groups have recently engineered CBEs derived from the conventional TadA adenosine deaminase domain).A / LS

[0122] In some cases, a fusion partner provides for subcellular localization, i.e. , the heterologous polypeptide contains a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence to keep the fusion protein out of the nucleus, e.g., a nuclear export sequence (NES), a sequence to keep the fusion protein retained in the cytoplasm, a mitochondrial localization signal for targeting to the mitochondria, a chloroplast localization signal for targeting to a chloroplast, an endoplasmic reticulum (ER) retention signal, and the like). In some cases, a Nucleic acid-binding effector fusion polypeptide (e.g., a CRISPR-Cas fusion polypeptide) does not include an NLS (which can be advantageous, e.g., when the target nucleic acid is an RNA that is present in the cytosol). In some cases, the heterologous polypeptide can provide a tag (i.e., the heterologous polypeptide is a detectable label) for ease of tracking and / or purification (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), yellow fluorescent protein (YEP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, tdTomato, and the like; a histidine tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag; and the like).

[0123] In some cases, the one or more fusion partners include one or more nuclear localization signals (NLSs) (e.g., in some cases 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more NLSs). Thus, in some cases, a fusion polypeptide of the present disclosure includes, in addition to the internally inserted hiNLS module(s), one or more NLSs (e.g., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more NLSs). In some cases, one or more NLSs (2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more NLSs) are positioned at or near (e.g., within 50 amino acids of) the N-terminus and / or the C-terminus. In some cases, one or more NLSs (2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more NLSs) are positioned at or near (e.g., within 50 amino acids of) the N-terminus. In some cases, one or more NLSs (2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more) are positioned at or near (e.g., within 50 amino acids of) the C-terminus. In some cases, one or more NLSs (3 or more, 4 or more, 5 or more, 6 or more, 7 ormore, 8 or more, 9 or more, or 10 or more) are positioned at or near (e.g., within 50 amino acids of) both the N-terminus and the C-terminus. In some cases, an NLS is positioned at the N-terminus and an NLS is positioned at the C-terminus.

[0124] In some cases, the one or more fusion partners includes 1 to 10 NLSs (e.g., 1-9, 1-8, 1-7, 1-6, 1-5, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 4-10, 4-9, 4-8, 4-7, 5-10, 5-9, or 5-8 NLSs). In some cases, the one or more fusion partners includes from 2 to 5 NLSs (e.g., 2-4 NLSs, or 2-3 NLSs). In some cases, the one or more fusion partners includes about 4 NLSs. In some cases, the one or more fusion partners includes about 7 NLSs.

[0125] Non-limiting examples of NLSs include an NLS sequence derived from: the NLS of the SV40 virus large T-antigen, having the amino acid sequence PKKKRKV (SEQ ID NO:1); the NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 43)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 44) or RQRRNELKRSP (SEQ ID NO: 49); the hRNPAI M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 50); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 51) of the IBB domain from importin-alpha; the sequences VSRKRPRP (SEQ ID NO: 52) and PPKKARED (SEQ ID NO: 53) of the myoma T protein; the sequence PQPKKKPL (SEQ ID NO: 54) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 55) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO: 56) and PKQKKRK (SEQ ID NO: 57) of the influenza virus NS1 ; the sequence RKLKKKIKKL (SEQ ID NO: 58) of the Hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 59) of the mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 60) of the human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 61) of the steroid hormone receptors (human) glucocorticoid. In general, NLS (or multiple NLSs) are of sufficient strength to drive accumulation of a subject Cas9 fusion polypeptide in the nucleus of a eukaryotic cell. If desired, detection of accumulation in the nucleus may be performed by any suitable technique. For example, a detectable marker may be fused to the Cas9 fusion polypeptide such that location within a cell may be visualized. Cell nuclei may also be isolated from cells, the contents of which may then be analyzed by any suitable process for detecting protein, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus may also be determined indirectly.PTD

[0126] In some cases, the one or more fusion partners includes a "Protein Transduction Domain" or PTD (also known as a CPP - cell penetrating peptide), which refers to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversing a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule, which can range from a small polar molecule to a large macromolecule and / or a nanoparticle, facilitates the molecule traversing a membrane, for example going from extracellular space to intracellular space, or cytosol to within an organelle. In some cases, a PTD is covalently linked to the amino terminus of a Cas9 fusion polypeptide. In some cases, a PTD is covalently linked to the carboxyl terminus of a Cas9 fusion polypeptide. In some cases, a PTD is covalently linked to the N-terminus and C-terminus of a Cas9 fusion polypeptide. In some cases, a PTD includes a nuclear localization signal (NLS) (e.g., in some cases 2 or more, 3 or more, 4 or more, or 5 or more NLSs). Thus, in some cases, a CRISPR-Cas fusion polypeptide includes one or more NLSs (e.g., 2 or more, 3 or more, 4 or more, or 5 or more NLSs), e.g., in addition to the inserted hiNLS module(s). In some cases, a PTD is covalently linked to a nucleic acid (e.g., a Cas9 guide RNA, a polynucleotide encoding a Cas9 guide RNA, a donor polynucleotide, etc.). Examples of PTDs include but are not limited to a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT comprising YGRKKRRQRRR; SEQ ID NO: 86); a polyarginine sequence comprising a number of arginines sufficient to direct entry into a cell (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines); a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489- 96); a Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7): 1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21 :1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO: 87); Transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 88);KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 89); and RQIKIWFQNRRMKWKK (SEQ ID NO: 90). Exemplary PTDs include but are not limited to, YGRKKRRQRRR (SEQ ID NO: 91), RKKRRQRRR (SEQ ID NO: 92); an arginine homopolymer of from 3 arginine residues to 50 arginine residues; Exemplary PTD domain amino acid sequences include, but are not limited to, any of the following: YGRKKRRQRRR (SEQ ID NO: 91); RKKRRQRR (SEQ ID NO: 93); YARAAARQARA (SEQ ID NO: 94); THRLPRRRRRR (SEQ ID NO: 95); andGGRRARRRRRR (SEQ ID NO: 96). In some cases, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June; 1(5-6): 371-381). ACPPs comprise a polycationic CPP (e.g., Arg9 or “R9”) connected via a cleavable linker to a matching polyanion (e.g., Glu9 or “E9”), which reduces the net charge to nearly zero and thereby inhibits adhesion and uptake into cells. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesiveness, thus “activating” the ACPP to traverse the membrane.

[0127] For examples of some of the above fusion partners (and more) used in the context of fusions with Cas9, Zinc Finger, and / or TALE proteins (for site specific target nucleic modification, modulation of transcription, and / or target protein modification, e.g., histone modification), see, e.g.: Nomura et al., J Am Chem Soc. 2007 Jul 18;129(28):8676-7; Rivenbark et al., Epigenetics. 2012 Apr;7(4):350-60; Nucleic Acids Res. 2016 Jul 8;44(12):5615-28; Gilbert et al., Cell. 2013 Jul 18;154(2):442-51 ; Kearns et al., Nat Methods. 2015 May;12(5):401-3; Mendenhall et al., Nat Biotechnol. 2013 Dec;31 (12): 1133-6; Hilton et al., Nat Biotechnol. 2015 May;33(5):510-7; Gordley et al., Proc Natl Acad Sci U S A. 2009 Mar 31 ;106(13):5053-8; Akopian et al., Proc Natl Acad Sci U S A. 2003 Jul 22;100(15):8688-91; Tan et al., J Virol. 2006 Feb;80(4):1939-48; Tan et al., Proc Natl Acad Sci U S A. 2003 Oct 14; 100(21): 11997- 2002; Papworth et al., Proc Natl Acad Sci U S A. 2003 Feb 18; 100(4): 1621 -6; Sanjana et al., Nat Protoc. 2012 Jan 5;7(1):171-92; Beerli et al., Proc Natl Acad Sci U S A. 1998 Dec 8; 95(25): 14628-33; Snowden et al., Curr Biol. 2002 Dec 23;12(24):2159-66; Xu et al., Cell Discov. 2016 May 3;2:16009; Komor et al., Nature. 2016 Apr 20;533(7603):420-4; Chaikind et al., Nucleic Acids Res. 2016 Aug 11 ; Choudhury et al., Oncotarget. 2016 Jun 23; Du et al., Cold Spring Harb Protoc. 2016 Jan 4; Pham et al., Methods Mol Biol. 2016;1358:43-57; Balboa et al., Stem Cell Reports. 2015 Sep 8;5(3):448-59; Hara et al., Sci Rep. 2015 Jun 9;5: 11221 ; Piatek et al., Plant Biotechnol J. 2015 May;13(4):578-89; Hu et al., Nucleic Acids Res. 2014 Apr;42(7):4375-90; Cheng et al., Cell Res. 2013 Oct;23(10):1163-71; and Maeder et al., Nat Methods. 2013 Oct;10(10):977-9. Examples of various suitable fusion partners include, but are not limited to those described in the following patents and applications (which disclose Cas9 proteins fused with various fusion partners): PCT patent applications: W02010075303, WO2012068627, and WO2013155555; U.S. patents: 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871 ,445; 8,865,406; 8,795,965;8,771 ,945; and 8,697,359; and U.S. patent applications: 20160304846, 20160215276, 20150166980, 20150071898, 20140068797; 20140170753; 20140179006;20140179770; 20140186843; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046; 20140273037; 20140273226; 20140273230; 20140273231 ; 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557; 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063; 20140335620; 20140342456; 20140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; all of which are hereby incorporated by reference in their entirety.Linkers

[0128] In some embodiments, a subject Cas9 fusion polypeptide can include a Cas9 protein that is linked to one or more heterologous polypeptides (one or more fusion partners) via a linker polypeptide (e.g., one or more linker polypeptides). In some embodiments, a Cas9 fusion polypeptide can be linked at the C-terminal and / or N-terminal end to a heterologous polypeptide (fusion partner) via a linker polypeptide (e.g., one or more linker polypeptides).

[0129] The linker polypeptide may have any of a variety of amino acid sequences. Proteins can be joined by a spacer peptide, generally of a flexible nature, although other chemical linkages are not excluded. Suitable linkers include polypeptides of between 4 amino acids and 40 amino acids in length, or between 4 amino acids and 25 amino acids in length. These linkers are generally produced by using synthetic, linkerencoding oligonucleotides to couple the proteins. Peptide linkers with a degree of flexibility can be used. The linking peptides may have virtually any amino acid sequence, bearing in mind that the preferred linkers will have a sequence that results in a generally flexible peptide. The use of small amino acids, such as glycine and alanine, are of use in creating a flexible peptide. The creation of such sequences is routine to those of skill in the art. A variety of different linkers are commercially available and are considered suitable for use.

[0130] Example linker polypeptides include glycine polymers (G)n, glycine-serine polymers (including, for example, (GS)n, GSGGSn (SEQ ID NO: 82), GGSGGSn (SEQ ID NO: 83), and GGGSn (SEQ ID NO: 84), where n is an integer of at least one), glycinealanine polymers, alanine-serine polymers. Example linkers can comprise amino acid sequences including, but not limited to, GGSG (SEQ ID NO: 77), GGSGG (SEQ IDNO: 78), GSGSG (SEQ ID NO: 76), GSGGG (SEQ ID NO: 79), GGGSG (SEQ ID NO: 80), GSSSG (SEQ ID NO: 81), and the like. The ordinarily skilled artisan will recognize that design of a peptide conjugated to any elements described above can include linkers that are all or partially flexible, such that the linker can include a flexible linker as well as one or more portions that confer less flexible structure.Cas9 Guide RNA

[0131] A nucleic acid molecule that binds to a subject Cas9 fusion polypeptide, forming a ribonucleoprotein complex (RNP), and targets the complex to a specific location within a target nucleic acid is referred to herein as a “Cas9 guide RNA,” a “guide RNA,” or simply a “guide.” A Cas9 guide RNA can be used to guide the protein to the target sequence. It is to be understood that in some cases, a hybrid DNA / RNA can be made such that a guide RNA includes DNA bases in addition to RNA bases - but the term “guide RNA” is still used herein to encompass such hybrid molecules. A subject guide RNA includes a guide sequence (also referred to as a “spacer”) (that hybridizes to target sequence of a target nucleic acid, e.g., target DNA) and a constant region (e.g., a region that is adjacent to the guide sequence and binds to the Cas9 fusion polypeptide). A “constant region” can also be referred to herein as a “protein-binding segment” or “scaffold” or “handle.” In many instances, the guide sequence and the constant region (scaffold) of a Cas9 guide RNA are heterologous to one another; i.e., the guide sequence and the constant region do not occur together in nature in a guide RNA (e.g., in some cases the guide sequence hybridizes to a target sequence of a eukaryotic cell and that target sequence is not targeted by the wild type system in nature).

[0132] As would be understood to one of ordinary skill in the art, the guide RNA can be introduced into a cell as an RNA (or as a DNA / RNA hybrid) or can be introduced as a nucleic acid encoding the RNA (e.g., a DNA such as an expression vector such as a viral, plasmid, or minicircle DNA), in which case the cell transcribes the RNA from the introduced DNA. In some cases, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., a Pol III promoter such as U6 or H1). In some cases, one or more guide RNAs (e.g., 1 , 2, 3, 4, 5, 6, 1-10, 1-8, 1-6, 1-5, 1-4, 1-3, 2- 10, 2-8, 2-6, 2-5, 2-4, 3-10, 3-8, 3-6, 3-5, two or more, three or more, four or more, or five or more) (or nucleotide sequences that encode said guide RNAs) can be introduced into the same cell (e.g., to target different sequences of the same target nucleic, to target different target nucleic acids, etc.).

[0133] A guide RNA can be said to include two segments, a targeting segment and a proteinbinding segment. The targeting segment of a guide RNA includes a nucleotide sequence (a guide sequence) that is complementary to (and therefore hybridizes with) a specific sequence (a target site) within a target nucleic acid (e.g., a target ssRNA, a target ssDNA, the complementary strand of a double stranded target DNA, etc.). The protein-binding segment (or “protein-binding sequence”) interacts with (binds to) a subject Cas9 fusion polypeptide. The protein-binding segment of a subject guide RNA includes two complementary stretches of nucleotides that hybridize to one another to form a double stranded RNA duplex (dsRNA duplex). Site-specific binding and / or cleavage of a target nucleic acid (e.g., genomic DNA) can occur at locations (e.g., target sequence of a target locus) determined by base-pairing complementarity between the guide RNA (the guide sequence of the guide RNA) and the target nucleic acid.

[0134] A Cas9 guide RNA and a Cas9 fusion polypeptide form a complex (e.g., bind via non- covalent interactions). The Cas9 guide RNA provides target specificity to the complex by including a targeting segment, which includes a guide sequence (a nucleotide sequence that is complementary to a sequence of a target nucleic acid). The Cas9 fusion polypeptide of the complex provides the site-specific activity (e.g., cleavage activity or an activity provided by a fusion partner). In other words, the Cas9 fusion polypeptide is guided to a target nucleic acid sequence (e.g. a target sequence in a chromosomal nucleic acid, e.g., a chromosome; a target sequence in an extrachromosomal nucleic acid, e.g. an episomal nucleic acid, a minicircle, an ssRNA, an ssDNA, etc.; a target sequence in a mitochondrial nucleic acid; a target sequence in a chloroplast nucleic acid; a target sequence in a plasmid; a target sequence in a viral nucleic acid; etc.) by virtue of its association with the Cas9 guide RNA.

[0135] The “guide sequence” also referred to as the “targeting sequence” of a Cas9 guide RNA can be modified so that the Cas9 guide RNA can target a Cas9 fusion polypeptide, to any desired sequence of any desired target nucleic acid, with the exception (e.g., as described herein) that the PAM sequence can be taken into account. Thus, for example, a Cas9 guide RNA can have a targeting segment with a guide sequence that has complementarity with (e.g., can hybridize to) a sequence in a nucleic acid in a eukaryotic cell, e.g., a viral nucleic acid, a eukaryotic nucleic acid (e.g., a eukaryotic chromosome, chromosomal sequence, a eukaryotic RNA, etc.), and the like.

[0136] In subject guide RNA can also be said to include an “activator” (e.g., a tracrRNA) and a “targeter” (e.g., a crRNA) (e.g., an “activator-RNA” and a “targeter-RNA”, respectively). When the “activator” and a “targeter” are two separate molecules the guide RNA is referred to herein as a “dual guide RNA”, a “dgRNA,” a “doublemolecule guide RNA”, or a “two-molecule guide RNA.” In some embodiments, the activator and targeter are covalently linked to one another (e.g., via intervening nucleotides) and the guide RNA is referred to herein as a “single guide RNA”, an “sgRNA,” a “single-molecule guide RNA,” or a “one-molecule guide RNA”. Thus, a subject single guide RNA comprises a targeter (e.g., targeter-RNA) and an activator (e.g., activator-RNA) that are linked to one another (e.g., by intervening nucleotides, or can be linked via a non-nucleic acid chemical linker), and hybridize to one another to form the double stranded RNA duplex (dsRNA duplex) of the protein-binding segment of the guide RNA, thus resulting in a stem-loop structure. Thus, the targeter and the activator each have a duplex-forming segment, where the duplex forming segment of the targeter and the duplex-forming segment of the activator have complementarity with one another and hybridize to one another.

[0137] In some embodiments, the linker of a single guide RNA is a stretch of nucleotides. In some cases, the targeter and activator of a single guide RNA are linked to one another by intervening nucleotides and the linker can have a length of from 3 to 20 nucleotides (nt) (e.g., from 3 to 15, 3 to 12, 3 to 10, 3 to 8, 3 to 6, 3 to 5, 3 to 4, 4 to 20, 4 to 15, 4 to 12, 4 to 10, 4 to 8, 4 to 6, or 4 to 5 nt). In some embodiments, the linker of a single guide RNA can have a length of from 3 to 100 nucleotides (nt) (e.g., from 3 to 80, 3 to 50, 3 to 30, 3 to 25, 3 to 20, 3 to 15, 3 to 12, 3 to 10, 3 to 8, 3 to 6, 3 to 5, 3 to 4, 4 to 100, 4 to 80, 4 to 50, 4 to 30, 4 to 25, 4 to 20, 4 to 15, 4 to 12, 4 to 10, 4 to 8, 4 to 6, or 4 to 5 nt). In some embodiments, the linker of a single guide RNA can have a length of from 3 to 10 nucleotides (nt) (e.g., from 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 to 4, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, or 4 to 5 nt). In some embodiments, the linker is 4 nt.Guide sequence of a guide RNA

[0138] The targeting segment of a subject guide RNA includes a guide sequence (i.e. , a targeting sequence), which is a nucleotide sequence that is complementary to a sequence (a target site) in a target nucleic acid. In other words, the targeting segment of a guide RNA can interact with a target nucleic acid (e.g., double stranded DNA (dsDNA), single stranded DNA (ssDNA), single stranded RNA (ssRNA), or doublestranded RNA (dsRNA)) in a sequence-specific manner via hybridization (i.e., base pairing). The guide sequence of a guide RNA can be modified (e.g., by genetic engineering) / designed to hybridize to any desired target sequence (e.g., while taking the PAM into account, e.g., when targeting a dsDNA target) within a target nucleic acid (e.g., a eukaryotic target nucleic acid such as genomic DNA).

[0139] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or more (e.g., 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100%.

[0140] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over the seven contiguous 3’-most nucleotides of the target site of the target nucleic acid (i.e., the nucleotides on the PAM-side of the target DNA).

[0141] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides.

[0142] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or more (e.g., 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or more (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or more (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19-25 contiguous nucleotides.

[0143] In some cases, the guide sequence has a length of 17-30 nucleotides (nt) (e.g., 17- 28, 17-25, 17-23, 17-22, 17-21 , 17-20, 17-19, 17-18, 18-30, 18-28, 18-25, 18-23, 18-22, 18-21 , 18-20, 18-19, 19-30, 19-28, 19-25, 19-23, 19-22, 19-21 , 19-20, 20-30, 20-28, 20-25, 20-23, 20-22, 20-21 , 21-30, 21-28, 21-25, 21-23, or 21-22 nt). In some cases, the guide sequence has a length of 19-25 nt (e.g., 19-22, 19-20, 20-25, 20-25, or 20-22 nt). In some cases, the guide sequence has a length of 19 or more nt (e.g.,20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 17-18 nt. In some cases, the guide sequence has a length of 19-25 nt. In some cases, the guide sequence has a length of 19-22 nt. In some cases, the guide sequence has a length of 17-22 nt. In some cases, the guide sequence has a length of 19 nt. In some cases, the guide sequence has a length of 20 nt. In some cases, the guide sequence has a length of21 nt. In some cases, the guide sequence has a length of 22 nt. In some cases, the guide sequence has a length of 23 nt.Protein-binding segment of a guide RNA

[0144] The protein-binding segment of a subject guide RNA interacts with a Cas9 fusion polypeptide. The guide RNA guides the bound Cas9 fusion polypeptide (as an RNP) to a specific nucleotide sequence within target nucleic acid via the above mentioned guide sequence. The protein-binding segment of a guide RNA comprises two stretches of nucleotides (the duplex-forming segment of the activator and the duplexforming segment of the targeter) that are complementary to one another and hybridizeto form a double stranded RNA duplex (dsRNA duplex). Thus, the protein-binding segment includes a dsRNA duplex.

[0145] In some cases, the dsRNA duplex region formed between the activator (activator- RNA) and targeter (targeter-RNA) (i.e., the activator / targeter dsRNA duplex) (e.g., in dual or single guide RNA format) includes 8-25 base pairs (bp) (e.g., from 8-22, 8-18, 8-15, 8-12, 12-25, 12-22, 12-18, 12-15, 13-25, 13-22, 13-18, 13-15, 14-25, 14-22, 14- 18, 14-15, 15-25, 15-22, 15-18, 17-25, 17-22, or 17-18 bp, e.g., 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, 20 bp, 21 bp, etc.). In some cases, the duplex region (e.g., in dual or single guide RNA format) includes 8 or more bp (e.g., 10 or more, 12 or more, 15 or more, or 17 or more bp). As would be known to one of ordinary skill in the art, not all nucleotides of the duplex region need to be paired, and therefore the duplex forming region can include a bulge. The term “bulge” herein is used to mean a stretch of nucleotides (which can be one nucleotide) that do not contribute to a double stranded duplex, but which are surround 5’ and 3’ by nucleotides that do contribute, and as such a bulge is considered part of the duplex region. In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) includes 1 bulge. In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) includes 1 or more bulges (e.g., 2 or more, 3 or more, 4 or more bulges). In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) includes 2 or more bulges (e.g., 3 or more, 4 or more bulges). In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) includes 1-5 bulges (e.g., 1-4, 1-3, 2-5, 2-4, or 2-3 bulges).

[0146] The duplex region of a subject guide RNA (in dual guide or single guide RNA format) can include one or more (1 , 2, 3, 4, 5, etc.) mutations relative to a naturally occurring duplex region. For example, in some cases a base pair can be maintained while the nucleotides contributing to the base pair from each segment (targeter and activator) can be different. In some cases, the duplex region of a subject guide RNA includes more paired bases, less paired bases, a smaller bulge, a larger bulge, fewer bulges, more bulges, or any convenient combination thereof, as compared to a naturally occurring duplex region (of a naturally occurring guide RNA).

[0147] In some cases, the activator (activator- RNA) of a subject guide RNA (in dual or single guide RNA format) includes one or more (in some cases 2, in some cases 3) internal RNA duplexes in addition to the activator / targeter dsRNA duplex). Internal RNAduplexes (hairpins) of the activator can be positioned 3’ of the activator / targeter dsRNA duplex (meaning - 3’ of the activator sequence). In some cases, the activator includes one hairpin positioned 3’ of the activator / targeter dsRNA duplex. In some cases, the activator includes two hairpins positioned 3’ of the activator / targeter dsRNA duplex. In some cases, the activator includes three hairpins positioned 3’ of the activator / targeter dsRNA duplex. In some cases, the activator includes two or more hairpins (e.g., 3 or more or 4 or more hairpins) positioned 3’ of the activator / targeter dsRNA duplex. In some cases, the activator includes 2 to 5 hairpins (e.g., 2 to 4, or 2 to 3 hairpins) positioned 3’ of the activator / targeter dsRNA duplex. In some cases, the activator-RNA (e.g., in dual or single guide RNA format) comprises at least 2 nucleotides (nt) (e.g., at least 3 or at least 4 nt) 3’ of the 3’-most hairpin stem. In some cases, the activator-RNA (e.g., in dual or single guide RNA format) comprises at least 4 nt 3’ of the 3’-most hairpin stem.

[0148] Internal RNA duplexes (hairpins) (e.g., one or more, two or more, one, two, or three) can also be positioned 5’ of the guide sequence (and there can be considered 5’ of the targeter (targeter-RNA / crRNA). In some cases, a guide RNA includes 1-4 hairpins positioned 5’ of the guide sequence. In some cases, a guide RNA includes 1- 2 hairpins positioned 5’ of the guide sequence. In some cases, a guide RNA includes 1 hairpin positioned 5’ of the guide sequence.

[0149] In some cases, the activator-RNA (e.g., in dual or single guide format) has a length of 65 nucleotides (nt) or more (e.g., 66 or more, 67 or more, 68 or more, 69 or more, 70 or more, or 75 or more nt). In some cases, the activator-RNA (e.g., in dual or single guide format) has a length of 66 nt or more (e.g., 67 or more, 68 or more, 69 or more, 70 or more, or 75 or more nt). In some cases, the activator-RNA (e.g., in dual or single guide format) has a length of 67 nt or more (e.g., 68 or more, 69 or more, 70 or more, or 75 or more nt). In some cases, the activator-RNA has a length of from 80 nt to 100 nt. In some cases, the activator-RNA has a length of 80 nt, 81 nt, 82 nt, 83 nt, 84 nt, 85 nt, 86 nt, 87 nt, 88 nt, 89 nt, 90 nt, 91 nt, 92 nt, 93 nt, 94 nt, 95 nt, 96 nt, 97 nt, 98 nt, 99 nt, or 100 nt (or more than 100 nt).

[0150] In some cases, the Cas9 guide RNA is a sgRNA that is 100-200 nt long (e.g., 100- 180, 100-160, 100-140, 100-120, 110-200, 110-180, 110-160, 110-140, 110-120,120- 200, 120-180, 120-160, 120-140, 130-200, 130-180, 130-160, or 130-140).

[0151] In some cases, the activator-RNA (e.g., in dual or single guide format) includes 45 or more nucleotides (nt) (e.g., 46 or more, 47 or more, 48 or more, 49 or more, 50 or more, 51 or more, 52 or more, 53 or more, 54 or more, or 55 or more nt) 3’ of thedsRNA duplex formed between the activator and the targeter (the activator / targeter dsRNA duplex). In some cases, the activator is truncated at the 5’ end relative to a naturally occurring activator. In some cases, the activator is extended at the 5’ end relative to a naturally occurring activator. In some cases, the activator is truncated at the 3’ end relative to a naturally occurring activator. In some cases, the activator is extended at the 3’ end relative to a naturally occurring activator.

[0152] The term “activator” or “activator RNA” is used herein to mean a tracrRNA-like molecule (tracrRNA: “trans-acting CRISPR RNA”) of a dual guide RNA (and therefore of a single guide RNA when the “activator” and the “targeter” are linked together by, e.g., intervening nucleotides). Thus, for example, a guide RNA (dgRNA or sgRNA) comprises an activator sequence (e.g., a tracrRNA sequence). A tracr molecule (a tracrRNA) is a naturally existing molecule that hybridizes with a CRISPR RNA molecule (a crRNA) to form a dual guide RNA. The term “activator” is used herein to encompass naturally existing tracrRNAs, but also to encompass tracrRNAs with modifications (e.g., truncations, extensions, sequence variations, base modifications, backbone modifications, linkage modifications, etc.) where the activator retains at least one function of a tracrRNA (e.g., contributes to the dsRNA duplex to which protein binds). In some cases, the activator provides one or more stem loops that can interact with protein. An activator can be referred to as having a tracr sequence (tracrRNA sequence) and in some cases is a tracrRNA, but the term “activator” is not limited to naturally existing tracrRNAs.

[0153] The term “targeter” or “targeter RNA” is used herein to refer to a crRNA-like molecule (crRNA: “CRISPR RNA”) of a dual guide RNA (and therefore of a single guide RNA when the “activator” and the “targeter” are linked together, e.g., by intervening nucleotides). Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a guide sequences and a duplex-forming segment (e.g., a duplex forming segment of a crRNA, which can also be referred to as a crRNA repeat). Because the sequence of a targeting segment (the segment that hybridizes with a target sequence of a target nucleic acid) of a targeter is modified by a user to hybridize with a desired target nucleic acid, the sequence of a targeter will often be a non-naturally occurring sequence. However, the duplex-forming segment of a targeter (described in more detail herein), which hybridizes with the duplex-forming segment of an activator, can include a naturally existing sequence (e.g., can include the sequence of a duplexforming segment of a naturally existing crRNA, which can also be referred to as a crRNA repeat). Thus, the term targeter is used herein to distinguish from naturallyoccurring crRNAs, despite the fact that part of a targeter (e.g., the duplex-forming segment) often includes a naturally occurring sequence from a crRNA. However, the term “targeter” encompasses naturally occurring crRNAs.

[0154] As noted above, a targeter comprises both the guide sequence of the guide RNA and a stretch (a “duplex-forming segment”) of nucleotides that forms one half of the dsRNA duplex of the protein-binding segment of the guide RNA. A corresponding tracrRNA-like molecule (activator) comprises a stretch of nucleotides (a duplexforming segment) that forms the other half of the dsRNA duplex of the protein-binding segment of the guide RNA. In other words, a stretch of nucleotides of the targeter is complementary to and hybridizes with a stretch of nucleotides of the activator to form the dsRNA duplex of the protein-binding segment of a guide RNA. As such, each targeter can be said to have a corresponding activator (which has a region that hybridizes with the targeter). The targeter molecule additionally provides the guide sequence. Thus, a targeter and an activator (as a corresponding pair) hybridize to form a guide RNA. The particular sequence of a given naturally existing crRNA or tracrRNA molecule can be characteristic of the species in which the RNA molecules are found.

[0155] Guide RNA sequences (e.g., for the scaffold) for many different Cas9 proteins are known, see, e.g., Gasiunas et al., Nat Commun. 2020 Nov 2;11 (1):5512: “A catalogue of biochemically diverse CRISPR-Cas9 orthologs.” A subject variant Cas9 protein can use the same guide RNA scaffolds that can be used with the corresponding wild type Cas9. As an illustrative example, in some cases, a subject Cas9 fusion polypeptide having an NmeCas9 protein is used with NmeCas9 guide RNA scaffolds, while a subject Cas9 fusion polypeptide having a SpyCas9 protein (S. pyogenes Cas9) is used with SpyCas9 guide RNA scaffolds, and a subject Cas9 fusion polypeptide having a SaCas9 protein (S. aureus Cas9) is used with SaCas9 guide RNA scaffolds. For example, in some cases, the portion of the targeter-RNA (e.g., crRNA) that contributes to the scaffold (i.e., is 3’ of the guide sequence) (e.g., when using an S. pyogenes Cas9 protein) includes: 5’-GUUUUAGAGCUAUGCUGUUUUG-3' (SEQ ID NO: 101). In some cases, it includes: 5 -GUUUUAGAGCUA-3' (SEQ ID NO: 102). in some cases, the activator-RNA (e.g., tracrRNA) (e.g., when using an S. pyogenes Cas9 protein) includes: ’5- AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUG GCACCGAGUCGGUGCUU-3' (SEQ ID NO: 103). In some cases (e.g., when usingan S. pyogenes Cas9 protein) a sgRNA includes 5’-GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCG-3' (SEQ ID NO: 104). In some cases (e.g., when using an S. pyogenes Cas9 protein) a sgRNA includes 5?-GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGCUU-3' (SEQ ID NO: 105). Mutations / variants of the above sequences can also be used and many suitable examples will be known to one of ordinary skill in the art.Nucleic Acid Modifications

[0156] As would be understood to one of ordinary skill in the art, a guide RNA (and / or a donor polynucleotide) can have one or more modifications, e.g., a base modification, a backbone modification, a sugar modification, etc., to provide the nucleic acid with a new or enhanced feature (e.g., improved stability).

[0157] In some cases, a guide RNA (and / or a donor polynucleotide) comprises one or more modifications, e.g., a base modification, a backbone modification, a sugar modification, etc., to provide the nucleic acid with a new or enhanced feature (e.g., improved stability). As is known in the art, a nucleoside is a base-sugar combination. The base portion of the nucleoside is normally a heterocyclic base. The two most common classes of such heterocyclic bases are the purines and the pyrimidines. Nucleotides are nucleosides that further include a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include a pentofuranosyl sugar, the phosphate group can be linked to the 2', the 3', or the 5' hydroxyl moiety of the sugar. In forming oligonucleotides, the phosphate groups covalently link adjacent nucleosides to one another to form a linear polymeric compound. In turn, the respective ends of this linear polymeric compound can be further joined to form a circular compound, however, linear compounds are generally suitable. In addition, linear compounds may have internal nucleotide base complementarity and may therefore fold in a manner as to produce a fully or partially double-stranded compound. Within oligonucleotides, the phosphate groups are commonly referred to as forming the internucleoside backbone of the oligonucleotide. The normal linkage or backbone of RNA and DNA is a 3' to 5' phosphodiester linkage.Modified backbones and modified internucleoside linkages

[0158] Examples of suitable guide RNA (and / or a donor polynucleotide) modifications include modified nucleic acid backbones and non-natural internucleoside linkages. Nucleic acids having modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.

[0159] Suitable modified oligonucleotide backbones containing a phosphorus atom therein include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates including 3'-alkylene phosphonates, 5'-alkylene phosphonates and chiral phosphonates, phosphinates, phosphoramidates including 3'-amino phosphoramidate and aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates and boranophosphates having normal 3'-5' linkages, 2'-5' linked analogs of these, and those having inverted polarity wherein one or more internucleotide linkages is a 3' to 3', 5' to 5' or 2' to 2' linkage. Suitable oligonucleotides having inverted polarity comprise a single 3' to 3' linkage at the 3'- most internucleotide linkage i.e. a single inverted nucleoside residue which may be a basic (the nucleobase is missing or has a hydroxyl group in place thereof). Various salts (such as, for example, potassium or sodium), mixed salts and free acid forms are also included.

[0160] In some cases, a guide RNA (and / or a donor polynucleotide) comprises one or more phosphorothioate and / or heteroatom internucleoside linkages, in particular -CH2-NH- O-CH2-, -CH2-N(CH3)-O-CH2- (known as a methylene (methylimino) or MMI backbone), -CH2-O-N(CH3)-CH2-, -CH2-N(CH3)-N(CH3)-CH2- and -O-N(CH3)-CH2-CH2- (wherein the native phosphodiester internucleotide linkage is represented as -O- P(=O)(OH)-O-CH2-). MMI type internucleoside linkages are disclosed in the above referenced U.S. Pat. No. 5,489,677. Suitable amide internucleoside linkages are disclosed in U.S. Pat. No. 5,602,240.

[0161] Also suitable are nucleic acids having morpholino backbone structures as described in, e.g., U.S. Pat. No. 5,034,506. For example, in some cases, a guide RNA (and / or a donor polynucleotide) comprises a 6-membered morpholino ring in place of a ribose ring. In some cases, a phosphorodiamidate or other non-phosphodiester internucleoside linkage replaces a phosphodiester linkage.

[0162] Suitable modified polynucleotide backbones that do not include a phosphorus atom therein have backbones that are formed by short chain alkyl or cycloalkylinternucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatomic or heterocyclic internucleoside linkages. These include those having morpholino linkages (formed in part from the sugar portion of a nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methylene formacetyl and thioformacetyl backbones; riboacetyl backbones; alkene containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2component parts.Mimetics

[0163] A guide RNA (and / or a donor polynucleotide) can be a nucleic acid mimetic. The term "mimetic" as it is applied to polynucleotides is intended to include polynucleotides wherein only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups, replacement of only the furanose ring is also referred to in the art as being a sugar surrogate. The heterocyclic base moiety or a modified heterocyclic base moiety is maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid, a polynucleotide mimetic that has been shown to have excellent hybridization properties, is referred to as a peptide nucleic acid (PNA). In PNA, the sugar-backbone of a polynucleotide is replaced with an amide containing backbone, in particular an aminoethylglycine backbone. The nucleotides are retained and are bound directly or indirectly to aza nitrogen atoms of the amide portion of the backbone.

[0164] One polynucleotide mimetic that has been reported to have excellent hybridization properties is a peptide nucleic acid (PNA). The backbone in PNA compounds is two or more linked aminoethylglycine units which gives PNA an amide containing backbone. The heterocyclic base moieties are bound directly or indirectly to aza nitrogen atoms of the amide portion of the backbone. Representative U.S. patents that describe the preparation of PNA compounds include, but are not limited to: U.S. Pat. Nos. 5,539,082; 5,714,331; and 5,719,262.

[0165] Another class of polynucleotide mimetic that has been studied is based on linked morpholino units (morpholino nucleic acid) having heterocyclic bases attached to the morpholino ring. A number of linking groups have been reported that link the morpholino monomeric units in a morpholino nucleic acid. One class of linking groups has been selected to give a non-ionic oligomeric compound. The non-ionicmorpholino-based oligomeric compounds are less likely to have undesired interactions with cellular proteins. Morpholino-based polynucleotides are non-ionic mimics of oligonucleotides which are less likely to form undesired interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholino-based polynucleotides are disclosed in U.S. Pat. No. 5,034,506. A variety of compounds within the morpholino class of polynucleotides have been prepared, having a variety of different linking groups joining the monomeric subunits.

[0166] A further class of polynucleotide mimetic is referred to as cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in a DNA / RNA molecule is replaced with a cyclohexenyl ring. CeNA DMT protected phosphoramidite monomers have been prepared and used for oligomeric compound synthesis following classical phosphoramidite chemistry. Fully modified CeNA oligomeric compounds and oligonucleotides having specific positions modified with CeNA have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, 122, 8595-8602). In general, the incorporation of CeNA monomers into a DNA chain increases its stability of a DNA / RNA hybrid. CeNA oligoadenylates formed complexes with RNA and DNA complements with similar stability to the native complexes. The study of incorporating CeNA structures into natural nucleic acid structures was shown by NMR and circular dichroism to proceed with easy conformational adaptation.

[0167] A further modification includes Locked Nucleic Acids (LNAs) in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring thereby forming a 2'-C,4'-C- oxymethylene linkage thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-), group bridging the 2' oxygen atom and the 4' carbon atom wherein n is 1 or 2 (Singh et al., Chem. Commun., 1998, 4, 455-456). LNA and LNA analogs display very high duplex thermal stabilities with complementary DNA and RNA (Tm=+3 to +10° C), stability towards 3'-exonucleolytic degradation and good solubility properties. Potent and nontoxic antisense oligonucleotides containing LNAs have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 2000, 97, 5633- 5638).

[0168] The synthesis and preparation of the LNA monomers adenine, cytosine, guanine, 5- methyl-cytosine, thymine and uracil, along with their oligomerization, and nucleic acid recognition properties have been described (Koshkin et al., Tetrahedron, 1998, 54, 3607-3630). LNAs and preparation thereof are also described in WO 98 / 39352 and WO 99 / 14226.Modified sugar moieties

[0169] A guide RNA (and / or a donor polynucleotide) can also include one or more substituted sugar moieties. Suitable polynucleotides comprise a sugar substituent group selected from: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S- or N- alkynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl and alkynyl may be substituted or unsubstituted C.sub.1 to C10 alkyl or C2 to C10 alkenyl and alkynyl. Particularly suitable are O((CH2)nO)mCH3, O(CH2)nOCH3, O(CH2)nNH2, O(CH2)nCH3, O(CH2)nONH2, and O(CH2)nON((CH2)nCH3)2, where n and m are from 1 to about 10. Other suitable polynucleotides comprise a sugar substituent group selected from: Ci to C10 lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, an RNA cleaving group, a reporter group, an intercalate^ a group for improving the pharmacokinetic properties of an oligonucleotide, or a group for improving the pharmacodynamic properties of an oligonucleotide, and other substituents having similar properties. A suitable modification includes 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-0-(2-methoxyethyl) or 2'-MOE) (Martin et al., Helv. Chim. Acta, 1995, 78, 486-504) i.e., an alkoxyalkoxy group. A further suitable modification includes 2'-dimethylaminooxyethoxy, i.e., a O(CH2)2ON(CH3)2group, also known as 2 -DMAOE, as described in examples hereinbelow, and 2'- dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethyl-amino-ethoxy- ethyl or 2 -DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2.

[0170] Other suitable sugar substituent groups include methoxy (-O-CH3), aminopropoxy (-0 CH2CH2CH2NH2), allyl (-CH2-CH=CH2), -O-allyl (-O- CH2— CH=CH2) and fluoro (F). 2'-sugar substituent groups may be in the arabino (up) position or ribo (down) position. A suitable 2'-arabino modification is 2'-F. Similar modifications may also be made at other positions on the oligomeric compound, particularly the 3' position of the sugar on the 3' terminal nucleoside or in 2'-5' linked oligonucleotides and the 5' position of 5' terminal nucleotide. Oligomeric compounds may also have sugar mimetics such as cyclobutyl moieties in place of the pentofuranosyl sugar.Base modifications and substitutions

[0171] A guide RNA (and / or a donor polynucleotide) may also include nucleobase (often referred to in the art simply as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include the purine bases adenine (A) andguanine (G), and the pyrimidine bases thymine (T), cytosine (C) and uracil (U). Modified nucleobases include other synthetic and natural nucleobases such as 5- methylcytosine (5-me-C), 5-hydroxymethyl cytosine, xanthine, hypoxanthine, 2- aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2- thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo uracil, cytosine and thymine, 5- uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo particularly 5-bromo, 5- trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7- methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7- deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine. Further modified nucleobases include tricyclic pyrimidines such as phenoxazine cytidine(1 H-pyrimido(5,4-b)(1 ,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1 H- pyrimido(5,4-b)(1 ,4)benzothiazin-2(3H)-one), G-clamps such as a substituted phenoxazine cytidine (e.g. 9-(2-aminoethoxy)-H-pyrimido(5,4-(b) (1 ,4)benzoxazin- 2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).

[0172] Heterocyclic base moieties may also include those in which the purine or pyrimidine base is replaced with other heterocycles, for example 7-deaza-adenine, 7- deazaguanosine, 2-aminopyridine and 2-pyridone. Further nucleobases include those disclosed in U.S. Pat. No. 3,687,808, those disclosed in The Concise Encyclopedia Of Polymer Science And Engineering, pages 858-859, Kroschwitz, J. I., ed. John Wiley & Sons, 1990, those disclosed by Englisch et al., Angewandte Chemie, International Edition, 1991 , 30, 613, and those disclosed by Sanghvi, Y. S., Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, S. T. and Lebleu, B., ed., CRC Press, 1993. Certain of these nucleobases are useful for increasing the binding affinity of an oligomeric compound. These include 5-substituted pyrimidines, 6- azapyrimidines and N-2, N-6 and O-6 substituted purines, including 2- aminopropyladenine, 5-propynyluracil and 5-propynylcytosine. 5-methylcytosine substitutions have been shown to increase nucleic acid duplex stability by 0.6-1.2° C. (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278) and are suitable base substitutions, e.g., when combined with 2'- O-methoxyethyl sugar modifications.Protospacer adjacent motif (PAM)

[0173] A wild type CRISPR / Cas protein (e.g., Cas9 protein) normally has nuclease activity that cleaves a target nucleic acid (e.g., a double stranded DNA (dsDNA)) at a target site defined by the region of complementarity between the guide sequence of the guide RNA and the target nucleic acid. In some cases, site-specific targeting to the target nucleic acid occurs at locations determined by both (i) base-pairing complementarity between the guide nucleic acid and the target nucleic acid; and (ii) a short motif referred to as the “protospacer adjacent motif” (PAM) in the target nucleic acid. For example, when a Cas9 protein binds to (in some cases cleaves) a dsDNA target nucleic acid, the PAM sequence that is recognized (bound) by the Cas9 protein is present on the non-complementary strand (the strand that does not hybridize with the targeting segment of the guide nucleic acid) of the target DNA. In some cases, a PAM sequence has a length in a range of from 1 nt to 15 nt (e.g., 1 nt to 14 nt, 1 nt to 13 nt, 1 nt to 12 nt, 1 nt to 11 nt, 1 nt to 10 nt, 1 nt to 9 nt, 1 nt to 9 nt, 1 nt to 8 nt, 1 nt to 7 nt, 1 nt to 6 nt, 1 nt to 5 nt, 1 nt to 4 nt, 1 nt to 3 nt, 2 nt to 15 nt, 2 nt to 14 nt, 2 nt to 13 nt, 2 nt to 12 nt, 2 nt to 11 nt, 2 nt to 10 nt, 2 nt to 9 nt, 2 nt to 8 nt, 2 nt to 7 nt, 2 nt to 6 nt, 2 nt to 5 nt, 2 nt to 4 nt, 2 nt to 3 nt, 2 nt, or 3 nt).

[0174] CRISRPR / Cas (e.g., Cas9) proteins from different species can have different PAM sequence requirements. For example, in some embodiments (e.g., when the Cas9 protein is derived from S. pyogenes or a closely related Cas9 is used; see for example, Chylinski et al., RNA Biol. 2013 May;10(5):726-37; and Jinek et al., Science. 2012 Aug 17;337(6096):816-21 ; both of which are hereby incorporated by reference in their entirety), the PAM sequence can be NRG because the S. pyogenes Cas9 PAM (PAM sequence) is NAG or NGG (or NRG where “R” is A or G). For example, a Cas9 PAM sequence for S. pyogenes Cas9 can be: NGG, NAG, AGG, CGG, GGG, TGG, AAG, CAG, GAG, and TAG. In some cases, the PAM is NGG.

[0175] In some cases (e.g., when a Cas9 protein is derived from the Cas9 protein of Neisseria meningitidis or a closely related Cas9 is used), the PAM sequence (e.g., of a target nucleic acid) can be 5’-NNNNGANN-3’, 5’-NNNNGTTN-3’, 5’-NNNNGNNT-3’, 5’-NNNNGTNN-3’, 5’-NNNNGNTN-3’, or 5’-NNNNGATT-3’, where N is any nucleotide. In some embodiments (e.g., when a Cas9 protein is derived from Streptococcus thermophilus #1 or a closely related Cas9 is used), the PAM sequence (e.g., of a target nucleic acid) can be 5’-NNAGAA-3’, 5’-NNAGGA-3’, 5’-NNGGAA-3’, 5’-NNANAA-3’, or 5’-NNGGGA-3’ where N is any nucleotide. In some embodiments (e.g., when a Cas9 protein is derived from Treponema denticola (TD) or a closelyrelated Cas9 is used), the PAM sequence (e g., of a target nucleic acid) can be 5’-NAAAAN-3’, 5’-NAAAAC-3’, 5’-NAAANC-3’, 5’-NANAAC-3’, or 5’-NNAAAC-3’, where N is any nucleotide. As would be known by one of ordinary skill in the art, additional PAM sequences for other Cas9 proteins are known in the art and / or can readily be determined using bioinformatic analysis (e.g., analysis of genomic sequencing data), and / or routine experimentation. See Esvelt et al., Nat Methods. 2013 Nov;10(11):1116-21 , for additional information. In some cases, the PAM sequence for S. aureus Cas9 is 5’-NNGRR(N)-3’ (in some cases 5’-NNGRRT-3’).Additional resources

[0176] More information (including examples) related to various Cas9 guide RNAs, Cas9 proteins, and Cas9 PAMs can be found in the art, for example, see Jinek et al., Science. 2012 Aug 17;337(6096):816-21; Chylinski et al., RNA Biol. 2013 May;10(5):726-37; Ma et al., Biomed Res Int. 2013;2013:270805; Hou et al., Proc Natl Acad Sci U S A. 2013 Sep 24; 110(39): 15644-9; Jinek et al., Elife. 2013;2:e00471; Pattanayak et al., Nat Biotechnol. 2013 Sep;31(9):839-43; Qi et al., Cell. 2013 Feb 28; 152(5): 1173-83; Wang et al., Cell. 2013 May 9;153(4):910-8; Auer et al., Genome Res. 2013 Oct 31; Chen et al., Nucleic Acids Res. 2013 Nov 1 ;41(20):e19; Cheng et al., Cell Res. 2013 Oct;23(10):1163-71 ; Cho et al., Genetics. 2013 Nov; 195(3): 1177- 80; DiCarlo et al., Nucleic Acids Res. 2013 Apr;41(7):4336-43; Dickinson et al., Nat Methods. 2013 Oct; 10(10): 1028-34; Ebina et al., Sci Rep. 2013;3:2510; Fujii et al., Nucleic Acids Res. 2013 Nov 1 ;41(20):e187; Hu et al., Cell Res. 2013 Nov;23(11):1322-5; Jiang et al., Nucleic Acids Res. 2013 Nov 1 ;41(20):e188; Larson et al., Nat Protoc. 2013 Nov;8(11):2180-96; Mali et al., Nat Methods. 2013 Oct; 10(10): 957-63; Nakayama et al., Genesis. 2013 Dec;51(12):835-43; Ran et al., Nat Protoc. 2013 Nov;8(11):2281-308; Ran et al., Cell. 2013 Sep 12; 154(6): 1380-9; Upadhyay et al., G3 (Bethesda). 2013 Dec 9;3(12):2233-8; Walsh et al., Proc Natl Acad Sci U S A. 2013 Sep 24;110(39): 15514-5; Xie et al., Mol Plant. 2013 Oct 9; Yang et al., Cell. 2013 Sep 12;154(6):1370-9; Briner et al., Mol Cell. 2014 Oct 23;56(2):333-9; and U.S. patents and patent applications: 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871 ,445; 8,865,406; 8,795,965; 8,771 ,945; 8,697,359; 20160304846, 20160215276, 20150166980, 20150071898, 20140068797; 20140170753; 20140179006; 20140179770; 20140186843; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046;20140273037; 20140273226; 20140273230; 20140273231; 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557; 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063; 20140335620; 20140342456; 20140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; all of which are hereby incorporated by reference in their entirety.Nucleic Acids

[0177] The present disclosure provides a nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). The present disclosure provides a nucleic acid / protein complex comprising: a) a Cas9 fusion polypeptide; and b) a guide RNA. The present disclosure provides a nucleic acid / protein complex comprising: a) a fusion Cas9 protein of the present disclosure; and b) a guide RNA.

[0178] The present disclosure provides one or more nucleic acids (e.g., RNA and / or DNA) comprising one or more of: a donor polynucleotide sequence, a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide), a Cas9 guide RNA (which can include two separate nucleotide sequences in the case of dual guide RNA format or which can include a single nucleotide sequence in the case of single guide RNA format), and a nucleotide sequence encoding a Cas9 guide RNA.

[0179] The present disclosure provides a nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). The present disclosure provides a recombinant expression vector that comprises a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). The present disclosure provides a recombinant expression vector that comprises a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). The present disclosure provides a recombinant expression vector that comprises: a) a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide); and b) a nucleotide sequence encoding a Cas9 guide RNA(s). The present disclosure provides a recombinant expression vector that comprises: a) a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide); and b) anucleotide sequence encoding a Cas9 guide RNA(s). In some cases, the nucleotide sequence encoding the protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) and / or the nucleotide sequence encoding the Cas9 guide RNA is operably linked to a promoter that is operable in a cell type of choice (e.g., a prokaryotic cell, a eukaryotic cell, a plant cell, an animal cell, a mammalian cell, a primate cell, a rodent cell, a human cell, etc.).

[0180] In some cases, a nucleotide sequence encoding an Cas9 fusion polypeptide of the present disclosure is codon optimized. This type of optimization can entail a mutation of a protein-coding (e.g., Cas9-encoding) nucleotide sequence to mimic the codon preferences of the intended host organism or cell while encoding the same protein. Thus, the codons can be changed, but the encoded protein remains unchanged. For example, if the intended target cell was a human cell, a human codon-optimized protein-encoding (e.g., Cas9-encoding) nucleotide sequence could be used. As another non-limiting example, if the intended host cell were a mouse cell, then a mouse codon-optimized protein-encoding (e.g., Cas9-encoding) nucleotide sequence could be generated. As another non-limiting example, if the intended host cell were a plant cell, then a plant codon-optimized protein-encoding (e.g., Cas9-encoding) nucleotide sequence could be generated. As another non-limiting example, if the intended host cell were an insect cell, then an insect codon-optimized proteinencoding (e.g., Cas9-encoding) nucleotide sequence could be generated.

[0181] The present disclosure provides one or more recombinant expression vectors that include (in different recombinant expression vectors in some cases, and in the same recombinant expression vector in some cases): (i) a nucleotide sequence of a donor template nucleic acid (where the donor template comprises a nucleotide sequence having homology to a target sequence of a target nucleic acid (e.g., a target genome)); (ii) a nucleotide sequence that encodes a Cas9 guide RNA that hybridizes to a target sequence of the target locus of the targeted genome (e.g., a single or dual guide RNA) (e.g., operably linked to a promoter that is operable in a target cell such as a eukaryotic cell); and (iii) a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) (e.g., operably linked to a promoter that is operable in a target cell such as a eukaryotic cell). The present disclosure provides one or more recombinant expression vectors that include (in different recombinant expression vectors in some cases, and in the same recombinant expression vector in some cases): (i) a nucleotide sequence of a donor template nucleic acid (where the donor template comprises a nucleotide sequence havinghomology to a target sequence of a target nucleic acid (e g., a target genome)); and (ii) a nucleotide sequence that encodes a Cas9 guide RNA that hybridizes to a target sequence of the target locus of the targeted genome (e.g., a single or dual guide RNA) (e.g., operably linked to a promoter that is operable in a target cell such as a eukaryotic cell). The present disclosure provides one or more recombinant expression vectors that include (in different recombinant expression vectors in some cases, and in the same recombinant expression vector in some cases): (i) a nucleotide sequence that encodes a Cas9 guide RNA that hybridizes to a target sequence of the target locus of the targeted genome (e.g., a single or dual guide RNA) (e.g., operably linked to a promoter that is operable in a target cell such as a eukaryotic cell); and (ii) a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) (e.g., operably linked to a promoter that is operable in a target cell such as a eukaryotic cell).

[0182] Suitable expression vectors include viral expression vectors (e.g. viral vectors based on vaccinia virus; poliovirus; adenovirus (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:25432549, 1994; Borras et al., Gene Ther 6:515 524, 1999; Li and Davidson, PNAS 92:7700 7704, 1995; Sakamoto et al., H Gene Ther 5:1088 1097, 1999; WO 94 / 12649, WO 93 / 03769; WO 93 / 19191 ; WO 94 / 28938; WO 95 / 11984 and WO 95 / 00655); adeno-associated virus (AAV) (see, e.g., Ali et al., Hum Gene Ther 9:81 86, 1998, Flannery et al., PNAS 94:69166921 , 1997; Bennett et al., Invest Opthalmol Vis Sci 38:28572863, 1997; Jomary et al., Gene Ther 4:683 690, 1997, Rolling et al., Hum Gene Ther 10:641 648, 1999; Ali et al., Hum Mol Genet 5:591 594, 1996; Srivastava in WO 93 / 09239, Samulski et al., J. Vir. (1989) 63:3822-3828; Mendelson et al., Virol. (1988) 166:154-165; and Flotte et al., PNAS (1993) 90:10613-10617); SV40; herpes simplex virus; human immunodeficiency virus (see, e.g., Miyoshi et al., PNAS 94:10319 23, 1997; Takahashi et al., J Virol 73:7812 7816, 1999); a retroviral vector (e.g., Murine Leukemia Virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus); and the like. In some cases, a recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some cases, a recombinant expression vector of the present disclosure is a recombinant lentivirus vector. In some cases, a recombinant expression vector of the present disclosure is a recombinant retroviral vector.

[0183] Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. may be used in the expression vector.

[0184] In some embodiments, a nucleotide sequence encoding a Cas9 guide RNA is operably linked to a control element, e.g., a transcription control element, such as a promoter. In some embodiments, a nucleotide sequence encoding a Cas9 fusion polypeptide is operably linked to a control element, e.g., a transcription control element, such as a promoter.

[0185] The transcription control element can be a promoter. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell type-specific promoter. In some cases, the transcription control element (e.g., the promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcription control element can be functional in eukaryotic cells, e.g., hematopoietic stem cells (e.g., mobilized peripheral blood (mPB) CD34(+) cell, bone marrow (BM) CD34(+) cell, etc.).

[0186] Non-limiting examples of eukaryotic promoters (promoters functional in a eukaryotic cell) include EF1a, those from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) from retrovirus, and mouse metallothionein-l. Selection of the appropriate vector and promoter is well within the level of ordinary skill in the art. The expression vector may also contain a ribosome binding site for translation initiation and a transcription terminator. The expression vector may also include appropriate sequences for amplifying expression. The expression vector may also include nucleotide sequences encoding protein tags (e.g., 6xHis tag, hemagglutinin tag, fluorescent protein, etc.) that can be fused to a subject protein, thus resulting in a chimeric polypeptide.

[0187] In some embodiments, a nucleotide sequence encoding a subject Cas9 fusion polypeptide is operably linked to an inducible promoter. In some embodiments, a nucleotide sequence encoding a subject Cas9 fusion polypeptide is operably linked to a constitutive promoter.

[0188] A promoter can be a constitutively active promoter (i.e. , a promoter that is constitutively in an active / ”ON” state), it may be an inducible promoter (i.e., a promoterwhose state, active / ”ON” or inactive / “OFF”, is controlled by an external stimulus, e.g., the presence of a particular temperature, compound, or protein.), it may be a spatially restricted promoter (i.e., transcription control element, enhancer, etc.)(e.g., tissue specific promoter, cell type specific promoter, etc.), and it may be a temporally restricted promoter (i.e., the promoter is in the “ON” state or “OFF” state during specific stages of embryonic development or during specific stages of a biological process, e.g., hair follicle cycle in mice).

[0189] Promoters can be derived from viruses and can therefore be referred to as viral promoters, or they can be derived from any organism, including prokaryotic or eukaryotic organisms. Promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Exemplary promoters include, but are not limited to the SV40 early promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); a herpes simplex virus (HSV) promoter, a cytomegalovirus (CMV) promoter such as the CMV immediate early promoter region (CMVIE), a rous sarcoma virus (RSV) promoter, and the like. Example Pol III promoters (e.g., for expressing a guide RNA) include, but are not necessarily limited to: a human U6 small nuclear promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497 - 500 (2002)), an enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1 ;31 (17)), a human H1 promoter (H1), and the like.

[0190] In some cases, a nucleotide sequence encoding a Cas9 guide RNA is operably linked to (under the control of) a promoter operable in a eukaryotic cell (e.g., a Pol III promoter such as a U6 promoter, an enhanced U6 promoter, an H1 promoter, and the like). In some cases, a nucleotide sequence encoding a subject Cas9 fusion polypeptide is operably linked to a promoter operable in a eukaryotic cell (e.g., a CMV promoter, a beta-actin promoter, an EF1a promoter, an estrogen receptor-regulated promoter, and the like).

[0191] Examples of inducible promoters include, but are not limited toT7 RNA polymerase promoter, T3 RNA polymerase promoter, Isopropyl-beta-D-thiogalactopyranoside (IPTG)-regulated promoter, lactose induced promoter, heat shock promoter, Tetracycline-regulated promoter, Steroid-regulated promoter, Metal-regulated promoter, estrogen receptor-regulated promoter, etc. Inducible promoters can therefore be regulated by molecules including, but not limited to, doxycycline; estrogen and / or an estrogen analog; IPTG; etc.

[0192] Inducible promoters suitable for use include any inducible promoter described herein or known to one of ordinary skill in the art. Examples of inducible promoters include,without limitation, chemically / biochemically-regulated and physically-regulated promoters such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)-responsive promoters and other tetracyclineresponsive promoter systems, which include a tetracycline repressor protein (tetR), a tetracycline operator sequence (tetO) and a tetracycline transactivator fusion protein (tTA)), steroid-regulated promoters (e.g., promoters based on the rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptors, and promoters from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters derived from metallothionein (proteins that bind and sequester metal ions) genes from yeast, mouse and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene or benzothiadiazole (BTH)), temperature / heat- inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light responsive promoters from plant cells).

[0193] In some cases, the promoter is a spatially restricted promoter (i.e., cell type specific promoter, tissue specific promoter, etc.) such that in a multi-cellular organism, the promoter is active (i.e., “ON”) in a subset of specific cells. Spatially restricted promoters may also be referred to as enhancers, transcription control elements, control sequences, etc. Any convenient spatially restricted promoter may be used as long as the promoter is functional in the targeted host cell (e.g., eukaryotic cell; prokaryotic cell).

[0194] In some cases, the promoter is a reversible promoter. Suitable reversible promoters, including reversible inducible promoters are known in the art. Such reversible promoters may be isolated and derived from many organisms, e.g., eukaryotes and prokaryotes. Modification of reversible promoters derived from a first organism for use in a second organism, e.g., a first prokaryote and a second a eukaryote, a first eukaryote and a second a prokaryote, etc., is well known in the art. Such reversible promoters, and systems based on such reversible promoters but also comprising additional control proteins, include, but are not limited to, alcohol regulated promoters (e.g., alcohol dehydrogenase I (alcA) gene promoter, promoters responsive to alcohol transactivator proteins (AlcR), etc.), tetracycline regulated promoters, (e.g., promoter systems including TetActivators, TetON, TetOFF, etc.), steroid regulated promoters (e.g., rat glucocorticoid receptor promoter systems, human estrogen receptor promoter systems, retinoid promoter systems, thyroid promoter systems, ecdysone promoter systems, mifepristone promoter systems, etc.), metal regulated promoters (e.g., metallothionein promoter systems, etc.), pathogenesis-related regulatedpromoters (e.g., salicylic acid regulated promoters, ethylene regulated promoters, benzothiadiazole regulated promoters, etc.), temperature regulated promoters (e.g., heat shock inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter, etc.), light regulated promoters, synthetic inducible promoters, and the like.

[0195] Methods of introducing a nucleic acid (e.g., a nucleic acid comprising a donor polynucleotide sequence, a subject Cas9 fusion polypeptide, and / or a Cas9 guide RNA) into a host cell are known in the art, and any convenient method can be used to introduce a nucleic acid (e.g., an expression construct) into a cell. Suitable methods include e.g., viral infection, transfection, lipofection, nucleofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle- mediated nucleic acid delivery, lipid nanoparticle mediated delivery, and the like.

[0196] Introducing the recombinant expression vector into cells can occur in vivo or can occur in any culture media and under any culture conditions that promote the survival of the cells. Introducing the recombinant expression vector into a target cell can be carried out in vivo or ex vivo or in vitro. Introducing the recombinant expression vector into a target cell can be carried out in vitro.

[0197] In some embodiments, a subject Cas9 fusion polypeptide is provided to a cell as RNA (e.g., as mRNA that is translated into protein within a cell). In some embodiments, a Cas9 guide RNA is provided as RNA. The RNA can be provided by direct chemical synthesis or may be transcribed in vitro from a DNA (e.g., encoding the protein). Once synthesized, the RNA may be introduced into a cell by any of the well-known techniques for introducing nucleic acids into cells (e.g., microinjection, electroporation, transfection, lipid nanoparticle delivery, etc.).

[0198] Nucleic acids may be provided to the cells using well-developed transfection techniques; see, e.g. Angel and Yanik (2010) PLoS ONE 5(7): e11756, and the commercially available TransMessenger® reagents from Qiagen, Stemfect™ RNA Transfection Kit from Stemgent, TranslT®-mRNA Transfection Kit from Mirus Bio LLC, nucleofection, and the like. See also Beumer et al. (2008) PNAS 105(50): 19821- 19826.

[0199] Vectors may be provided directly to a target host cell. In other words, the cells can be contacted with vectors comprising the subject nucleic acids such that the vectors are taken up by the cells. Methods for contacting cells with nucleic acid vectors that are plasmids include electroporation, calcium chloride transfection, microinjection, andlipofection are well known in the art. For viral vector delivery, cells can be contacted with viral particles comprising the subject viral expression vectors.

[0200] Retroviruses, for example, lentiviruses, are suitable for use in methods of the present disclosure. Commonly used retroviral vectors are “defective”, i.e. unable to produce viral proteins required for productive infection. Rather, replication of the vector requires growth in a packaging cell line. To generate viral particles comprising nucleic acids of interest, the retroviral nucleic acids comprising the nucleic acid are packaged into viral capsids by a packaging cell line. Different packaging cell lines provide a different envelope protein (ecotropic, amphotropic or xenotropic) to be incorporated into the capsid, this envelope protein determining the specificity of the viral particle for the cells (ecotropic for murine and rat; amphotropic for most mammalian cell types including human, dog and mouse; and xenotropic for most mammalian cell types except murine cells). The appropriate packaging cell line may be used to ensure that the cells are targeted by the packaged viral particles. Methods of introducing subject vector expression vectors into packaging cell lines and of collecting the viral particles that are generated by the packaging lines are well known in the art. Nucleic acids can also introduced by direct micro-injection (e.g., injection of RNA).

[0201] Vectors used for providing the nucleic acids encoding Cas9 guide RNA and / or a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) to a target host cell can include suitable promoters for driving the expression, that is, transcription activation, of the nucleic acid of interest. In other words, in some cases, the nucleic acid of interest will be operably linked to a promoter. This may include ubiquitously acting promoters, for example, the CMV-p-actin promoter, or inducible promoters, such as promoters that are active in particular cell populations or that respond to the presence of drugs such as tetracycline. By transcription activation, it is intended that transcription will be increased above basal levels in the target cell by 10 fold, by 100 fold, more usually by 1000 fold. In addition, vectors used for providing a nucleic acid encoding a Cas9 guide RNA and / or a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) to a cell may include nucleic acid sequences that encode for selectable markers in the target cells, so as to identify cells that have taken up the Cas9 guide RNA and / or subject protein.

[0202] A nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide), is in some cases an RNA. Thus, a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be introduced into cells as RNA. Methods of introducing RNA into cells are known in theart and may include, for example, direct injection, transfection, or any other method used for the introduction of DNA. A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) may instead be provided to cells as a polypeptide. Such a polypeptide may optionally be fused to a polypeptide domain that increases solubility of the product. The domain may be linked to the polypeptide through a defined protease cleavage site, e.g. a TEV sequence, which is cleaved by TEV protease. The linker may also include one or more flexible sequences, e.g. from 1 to 10 glycine residues. In some embodiments, the cleavage of the fusion protein is performed in a buffer that maintains solubility of the product, e.g. in the presence of from 0.5 to 2 M urea, in the presence of polypeptides and / or polynucleotides that increase solubility, and the like. Domains of interest include endosomolytic domains, e.g. influenza HA domain; and other polypeptides that aid in production, e.g. IF2 domain, GST domain, GRPE domain, and the like. The polypeptide may be formulated for improved stability. For example, the peptides may be PEGylated, where the polyethyleneoxy group provides for enhanced lifetime in the blood stream.

[0203] Additionally or alternatively, a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be introduced as protein into a cell (e.g., in some cases Fused to a polypeptide permeant domain to promote uptake by the cell). In some cases, the Cas9 fusion polypeptide is introduced to a cell as an RNP (i.e., pre-complexed with a Cas9 guide RNA). In some cases, a RNP is delivered using peptide-enabled RNP delivery for CRISPR engineering (PERC) (see, e.g., Foss et al., Nature Biomedical Engineering, 2023: p. 1-14). In some cases, a RNP is delivered using a nanoparticle formulation. In some cases, a RNP is delivered using a lipid nanoparticle formulation.

[0204] A number of permeant domains are known in the art and may be used in the nonintegrating polypeptides of the present disclosure, including peptides, peptidomimetics, and non-peptide carriers. For example, a permeant peptide may be derived from the third alpha helix of Drosophila melanogaster transcription factor Antennapaedia, referred to as penetratin, which comprises the amino acid sequence RQIKIWFQNRRMKWKK (SEQ ID NO: 946). As another example, the permeant peptide comprises the HIV-1 tat basic region amino acid sequence, which may include, for example, amino acids 49-57 of naturally-occurring tat protein. Other permeant domains include poly-arginine motifs, for example, the region of amino acids 34-56 of HIV-1 rev protein, nona-arginine, octa-arginine, and the like. (See, for example, Futaki et al. (2003) Curr Protein Pept Sci. 2003 Apr; 4(2): 87-9 and 446; and Wender et al. (2000) Proc. Natl. Acad. Sci. U.S.A 2000 Nov. 21 ; 97(24): 13003-8;published U.S. Patent applications 20030220334; 20030083256; 20030032593; and 20030022831 , herein specifically incorporated by reference for the teachings of translocation peptides and peptoids). The nona-arginine (R9) sequence is one of the more efficient PTDs that have been characterized (Wender et al. 2000; Uemura et al. 2002). The site at which the fusion is made may be selected in order to optimize the biological activity, secretion or binding characteristics of the polypeptide. The optimal site will be determined by routine experimentation.

[0205] A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) may be produced in vitro or by eukaryotic cells or by prokaryotic cells, and it may be further processed by unfolding, e.g. heat denaturation, dithiothreitol reduction, etc. and may be further refolded, using methods known in the art.

[0206] Modifications of interest that do not alter primary sequence include chemical derivatization of polypeptides, e.g., acylation, acetylation, carboxylation, amidation, etc. Also included are modifications of glycosylation, e.g. those made by modifying the glycosylation patterns of a polypeptide during its synthesis and processing or in further processing steps; e.g. by exposing the polypeptide to enzymes which affect glycosylation, such as mammalian glycosylating or deglycosylating enzymes. Also embraced are sequences that have phosphorylated amino acid residues, e.g. phosphotyrosine, phosphoserine, or phosphothreonine.

[0207] Also suitable for inclusion in embodiments of the present disclosure are nucleic acids (e.g., encoding a Cas9 guide RNA, encoding a protein of the present disclosure, e.g., a subject Cas9 fusion polypeptide, and the like), and proteins of the present disclosure (e.g., a Cas9 fusion polypeptide) that have been modified using ordinary molecular biological techniques and synthetic chemistry so as to improve their resistance to proteolytic degradation, to change the target sequence specificity, to optimize solubility properties, to alter protein activity (e.g., transcription modulatory activity, enzymatic activity, etc.) or to render them more suitable. Analogs of such polypeptides include those containing residues other than naturally occurring L-amino acids, e.g. D-amino acids or non-naturally occurring synthetic amino acids. D-amino acids may be substituted for some or all of the amino acid residues.

[0208] A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) may be prepared by in vitro synthesis, using conventional methods as known in the art. Various commercial synthetic apparatuses are available, for example, automated synthesizers by Applied Biosystems, Inc., Beckman, etc. By using synthesizers, naturally occurring amino acids may be substituted with unnatural amino acids. Theparticular sequence and the manner of preparation will be determined by convenience, economics, purity required, and the like.

[0209] If desired, various groups may be introduced into the peptide during synthesis or during expression, which allow for linking to other molecules or to a surface. Thus cysteines can be used to make thioethers, histidines for linking to a metal ion complex, carboxyl groups for forming amides or esters, amino groups for forming amides, and the like.

[0210] A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) may also be isolated and purified in accordance with conventional methods of recombinant synthesis. A lysate may be prepared of the expression host and the lysate purified using high performance liquid chromatography (HPLC), exclusion chromatography, gel electrophoresis, affinity chromatography, or other purification technique. For the most part, the compositions which are used will comprise 20% or more by weight of the desired product, more usually 75% or more by weight, preferably 95% or more by weight, and for therapeutic purposes, usually 99.5% or more by weight, in relation to contaminants related to the method of preparation of the product and its purification. Usually, the percentages will be based upon total protein. Thus, in some cases, a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) is at least 80% pure, at least 85% pure, at least 90% pure, at least 95% pure, at least 98% pure, or at least 99% pure (e.g., free of contaminants, non-desired proteins or other macromolecules, etc.).

[0211] In cases in which two or more different targeting complexes are provided to the cell (e.g., two different Cas9 guide RNAs that are complementary to different sequences within the same or different target nucleic acid), the complexes may be provided simultaneously (e.g. as two polypeptides and / or nucleic acids), or delivered simultaneously. Alternatively, they may be provided consecutively, e.g. the targeting complex being provided first, followed by the second targeting complex, etc. or vice versa.

[0212] To improve the delivery of a DNA vector into a target cell, the DNA can be protected from damage and its entry into the cell facilitated, for example, by using lipoplexes and polyplexes. Thus, in some cases, a nucleic acid of the present disclosure (e.g., a recombinant expression vector of the present disclosure) can be covered with lipids in an organized structure like a micelle or a liposome. When the organized structure is complexed with DNA it is called a lipoplex. There are three types of lipids, anionic (negatively-charged), neutral, or cationic (positively-charged). Lipoplexes that utilizecationic lipids have proven utility for gene transfer. Cationic lipids, due to their positive charge, naturally complex with the negatively charged DNA. Also as a result of their charge, they interact with the cell membrane. Endocytosis of the lipoplex then occurs, and the DNA is released into the cytoplasm. The cationic lipids also protect against degradation of the DNA by the cell.

[0213] Complexes of polymers with DNA are called polyplexes. Most polyplexes consist of cationic polymers and their production is regulated by ionic interactions. One large difference between the methods of action of polyplexes and lipoplexes is that polyplexes cannot release their DNA load into the cytoplasm, so to this end, cotransfection with endosome-lytic agents (to lyse the endosome that is made during endocytosis) such as inactivated adenovirus must occur. However, this is not always the case; polymers such as polyethylenimine have their own method of endosome disruption as does chitosan and trimethylchitosan.

[0214] Dendrimers, a highly branched macromolecule with a spherical shape, may be also be used to genetically modify stem cells. The surface of the dendrimer particle may be functionalized to alter its properties. In particular, it is possible to construct a cationic dendrimer (i.e. , one with a positive surface charge). When in the presence of genetic material such as a DNA plasmid, charge complementarity leads to a temporary association of the nucleic acid with the cationic dendrimer. On reaching its destination, the dendrimer-nucleic acid complex can be taken up into a cell by endocytosis.

[0215] In some cases, a nucleic acid of the disclosure (e.g., an expression vector) includes an insertion site for a guide sequence of interest. For example, a nucleic acid can include an insertion site for a guide sequence of interest, where the insertion site is immediately adjacent to a nucleotide sequence encoding the portion of a Cas9 guide RNA that does not change when the guide sequence is changed to hybridized to a desired target sequence (e.g., sequences that contribute to the Cas9 binding aspect of the guide RNA, e.g, the sequences that contribute to the dsRNA duplex(es) of the Cas9 guide RNA - this portion of the guide RNA can also be referred to as the ‘scaffold’ or 'constant region’ of the guide RNA). Thus, in some cases, a subject nucleic acid (e.g., an expression vector) includes a nucleotide sequence encoding a Cas9 guide RNA, except that the portion encoding the guide sequence portion of the guide RNA is an insertion sequence (an insertion site). An insertion site is any nucleotide sequence used for the insertion of the desired sequence. “Insertion sites” for use with various technologies are known to those of ordinary skill in the art and any convenient insertion site can be used. An insertion site can be for any method formanipulating nucleic acid sequences. For example, In some cases, the insertion site is a multiple cloning site (MCS) (e.g., a site including one or more restriction enzyme recognition sequences), a site for ligation independent cloning, a site for recombination based cloning (e.g., recombination based on att sites), a nucleotide sequence recognized by a CRISPR / Cas (e.g. Cas9) based technology, and the like.

[0216] An insertion site can be any desirable length, and can depend on the type of insertion site (e.g., can depend on whether (and how many) the site includes one or more restriction enzyme recognition sequences, whether the site includes a target site for a CRISPR / Cas protein, etc.). In some cases, an insertion site of a subject nucleic acid is 3 or more nucleotides (nt) in length (e.g., 5 or more, 8 or more, 10 or more, 15 or more, 17 or more, 18 or more, 19 or more, 20 or more or 25 or more, or 30 or more nt in length). In some cases, the length of an insertion site of a subject nucleic acid has a length in a range of from 2 to 50 nucleotides (nt) (e.g., from 2 to 40 nt, from 2 to 30 nt, from 2 to 25 nt, from 2 to 20 nt, from 5 to 50 nt, from 5 to 40 nt, from 5 to 30 nt, from 5 to 25 nt, from 5 to 20 nt, from 10 to 50 nt, from 10 to 40 nt, from 10 to 30 nt, from 10 to 25 nt, from 10 to 20 nt, from 17 to 50 nt, from 17 to 40 nt, from 17 to 30 nt, from 17 to 25 nt). In some cases, the length of an insertion site of a subject nucleic acid has a length in a range of from 5 to 40 nt.Introducing components into a target cell

[0217] A Cas9 guide RNA (or a nucleic acid comprising a nucleotide sequence encoding same), and / or a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) (or a nucleic acid comprising a nucleotide sequence encoding same) and / or a donor template polynucleotide can be introduced into a host cell by any of a variety of well-known methods. A Cas9 guide RNA can be provided directly (e.g., as one or more RNA molecules), or can be provided as a DNA encoding the guide RNA.

[0218] Methods of introducing a nucleic acid (mRNA and / or DNA) into a cell (e.g., prokaryotic cell, eukaryotic cell, plant cell, vertebrate cell, mammalian cell, primate cell, nonhuman primate cell, human cell, and the like) are known in the art, and any convenient method can be used to introduce a nucleic acid (e.g., an expression construct) into a target cell. Suitable methods include, e.g., viral infection, transfection, conjugation, protoplast fusion, lipofection, nucleofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE- dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, agrobacterium-mediated transformation,direct micro injection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et., al Adv Drug Deliv Rev. 2012 Sep 13. pii: S0169-409X(12)00283-9. doi: 10.1016 / j.addr.2012.09.023 ), and the like. Any or all of the components can be introduced into a cell as a composition [e.g., including any convenient combination of: a subject Cas9 fusion protein (or a nucleic acid comprising a nucleotide sequence encoding same), a Cas9 guide RNA (or a DNA comprising a nucleotide sequence encoding same), a donor polynucleotide, etc ] using known methods, e.g., such as nucleofection.

[0219] In some cases, a protein of the present disclosure (e.g., a subject Cas9 fusion polypeptide) is provided as a nucleic acid (e.g., an mRNA, a DNA, a plasmid, an expression vector, a viral vector, etc.) that encodes protein. In some cases, a protein of the present disclosure (e.g., a subject Cas9 fusion polypeptide) is provided as a protein.

[0220] A protein of the present disclosure (e.g., a subject Cas9 fusion polypeptide) can be introduced into a cell (provided to the cell) by any convenient method; such methods are known to those of ordinary skill in the art. As an illustrative example, a protein of the present disclosure (e.g., a subject Cas9 fusion polypeptide) can be injected directly into a cell (e.g., with or without a Cas9 guide RNA, with or without a nucleic acid encoding a Cas9 guide RNA, and with or without a donor polynucleotide). As another example, a preformed complex of a protein of the present disclosure (e.g., a subject Cas9 fusion polypeptide) plus a Cas9 guide RNA is referred to as a ribonucleoprotein (RNP) complex, can be introduced into a cell (e.g., via nucleofection; via injection; via a protein transduction domain (PTD) conjugated to one or more components, e.g., conjugated to the protein, conjugated to a guide RNA, and the like). As noted above, a Cas9 guide RNA can be introduced into a cell as RNA or as DNA encoding the RNA (e.g., an expression vector encoding the Cas9 guide RNA).Host Cells (modified cells, e.g., genetically modified cells)

[0221] The present disclosure provides host cells comprising (e.g., genetically modified to comprise) a protein and / or nucleic acid of the present disclosure. The present disclosure provides host cells comprising (e.g., modified to comprise, e.g., genetically modified to comprise) a recombinant vector of the present disclosure. The present disclosure provides host cells comprising a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). The present disclosure provides host cellscomprising a nucleic acid molecule that includes a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). In some cases, the nucleotide sequence encoding a protein of the present disclosure is integrated into the genome of the host cell. In some cases, the nucleotide sequence encoding a protein of the present disclosure is not integrated into the genome of the host cell (e.g., in some cases it is maintained episomally). The modified host cells are in some cases in vitro host cells. In some cases, the modified host cells are in vivo host cells.

[0222] Suitable host cells include, e.g. a bacterial cell; an archaeal cell; a cell of a single-cell eukaryotic organism; a plant cell; an algal cell, e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like; a fungal cell (e.g., a yeast cell); an animal cell; an invertebrate animal cell (e.g. fruit fly, cnidarian, echinoderm, nematode, etc.); a vertebrate animal cell (e.g., fish, amphibian, reptile, bird, mammal); a mammalian cell (e.g., rodent, mouse, rat, primate, non-human primate, human, etc.); and the like.

[0223] A suitable host cell can be a stem cell (e.g. an embryonic stem (ES) cell, an induced pluripotent stem (iPS) cell); a germ cell; a somatic cell, e.g. a fibroblast, a hematopoietic cell, a neuron, a muscle cell, a bone cell, a hepatocyte, a pancreatic cell; an in vitro or in vivo embryonic cell of an embryo at any stage, e.g., a 1-cell, 2- cell, 4-cell, 8-cell, etc. stage zebrafish embryo; etc.). Cells may be from established cell lines or they may be primary cells, where “primary cells”, “primary cell lines”, and “primary cultures” are used interchangeably herein to refer to cells and cells cultures that have been derived from a subject and allowed to grow in vitro for a limited number of passages, i.e. splittings, of the culture. For example, primary cultures include cultures that may have been passaged 0 times, 1 time, 2 times, 4 times, 5 times, 10 times, or 15 times, but not enough times go through the crisis stage. Primary cell lines can be maintained for fewer than 10 passages in vitro. Host cells are in many embodiments unicellular organisms, or are grown in culture.

[0224] If the cells are primary cells, they may be harvest from an organism (e.g., an individual) by any convenient method. For example, leukocytes may be conveniently harvested by apheresis, leukocytapheresis, density gradient separation, etc., while cells from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc. are most conveniently harvested by biopsy. An appropriate solution may be used for dispersion or suspension of the harvested cells. Such solution will generally be a balanced salt solution, e.g. normal saline, phosphate-buffered saline (PBS), Hank’s balanced salt solution, etc., conveniently supplemented with fetal calf serum or other naturally occurring factors, in conjunction with an acceptable buffer at low concentration, e.g., from 5-25 mM. Convenient buffers include HEPES, phosphate buffers, lactate buffers, etc. The cells may be used immediately, or they may be stored, frozen, for long periods of time, being thawed and capable of being reused. In such cases, the cells can be frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered medium, or some other such solution as is commonly used in the art to preserve cells at such freezing temperatures, and thawed in a manner as commonly known in the art for thawing frozen cultured cells.

[0225] In some embodiments, a subject genetically modified host cell is in vitro. In some embodiments, a subject genetically modified host cell is ex vivo. In some embodiments, a subject genetically modified host cell is in vivo. In some embodiments, a subject genetically modified host cell is a prokaryotic cell or is derived from a prokaryotic cell. In some embodiments, a subject genetically modified host cell is a bacterial cell or is derived from a bacterial cell. In some embodiments, a subject genetically modified host cell is an archaeal cell or is derived from an archaeal cell. In some embodiments, a subject genetically modified host cell is a eukaryotic cell or is derived from a eukaryotic cell. In some embodiments, a subject genetically modified host cell is a plant cell or is derived from a plant cell. In some embodiments, a subject genetically modified host cell is an animal cell or is derived from an animal cell. In some embodiments, a subject genetically modified host cell is an invertebrate cell or is derived from an invertebrate cell. In some embodiments, a subject genetically modified host cell is a vertebrate cell or is derived from a vertebrate cell. In some embodiments, a subject genetically modified host cell is a mammalian cell or is derived from a mammalian cell. In some embodiments, a subject genetically modified host cell is a rodent cell or is derived from a rodent cell. In some embodiments, a subject genetically modified host cell is a human cell or is derived from a human cell. In some embodiments, a subject genetically modified host cell is a non-human mammalian cell or is derived from a non-human mammalian cell. In some embodiments, a subject genetically modified host cell is a non-human primate cell or is derived from a non-human primate cell. In some embodiments, a subject genetically modified host cell is an insect cell or is derived from an insect cell. In some embodiments, a subject genetically modified host cell is an arachnid cell or is derived from an arachnid cell.

[0226] The present disclosure further provides progeny of a subject genetically modified cell, where the progeny can comprise the same exogenous nucleic acid or polypeptide as the subject genetically modified cell from which it was derived. The present disclosure further provides a composition comprising a subject genetically modified host cell. Donor Polynucleotide (donor template)

[0227] In some cases, the contacting occurs under conditions that are permissive for nonhomologous end joining or homology-directed repair. In some cases, the method further comprises contacting the target DNA with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide integrates into the target DNA (e.g., a sequence of the donor polynucleotide, e.g., DNA, is incorporated into the target DNA, e.g., genomic DNA). In some cases, the method does not comprise contacting a cell with a donor polynucleotide, and the target DNA is modified such that nucleotides within the target DNA are deleted.

[0228] In some cases, Cas9 guide RNA and a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) are coadministered (e.g., contacted with a target nucleic acid, administered to cells, etc.) with a donor polynucleotide sequence that includes at least a segment with homology to the target DNA sequence, the subject methods may be used to add, i.e. insert or replace, nucleic acid material to a target DNA sequence (e.g. to “knock in” a nucleic acid that encodes for a protein, an siRNA, an miRNA, etc.), to add a tag (e.g., 6xHis, a fluorescent protein (e.g., a green fluorescent protein; a yellow fluorescent protein, etc.), hemagglutinin (HA), FLAG, etc.), to add a regulatory sequence to a gene (e.g. promoter, polyadenylation signal, internal ribosome entry sequence (IRES), 2A peptide, start codon, stop codon, splice signal, localization signal, etc.), to modify a nucleic acid sequence (e.g., introduce a mutation or to correct a mutation to wild type), and the like. As such, a complex comprising a Cas9 guide RNA and a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) is useful in any in vitro or in vivo application in which it is desirable to modify DNA in a site-specific, i.e. “targeted”, way, for example gene knock-out, gene knock-in, gene editing, gene tagging, etc., as used in, for example, gene therapy, e.g. to treat a disease or as an antiviral, antipathogenic, or anticancer therapeutic, the production of genetically modified organisms in agriculture, the large scale production of proteins by cells for therapeutic, diagnostic, or research purposes, the induction ofiPS cells, biological research, the targeting of genes of pathogens for deletion or replacement, etc.

[0229] In applications in which it is desirable to insert a polynucleotide sequence into a target DNA sequence, a polynucleotide comprising a donor sequence to be inserted can also provided to the cell. A donor polynucleotide (also referred to as a donor template) will include a “donor sequence,” which is a nucleic acid sequence to be inserted at the cleavage site induced by a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). The donor polynucleotide will contain sufficient homology to a genomic sequence at the cleavage site, e.g. 70%, 80%, 85%, 90%, 95%, or 100% homology with the nucleotide sequences flanking the cleavage site, e.g. within about 50 bases or less of the cleavage site, e.g. within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or immediately flanking the cleavage site, to support homology-directed repair between it and the genomic sequence to which it bears homology. Approximately 25, 50, 100, or 200 nucleotides, or more than 200 nucleotides, of sequence homology between a donor and a genomic sequence (or any integral value between 10 and 200 nucleotides, or more) will support homology- directed repair. Donor sequences can be of any length, e.g. 10 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more, etc.

[0230] The donor sequence is typically not identical to the genomic sequence that it replaces. Rather, the donor sequence may contain at least one or more single base changes, insertions, deletions, inversions or rearrangements with respect to the genomic sequence, so long as sufficient homology is present to support homology- directed repair. In some embodiments, the donor sequence comprises a non- homologous sequence flanked by two regions of homology, such that homology- directed repair between the target DNA region and the two flanking sequences results in insertion of the non-homologous sequence at the target region. Donor sequences may also comprise a vector backbone containing sequences that are not homologous to the DNA region of interest and that are not intended for insertion into the DNA region of interest. Generally, the homologous region(s) of a donor sequence will have at least 50% sequence identity to a genomic sequence with which recombination is desired. In certain embodiments, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.9% sequence identity is present. Any value between 1% and 100% sequence identity can be present, depending upon the length of the donor polynucleotide.

[0231] The donor sequence may comprise certain sequence differences as compared to the genomic sequence, e.g. restriction sites, nucleotide polymorphisms, selectable markers (e.g., drug resistance genes, fluorescent proteins, enzymes etc.), etc., which may be used to assess for successful insertion of the donor sequence at the cleavage site or in some cases may be used for other purposes (e.g., to signify expression at the targeted genomic locus). In some cases, if located in a coding region, such nucleotide sequence differences will not change the amino acid sequence, or will make silent amino acid changes (i.e. , changes which do not affect the structure or function of the protein). Alternatively, these sequences differences may include flanking recombination sequences such as FLPs, loxP sequences, or the like, that can be activated at a later time for removal of the marker sequence. In some cases, the nucleotide difference are intended to change the sequence of the target nucleic acid (e.g., DNA).

[0232] The donor sequence may be provided to the cell as single-stranded DNA, singlestranded RNA (e.g., as part of the guide RNA as used in ‘prime editing’, doublestranded DNA, or double-stranded RNA. It may be introduced into a cell in linear or circular form. If introduced in linear form, the ends of the donor sequence may be protected (e.g., from exonucleolytic degradation) by methods known to those of skill in the art. For example, one or more dideoxynucleotide residues are added to the 3' terminus of a linear molecule and / or self-complementary oligonucleotides are ligated to one or both ends. See, for example, Chang et al. (1987) Proc. Natl. Acad Sci USA 84:4959-4963; Nehls et al. (1996) Science 272:886-889. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, addition of terminal amino group(s) and the use of modified internucleotide linkages such as, for example, phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues. As an alternative to protecting the termini of a linear donor sequence, additional lengths of sequence may be included outside of the regions of homology that can be degraded without impacting recombination. A donor sequence can be introduced into a cell as part of a vector molecule having additional sequences such as, for example, replication origins, promoters and genes encoding antibiotic resistance. Moreover, donor sequences can be introduced as naked nucleic acid, as nucleic acid complexed with an agent such as a liposome or poloxamer, or can be delivered by viruses (e.g., adenovirus, AAV), as described above for nucleic acids encoding a Cas9 guide RNA and / or a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) and / or donor polynucleotide.

[0233] Because of the presence of a PAM sequence, it is contemplated herein that Cas9 could potentially cleave inserted sequence either prior to or following homology directed repair (e.g., homologous recombination), resulting in a possible non- homologous-end-joining event and further DNA sequence mutation at a target sequence (e.g., chromosomal locus) of interest. Therefore, to avoid cleavage of the donor sequence before and / or after Cas9-mediated homology directed repair, in some embodiments, alternate versions of the donor sequence may be used where mutations (e.g., silent mutations) are introduced. These mutations (e.g., silent mutations) may disrupt Cas9 binding and cleavage, but not disrupt the amino acid sequence of the repaired gene. For example, a donor sequence can in some cases include a mutated PAM (e.g., insertion of the donor sequence can result in mutation of the PAM sequence to reduce / eliminate additional cleavage events).Non-Human Modified Organisms

[0234] If a modified cell (e.g., genetically modified cell, as described above) is a eukaryotic single-cell organism, then the modified cell can be considered a genetically modified organism. In some embodiments, a subject non-human modified organism (e.g., nonhuman genetically modified organism) is a transgenic multicellular organism, in which a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) is integrated into the genome of the organism.

[0235] In some embodiments, a subject genetically modified non-human host cell (e.g., a cell that has been genetically modified with an exogenous nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can generate a subject genetically modified non-human organism (e.g., a rodent, a mouse, a fish, an amphibian, a frog, a reptile, an ungulate, a fly, a worm, an insect, an arachnid, an annelid, a plant, etc.). For example, if the genetically modified host cell is a pluripotent stem cell (i.e., PSC) or a germ cell (e.g., sperm, oocyte, etc.), an entire genetically modified organism can be derived from the genetically modified host cell. In some embodiments, the genetically modified host cell is a pluripotent stem cell (e.g., ESC, iPSC, pluripotent plant stem cell, etc.) or a germ cell (e.g., sperm cell, oocyte, etc.), either in vivo or in vitro, that can give rise to a genetically modified organism. In some embodiments the genetically modified host cell is a vertebrate PSC (e.g., ESC, iPSC, etc.) and is used to generate a genetically modified organism (e.g. by injecting a PSC into a blastocyst to produce a chimeric / mosaic animal, which could then be mated to generate non-chimeric / non-mosaic genetically modified organisms; grafting in the case of plants; etc.). Any convenient method / protocol for producing a genetically modified organism is suitable for producing a genetically modified host cell comprising an exogenous nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). Methods of producing genetically modified organisms are known in the art. For example, see Cho et al., Curr Protoc Cell Biol. 2009 Mar;Chapter 19: Unit 19.11: Generation of transgenic mice; Gama et al., Brain Struct Funct. 2010 Mar;214(2-3):91-109. Epub 2009 Nov 25: Animal transgenesis: an overview; Husaini et al., GM Crops. 2011 Jun-Dec;2(3): 150-62. Epub 2011 Jun 1: Approaches for gene targeting and targeted gene expression in plants.

[0236] In some embodiments, a genetically modified organism comprises a target cell for methods of the disclosure, and thus can be considered a source for target cells. For example, if a genetically modified cell comprising one or more exogenous nucleic acids comprising nucleotide sequences encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) is used to generate a genetically modified organism, then the cells of the genetically modified organism comprise the one or more exogenous nucleic acids comprising nucleotide sequences encoding the protein. In some such embodiments, the DNA of a cell or cells of the genetically modified organism can be targeted for modification by introducing into the cell or cells a Cas9 guide RNA (e.g., a truncated Cas9 guide RNA) (or a nucleic acid encoding the Cas9 guide RNA). For example, the introduction of a Cas9 guide RNA (or a DNA encoding the same) into a subset of cells (e.g., brain cells, intestinal cells, kidney cells, lung cells, blood cells, etc.) of the genetically modified organism can target the DNA of such cells for modification, the genomic location of which will depend on the targeting sequence of the introduced Cas9 guide RNA.

[0237] In some cases, a genetically modified organism is a source of target cells for methods of the disclosure. For example, a genetically modified organism comprising cells that are genetically modified with an exogenous nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can provide a source of genetically modified cells, for example PSCs (e.g., ESCs, iPSCs, sperm, oocytes, etc.), neurons, progenitor cells, cardiomyocytes, etc.

[0238] In some cases, a genetically modified cell is a PSC comprising an exogenous nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). As such, the PSC can be a target cell suchthat the DNA of the PSC can be targeted for modification by introducing into the PSC a Cas9 guide RNA (or a nucleic acid encoding the Cas9 guide RNA), and optionally a donor nucleic acid (donor polynucleotide), and the genomic location of the modification will depend on the targeting sequence of the introduced Cas9 guide RNA. Thus, in some embodiments, the methods described herein can be used to modify the DNA (e.g., delete and / or replace any desired genomic location) of PSCs derived from a subject genetically modified organism. Such modified PSCs can then be used to generate organisms having both (i) an exogenous nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) and (ii) a DNA modification that was introduced into the PSC.

[0239] An exogenous nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be under the control of (i.e. , operably linked to) an unknown promoter (e.g., when the nucleic acid randomly integrates into a host cell genome) or can be under the control of (i.e., operably linked to) a known promoter. Suitable known promoters can be any known promoter and include constitutively active promoters (e.g., CMV promoter, EF1a promoter), inducible promoters (e.g., heat shock promoter, Tetracycline-regulated promoter, Steroid-regulated promoter, Metal-regulated promoter, estrogen receptor- regulated promoter, etc.), spatially restricted and / or temporally restricted promoters (e.g., a tissue specific promoter, a cell type specific promoter, etc.), etc.

[0240] A subject genetically modified non-human organism can be any organism other than a human, including for example, a plant (e.g., a tobacco plant, a fruiting tree, a dicot, a monocot, a soy plant, a legume, a wheat plant, a barley plant, a rice plant, a tomato plant, a corn plant, a crop plant, and the like); algae; an invertebrate (e.g., a cnidarian, an echinoderm, a worm, an insect, an arachnid, an annelid, a fly, etc.); a vertebrate (e.g., a fish (e.g., zebrafish, puffer fish, gold fish, etc.), an amphibian (e.g., salamander, frog, etc.), a reptile, a bird, a mammal, etc.); an ungulate (e.g., a goat, a pig, a sheep, a cow, etc.); a rodent (e.g., a mouse, a rat, a hamster, a guinea pig); a lagomorpha (e.g., a rabbit); etc.Transgenic non-human animals

[0241] As described above, in some embodiments, a subject nucleic acid (e.g., one or more nucleic acids comprising nucleotide sequences encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide), e.g., a recombinant expressionvector, is used as a transgene to generate a transgenic animal that produces a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). Thus, the present disclosure further provides a transgenic non-human animal, which animal comprises a transgene comprising a subject nucleic acid comprising a nucleotide sequence encoding protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). In some embodiments, the genome of the transgenic non-human animal comprises a subject nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). In some embodiments, the transgenic non- human animal is homozygous for the genetic modification. In some embodiments, the transgenic non-human animal is heterozygous for the genetic modification. In some embodiments, the transgenic non-human animal is a vertebrate, for example, a fish (e.g., zebra fish, gold fish, puffer fish, cave fish, etc.), an amphibian (frog, salamander, etc.), a bird (e.g., chicken, turkey, etc.), a reptile (e.g., snake, lizard, etc.), a mammal (e.g., an ungulate, e.g., a pig, a cow, a goat, a sheep, etc.; a lagomorph (e.g., a rabbit); a rodent (e.g., a rat, a mouse); a non-human primate; etc.), etc. In some embodiments, the transgenic non-human animal is an invertebrate (e.g., an insect, an arachnid, etc.).

[0242] Nucleotide sequences encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be under the control of (i.e. , operably linked to) an unknown promoter (e.g., when the nucleic acid randomly integrates into a host cell genome) or can be under the control of (i.e., operably linked to) a known promoter. Suitable known promoters can be any known promoter and include constitutively active promoters (e.g., CMV promoter, EF1a, and the like), inducible promoters (e.g., heat shock promoter, Tetracycline-regulated promoter, Steroid-regulated promoter, Metal-regulated promoter, estrogen receptor-regulated promoter, etc.), spatially restricted and / or temporally restricted promoters (e.g., a tissue specific promoter, a cell type specific promoter, etc.), etc.Transgenic plants

[0243] As described above, in some embodiments, a subject nucleic acid (e.g., one or more nucleic acids comprising nucleotide sequences encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide), e.g., a recombinant expression vector, is used as a transgene to generate a transgenic plant that produces a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). Thus, the present disclosure further provides a transgenic plant, which plant comprises a transgenecomprising a subject nucleic acid comprising a nucleotide sequence encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). In some embodiments, the genome of the transgenic plant comprises a subject nucleic acid. In some embodiments, the transgenic plant is homozygous for the genetic modification. In some embodiments, the transgenic plant is heterozygous for the genetic modification.

[0244] Methods of introducing exogenous nucleic acids into plant cells are well known in the art. Such plant cells are considered “transformed,” as defined above. Suitable methods include viral infection (such as double stranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whiskers technology, Agrobacterium-mediated transformation and the like. The choice of method is generally dependent on the type of cell being transformed and the circumstances under which the transformation is taking place (i.e. in vitro, ex vivo, or in vivo).

[0245] Transformation methods based upon the soil bacterium Agrobacterium tumefaciens are particularly useful for introducing an exogenous nucleic acid molecule into a vascular plant. The wild type form of Agrobacterium contains a Ti (tumor-inducing) plasmid that directs production of tumorigenic crown gall growth on host plants. T ransfer of the tumor-inducing T-DNA region of the Ti plasmid to a plant genome requires the Ti plasmid-encoded virulence genes as well as T-DNA borders, which are a set of direct DNA repeats that delineate the region to be transferred. An Agrobacterium- based vector is a modified form of a Ti plasmid, in which the tumor inducing functions are replaced by the nucleic acid sequence of interest to be introduced into the plant host.

[0246] Agrobacterium-mediated transformation generally employs cointegrate vectors or binary vector systems, in which the components of the Ti plasmid are divided between a helper vector, which resides permanently in the Agrobacterium host and carries the virulence genes, and a shuttle vector, which contains the gene of interest bounded by T-DNA sequences. A variety of binary vectors is well known in the art and are commercially available, for example, from Clontech (Palo Alto, Calif.). Methods of coculturing Agrobacterium with cultured plant cells or wounded tissue such as leaf tissue, root explants, hypocotyledons, stem pieces or tubers, for example, also are well known in the art. See, e.g., Glick and Thompson, (eds.), Methods in Plant Molecular Biology and Biotechnology, Boca Raton, Fla.: CRC Press (1993).

[0247] Microprojectile-mediated transformation also can be used to produce a subject transgenic plant. This method, first described by Klein et al. (Nature 327:70-73 (1987)), relies on microprojectiles such as gold or tungsten that are coated with the desired nucleic acid molecule by precipitation with calcium chloride, spermidine or polyethylene glycol. The microprojectile particles are accelerated at high speed into an angiosperm tissue using a device such as the BIOLISTIC PD-1000 (Biorad; Hercules Calif.).

[0248] A subject nucleic acid may be introduced into a plant in a manner such that the nucleic acid is able to enter a plant cell(s), e.g., via an in vivo or ex vivo protocol. By "in vivo,” it is meant in the nucleic acid is administered to a living body of a plant e.g. infiltration. By “ex vivo” it is meant that cells or explants are modified outside of the plant, and then such cells or organs are regenerated to a plant. A number of vectors suitable for stable transformation of plant cells or for the establishment of transgenic plants have been described, including those described in Weissbach and Weissbach, (1989) Methods for Plant Molecular Biology Academic Press, and Gelvin et al., (1990) Plant Molecular Biology Manual, Kluwer Academic Publishers. Specific examples include those derived from a Ti plasmid of Agrobacterium tumefaciens, as well as those disclosed by Herrera- Estrella et al. (1983) Nature 303: 209, Bevan (1984) Nucl Acid Res. 12: 8711-8721, Klee (1985) Bio / Technolo 3: 637-642. Alternatively, non-Ti vectors can be used to transfer the DNA into plants and cells by using free DNA delivery techniques. By using these methods transgenic plants such as wheat, rice (Christou (1991) Bio / Technology 9:957-9 and 4462) and corn (Gordon-Kamm (1990) Plant Cell 2: 603-618) can be produced. An immature embryo can also be a good target tissue for monocots for direct DNA delivery techniques by using the particle gun (Weeks et al. (1993) Plant Physiol 102: 1077-1084; Vasil (1993) Bio / Technolo 10: 667-674; Wan and Lemeaux (1994) Plant Physiol 104: 37-48 and for Agrobacterium- mediated DNA transfer (Ishida et al. (1996) Nature Biotech 14: 745-750). Exemplary methods for introduction of DNA into chloroplasts are biolistic bombardment, polyethylene glycol transformation of protoplasts, and microinjection (Danieli et al Nat. Biotechnol 16:345-348, 1998; Staub et al Nat. Biotechnol 18: 333-338, 2000; O’Neill et al Plant J. 3:729-738, 1993; Knoblauch et al Nat. Biotechnol 17: 906-909; U.S. Pat. Nos. 5,451,513, 5,545,817, 5,545,818, and 5,576,198; in Inti. Application No. WO 95 / 16783; and in Boynton et al., Methods in Enzymology 217: 510-536 (1993), Svab et al., Proc. Natl. Acad. Sci. USA 90: 913-917 (1993), and McBride et al., Proc. Natl. Acad. Sci. USA 91 : 7301-7305 (1994)). Any vector suitable for the methods of biolisticbombardment, polyethylene glycol transformation of protoplasts and microinjection will be suitable as a targeting vector for chloroplast transformation. Any double stranded DNA vector may be used as a transformation vector, especially when the method of introduction does not utilize Agrobacterium.

[0249] Plants which can be genetically modified include grains, forage crops, fruits, vegetables, oil seed crops, palms, forestry, and vines. Specific examples of plants which can be modified follow: maize, banana, peanut, field peas, sunflower, tomato, canola, tobacco, wheat, barley, oats, potato, soybeans, cotton, carnations, sorghum, lupin and rice.

[0250] Also provided by the subject disclosure are transformed plant cells, tissues, plants and products that contain the transformed plant cells. A feature of the subject transformed cells, and tissues and products that include the same is the presence of a subject nucleic acid integrated into the genome, and production by plant cells of a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). Recombinant plant cells of the present invention are useful as populations of recombinant cells, or as a tissue, seed, whole plant, stem, fruit, leaf, root, flower, stem, tuber, grain, animal feed, a field of plants, and the like.

[0251] Nucleotide sequences encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be under the control of (i.e. , operably linked to) an unknown promoter (e.g., when the nucleic acid randomly integrates into a host cell genome) or can be under the control of (i.e., operably linked to) a known promoter. Suitable known promoters can be any known promoter and include constitutively active promoters, inducible promoters, spatially restricted and / or temporally restricted promoters, etc.Methods

[0252] A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) finds use in a variety of methods. A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be used in any method that a Cas9 protein can be used. For example, a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be used to (i) modify (e.g., cleave, e.g., nick; methylate; etc.) target nucleic acid (DNA or RNA; single stranded or double stranded); (ii) modulate transcription of a target nucleic acid; (iii) label a target nucleic acid; (iv) bind a target nucleic acid (e.g., for purposes of isolation, labeling, imaging, tracking, etc.); (v) modify a polypeptide(e.g., a histone) associated with a target nucleic acid; and the like. Because a method that uses a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) includes binding of the polypeptide to a particular region in a target nucleic acid (by virtue of being targeted there by an associated Cas9 guide RNA), the methods are generally referred to herein as methods of binding (e.g., a method of binding a target nucleic acid). However, it is to be understood that in some cases, while a method of binding may result in nothing more than binding of the target nucleic acid, in other cases, the method can have different final results (e.g., the method can result in modification of the target nucleic acid, e.g., cleavage / methylation / etc., modulation of transcription from the target nucleic acid, modulation of translation of the target nucleic acid, genome editing, modulation of a protein associated with the target nucleic acid, isolation of the target nucleic acid, etc.). For examples of suitable methods, see, for example, Jinek et al., Science. 2012 Aug 17;337(6096):816-21 ; Chylinski et al., RNA Biol. 2013 May;10(5):726-37; Ma et al., Biomed Res Int.2013;2013:270805; Hou et al., Proc Natl Acad Sci U S A. 2013 Sep24; 110(39): 15644-9; Jinek et al., Elife. 2013;2:e00471; Pattanayak et al., Nat Biotechnol. 2013 Sep;31(9):839-43; Qi et al, Cell. 2013 Feb 28; 152(5): 1173-83; Wang et al., Cell. 2013 May 9;153(4):910-8; Auer et al., Genome Res. 2013 Oct 31 ; Chen et al., Nucleic Acids Res. 2013 Nov 1 ;41(20):e19; Cheng et al., Cell Res. 2013 Oct;23(10): 1163-71; Cho et al., Genetics. 2013 Nov; 195(3): 1177-80; DiCarlo et al., Nucleic Acids Res. 2013 Apr;41(7):4336-43; Dickinson et al., Nat Methods. 2013 Cct;10(10):1028-34; Ebina et al., Sci Rep. 2013;3:2510; Fujii et al., Nucleic Acids Res. 2013 Nov 1 ;41 (20):e187; Hu et al., Cell Res. 2013 Nov;23(11):1322-5; Jiang et al., Nucleic Acids Res. 2013 Nov 1 ;41(20):e188; Larson et al., Nat Protoc. 2013 Nov;8(11):2180-96; Mali et al. Nat Methods. 2013 Oct;10(10):957-63; Nakayama et al., Genesis. 2013 Dec;51(12):835-43; Ran et al., Nat Protoc. 2013 Nov;8(11):2281- 308; Ran et al., Cell. 2013 Sep 12;154(6):1380-9; Upadhyay et al., G3 (Bethesda). 2013 Dec 9;3(12):2233-8; Walsh et al., Proc Natl Acad Sci U S A. 2013 Sep 24; 110(39): 15514-5; Xie et al., Mol Plant. 2013 Oct 9; Yang et al., Cell. 2013 Sep 12;154(6):1370-9; and U.S. patents and patent applications: 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871 ,445; 8,865,406; 8,795,965; 8,771 ,945; 8,697,359;20160304846, 20160215276, 20150166980, 20150071898, 20140068797; 20140170753; 20140179006; 20140179770; 20140186843; 20140186919; 20140186958; 20140189896; 20140227787; 20140234972; 20140242664; 20140242699; 20140242700; 20140242702; 20140248702; 20140256046;20140273037; 20140273226; 20140273230; 20140273231; 20140273232; 20140273233; 20140273234; 20140273235; 20140287938; 20140295556; 20140295557; 20140298547; 20140304853; 20140309487; 20140310828; 20140310830; 20140315985; 20140335063; 20140335620; 20140342456; 20140342457; 20140342458; 20140349400; 20140349405; 20140356867; 20140356956; 20140356958; 20140356959; 20140357523; 20140357530; 20140364333; and 20140377868; all of which are hereby incorporated by reference in their entirety.

[0253] For example, the present disclosure provides (but is not limited to) methods of cleaving a target nucleic acid; methods of editing a target nucleic acid; methods of modulating transcription from a target nucleic acid; methods of increasing transcription from a target nucleic acid; methods of decreasing transcription from a target nucleic acid; methods of isolating a target nucleic acid; methods of binding a target nucleic acid; methods of imaging a target nucleic acid; methods of modifying a target nucleic acid; methods of editing a target nucleic acid; methods of gene editing; and the like.

[0254] As used herein, the terms / phrases “contact a target nucleic acid” and “contacting a target nucleic acid”, for example, with a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) encompass all methods for contacting the target nucleic acid. For example, a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) can be provided as protein, RNA (encoding the polypeptide), or DNA (encoding the polypeptide); while a Cas9 guide RNA can be provided as a guide RNA or as a nucleic acid encoding the guide RNA. As such, when, for example, performing a method in a cell (e.g., inside of a cell in vitro, inside of a cell in vivo, inside of a cell ex vivo), a method that includes contacting the target nucleic acid encompasses the introduction into the cell of any or all of the components in their active / final state (e.g., in the form of a protein(s) [for the protein], in the form of an RNA [for the guide RNA]), and also encompasses the introduction into the cell of one or more nucleic acids (DNA, RNA) encoding one or more of the components (e.g., nucleic acid(s) having nucleotide sequence(s) encoding a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide), nucleic acid(s) having nucleotide sequence(s) encoding guide RNA(s), and the like). In some cases, the protein and the guide RNA are introduced as an RNP (i.e. , pre-complexed). In some cases, a donor polynucleotide is also introduced into a cell. Because the methods can also be performed in vitro outside of a cell, a method that includes contacting a target nucleic acid, (unlessotherwise specified) encompasses contacting outside of a cell in vitro, inside of a cell in vitro (e.g., a cell line in culture), inside of a cell in vivo, inside of a cell ex vivo (e.g., a cultured primary cell from an individual), etc.

[0255] In some cases, a subject method is a method that includes contacting a target nucleic acid with a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide). In some cases, a subject method includes contacting a target nucleic acid with a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) and a Cas9 guide RNA. In some cases, a subject method includes contacting a target nucleic acid with a protein of the present disclosure (e.g., subject Cas9 fusion polypeptide) and a Cas9 guide RNA, and a donor polynucleotide.Target nucleic acids and target cells of interest

[0256] A protein of the present disclosure (e.g., subject Cas9 fusion polypeptide), when bound to a guide RNA, can bind to a target nucleic acid, and in some cases, can bind to and modify a target nucleic acid. A target nucleic acid can be any nucleic acid (e.g., DNA, RNA), can be double stranded or single stranded, can be any type of nucleic acid (e.g., a chromosome, derived from a chromosome, chromosomal, plasmid, viral, extracellular, intracellular, mitochondrial, chloroplast, linear, circular, etc.) and can be from any organism (e.g., as long as the Cas9 guide RNA can hybridize to a target sequence in a target nucleic acid, that target nucleic acid can be targeted - taking the PAM into account for double stranded target nucleic acids, as would be understood by one of ordinary skill in the art and as described elsewhere herein).

[0257] A target nucleic acid can be DNA or RNA. A target nucleic acid can be double stranded (e.g., dsDNA, dsRNA) or single stranded (e.g., ssRNA, ssDNA). In some cases, a target nucleic acid is single stranded. In some cases, a target nucleic acid is a single stranded RNA (ssRNA). In some cases, a target ssRNA (e.g., a target cell ssRNA, a viral ssRNA, etc.) is selected from: mRNA, rRNA, tRNA, non-coding RNA (ncRNA), long non-coding RNA (IncRNA), and microRNA (miRNA). In some cases, a target nucleic acid is a single stranded DNA (ssDNA) (e.g., a viral DNA). In some cases, a target nucleic acid is a double stranded DNA (dsDNA).

[0258] A target nucleic acid can be located anywhere, for example, outside of a cell in vitro, inside of a cell in vitro (e.g., in a cell in culture), inside of a cell in vivo, inside of a cell ex vivo (e.g., a cell recently isolated from an individual); or inside of an organelle (e.g., mitochondrion; nucleus; etc.) within a cell that is in vitro, in vivo, or ex vivo. Suitabletarget cells (which can comprise target nucleic acids such as genomic DNA) include, but are not limited to: a bacterial cell; an archaeal cell; a cell of a single-cell eukaryotic organism; a plant cell; an algal cell, e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like; a fungal cell (e.g., a yeast cell); an animal cell; a cell from an invertebrate animal (e.g. fruit fly, a cnidarian, an echinoderm, a nematode, etc.); a cell of an insect (e.g., a mosquito; a bee; an agricultural pest; etc ); a cell of an arachnid (e.g., a spider; a tick; etc.); a cell from a vertebrate animal (e.g., a fish, an amphibian, a reptile, a bird, a mammal); a cell from a mammal (e.g., a cell from a rodent; a cell from a human; a cell of a non-human mammal; a cell of a rodent (e.g., a mouse, a rat); a cell of a lagomorph (e.g., a rabbit); a cell of an ungulate (e.g., a cow, a horse, a camel, a llama, a vicuna, a sheep, a goat, etc.); a cell of a marine mammal (e.g., a whale, a seal, an elephant seal, a dolphin, a sea lion; etc.) and the like.

[0259] Any type of cell may be of interest (e.g. a stem cell, e.g. an embryonic stem (ES) cell, an induced pluripotent stem (iPS) cell, a germ cell (e.g., an oocyte, a sperm, an oogonia, a spermatogonia, etc.), a somatic cell, e.g. a fibroblast, a hematopoietic cell, a neuron, a muscle cell, a bone cell, a hepatocyte, a pancreatic cell; an in vitro or in vivo embryonic cell of an embryo at any stage, e.g., a 1-cell, 2-cell, 4-cell, 8-cell, etc. stage zebrafish embryo; etc.).

[0260] Cells may be from established cell lines (e.g., in vitro cell culture) or they may be primary cells (e.g., ex vivo cells in culture), where “primary cells”, “primary cell lines”, and “primary cultures” are used interchangeably herein to refer to cells and cells cultures that have been derived from a subject and allowed to grow in vitro for a limited number of passages, i.e. splittings, of the culture. For example, primary cultures are cultures that may have been passaged 0 times, 1 time, 2 times, 4 times, 5 times, 10 times, or 15 times, but not enough times go through the crisis stage. Typically, the primary cell lines are maintained for fewer than 10 passages in vitro. Target cells can be unicellular organisms and / or can be grown in culture. If the cells are primary cells, they may be harvest from an individual by any convenient method. For example, leukocytes may be conveniently harvested by apheresis, leukocytapheresis, density gradient separation, etc., while cells from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc. can be conveniently harvested by biopsy.

[0261] In some of the above applications, the subject methods may be employed to induce target nucleic acid cleavage, target nucleic acid modification, and / or to bind targetnucleic acids (e.g., for visualization, for collecting and / or analyzing, etc.) in cells in vivo and / or ex vivo and / or in vitro (e.g., to disrupt production of a protein encoded by a targeted mRNA). Because the guide RNA provides specificity by hybridizing to target nucleic acid, a cell of interest in the disclosed methods may include a cell from any organism (e.g. a bacterial cell, an archaeal cell, a cell of a single-cell eukaryotic organism, a plant cell, an algal cell, e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like, a fungal cell (e.g., a yeast cell), an animal cell, a cell from an invertebrate animal (e.g. fruit fly, cnidarian, echinoderm, nematode, etc.), a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), a cell from a mammal, a cell from a rodent, a cell from a human, etc.). In some cases, a Cas9 fusion polypeptide of the present disclosure (and / or nucleic acid encoding the protein such as DNA and / or RNA), and / or Cas9 guide RNA (and / or a DNA encoding the guide RNA), and / or donor template (donor polynucleotide), and / or RNP can be introduced into an individual (i.e. , the target cell can be in vivo) (e.g., a mammal, a rat, a mouse, a pig, a primate, a non-human primate, a human, etc.). In some case, such an administration can be for the purpose of treating and / or preventing a disease, e.g., by editing the genome of targeted cells.

[0262] Plant cells include cells of a monocotyledon, and cells of a dicotyledon. The cells can be root cells, leaf cells, cells of the xylem, cells of the phloem, cells of the cambium, apical meristem cells, parenchyma cells, collenchyma cells, sclerenchyma cells, and the like. Plant cells include cells of agricultural crops such as wheat, corn, rice, sorghum, millet, soybean, etc. Plant cells include cells of agricultural fruit and nut plants, e.g., plant that produce apricots, oranges, lemons, apples, plums, pears, almonds, etc.

[0263] Additional examples of target cells are listed above in the section titled “Modified cells.” Non-limiting examples of cells (target cells) include: a prokaryotic cell, eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a single-cell eukaryotic organism, a protozoa cell, a cell from a plant (e.g., cells from plant crops, fruits, vegetables, grains, soy bean, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, angiosperms, ferns, clubmosses, hornworts, liverworts, mosses, dicotyledons, monocotyledons, etc.), an algal cell, (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, and the like), seaweeds (e.g. kelp) afungal cell (e.g., a yeast cell, a cell from a mushroom), an animal cell, a cell from an invertebrate animal (e.g., fruit fly, cnidarian, echinoderm, nematode, etc.), a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), a cell from a mammal (e.g., an ungulate (e.g., a pig, a cow, a goat, a sheep); a rodent (e.g., a rat, a mouse); a non-human primate; a human; a feline (e.g., a cat); a canine (e.g., a dog); etc.), and the like. In some cases, the cell is a cell that does not originate from a natural organism (e.g., the cell can be a synthetically made cell; also referred to as an artificial cell).

[0264] A cell can be an in vitro cell (e.g., established cultured cell line). A cell can be an ex vivo cell (cultured cell from an individual). A cell can be an in vivo cell (e.g., a cell in an individual). A cell can be an isolated cell. A cell can be a cell inside of an organism. A cell can be an organism. A cell can be a cell in a cell culture (e.g., in vitro cell culture). A cell can be one of a collection of cells. A cell can be a prokaryotic cell or derived from a prokaryotic cell. A cell can be a bacterial cell or can be derived from a bacterial cell. A cell can be an archaeal cell or derived from an archaeal cell. A cell can be a eukaryotic cell or derived from a eukaryotic cell. A cell can be a plant cell or derived from a plant cell. A cell can be an animal cell or derived from an animal cell. A cell can be an invertebrate cell or derived from an invertebrate cell. A cell can be a vertebrate cell or derived from a vertebrate cell. A cell can be a mammalian cell or derived from a mammalian cell. A cell can be a rodent cell or derived from a rodent cell. A cell can be a human cell or derived from a human cell. A cell can be a microbe cell or derived from a microbe cell. A cell can be a fungi cell or derived from a fungi cell. A cell can be an insect cell. A cell can be an arthropod cell. A cell can be a protozoan cell. A cell can be a helminth cell.

[0265] Suitable cells include a stem cell (e.g. an embryonic stem (ES) cell, an induced pluripotent stem (iPS) cell; a germ cell (e.g., an oocyte, a sperm, an oogonia, a spermatogonia, etc.); a somatic cell, e.g. a fibroblast, an oligodendrocyte, a glial cell, a hematopoietic cell, a neuron, a muscle cell, a bone cell, a hepatocyte, a pancreatic cell, etc.

[0266] Suitable cells include human embryonic stem cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, blood stem cells, myoblasts, adult stem cells, bone marrow cells, mesenchymal cells, embryonic stem cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, neurons, astrocytes, islet cells, alpha cells, beta cells delta cells, osteoblasts, chondrocytes, exogenous cells, endogenouscells, stem cells, hematopoietic stem cells, bone-marrow derived progenitor cells, myocardial cells, skeletal cells, fetal cells, undifferentiated cells, multi-potent progenitor cells, unipotent progenitor cells, monocytes, cardiac myoblasts, skeletal myoblasts, macrophages, capillary endothelial cells, xenogenic cells, allogenic cells, and post-natal stem cells.

[0267] In some cases, the cell is an immune cell, a neuron, an epithelial cell, and endothelial cell, or a stem cell. In some cases, the immune cell is a T cell, a B cell, a monocyte, a natural killer cell, a dendritic cell, or a macrophage. In some cases, the immune cell is a cytotoxic T cell. In some cases, the immune cell is a helper T cell. In some cases, the immune cell is a regulatory T cell (Treg).

[0268] In some cases, the cell is a stem cell. Stem cells include adult stem cells. Adult stem cells are also referred to as somatic stem cells.

[0269] Adult stem cells are resident in differentiated tissue, but retain the properties of selfrenewal and ability to give rise to multiple cell types, usually cell types typical of the tissue in which the stem cells are found. Numerous examples of somatic stem cells are known to those of skill in the art, including muscle stem cells; hematopoietic stem cells; epithelial stem cells; neural stem cells; mesenchymal stem cells; mammary stem cells; intestinal stem cells; mesodermal stem cells; endothelial stem cells; olfactory stem cells; neural crest stem cells; and the like.

[0270] Stem cells of interest include mammalian stem cells, where the term “mammalian” refers to any animal classified as a mammal, including humans; non-human primates; domestic and farm animals; and zoo, laboratory, sports, or pet animals, such as dogs, horses, cats, cows, mice, rats, rabbits, etc. In some cases, the stem cell is a human stem cell. In some cases, the stem cell is a rodent (e.g., a mouse; a rat) stem cell. In some cases, the stem cell is a non-human primate stem cell.

[0271] Stem cells can express one or more stem cell markers, e.g., SOX9, KRT19, KRT7, LGR5, CA9, FXYD2, CDH6, CLDN18, TSPAN8, BPIFB1 , OLFM4, CDH17, and PPARGC1A.

[0272] In some cases, the stem cell is a hematopoietic stem cell (HSC). HSCs are mesoderm-derived cells that can be isolated from bone marrow, blood, cord blood, fetal liver and yolk sac. HSCs are characterized as CD34+ and CD3-. HSCs can repopulate the erythroid, neutrophil-macrophage, megakaryocyte and lymphoid hematopoietic cell lineages in vivo. In vitro, HSCs can be induced to undergo at least some self-renewing cell divisions and can be induced to differentiate to the same lineages as is seen in vivo. As such, HSCs can be induced to differentiate into one ormore of erythroid cells, megakaryocytes, neutrophils, macrophages, and lymphoid cells.

[0273] In other cases, the stem cell is a neural stem cell (NSC). Neural stem cells (NSCs) are capable of differentiating into neurons, and glia (including oligodendrocytes, and astrocytes). A neural stem cell is a multipotent stem cell which is capable of multiple divisions, and under specific conditions can produce daughter cells which are neural stem cells, or neural progenitor cells that can be neuroblasts or glioblasts, e.g., cells committed to become one or more types of neurons and glial cells respectively. Methods of obtaining NSCs are known in the art.

[0274] In other cases, the stem cell is a mesenchymal stem cell (MSC). MSCs originally derived from the embryonal mesoderm and isolated from adult bone marrow, can differentiate to form muscle, bone, cartilage, fat, marrow stroma, and tendon. Methods of isolating MSC are known in the art; and any known method can be used to obtain MSC. See, e.g., U.S. Pat. No. 5,736,396, which describes isolation of human MSC.

[0275] A cell is in some cases a plant cell. A plant cell can be a cell of a monocotyledon. A cell can be a cell of a dicotyledon.

[0276] In some cases, the cell is a plant cell. For example, the cell can be a cell of a major agricultural plant, e.g., Barley, Beans (Dry Edible), Canola, Corn, Cotton (Pima), Cotton (Upland), Flaxseed, Hay (Alfalfa), Hay (Non-Alfalfa), Oats, Peanuts, Rice, Sorghum, Soybeans, Sugarbeets, Sugarcane, Sunflowers (Oil), Sunflowers (Non-Oil), Sweet Potatoes, Tobacco (Burley), Tobacco (Flue-cured), Tomatoes, Wheat (Durum), Wheat (Spring), Wheat (Winter), and the like. As another example, the cell is a cell of a vegetable crops which include but are not limited to, e.g., alfalfa sprouts, aloe leaves, arrow root, arrowhead, artichokes, asparagus, bamboo shoots, banana flowers, bean sprouts, beans, beet tops, beets, bittermelon, bok choy, broccoli, broccoli rabe (rappini), brussels sprouts, cabbage, cabbage sprouts, cactus leaf (nopales), calabaza, cardoon, carrots, cauliflower, celery, chayote, Chinese artichoke (crosnes), Chinese cabbage, Chinese celery, Chinese chives, choy sum, chrysanthemum leaves (tung ho), collard greens, corn stalks, corn-sweet, cucumbers, daikon, dandelion greens, dasheen, dau mue (pea tips), donqua (winter melon), eggplant, endive, escarole, fiddle head ferns, field cress, frisee, gai choy (Chinese mustard), gailon, galanga (siam, thai ginger), garlic, ginger root, gobo, greens, hanover salad greens, huauzontle, Jerusalem artichokes, jicama, kale greens, kohlrabi, lamb's quarters (quilete), lettuce (bibb), lettuce (boston), lettuce (boston red),lettuce (green leaf), lettuce (iceberg), lettuce (lolla rossa), lettuce (oak leaf - green), lettuce (oak leaf - red), lettuce (processed), lettuce (red leaf), lettuce (romaine), lettuce (ruby romaine), lettuce (russian red mustard), linkok, Io bok, long beans, lotus root, mache, maguey (agave) leaves, malanga, mesculin mix, mizuna, moap (smooth luffa), moo, moqua (fuzzy squash), mushrooms, mustard, nagaimo, okra, ong choy, onions green, opo (long squash), ornamental corn, ornamental gourds, parsley, parsnips, peas, peppers (bell type), peppers, pumpkins, radicchio, radish sprouts, radishes, rape greens, rape greens, rhubarb, romaine (baby red), rutabagas, salicornia (sea bean), sinqua (angled / ridged luffa), spinach, squash, straw bales, sugarcane, sweet potatoes, swiss chard, tamarindo, taro, taro leaf, taro shoots, tatsoi, tepeguaje (guaje), tindora, tomatillos, tomatoes, tomatoes (cherry), tomatoes (grape type), tomatoes (plum type), tumeric, turnip tops greens, turnips, water chestnuts, yampi, yams (names), yu choy, yuca (cassava), and the like.

[0277] A cell is in some cases an arthropod cell. For example, the cell can be a cell of a suborder, a family, a sub-family, a group, a sub-group, or a species of, e.g., Chelicerata, Myriapodia, Hexipodia, Arachnida, Insecta, Archaeognatha, Thysanura, Palaeoptera, Ephemeroptera, Odonata, Anisoptera, Zygoptera, Neoptera, Exopterygota, Plecoptera, Embioptera, Orthoptera, Zoraptera, Dermaptera, Dictyoptera, Notoptera, Grylloblattidae, Mantophasmatidae, Phasmatodea, Blattaria, Isoptera, Mantodea, Parapneuroptera, Psocoptera, Thysanoptera, Phthiraptera, Hemiptera, Endopterygota or Holometabola, Hymenoptera, Coleoptera, Strepsiptera, Raphidioptera, Megaloptera, Neuroptera, Mecoptera, Siphonaptera, Diptera, Trichoptera, or Lepidoptera.

[0278] A cell is in some cases an insect cell. For example, in some cases, the cell is a cell of a mosquito, a grasshopper, a true bug, a fly, a flea, a bee, a wasp, an ant, a louse, a moth, or a beetle.Compositions

[0279] The present disclosure provides compositions, including pharmaceutical compositions, comprising a Cas9 fusion polypeptide of the present disclosure (and / or a nucleic acid encoding it). In some cases, a composition also includes a Cas9 guide RNA and / or a nucleic acid encoding it (e.g., in some cases as a Cas9 fusion polypeptide / guide RNA RNP). In some cases, a composition includes a donor polynucleotide. In some cases, a composition includes a nucleic acid of the presentdisclosure, a recombinant expression vector of the present disclosure, an RNP of the present disclosure, or a system of the present disclosure.

[0280] Pharmaceutical compositions can include, depending on the formulation desired, pharmaceutically-acceptable, non-toxic carriers of diluents, which are defined as vehicles commonly used to formulate pharmaceutical compositions for animal or human administration. The diluent is selected so as not to affect the biological activity of the combination. Examples of such diluents are distilled water, buffered water, physiological saline, phosphate buffered saline (PBS), Ringer's solution, dextrose solution, and Hank's solution. In addition, the pharmaceutical composition or formulation can include other carriers, adjuvants, or non-toxic, nontherapeutic, nonimmunogenic stabilizers, excipients and the like. The compositions can also include additional substances to approximate physiological conditions, such as pH adjusting and buffering agents, toxicity adjusting agents, wetting agents and detergents.

[0281] The composition can also include any of a variety of stabilizing agents, such as an antioxidant for example. When the pharmaceutical composition includes a polypeptide, the polypeptide can be complexed with various well-known compounds that enhance the in vivo stability of the polypeptide, or otherwise enhance its pharmacological properties (e.g., increase the half-life of the polypeptide, reduce its toxicity, enhance solubility or uptake). Examples of such modifications or complexing agents include sulfate, gluconate, citrate and phosphate. The nucleic acids or polypeptides of a composition can also be complexed with molecules that enhance their in vivo attributes. Such molecules include, for example, carbohydrates, polyamines, amino acids, other peptides, ions (e.g., sodium, potassium, calcium, magnesium, manganese), and lipids.

[0282] Further guidance regarding formulations that are suitable for various types of administration can be found in Remington's Pharmaceutical Sciences, Mace Publishing Company, Philadelphia, Pa., 17th ed. (1985). For a brief review of methods for drug delivery, see, Langer, Science 249:1527-1533 (1990).

[0283] The data obtained from cell culture and / or animal studies can be used in formulating a range of dosages for humans. The dosage of the active ingredient typically lines within a range of circulating concentrations that include the ED50 with low toxicity. The dosage can vary within this range depending upon the dosage form employed and the route of administration utilized.

[0284] The components used to formulate the pharmaceutical compositions are generally of high purity and are substantially free of potentially harmful contaminants (e.g., at least National Food (NF) grade, generally at least analytical grade, and more typically at least pharmaceutical grade). Moreover, compositions intended for in vivo use are usually sterile. To the extent that a given compound must be synthesized prior to use, the resulting product is typically substantially free of any potentially toxic agents, particularly any endotoxins, which may be present during the synthesis or purification process. Compositions for parental administration are also sterile, substantially isotonic and made under GMP conditions.Kits

[0285] Provided are kits / systems for carrying out a subject method. Such kits comprise various combinations of components useful in any of the methods described elsewhere herein. As an example, in some embodiments a subject kit includes a subject Cas9 fusion polypeptide or a nucleic acid encoding it. In some cases, a subject kit further includes a guide RNA or a nucleic acid encoding it.

[0286] A kit can further include one or more additional reagents, where such additional reagents can be any convenient reagent. Components of a subject kit can be in separate containers; or can be combined in a single container. In some cases one or more of a kit’s components are pharmaceutically formulated for administration to a human.

[0287] In addition to above-mentioned components, a subject kit can further include instructions for using the components of the kit to practice the subject methods (e.g., dosing instructions, instructions to administer the component(s) to an individual. The instructions for practicing the subject methods are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e. , associated with the packaging or subpackaging) etc. In some embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g. CD-ROM, diskette, flash drive, etc. In some embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g. via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. Aswith the instructions, this means for obtaining the instructions is recorded on a suitable substrate.EXEMPLARY NON-LIMITING ASPECTS OF THE DISCLOSURE

[0288] Aspects, including embodiments, of the present subject matter described above may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure are provided below. As will be apparent to those of ordinary skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below. It will be apparent to one of ordinary skill in the art that various changes and modifications can be made without departing from the spirit or scope of the invention.1. A Cas9 fusion polypeptide, comprising: (a) a Cas9 protein, and (b) one or more linker-NLS1-linker-NLS2-linker (hiNLS) modules inserted internally within the Cas9 protein, wherein NLS1 is a first nuclear localization signal (NLS) and NLS2 is a second NLS, and wherein NLS1 and NLS2 can be the same or different, wherein the one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248, of SEQ ID NO: 1, or any combination thereof.2. The Cas9 fusion polypeptide of 1 , wherein the Cas9 fusion polypeptide further comprises at least one N-terminal NLS and / or at least one C-terminal NLS.3. The Cas9 fusion polypeptide of 1 , wherein the Cas9 fusion polypeptide further comprises at one C-terminal NLS.4. The Cas9 fusion polypeptide of 1 , wherein the Cas9 fusion polypeptide further comprises one N-terminal NLS and two C-terminal NLSs.5. The Cas9 fusion polypeptide of any one of 1-4, wherein the NLS1 and / or NLS2 of at least one of the one or more hiNLS modules comprises: PKKKRKV (SEQ ID NO: 41), PAAKKKKLD (SEQ ID NO: 42), KRPAATKKAGQAKKKK (SEQ ID NO: 43), PAAKRVKLD (SEQ ID NO: 44), KRTADGSEFESPKKKRKVE (SEQ ID NO: 45), KRTADGSEFESPKKARKVE (SEQ ID NO: 46),KRTADGSEFESPKKKAKVE (SEQ ID NO: 47), or KR(X1 )5-I5KK(X2)(X3)KV (SEQ ID NO: 48), wherein X1 is any amino acid, X2 is lysine or alanine, and X3 is lysine, arginine, or alanine.6. The Cas9 fusion polypeptide of any one of 1-5, wherein the NLS1 and / or the NLS2 of at least one of the one or more hiNLS modules is PKKKRKV (SEQ ID NO: 41).7. The Cas9 fusion polypeptide of any one of 1-6, wherein the NLS1 and / or the NLS2 of at least one of the one or more hiNLS modules is PAAKKKKLD (SEQ ID NO: 42).8. The Cas9 fusion polypeptide of any one of 1-4, wherein at least one of the one or more hiNLS modules comprises the sequence GSGSGPKKKRKVGSGSGPKKKRKVGSGSG (SEQ ID NO: 66).9. The Cas9 fusion polypeptide of any one of 1-4, wherein at least one of the one or more hiNLS modules comprises the sequence GSGSGPAAKKKKLDGSGSGPAAKKKKLDGSGSG (SEQ ID NO: 67).10. The Cas9 fusion polypeptide of any one of 1-9, wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205 of SEQ ID NO: 1.11. The Cas9 fusion polypeptide of any one of 1-10, wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to K468 of SEQ ID NO: 1.12. The Cas9 fusion polypeptide of any one of 1-11 , wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to 11022 of SEQ ID NO: 1.13. The Cas9 fusion polypeptide of any one of 1-12, wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to S1248 of SEQ ID NO: 1.14. The Cas9 fusion polypeptide of any one of 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that provides transcription modulation activity, DNA-modifying activity, and / or protein modifying activity.15. The Cas9 fusion polypeptide of any one of 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that provides a DNA- modifying activity selected from: nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deaminationactivity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity and glycosylase activity.16. The Cas9 fusion polypeptide of any one of 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous deaminase polypeptide.17. The Cas9 fusion polypeptide of any one of 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that provides reverse transcriptase activity.18. The Cas9 fusion polypeptide of any one of 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that exhibits histone modification activity.19. The Cas9 fusion polypeptide of any one of 1-18, wherein the Cas9 protein lacks a catalytically active RuvC domain and / or lacks a catalytically active HNH domain.20. The Cas9 fusion polypeptide of 19, wherein the Cas9 protein has nickase activity.21 . The Cas9 fusion polypeptide of 19, wherein the Cas9 protein lacks a catalytically active RuvC domain and lacks a catalytically active HNH domain.22. The Cas9 fusion polypeptide of any one of 1-21 , wherein the Cas9 protein is an S. pyogenes, S. aureus, S. thermophilus, orN. meningitidis Cas9.23. The Cas9 fusion polypeptide of any one of 1-21 , wherein the Cas9 protein, when considered in the absence of the inserted one or more hiNLS modules, comprises an amino acid sequence that is 80% or more identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 1.24. The Cas9 fusion polypeptide of any one of 1-21 , wherein the Cas9 protein, when considered in the absence of the inserted one or more hiNLS modules, comprises an amino acid sequence that is 80% or more identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-4.25. The Cas9 fusion polypeptide of any one of 1-21 , wherein the Cas9 protein, when considered in the absence of the inserted one or more hiNLS modules, comprises an amino acid sequence that is 80% or more identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-33.26. The Cas9 fusion polypeptide of any one of 1-25, wherein the linker is any one of SEQ ID NOs: 76-81.27. The Cas9 fusion polypeptide of any one of 1-25, wherein the linker is GSGSG (SEQ ID NO: 76).28. A nucleic acid comprising a nucleotide sequence encoding the Cas9 fusion polypeptide of any one of 1-27.29. The nucleic acid of 28, wherein said nucleotide sequence is operably linked to a promoter.30. The nucleic acid of 29, wherein the promoter is operable in a eukaryotic cell.31 . The nucleic acid of any one of 28-30, wherein said nucleic acid is a recombinant expression vector.32. The nucleic acid of 31 , wherein the recombinant expression vector is a linear expression vector, a circular expression vector, a plasmid, or a viral expression vector.33. The nucleic acid of 28, wherein said nucleic acid is an mRNA.34. A cell, comprising the Cas9 fusion polypeptide of any one of 1-27 and / or the nucleic acid of any one of 28-33.35. The cell of 34, further comprising a Cas9 guide RNA or a nucleic acid encoding the Cas9 guide RNA.36. The cell of 34 or 35, wherein the cell is a eukaryotic cell.37. The cell of 34 or 35, wherein the cell is a plant cell, an invertebrate cell, a vertebrate cell, a fish cell, an amphibian cell, an avian cell, a mammalian cell, a rodent cell, a mouse cell, a rat cell, a pig cell, a non-human primate cell, a primate cell, or a human cell.38. The cell of 34 or 35, wherein the cell is a prokaryotic cell.39. A composition comprising:(a) the Cas9 fusion polypeptide of any one of 1-27 or the nucleic acid of any one of 28-33; and(b) a Cas9 guide RNA or a nucleic acid encoding the Cas9 guide RNA.40. The composition of 39, comprising a ribonucleoprotein (RNP) complex comprising the Cas9 fusion polypeptide and the Cas9 guide RNA.41 . A method of binding, modifying, and / or modulating transcription of a target DNA, the method comprising: contacting a target DNA inside of a cell with the Cas9 fusion polypeptide of any one of 1-27 and a Cas9 guide RNA.42. The method of 41, wherein said contacting causes cleavage of the target DNA, thereby modifying the target DNA.43. The method of 41 or 42, wherein said contacting results in editing of the target DNA.44. The method of 41, wherein the Cas9 fusion polypeptide is fused to a heterologous protein that is a transcription activator, a transcription repressor, a histone modifier, a deaminase, or a reverse transcriptase.45. The method of any one of 41-44, wherein the target DNA is genomic DNA.46. The method of any one of 41-45, wherein the cell is a eukaryotic cell.47. The method of 46, wherein the eukaryotic cell is in vivo.48. The method of 46, wherein the eukaryotic cell is in vitro or ex vivo.49. The method of any one of 41-48, wherein said contacting comprises introducing into the cell: (a) the Cas9 fusion polypeptide or a nucleic acid encoding the Cas9 fusion polypeptide; and (b) the Cas9 guide RNA or a nucleic acid encoding the Cas9 guide RNA.50. The method of any one of 41-49, wherein the method comprises introducing a donor polynucleotide into the cell.EXPERIMENTAL EXAMPLES

[0289] The following examples are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way be construed as being limited to the following examples, but rather should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.

[0290] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative examples, make and utilize the present invention and practice the claimed methods. The following working examples therefore are not to be construed as limiting in any way the remainder of the disclosure.

[0291] General methods in molecular and cellular biochemistry can be found in such standard textbooks as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., HaRBor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995);Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference. Reagents, cloning vectors, cells, and kits for methods referred to in, or related to, this disclosure are available from commercial vendors such as BioRad, Agilent Technologies, Thermo Fisher Scientific, Sigma-Aldrich, New England Biolabs (NEB), Takara Bio USA, Inc., and the like, as well as repositories such as e.g., Addgene, Inc., American Type Culture Collection (ATCC), and the like.Example 1: Hairpin Internal Nuclear Localization Signals (hiNLS) in Cas9 Enhance Editing in Primary Human Lymphocytes

[0292] The approach presented here enhances the editing efficiency of Cas9 - as demonstrated by increasing the editing efficiency of Cas9 in cultured (in vitro) primary human CD4+T cells through the development of Cas9 variants featuring one or more hairpin internal NLS (hiNLS) modules, each of which comprises a two-NLS “hairpin” seguence introduced at a site within the native Cas9 backbone. A systematic approach was employed to design and characterization Cas9 variants with hiNLS motifs, culminating in an array of new constructs. When delivered via electroporation or via PERC, hiNLS Cas9 constructs exhibited higher editing efficiencies than a previously reported NLS-rich construct with an appealing array of properties. Furthermore, hiNLS Cas9 constructs can be produced with high yield, suggesting that they will likely be suitable for widespread use in research, biotechnology, and therapeutics.RESULTS AND DISCUSSION

[0293] hiNLS Cas9 design. Cas9 was engineered for efficient nuclear import by introducing NLS sequences in the protein’s backbone, in hopes that new variants could be discovered with robust activity and desirable properties that other NLS-rich constructs lack. This approach represents a departure from previous efforts, which introduced NLS sequences at one or both termini of the engineered CRISPR proteins. Internal NLS sequences have been introduced within the backbone of a CRISPR enzyme in the context of base editor engineering (with NLS sequence fused adjacent to an inserted deaminase domain). However, it was tested here whether NLS sequencescould be inserted at one or more backbone sites within S. pyogenes Cas9 to confer efficient nuclear localization, and thus efficient genome editing efficiency.

[0294] Because the N- and C-termini of Cas9 are proximal to one another in the 3D structure of Cas9, sites were selected for NLS insertion that were more evenly distributed across the surface of the enzyme.

[0295] Previously, randomized insertional mutagenesis experiments studying Cas9 revealed regions capable of tolerating insertions of a PDZ domain without disruption of the enzyme’s binding and cleavage functions (Oakes et al., Nature biotechnology, 2016. 34(6): p. 646-651). As part of the work described herein, local clusters of amino acids were explored that demonstrated a high tolerance to insertions, often around flexible loops, the ends of helices, and at solvent-exposed residues. Specifically, regions of high plasticity were profiled at six clusters within the helical recognition (REC) lobe, the linker between the REC and nuclease lobes, the HNH domain, three extended sites in the RuvCIII region and throughout the PAM interacting (C-terminal) domain. Another study demonstrated that catalytically-dead Cas9 (dCas9) can tolerate large single deletions to the REC2, REC3, HNH, and RuvC domains (Shams et al., Nature Communications, 2021. 12(1): p. 5664). Thus, it is possible to insert an entire exogenous domain at numerous sites within Cas9’s primary sequence while maintaining near-native levels of activity.

[0296] It was hypothesized that installing hiNLS motifs (Fig. 3) at the designated locations would be least likely to perturb the enzyme’s native function. Based on this evidence, four regions within the Cas9 backbone (referred to as sites 1-4) (see Fig. 1) were selected to introduce hiNLS modules either alone or in combination. Next, the three- dimensional structure of Cas9 was assessed to identify available surface residues that overlapped with regions of high plasticity that were previously observed to tolerate a synthetic domain insertion while retaining RNA-guided DNA-binding activity. hiNLS were installed immediately following four amino acid positions in Cas9: Gly-205 (helical domain), Lys-468 (helical domain), lle-1022 (RuvC-lll), and Ser-1248 (CTD) (Fig.1a, b).

[0297] 15 hiNLS Cas9 variants were engineered (Fig. 1c) based on two cysteine-free Cas9 “backbone” constructs distinguished by terminal NLS positions: s-Cas9 and t-Cas9 (Fig. 1b). The s-Cas9 backbone possesses a single monopartite SV40 motif at its C- terminus (akin to the 1 * NLS Cas9 reported previously), and the t-Cas9 backbone (Addgene ID# 196245; referred to as 3* NLS-Cas9 when originally reported and later as Cas9-triNLS) bears three different terminal NLS motifs: N-terminal c-Myc (M) aswell as C-terminal SV40 (S) and nucleoplasmin sequences, respectively. Each hiNLS module features a “linker-NLS-linker-NLS-linker” layout, selected to increase local NLS valency and allow each NLS to adopt the extended, linear orientation that is optimal for binding to importin proteins. Constructs were named based on the backbone type (either “s” or “t”) as well as the hiNLS type (an “S” or “M” pair) and location at position 1 (Gly-205), position 2 (Lys-468), position 3 (He- 1022), and / or position 4 (Ser-1248) (Fig. 1b).

[0298] Single hiNLS installation in Cas9 impacts B2M knockout. In effort to enhance genome editing efficiencies in primary human CD4+T cells, the impact of strategically inserting a single hiNLS at four distinct internal sites within Cas9 was investigated. To evaluate the potential improvements in editing efficacy, hiNLS motifs were incorporated at each of the specified internal sites 1-4 within the Cas9 framework, as delineated in Figure 1 b. The efficacy of the hiNLS Cas9 S2 / W-RNP variants targeting B2M in primary human CD4+T cells was assessed. Knockout efficiencies were evaluated four days post-delivery by either electroporation or PERC, a peptide- enabled RNP delivery for CRISPR engineering. In contrast to electroporation, PERC is a hardware-free approach to CRISPR delivery that could conceivably be used in vivo, and can be considered a proxy for other in vivo delivery technologies - such as virus-like-particle or lipid nanoparticle - that will typically introduce a relatively low “dose” of CRISPR RNP enzyme into cells, heightening the need for potent NLS activity. Low pmol conditions for PERC were used to enhance the dynamic range of the experimental setup, allowing for the sensitive detection of subtle variations, and enabling the identification of pronounced differences among distinct genetic constructs. This approach not only heightened the sensitivity of the measurements but also contributed to a more comprehensive understanding of the intricate biological processes involved in hiNLS editing. Knockout efficiencies were similar in all hiNLS Cas9 variants but one when compared to backbone constructs s-Cas9 and t-Cas9 (Fig. 1), without substantial impact on viability (Fig. 3a). Insertion of a single hiNLS at Cas9 residues Gly-205, lle-1022, or Ser-1248 supported full efficiency of editing at the B2M locus via electroporation. Following PERC, however, three hiNLS variants exhibited editing efficiency superior to that attained by s-Cas9. Notably, novel Cas9 variant s-M4 achieved 21% B2M knockout via PERC compared to 26% by the high performing t-Cas9. The s-M2 construct performed poorly, perhaps due to the insertion site’s proximity to the bridge helix essential for the allosteric signal transduction that underlies the nuclease activity of Cas9.

[0299] Multiplex hiNLS Cas9 boosts CD3 and TRAC knockout. The introduction of three different NLS motifs in Cas9 has proven to improve editing efficiency beyond the enzyme’s native ability. This basis and the observation that the addition of a single hiNLS (M3 and M4) to the Cas9-triNLS construct (t-Cas9) resulted in improved editing outcomes (Fig. 2), the potential synergistic effects of incorporating multiple hiNLS motifs inside the t-Cas9 architecture was explored. To address whether the potency of t-Cas9 may not solely arise from the sheer number of NLS motifs, a comparative analysis was performed to evaluate S-M1M4 against s-S1S4 and s-S1S3S4. t- 1 3S4 was also examined, aiming to discern whether diversity in NLS motifs surpasses considerations of NLS count or charge in influencing editing efficiency. To test whether combining multiple hiNLS in Cas9 would improve editing, multiplex hiNLS motifs were engineered inside Cas9 and CD3 and TRAC knockout efficiencies were evaluated in primary human CD4+T cells (Fig. 2 and Fig. 3). Delivered via electroporation, hiNLS Cas9 CD3-RNP variants s-M3, s-M4, S-M1 M4, S-M1M3M4, and t-S4 exhibited higher CD3 knockout percentages than t-Cas9. Further, variants s- M1 M4, S-M1 3M4, t-S4, t-M4, t-M1M4, and t-M1M3M4 achieved higher CD3 knockout percentages than t-Cas9 using PERC. A slightly less potent gRNA was tested in hopes of observing differences between constructs following delivery via electroporation. In contrast to targeting B2M, CD3 editing efficiencies dramatically change based on different hiNLS constructs (with some out-performing t-Cas9) perhaps demonstrating efficient gRNA can promote maximal editing efficiency regardless of construct potency. To showcase the robust and versatile application of hiNLS Cas9 variants across a variety of target genes, TRAC was targeted in T cells. Electroporated cells treated with hiNLS Cas9 7RAC-RNP variants S-M1M4, s- M1M3M4, t-S4, t-M4, t-M1M4, and t-M1M3M4, and s-M4 exhibited higher TRAC knockout percentages than t-Cas9 without substantial impact on viability (Fig. 3; Fig. 4).DISCUSSION

[0300] The use of Cas9 variants with terminally-fused NLS has been a prevailing strategy in genome editing applications. The engineering of Cas9 variants equipped with hiNLS improves genome editing, e.g., as demonstrated in human lymphocytes. Most notably, protein recovery efforts from prior recombinant high-NLS Cas9 constructs (e.g. 6 x SV40) have commonly produced poor yields incurring 0.6 to 1.5 mg / L, an order of magnitude less than native Cas9 expression. Comparably, hiNLS Cas9constructs reliably produced between 3 to 9 mg / L. Further, strategic positioning of hiNLS within Cas9 constructs has yielded enhanced knockout efficiencies in critical genes such as B2M, CD3, and TRAC, with preservation of cell viability and maintenance of high yields via recombinant expression (Fig. 1c). The versatility demonstrated across different target genes highlights the robustness and potential translational impact of hiNLS Cas9 variants. The innovative strategy presented herein holds promise for advancing T cell-based therapies and underscores the significance of hiNLS in shaping the next generation of Cas9 engineering.

[0301] The findings reported here suggest that the potency of hiNLS in boosting editing efficiency may not be solely dictated by the sheer number of charges or NLS motifs. This study demonstrates that hiNLS constructs yield higher editing efficiencies compared to terminally located NLS Cas9 constructs, offering a promising avenue to load Cas9 with NLS without compromising protein yield. These findings not only contribute to understanding of the factors influencing Cas9 editing but also provide Cas9 variants with enhanced gene editing capabilities.

[0302] This study also demonstrates that incorporating hiNLS within the Cas9 open reading frame surpasses the limitations of traditional terminal NLS tags. The enhanced protein yield with hiNLS is likely multifactorial. Firstly, the mRNA surrounding terminal NLS sequences may indeed fold into non-translatable secondary structures, especially with larger, repetitive tag sequences. Additionally, the highly positively charged nature of these motifs could deter ribosome binding and translation initiation, further amplified with more NLS repeats. Conversely, hiNLS appears to seamlessly embed NLS functionality within the Cas9 structure, minimizing disruption to mRNA folding and ribosome interaction. Furthermore, the internal location likely mitigates electrostatic repulsion from the concentrated positive charge, potentially reducing protein aggregation and improving stability. This suggests that hiNLS not only bypasses translation impediments but also optimizes protein folding and solubility, ultimately leading to significantly higher Cas9 production. Thus, hiNLS not only unlocks higher protein yields similar to native Cas9 but also paves the way for further Cas9 optimization through fine-tuning other internal nucleus-targeting sequences and even removing terminally-fused NLS sequences altogether.

[0303] With hiNLS Cas9, new heights in NLS density and potency are achieved without sacrificing yield. hiNLS marks a crucial shift in Cas9 protein engineering, enabling efficient production of potent gene editing tools and pushing the boundaries of therapeutic protein design. These insights underscore the importance of consideringnot only the role of NLS motifs in enhancing gene editing but also their impact on translation efficiency and overall protein yield, informing the rational design of Cas9 variants for both heightened editing capabilities and robust recombinant protein production.MATERIALS AND METHODS

[0304] Cell Culture. Isolated T cells were thawed (on day -3 relative to delivery) and cultured overnight in X-VIVO 15 medium (Lonza) supplemented with 5% fetal bovine serum (FBS), 50 pM 2-mercaptoethanol and 10 mM A / -acetyl-L-cysteine. Cells were activated (on day -2 relative to delivery) and cultured at 1 x 106cells mL"1for 2 days in supplemented X-VIVO 15 medium with anti-human CD3 / CD28 magnetic Dynabeads (Gibco 40203D) at a bead-to-cell ratio of 1 :1 , 200 U mL-1IL-2 (Proleukin), 5 ng mL-1IL-7 (R&D Systems) and 5 ng mL-1IL-15 (R&D Systems). After 2 days of activation, Dynabeads were removed from the cell culture using an EasySep cell separation magnet (STEMCELL). For genome editing, 200 x 1O3T cells per well were suspended in 100 pl Opti-MEM (Gibco) (treatments detailed below). After 1 h of treatment, FBS-supplemented X-VIVO-15 recovery medium was added to cells. Unless specified otherwise, after genome editing (as described below), cells were cultured in supplemented X-VIVO 15 at 0.5 x 106cells mL'1with 300 U mL-1IL-2 (Proleukin) and split every 2-3 days. Cell viability was assessed using a CellTiter-Glo assay (Promega G7570) according to the manufacturer-provided instructions. Luminescence was measured using a Spark plate reader.

[0305] Peptides. TAT peptide was purchased from GenScript (GSCRPT-RP20256), and other peptides were procured via custom solid phase synthesis (CPC Scientific; 95% purity). All peptides were stored lyophilized or as 10 mM stocks in DMSO at -20 °C in a desiccator.

[0306] Cloning, overexpression, and protein purification. All cloning was conducted by Gibson Assembly Master Mix enzyme cloning methods (New England Biolabs). DNA templates were derived by PCR amplification and carried out using Q5 High Fidelity DNA Polymerase (New England Biolabs). All primers and gBIocks Gene Fragments used in this work were obtained from Integrated DNA Technologies. Vectors created were transformed into XL-Blue competent cells (Agilent Technologies) prepared by UC Berkeley MacroLab. All plasmids used in this work were freshly prepared from 5 mL of XL-Blue cell culture using QIAprep Spin Plasmid Miniprep (QIAGEN). Molecularbiology grade, DEPC-treated water was used in all assays, transfections, and PCR reactions to ensure exclusion of DNAse activity.

[0307] In the initial step of protein purification, Escherichia coli strain BL21 Star (DE3) cells expressing the His-tagged proteins were harvested and resuspended in lysis buffer composed of 20 mM HEPES (pH 7.5), 1 M NaCI, 10% (v / v) glycerol, 10 mM imidazole, and 1 mM phenylmethylsulfonyl fluoride (PMSF). The cells were then lysed by sonication, and the resulting lysate was clarified by centrifugation at 20,000 x g for 30 minutes at 4°C to remove cellular debris and recover the supernatant. The clarified lysate was loaded onto a nickel resin column, which had been equilibrated with wash buffer for the nickel column, containing 20 mM HEPES (pH 7.5), 1 M NaCI, 10% (v / v) glycerol, 10 mM imidazole. Subsequently, the column was washed with at least five column volumes of wash buffer to remove non-specifically bound proteins. The His- tagged proteins were then eluted from the nickel column using elution buffer containing 20 mM HEPES (pH 7.5), 100 mM NaCI, 10% (v / v) glycerol, and 300 mM imidazole. Fractions were collected and analyzed by SDS-PAGE to monitor protein purity and yield. The His-tag was removed from the purified Cas9 by an overnight incubation at 4°C with Tobacco Etch Virus (TEV) protease. TEV protease cleaves the His-tag, allowing for the isolation of the target protein in its native form. After TEV protease treatment, Cas9 proteins were subjected to an addition purification step using heparin affinity chromatography. A binding buffer for the heparin column, composed of 20 mM HEPES (pH 7.5), 300 mM NaCI, and 10% (v / v) glycerol, was used to load the fractions onto the column. Non-binding proteins were washed away, and the Cas9 was eluted from the heparin column using elution buffer, which contained 20 mM HEPES (pH 7.5), 1 M NaCI, and 10% (v / v) glycerol. The eluted fractions were collected for subsequent purification steps. The final purification step involved size exclusion chromatography (SEC). SEC was performed on a Akta Purifier using a HiLoad 16 / 60 S200 superdex column for Cas9 with gel filtration buffer (20 mM HEPES (pH 7.5), 150 mM NaCI, 10% (v / v) glycerol) with a flow rate of 1 mL / min. Protein was loaded in volumes no greater than 2 mL. Purified proteins were concentrated to ~50 pM in a buffer of 20 mM HEPES (pH 7.5), 150 mM NaCI and 10% (v / v) glycerol, and stored at -80 °C.

[0308] Cas9 sgRNA. S. pyogenes Cas9 single guide RNAs (sgRNAs) were purchased with manufacturer-recommended standard chemical modifications from Synthego and resuspended in water, or from IDT(Alt-R) and resuspended in IDT duplex buffer or diethyl pyrocarbonate (DEPC)-treated water. Before use, sgRNAs were diluted to30 pM in 20 mM HEPES pH 7.5 and 150 mM NaCI, then refolded by warming to 95 °C for 5 min and slow cooling to room temperature for 25 min.

[0309] RNP formation. For electroporation and peptide-mediated delivery experiments, Cas9 protein was diluted to 20 pM in ‘RNP buffer’ (20 mM HEPES pH 7.5, 150 mM NaCI, 10% glycerol and 2 mM MgCh). The sgRNA was diluted to 30 pM in 20 mM HEPES pH 7.5, 150 mM NaCI, 10% glycerol and 2 mM MgCh. The molar ratio of Cas9:sgRNA was 1:1.5 unless otherwise specified. Cas9 was mixed with sgRNA in equal volumes yielding 10 pM RNP complex, with the RNP concentration defined by the amount of Cas9 protein.

[0310] Peptide and RNP delivery formulations in T cells. Peptides (10 mM in 100% DMSO) were diluted in DEPC-treated water to 1 mM (resulting in a solution of 90% water and 10% DMSO) and added to RNP, resulting in a volume no greater than 20% of the eventual final volume (for example, <20 pl formulation for a well with 100 pl cells in Opti-MEM). The RNP / peptide mixture was added to a 96-well round-bottom plate, and 200 x 103cells in 100 pl Opti-MEM (Gibco) per well was added to the RNP / peptide mixture. The final dose of RNP during cell treatment was 50 pmol per well with a final peptide concentration of 10 pM, unless stated otherwise. In all cases of peptide-enabled delivery, the final concentration of DMSO was proportional to the peptide concentration: 0.1 % DMSO per 10 pM peptide; this concentration of DMSO was used for the ‘mock’ negative control conditions. After a 1 h incubation at 37 °C, 100 pl of treated volume was split in half into two plates, then 150 pl culture medium was added per well, thus diluting but not removing the treatment. The concentration of additives and stimulation cocktail in the recovery medium was such that the final concentrations matched the description above for each cell type.

[0311] RNP electroporation in T cells. In a 4D nucleofector (Lonza), 20 pmol of Cas9 RNP was electroporated into 200 x 1O3T cells resuspended in 20 pl of P3 buffer and supplement (Lonza V4XP-3032) using the EH-115 pulse code. Cells were incubated for 10 min at 37 °C, then rescued with 80 pl of pre-warmed culture medium before diluting for further cell culture as above.

[0312] Flow cytometry. Flow cytometry was performed on an Attune NxT flow cytometer with a 96-well autosampler (Thermo Fisher Scientific) or on an LSRFortessa X-50 flow cytometer (BD Biosciences). Cells were resuspended in FACS buffer (phosphate- buffered saline (PBS), 2% FBS and 1 mM EDTA) and stained with live-dead stain and surface marker-targeting antibodies according to manufacturer-provided instructions.Sampling was at defined volumes (60 pl per well) to quantify cell counts. Cytometry data were processed and analyzed using FlowJo software (BD Biosciences).Example 2: Hairpin Internal Nuclear Localization Signals in CRISPR-Cas9 Enhance Editing in Primary Human Lymphocytes

[0313] This example includes some information from the above example, but also includes additional information.

[0314] Incorporation of nuclear localization signal (NLS) sequences at one or both termini of CRISPR enzymes is a widely adopted strategy to facilitate genome editing. Engineered variants of CRISPR enzymes with diverse NLS sequences have demonstrated superior performance, promoting nuclear localization and efficient DNA editing. However, limiting NLS fusion to the CRISPR protein’s termini can negatively impact protein yield via recombinant expression. Here we present a distinct strategy involving the installation of hairpin internal NLS sequences (hiNLS) at rationally selected sites within the backbone of CRI...

Claims

CLAIMSWhat is claimed is:

1. A Cas9 fusion polypeptide, comprising: (a) a Cas9 protein, and (b) one or more linker- NLS1-linker-NLS2-linker (hiNLS) modules inserted internally within the Cas9 protein, wherein NLS1 is a first nuclear localization signal (NLS) and NLS2 is a second NLS, and wherein NLS1 and NLS2 can be the same or different, wherein the one or more hiNLS modules are inserted at a position(s) that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205, K468, 11022, or S1248, of SEQ ID NO: 1 , or any combination thereof.

2. The Cas9 fusion polypeptide of claim 1 , wherein the Cas9 fusion polypeptide further comprises at least one N-terminal NLS and / or at least one C-terminal NLS.

3. The Cas9 fusion polypeptide of claim 1 , wherein the Cas9 fusion polypeptide further comprises one C-terminal NLS.

4. The Cas9 fusion polypeptide of claim 1 , wherein the Cas9 fusion polypeptide further comprises one N-terminal NLS and two C-terminal NLSs.

5. The Cas9 fusion polypeptide of any one of claims 1-4, wherein the NLS1 and / or NLS2 of at least one of the one or more hiNLS modules comprises: PKKKRKV (SEQ ID NO: 41), PAAKKKKLD (SEQ ID NO: 42), KRPAATKKAGQAKKKK (SEQ ID NO: 43), PAAKRVKLD (SEQ ID NO: 44), KRTADGSEFESPKKKRKVE (SEQ ID NO: 45), KRTADGSEFESPKKARKVE (SEQ ID NO: 46), KRTADGSEFESPKKKAKVE (SEQ ID NO: 47), or KR(X1 )5-I5KK(X2)(X3)KV (SEQ ID NO: 48), wherein X1 is any amino acid, X2 is lysine or alanine, and X3 is lysine, arginine, or alanine.

6. The Cas9 fusion polypeptide of any one of claims 1-5, wherein the NLS1 and / or the NLS2 of at least one of the one or more hiNLS modules is PKKKRKV (SEQ ID NO: 41).

7. The Cas9 fusion polypeptide of any one of claims 1-6, wherein the NLS1 and / or the NLS2 of at least one of the one or more hiNLS modules is PAAKKKKLD (SEQ ID NO: 42).

8. The Cas9 fusion polypeptide of any one of claims 1-4, wherein at least one of the one or more hiNLS modules comprises the sequenceGSGSGPKKKRKVGSGSGPKKKRKVGSGSG (SEQ ID NO: 66).

9. The Cas9 fusion polypeptide of any one of claims 1-4, wherein at least one of the one or more hiNLS modules comprises the sequence GSGSGPAAKKKKLDGSGSGPAAKKKKLDGSGSG (SEQ ID NO: 67).

10. The Cas9 fusion polypeptide of any one of claims 1-9, wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to G205 of SEQ ID NO: 1.11 . The Cas9 fusion polypeptide of any one of claims 1-10, wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to K468 of SEQ ID NO: 1.

12. The Cas9 fusion polypeptide of any one of claims 1-11 , wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to 11022 of SEQ ID NO: 1.

13. The Cas9 fusion polypeptide of any one of claims 1-12, wherein one hiNLS module is inserted at a position that is immediately adjacent and C-terminal to an amino acid residue corresponding to S1248 of SEQ ID NO: 1.

14. The Cas9 fusion polypeptide of any one of claims 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that provides transcription modulation activity, DNA-modifying activity, and / or protein modifying activity.

15. The Cas9 fusion polypeptide of any one of claims 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that provides a DNA-modifying activity selected from: nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer forming activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity and glycosylase activity.

16. The Cas9 fusion polypeptide of any one of claims 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous deaminase polypeptide.

17. The Cas9 fusion polypeptide of any one of claims 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that provides reverse transcriptase activity.

18. The Cas9 fusion polypeptide of any one of claims 1-13, wherein the Cas9 fusion polypeptide further comprises a heterologous polypeptide that exhibits histone modification activity.

19. The Cas9 fusion polypeptide of any one of claims 1-18, wherein the Cas9 protein lacks a catalytically active RuvC domain and / or lacks a catalytically active HNH domain.

20. The Cas9 fusion polypeptide of claim 19, wherein the Cas9 protein has nickase activity.21 . The Cas9 fusion polypeptide of claim 19, wherein the Cas9 protein lacks a catalytically active RuvC domain and lacks a catalytically active HNH domain.

22. The Cas9 fusion polypeptide of any one of claims 1-21 , wherein the Cas9 protein is an S. pyogenes, S. aureus, S. thermophilus, or N. meningitidis Cas9.

23. The Cas9 fusion polypeptide of any one of claims 1-21 , wherein the Cas9 protein, when considered in the absence of the inserted one or more hiNLS modules, comprises an amino acid sequence that is 80% or more identical to the S. pyogenes amino acid sequence set forth in SEQ ID NO: 1.

24. The Cas9 fusion polypeptide of any one of claims 1-21 , wherein the Cas9 protein, when considered in the absence of the inserted one or more hiNLS modules, comprises an amino acid sequence that is 80% or more identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-4.

25. The Cas9 fusion polypeptide of any one of claims 1-21 , wherein the Cas9 protein, when considered in the absence of the inserted one or more hiNLS modules, comprises an aminoacid sequence that is 80% or more identical to the amino acid sequence set forth in any one of SEQ ID NOs: 1-33.

26. The Cas9 fusion polypeptide of any one of claims 1-25, wherein the linker is any one of SEQ ID NOs: 76-81.

27. The Cas9 fusion polypeptide of any one of claims 1-25, wherein the linker is GSGSG (SEQ ID NO: 76).

28. A nucleic acid comprising a nucleotide sequence encoding the Cas9 fusion polypeptide of any one of Claims 1-27.

29. The nucleic acid of claim 28, wherein said nucleotide sequence is operably linked to a promoter.

30. The nucleic acid of claim 29, wherein the promoter is operable in a eukaryotic cell.31 . The nucleic acid of any one of claims 28-30, wherein said nucleic acid is a recombinant expression vector.

32. The nucleic acid of claim 31 , wherein the recombinant expression vector is a linear expression vector, a circular expression vector, a plasmid, or a viral expression vector.

33. The nucleic acid of claim 28, wherein said nucleic acid is an mRNA.

34. A cell, comprising the Cas9 fusion polypeptide of any one of Claims 1-27 and / or the nucleic acid of any one of claims 28-33.

35. The cell of claim 34, further comprising a Cas9 guide RNA or a nucleic acid encoding the Cas9 guide RNA.

36. The cell of claim 34 or claim 35, wherein the cell is a eukaryotic cell.

37. The cell of claim 34 or claim 35, wherein the cell is a plant cell, an invertebrate cell, a vertebrate cell, a fish cell, an amphibian cell, an avian cell, a mammalian cell, a rodent cell, a mouse cell, a rat cell, a pig cell, a non-human primate cell, a primate cell, or a human cell.

38. The cell of claim 34 or claim 35, wherein the cell is a prokaryotic cell.

39. A composition comprising:(a) the Cas9 fusion polypeptide of any one of claims 1-27 or the nucleic acid of any one of claims 28-33; and(b) a Cas9 guide RNA or a nucleic acid encoding the Cas9 guide RNA.

40. The composition of claim 39, comprising a ribonucleoprotein (RNP) complex comprising the Cas9 fusion polypeptide and the Cas9 guide RNA.41 . A method of binding, modifying, and / or modulating transcription of a target DNA, the method comprising: contacting a target DNA inside of a cell with the Cas9 fusion polypeptide of any one of Claims 1-27 and a Cas9 guide RNA.

42. The method of claim 41 , wherein said contacting causes cleavage of the target DNA, thereby modifying the target DNA.

43. The method of claim 41 or claim 42, wherein said contacting results in editing of the target DNA.

44. The method of claim 41 , wherein the Cas9 fusion polypeptide is fused to a heterologous protein that is a transcription activator, a transcription repressor, a histone modifier, a deaminase, or a reverse transcriptase.

45. The method of any one of claims 41-44, wherein the target DNA is genomic DNA.

46. The method of any one of claims 41-45, wherein the cell is a eukaryotic cell.

47. The method of claim 46, wherein the eukaryotic cell is in vivo.

48. The method of claim 46, wherein the eukaryotic cell is in vitro or ex vivo.

49. The method of any one of claims 41-48, wherein said contacting comprises introducing into the cell: (a) the Cas9 fusion polypeptide or a nucleic acid encoding the Cas9 fusion polypeptide; and (b) the Cas9 guide RNA or a nucleic acid encoding the Cas9 guide RNA.

50. The method of any one of claims 41-49, wherein the method comprises introducing a donor polynucleotide into the cell.