RNA-guided nucleic acid modifying enzyme and method for using the same
CasX proteins and their guide RNAs enhance CRISPR-Cas systems by providing programmable DNA interference, addressing the need for additional Class 2 systems and improving genome engineering capabilities.
Patent Information
- Application Number
- JP2022199477
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-09-30
- Filing Date
- 2022-12-14
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2037-09-28
AI Technical Summary
Current CRISPR-Cas technology is primarily based on systems from cultured bacteria, and there is a need for additional Class 2 CRISPR/Cas systems, particularly those involving Cas protein and guide RNA combinations, to expand the range of applications in genome engineering.
The development of RNA-guided endonuclease polypeptides, referred to as CasX proteins, along with their associated guide RNAs, which can be used in various applications, including programmable DNA interference and editing.
CasX proteins provide enhanced programmable DNA interference capabilities, expanding the versatility of CRISPR-Cas systems and enabling efficient genome editing in diverse biological contexts.
Smart Images

Figure 0007808332000006 
Figure 0007808332000007 
Figure 0007808332000008
Abstract
Description
[Background technology]
[0001] An example of a pathway unknown to science before the era of DNA sequencing, the CRISPR-Cas system is now understood to confer acquired immunity to phages and viruses in bacteria and archaea. Intensive research over the past several decades has clarified the biochemistry of this system. CRISPR-Cas systems consist of Cas proteins, responsible for acquiring, targeting, and cleaving foreign DNA or RNA, and CRISPR arrays containing direct repeats flanked by short spacer sequences that guide the Cas proteins to their targets. Class 2 CRISPR-Cas systems are the latest evolution, in which a single RNA-bound Cas protein is responsible for binding to and cleaving the target sequence. The programmable nature of these minimal systems has facilitated their use as a versatile technology that is revolutionizing the field of genome engineering.
[0002] Current CRISPR-Cas technology is based on the system from cultured bacteria, and the majority of non-isolated organisms remain undeveloped.To date, only a few class 2 CRISPR / Cas systems have been discovered.There is a need in the art for additional class 2 CRISPR / Cas systems (for example, the combination of Cas protein + guide RNA). Summary of the Invention
[0003] The present disclosure provides RNA-guided endonuclease polypeptides, herein referred to as "CasX" polypeptides (also referred to as "CasX proteins"), nucleic acids encoding CasX polypeptides, and modified host cells containing CasX polypeptides and / or nucleic acids encoding same. CasX polypeptides are useful in a variety of applications, as provided.
[0004] The present disclosure provides guide RNAs that bind to and provide sequence specificity to CasX proteins (referred to herein as "CasX guide RNAs"), nucleic acids encoding CasX guide RNAs, and modified host cells containing the CasX guide RNAs and / or nucleic acids encoding same. The CasX guide RNAs are useful in a variety of applications, as provided.
[0005] The present disclosure provides archaeal Cas9 polypeptides and nucleic acids encoding them, as well as their associated guide RNAs (archaeal Cas9 guide RNAs) and nucleic acids encoding them. [Brief explanation of the drawings]
[0006] [Figure 1-1] Two naturally occurring CasX protein sequences are shown. [Figure 1-2] Two naturally occurring CasX protein sequences are shown. [Figure 2] An alignment of two identified naturally occurring CasX protein sequences is shown. [Figure 3-1] A shows a schematic domain representation of CasX and results from various studies attempting to identify CasX homologs. It also shows portions of CRISPR loci containing CasX identified from two different species. [Figure 3-2] B shows a schematic domain representation of CasX and results from various studies attempting to identify CasX homologs. It also shows portions of CRISPR loci containing CasX identified from two different species. [Figure 4-1] A and B show experiments performed to demonstrate plasmid interference by CasX expressed in Escherichia coli. [Figure 4-2] C shows experiments performed to demonstrate plasmid interference by CasX expressed in Escherichia coli. [Figure 5-1]A shows an experiment performed to determine the PAM sequence of CasX (PAM-dependent plasmid interference by CasX). [Figure 5-2] B and C show experiments performed to determine the PAM sequence of CasX (PAM-dependent plasmid interference by CasX). [Figure 6-1] A and B show experiments performed to determine that CasX is a dual-inducible CRISPR-Cas effector complex. [Figure 6-2] C shows experiments performed to determine that CasX is a dual-inducible CRISPR-Cas effector complex. [Figure 7] A schematic diagram of CasX RNA-induced DNA interference is presented. [Figure 8] A schematic of the experimental design for one embodiment to demonstrate editing in human cells using CasX is presented. [Figure 9] Data are presented showing the recombinant expression and purification of CasX. [Figure 10] Data are presented using a variety of different tracrRNA sequences (different sequence lengths) for cleavage activity. [Figure 11] Data on CasX function at room temperature versus 37°C are presented. [Figure 12-1] A and B present information about the archaeal Cas9 CRISPR system (ARMAN-1 type II CRISPR-cas system). [Figure 12-2] C presents information about the archaeal Cas9 CRISPR system (ARMAN-1 type II CRISPR-cas system). [Figure 12-3] D and E present information about the archaeal Cas9 CRISPR system (ARMAN-1 type II CRISPR-cas system). [Figure 13] Examples of archaeal Cas9 proteins (ARMAN-1 and ARMAN-4, SEQ ID NOs: 71 and 72, respectively) are provided. The catalytic residues corresponding to D10 and H840 in S. pyogenes (S. pyogenes) are bolded and underlined. [Figure 14-1] Examples of dual-guide (top panel) (top RNA SEQ ID NO: 73, bottom RNA SEQ ID NO: 77) and single-guide (bottom panel) (SEQ ID NO: 79) formats that can be used with archaeal Cas9 proteins (e.g., ARMAN-1 Cas9) are presented. [Figure 14-2] Examples of dual-guide (top panel) (top RNA SEQ ID NO: 73, bottom RNA SEQ ID NO: 77) and single-guide (bottom panel) (SEQ ID NO: 79) formats that can be used with archaeal Cas9 proteins (e.g., ARMAN-1 Cas9) are presented. [Figure 15] Examples of dual-guide (top panel) (top RNA SEQ ID NO: 74, bottom RNA SEQ ID NO: 78) and single-guide (bottom panel) (SEQ ID NO: 80) formats that can be used with archaeal Cas9 proteins (e.g., ARMAN-4 Cas9) are presented. [Figure 16] Two newly identified non-archaeal Cas9 proteins are presented. [Figure 17-1] We present (i) an alignment of two newly identified non-archaeal Cas9 proteins with the ARMAN-1 and ARMAN-4 Cas9 proteins, and (ii) an alignment of two closely related Cas9 proteins from uncultivated bacteria with the Cas9 proteins from ARMAN-1 and ARMAN-4 and the structurally solved Actinomyces naeslundii Cas9. [Figure 17-2] We present (i) an alignment of two newly identified non-archaeal Cas9 proteins with the ARMAN-1 and ARMAN-4 Cas9 proteins, and (ii) an alignment of two closely related Cas9 proteins from uncultivated bacteria with the Cas9 proteins from ARMAN-1 and ARMAN-4 and the structurally solved Actinomyces naeslundii Cas9. [Figure 17-3]We present (i) an alignment of two newly identified non-archaeal Cas9 proteins with the ARMAN-1 and ARMAN-4 Cas9 proteins, and (ii) an alignment of two closely related Cas9 proteins from uncultivated bacteria with the Cas9 proteins from ARMAN-1 and ARMAN-4 and the structurally solved Actinomyces naeslundii Cas9. [Figure 17-4] We present (i) an alignment of two newly identified non-archaeal Cas9 proteins with the ARMAN-1 and ARMAN-4 Cas9 proteins, and (ii) an alignment of two closely related Cas9 proteins from uncultivated bacteria with the Cas9 proteins from ARMAN-1 and ARMAN-4 and the structurally solved Actinomyces naeslundii Cas9. [Figure 17-5] We present (i) an alignment of two newly identified non-archaeal Cas9 proteins with the ARMAN-1 and ARMAN-4 Cas9 proteins, and (ii) an alignment of two closely related Cas9 proteins from uncultivated bacteria with the Cas9 proteins from ARMAN-1 and ARMAN-4 and the structurally solved Actinomyces naeslundii Cas9. [Figure 17-6] We present (i) an alignment of two newly identified non-archaeal Cas9 proteins with the ARMAN-1 and ARMAN-4 Cas9 proteins, and (ii) an alignment of two closely related Cas9 proteins from uncultivated bacteria with the Cas9 proteins from ARMAN-1 and ARMAN-4 and the structurally solved Actinomyces naeslundii Cas9. [Figure 18]Panels A and B present newly identified CRISPR-Cas systems from uncultivated organisms. Panel A shows the ratio of representative and unrepresented major lineages across all bacteria and archaea, based on data from Hug et al. These results highlight the extensive yet largely unexplored biology in these domains. Archaeal Cas9 and the novel CRISPR-CasY were found exclusively in unrepresented lineages. Panel B shows the locus organization of the newly discovered CRISPR-Cas systems. [Figure 19-1] A and B present the diversity of the ARMAN-1 CRISPR array and the identification of the ARMAN-1 Cas9 PAM sequence. A shows a CRISPR array reconstructed from 15 different AMD samples. White boxes indicate repeats, and colored diamonds indicate spacers (identical spacers are similarly colored, unique spacers are black). Conserved regions of the array are highlighted (right). The diversity of recently acquired spacers (left) indicates that the system is active. An analysis that also includes CRISPR fragments from the read data is presented in Figure 25. In B, a single putative viral contig reconstructed from AMD metagenomic data contains 56 protospacers (red vertical bar) from the ARMAN-1 CRISPR array. [Figure 19-2] In C, sequence analysis revealed a conserved “NGG” PAM motif downstream of the protospacer on the non-target strand. [Figure 20-1]Panels A and B present data demonstrating that CasX mediates programmable DNA interference in E. coli. In A, a diagram of the CasX plasmid interference assay is shown. E. coli expressing a minimal CasX locus are transformed with a plasmid containing a spacer matching a sequence in the CRISPR array (target) or a plasmid containing a non-matching spacer (non-target). After transformation, cultures are plated and colony-forming units (cfu) are quantified. In B, serial dilutions of E. coli expressing the Planctomyces CasX gene coordinate targeting spacer 1 (sX.1) and transformed with identified targets (sX1, CasX spacer 1; sX2, CasX spacer 2; NT, non-target) are shown. [Figure 20-2] C and D present data showing that CasX mediates programmable DNA interference in E. coli. C shows plasmid interference by Deltaproteobacterial CasX. Experiments were performed in triplicate. Mean ± standard deviation is shown. D shows PAM depletion assay for the Planctomyces CasX locus expressed in E. coli. WebLogos were generated using PAM sequences that were more than 30-fold depleted compared to the control library. [Figure 21] Figures A–C present data demonstrating that CasX is a dual-inducible CRISPR complex. Figure A shows mapping of environmental RNA sequences (metatranscriptome data) to the CasX CRISPR locus illustrated below (red arrow, putative tracrRNA; white box, repeat sequence; green diamond, spacer sequence). The inset shows a detailed view of the first repeat and spacer. Figure B shows a diagram of CasX double-stranded DNA interference. The site of RNA treatment is indicated by a black arrow. Figure C shows the results of a plasmid interference assay (T, targeting; NT, non-targeting) using the putative tracrRNA knocked out of the CasX locus and CasX coexpressed with crRNA alone, a truncated sgRNA, or a full-length sgRNA. Experiments were performed in triplicate. Mean ± standard deviation is shown. [Figure 22-1]A presents data showing that expression of the CasY locus in E. coli is sufficient for DNA interference, and shows a diagram of the CasY locus and adjacent proteins. [Figure 22-2] B and C present data showing that expression of the CasY locus in E. coli is sufficient for DNA interference. B shows a 5′ PAM sequence that depletes CasY more than threefold relative to the control library. c shows plasmid interference by E. coli expressing CasY.1 and transformed with targets containing the indicated PAM. Experiments were performed in triplicate. Mean ± standard deviation is shown. [Figure 23] Panels A and B present the newly identified CRISPR-Cas in the context of known systems. A shows a simplified phylogenetic tree of universal Cas1 proteins. Known CRISPR-type systems are marked with wedges and branches, while newly described systems are bold lines. A detailed Cas1 phylogeny is presented in Supplementary Data 2. Panel B shows a proposed evolutionary scenario in which archaeal type II systems arose as a result of recombination between type II-B and type II-C loci. [Figure 24-1] We present that archaeal Cas9 from ARMAN-4 is found on many contigs with degenerate CRISPR arrays. Cas9 from ARMAN-4 is highlighted in dark red on 16 different contigs. Proteins with putative domains or functions are labeled, but hypothetical proteins are not. Fifteen contigs contain two degenerate direct repeats (one bp mismatch) and a single conserved spacer. The remaining contigs contain only one direct repeat. Unlike ARMAN-1, no additional Cas proteins are found adjacent to Cas9 in ARMAN-4. [Figure 24-2]We present that archaeal Cas9 from ARMAN-4 is found on many contigs with degenerate CRISPR arrays. Cas9 from ARMAN-4 is highlighted in dark red on 16 different contigs. Proteins with putative domains or functions are labeled, but hypothetical proteins are not. Fifteen contigs contain two degenerate direct repeats (one bp mismatch) and a single conserved spacer. The remaining contigs contain only one direct repeat. Unlike ARMAN-1, no additional Cas proteins are found adjacent to Cas9 in ARMAN-4. [Figure 24-3] We present that archaeal Cas9 from ARMAN-4 is found on many contigs with degenerate CRISPR arrays. Cas9 from ARMAN-4 is highlighted in dark red on 16 different contigs. Proteins with putative domains or functions are labeled, but hypothetical proteins are not. Fifteen contigs contain two degenerate direct repeats (one bp mismatch) and a single conserved spacer. The remaining contigs contain only one direct repeat. Unlike ARMAN-1, no additional Cas proteins are found adjacent to Cas9 in ARMAN-4. [Figure 24-4] We present that archaeal Cas9 from ARMAN-4 is found on many contigs with degenerate CRISPR arrays. Cas9 from ARMAN-4 is highlighted in dark red on 16 different contigs. Proteins with putative domains or functions are labeled, but hypothetical proteins are not. Fifteen contigs contain two degenerate direct repeats (one bp mismatch) and a single conserved spacer. The remaining contigs contain only one direct repeat. Unlike ARMAN-1, no additional Cas proteins are found adjacent to Cas9 in ARMAN-4. [Figure 25-1]A complete reconstruction of the ARMAN-1 CRISPR array is presented. The CRISPR array reconstruction includes the assembled reference sequence as well as array segments reconstructed from short DNA reads. Green arrows indicate repeats, and colored arrows indicate CRISPR spacers (identical spacers are colored the same color, while unique spacers are colored black). In CRISPR systems, spacers are typically added unidirectionally, so the diverse spacers on the left are recent acquisitions. [Figure 25-2] A complete reconstruction of the ARMAN-1 CRISPR array is presented. The CRISPR array reconstruction includes the assembled reference sequence as well as array segments reconstructed from short DNA reads. Green arrows indicate repeats, and colored arrows indicate CRISPR spacers (identical spacers are colored the same color, while unique spacers are colored black). In CRISPR systems, spacers are typically added unidirectionally, so the diverse spacers on the left are recent acquisitions. [Figure 25-3] A complete reconstruction of the ARMAN-1 CRISPR array is presented. The CRISPR array reconstruction includes the assembled reference sequence as well as array segments reconstructed from short DNA reads. Green arrows indicate repeats, and colored arrows indicate CRISPR spacers (identical spacers are colored the same color, while unique spacers are colored black). In CRISPR systems, spacers are typically added unidirectionally, so the diverse spacers on the left are recent acquisitions. [Figure 25-4] A complete reconstruction of the ARMAN-1 CRISPR array is presented. The CRISPR array reconstruction includes the assembled reference sequence as well as array segments reconstructed from short DNA reads. Green arrows indicate repeats, and colored arrows indicate CRISPR spacers (identical spacers are colored the same color, while unique spacers are colored black). In CRISPR systems, spacers are typically added unidirectionally, so the diverse spacers on the left are recent acquisitions. [Figure 25-5] A complete reconstruction of the ARMAN-1 CRISPR array is presented. The CRISPR array reconstruction includes the assembled reference sequence as well as array segments reconstructed from short DNA reads. Green arrows indicate repeats, and colored arrows indicate CRISPR spacers (identical spacers are colored the same color, while unique spacers are colored black). In CRISPR systems, spacers are typically added unidirectionally, so the diverse spacers on the left are recent acquisitions. [Figure 25-6] A complete reconstruction of the ARMAN-1 CRISPR array is presented. The CRISPR array reconstruction includes the assembled reference sequence as well as array segments reconstructed from short DNA reads. Green arrows indicate repeats, and colored arrows indicate CRISPR spacers (identical spacers are colored the same color, while unique spacers are colored black). In CRISPR systems, spacers are typically added unidirectionally, so the diverse spacers on the left are recent acquisitions. [Figure 26] A and B show that ARMAN-1 spacers map to the genomes of archaeal community members. In A, protospacers from ARMAN-1 (red arrows) map to the genome of ARMAN-2, a nanoarchaeon from the same environment. Six protospacers map uniquely to a portion of the genome flanked by two long terminal repeats (LTRs), and two additional protospacers match perfectly within the LTRs (blue and green). This region may be a transposon, suggesting that the CRISPR-Cas system of ARMAN-1 plays a role in suppressing mobilization of this element. In B, the protospacers also map to a Thermoplasmales archaeon (I-plasma), another member of the Richmond Mine ecosystem found in the same sample as the ARMAN organism. The protospacers cluster within a region of the genome that encodes a short hypothetical protein, suggesting that this may also represent a mobile element. [Figure 27-1]A presents the predicted secondary structures of ARMAN-1 crRNA and tracrRNA. In A, CRISPR repeats and tracrRNA anti-repeats are shown in black, while spacer-derived sequences are shown as a series of green Ns. Because no clear termination signal can be predicted from the locus, three different tracrRNA lengths were tested based on their secondary structures: 69, 104, and 179 (red, blue, and pink, respectively). [Figure 27-2] B presents the predicted secondary structures of ARMAN-1 crRNA and tracrRNA. B shows the engineered single guide RNA corresponding to the dual guide in A. [Figure 27-3] C-E show the predicted secondary structures of ARMAN-1 crRNA and tracrRNA. C shows the ARMAN-4 Cas9 dual guide with two different hairpins (75 and 122) at the 3' end of tracrRNA. D shows the engineered single guide RNA corresponding to the dual guide in C. E shows the conditions tested in the E. coli in vivo targeting assay. [Figure 28] Panels A and B present a purification outline for in vitro biochemical studies. In A, ARMAN-1 (AR1) and ARMAN-4 (AR4) Cas9 were expressed and purified under various conditions outlined in the supplemental material. Proteins boxed in blue were tested for in vitro cleavage activity. In B, fractions of AR1-Cas9 and AR4-Cas9 purification were separated on a 10% SDS-PAGE gel. [Figure 29] We present newly identified CRISPR-Cas systems compared to known proteins. We show the similarity of CasX and CasY to known proteins based on the following searches: (1) a Blast search against the NCBI nonredundant (NR) protein database, (2) a hidden Markov model (HMM) search against the HMM database of all known proteins, and (3) a distant homology search using HHpred30. [Figure 30-1]Panels A and B present data on programmed DNA interference by CasX. Panel A, continuing from Panel C of Figure 20, shows a plasmid interference assay (sX1, CasX spacer 1; sX2, CasX spacer 2; NT, non-target) for CasX2 (Planctomyces) and CasX1 (Deltaproteobacteria). Experiments were performed in triplicate. Mean ± standard deviation is shown. Panel B, continuing from Panel B of Figure 20, shows serial dilutions of E. coli expressing the CasX locus and transformed with identified targets. [Figure 30-2] (C) Data on programmed DNA interference by CasX. (C) PAM depletion assay for deltaproteobacterial CasX. WebLogos were generated using PAM sequences that were depleted above the indicated PAM depletion threshold (PDVT) compared to a control library. [Figure 30-3] (D) Data on programmed DNA interference by CasX are presented. (D) PAM depletion assay for Planctomyces CasX expressed in E. coli. PAM sequences that were depleted above the indicated PAM depletion threshold (PDVT) compared to a control library were used to generate WebLogos. [Figure 30-4] E and F present data on programmed DNA interference by CasX. E is a diagram showing the location of the Northern blot probe for CasX.1, and F shows a Northern blot for CasX.1 tracrRNA in total RNA extracted from E. coli expressing the CasX.1 locus. [Figure 31] Figure 1 shows an evolutionary tree of Cas9 homologs. Maximum likelihood phylogenetic tree of the Cas9 protein, showing previously described systems colored based on their type (II-A blue, II-B green, and II-C purple). Archaeal Cas9 clusters with type II-C CRISPR-Cas systems, along with two newly described bacterial Cas9s from uncultured bacteria. [Figure 32] A table of cleavage conditions assayed for Cas9 from ARMAN-1 and ARMAN-4 is presented. DETAILED DESCRIPTION OF THE INVENTION
[0007] As used herein, "heterologous" refers to a nucleotide or polypeptide sequence not found in a naturally occurring nucleic acid or protein, respectively. For example, with respect to a CasX polypeptide, a heterologous polypeptide contains an amino acid sequence from a protein other than the CasX polypeptide. In some cases, a portion of a CasX protein from one species is fused to a portion of a CasX protein from a different species. CasX sequences from each species can therefore be considered heterologous to each other. As another example, a CasX protein (e.g., a dCasX protein) can be fused to an activity domain from a non-CasX protein (e.g., a histone deacetylase), and the sequence of the activity domain can be considered a heterologous polypeptide (it is heterologous to the CasX protein).
[0008] The terms "polynucleotide" and "nucleic acid," used interchangeably herein, refer to polymers of nucleotides of any length, either ribonucleotides or deoxynucleotides. Thus, the terms include, but are not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The terms "polynucleotide" and "nucleic acid" should be understood to include single-stranded (such as sense or antisense) and double-stranded polynucleotides, as applicable to the embodiments being described.
[0009] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymeric forms of amino acids of any length, which may include genetically encoded and non-genetically encoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones. This term also encompasses fusion proteins, including, but not limited to, fusion proteins with heterologous amino acid sequences, fusions with heterologous and homologous leader sequences, with or without an N-terminal methionine residue; immunologically tagged proteins; and the like.
[0010] As used herein, the term "naturally-occurring" when applied to a nucleic acid, protein, cell, or organism refers to a nucleic acid, cell, protein, or organism that is found in nature.
[0011] As used herein, the term "isolated" is meant to describe a polynucleotide, polypeptide, or cell that is in an environment that is different from the environment that the polynucleotide, polypeptide, or cell naturally occurs in. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.
[0012] As used herein, the term "exogenous nucleic acid" refers to a nucleic acid that is not normally or naturally found in and / or produced by a given bacterium, organism, or cell in nature. As used herein, the term "endogenous nucleic acid" refers to a nucleic acid that is not normally found in and / or produced by a given bacterium, organism, or cell in nature. "Endogenous nucleic acid" is also referred to as "natural nucleic acid" or a nucleic acid that is "native" to a given bacterium, organism, or cell.
[0013] As used herein, "recombinant" refers to the product of various combinations of cloning, restriction, and / or ligation steps that result in constructs in which a particular nucleic acid (DNA or RNA) has structural coding or non-coding sequences distinguishable from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers or from a series of synthetic oligonucleotides to provide synthetic nucleic acids expressible from recombinant transcription units contained in cells or cell-free transcription and translation systems. Such sequences can be provided in the form of open reading frames uninterrupted by internal non-translated sequences, or introns, typically present in eukaryotic genes. Genomic DNA containing relevant sequences can also be used in forming recombinant genes or transcription units. Sequences of non-translated DNA can be present 5' or 3' from the open reading frame, such that such sequences do not interfere with the manipulation or expression of the coding region and may actually act to modulate the production of a desired product by various mechanisms (see "DNA Regulatory Sequences," below).
[0014] Thus, for example, the terms "recombinant" polynucleotide or "recombinant" nucleic acid refer to those that do not occur in nature, e.g., are created by the artificial combination of two otherwise separated segments of sequence through human intervention. This artificial combination is often achieved either by chemical synthesis or by artificial manipulation of isolated segments of nucleic acid, e.g., by genetic engineering techniques. This is usually done to replace a codon with a redundant codon or a conserved amino acid that encodes it, typically while introducing or removing a sequence recognition site. Alternatively, it is done to join nucleic acid segments of desired functions together to generate a desired combination of functions. This artificial combination is often achieved either by chemical synthesis or by artificial manipulation of isolated segments of nucleic acid, e.g., by genetic engineering techniques.
[0015] Similarly, the term "recombinant" polypeptide refers to a polypeptide that does not occur in nature, e.g., that is made by the artificial combination of two otherwise separate segments of amino acid sequence, e.g., through human intervention. Thus, for example, a polypeptide comprising a heterologous amino acid sequence is recombinant.
[0016] By "construct" or "vector" is meant a recombinant nucleic acid, generally recombinant DNA, generated for the purpose of expressing and / or propagating a specific nucleotide sequence(s) or used in the construction of other recombinant nucleotide sequences.
[0017] The terms "DNA regulatory sequence," "control element," and "regulatory element," used interchangeably herein, refer to transcriptional and translational control sequences, such as promoters, enhancers, polyadenylation signals, terminators, proteolysis signals, and the like, that provide and / or regulate the expression of a coding sequence and / or the production of the encoded polypeptide in a host cell.
[0018] The term "transformation" is used interchangeably herein with "genetic modification" and refers to a permanent or transient genetic change induced in a cell after introducing new nucleic acid (DNA exogenous to the cell) into the cell. The genetic change ("modification") can be achieved either by integration of the new nucleic acid into the genome of the host cell or by transient or stable maintenance of the new nucleic acid as an episomal element. If the cell is eukaryotic, a permanent genetic change is generally achieved by introduction of the new DNA into the genome of the cell. In prokaryotic cells, permanent changes can be introduced into the chromosome or via extrachromosomal elements such as plasmids and expression vectors, which may contain one or more selectable markers to aid in their maintenance in the recombinant host cell. Suitable methods of genetic modification include viral infection, transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, and the like. The choice of method generally depends on the type of cell being transformed and the context in which the transformation is occurring (i.e., in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.
[0019] "Operably linked" refers to a juxtaposition, where the components so described are in a relationship permitting them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression. As used herein, the terms "heterologous promoter" and "heterologous regulatory region" refer to promoters and other regulatory regions that are not normally associated with a particular nucleic acid in nature. For example, a "transcriptional regulatory region" heterologous to a coding region is a transcriptional regulatory region that is not normally associated with the coding sequence in nature.
[0020] As used herein, "host cell" refers to a eukaryotic cell, a prokaryotic cell, or a cell from a multicellular organism cultured as a unicellular entity (e.g., a cell line) in vivo or in vitro; the eukaryotic or prokaryotic cell can be or has been used as a recipient of nucleic acid (e.g., an expression vector), and includes the progeny of the original cell that has been genetically modified with the nucleic acid. It is understood that the progeny of a single cell may not necessarily be completely identical in morphology or in genomic or total DNA complement to the original parent due to natural, inadvertent, or deliberate mutation. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, e.g., an expression vector, has been introduced. For example, a prokaryotic host cell of interest is a prokaryotic host cell (e.g., a bacterium) that has been genetically modified by the introduction into a suitable prokaryotic host cell of heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the prokaryotic host cell (not normally naturally found in the prokaryotic host cell), or a recombinant nucleic acid that is not normally found in the prokaryotic host cell; and a eukaryotic host cell of interest is a eukaryotic host cell that has been genetically modified by the introduction into a suitable eukaryotic host cell of heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the eukaryotic host cell, or a recombinant nucleic acid that is not normally found in the eukaryotic host cell.
[0021] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues in proteins that have similar side chains. For example, the group of amino acids with aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic-hydroxyl side chains consists of serine and threonine; the group of amino acids with amide-containing side chains consists of asparagine and glutamine; the group of amino acids with aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains consists of lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0022] A polynucleotide or polypeptide has a certain percentage of "sequence identity" with another polynucleotide or polypeptide, meaning that, when aligned, the percentage of bases or amino acids are identical and in the same relative positions when the two sequences are compared. Sequence similarity can be determined in several different ways. To determine sequence identity, sequences can be aligned using methods and computer programs, including BLAST, available at ncbi.nlm.nih.gov / BLAST. See, e.g., Altschul et al. (1990), J. Mol. Biol. 215:403-10. Another alignment algorithm is FASTA, available in the Genetics Computing Group (GCG) package (Madison, Wisconsin, USA), a wholly owned subsidiary of Oxford Molecular Group, Inc. Other techniques for alignment are described in Methods in Enzymology, vol. 266: Computer Methods for Macromolecular Sequence Analysis (1996), ed. Doolittle, Academic Press, Inc., a division of Harcourt Brace & Co., San Diego, California, USA. Alignment programs that allow gaps in sequences are of particular interest. Smith-Waterman is one type of algorithm that allows gaps in sequence alignment. See Meth. Mol. Biol. 70: 173-187 (1997). Sequences can also be aligned using the GAP program, which uses the Needleman-Wunsch alignment method. See J. Mol. Biol. 48: 443-453 (1970).
[0023] As used herein, the terms "treatment," "treating," and the like refer to obtaining a desired pharmacological and / or physiological effect. The effect may be prophylactic, in terms of completely or partially preventing the disease or condition, and / or therapeutic, in terms of partially or completely curing the disease and / or side effects caused by the disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal, e.g., a human, including (a) preventing the onset of the disease in a subject who is susceptible to, but has not yet been diagnosed with, the disease; (b) inhibiting the disease, i.e., preventing its development; and (c) relieving the disease, i.e., causing regression of the disease.
[0024] The terms "individual," "subject," "host," and "patient," used interchangeably herein, refer to individual organisms, e.g., mammals, including, but not limited to, mice, monkeys, humans, mammalian livestock animals, mammalian sport animals, and mammalian pets.
[0025] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0026] Where a range of values is provided, it is understood that each intervening value between the upper and lower limit of that range, to the tenth of the unit of the lower limit, and any other stated or intervening value in that stated range, is encompassed within the invention, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specific excluded limits in the stated range. Where the stated range includes one or both of its limits, ranges excluding either or both of those included limits are also encompassed within the invention.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0028] It should be noted that, as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to a "CasX polypeptide" includes a plurality of such polypeptides; a reference to a "guide RNA" includes a reference to one or more guide RNAs and equivalents thereof known to those skilled in the art; and so forth. It is further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a guideline precedent for use of exclusive terminology, such as "solely," "only," or "negative" limitations with regard to the recitation of claim elements.
[0029] It is understood that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of the embodiments related to the present invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination were individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such subcombination were individually and explicitly disclosed herein.
[0030] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein should be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0031] The present disclosure provides RNA-guided endonuclease polypeptides, herein referred to as "CasX" polypeptides (also referred to as "CasX proteins"), nucleic acids encoding CasX polypeptides, and modified host cells containing CasX polypeptides and / or nucleic acids encoding same. CasX polypeptides are useful in a variety of applications, as provided.
[0032] The present disclosure provides guide RNAs that bind to and provide sequence specificity to CasX proteins (referred to herein as "CasX guide RNAs"), nucleic acids encoding CasX guide RNAs, and modified host cells containing the CasX guide RNAs and / or nucleic acids encoding same. The CasX guide RNAs are useful in a variety of applications, as provided.
[0033] The present disclosure provides archaeal Cas9 polypeptides and nucleic acids encoding them, as well as their associated guide RNAs (archaeal Cas9 guide RNAs) and nucleic acids encoding them.
[0034] composition CRISPR / CASX proteins and guide RNAs A CRISPR / Cas endonuclease (e.g., a CasX protein) interacts with (binds to) a corresponding guide RNA (e.g., a CasX guide RNA) to form a ribonucleoprotein (RNP) complex that is targeted to a specific site in a target nucleic acid through base pairing between the guide RNA and the target sequence within the target nucleic acid molecule. The guide RNA contains a nucleotide sequence (guide sequence) that is complementary to the sequence of the target nucleic acid (target site). Thus, the CasX protein forms a complex with the CasX guide RNA, and the guide RNA provides sequence specificity to the RNP complex via the guide sequence. The CasX protein of the complex provides site-specific activity. In other words, the CasX protein is guided to (e.g., stabilized at) the target site within a target nucleic acid sequence (e.g., a chromosomal or extrachromosomal sequence, e.g., an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by its association with the guide RNA.
[0035] The present disclosure provides compositions comprising a CasX polypeptide (and / or a nucleic acid encoding a CasX polypeptide) (e.g., the CasX polypeptide can be a naturally occurring protein, a nickase CasX protein, a dCasX protein, a chimeric CasX protein, etc.). The present disclosure provides compositions comprising a CasX guide RNA (and / or a nucleic acid encoding a CasX guide RNA) (e.g., the CasX guide RNA can be in a dual- or single-guide format). The present disclosure provides compositions comprising (a) a CasX polypeptide (and / or a nucleic acid encoding a CasX polypeptide) (e.g., the CasX polypeptide can be a naturally occurring protein, a nickase CasX protein, a dCasX protein, a chimeric CasX protein, etc.) and (b) a CasX guide RNA (and / or a nucleic acid encoding a CasX guide RNA) (e.g., the CasX guide RNA can be in a dual- or single-guide format). The present disclosure provides nucleic acid / protein complexes (RNP complexes) comprising: (a) a CasX polypeptide of the present disclosure (e.g., the CasX polypeptide can be a naturally occurring protein, a nickase CasX protein, a dCasX protein, a chimeric CasX protein, etc.), and (b) a CasX guide RNA (e.g., the CasX guide RNA can be in a dual- or single-guide format).
[0036] CasX protein A CasX polypeptide (this term is used interchangeably with the term "CasX protein") can bind to and / or modify (e.g., cleave, nick, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with the target nucleic acid (e.g., methylate or acetylate a histone tail) (e.g., in some cases, the CasX protein includes an active fusion partner, and in other cases, the CasX protein provides nuclease activity). In some cases, the CasX protein is a naturally occurring protein (e.g., naturally occurring in a prokaryotic cell). In other cases, the CasX protein is not a naturally occurring polypeptide (e.g., the CasX protein is a mutant CasX protein, a chimeric protein, etc.).
[0037] The assay to determine whether a given protein interacts with a CasX guide RNA can be any convenient binding assay that tests for binding between a protein and a nucleic acid. Suitable binding assays (e.g., gel shift assays) will be known to those of skill in the art (e.g., assays involving adding a CasX guide RNA and a protein to a target nucleic acid). The assay to determine whether a protein has activity (e.g., determining whether the protein has nuclease activity and / or some heterologous activity that cleaves a target nucleic acid) can be any convenient assay (e.g., any convenient nucleic acid cleavage assay that tests for nucleic acid cleavage). Suitable assays (e.g., cleavage assays) will be known to those of skill in the art.
[0038] Naturally occurring CasX proteins function as endonucleases that catalyze double-strand cleavage at specific sequences in target double-stranded DNA (dsDNA). Sequence specificity is provided by associated guide RNAs that hybridize to target sequences within the target DNA. Naturally occurring guide RNAs include tracrRNAs hybridized to crRNAs, and the crRNAs contain guide sequences that hybridize to target sequences in the target DNA.
[0039] In some embodiments, the CasX protein of the subject methods and / or compositions is (or is derived from) a naturally occurring (wild-type) protein. Examples of naturally occurring CasX proteins are shown in FIG. 1 and set forth as SEQ ID NOS: 1-2. Examples of naturally occurring CasX proteins are shown in FIG. 1 and set forth as SEQ ID NOS: 1-3. An alignment of two naturally occurring CasX proteins is presented in FIG. 2 ("gwa2" is CasX1 and "gwc2" is CasX2). Partial DNA scaffolds of CRISPR loci assembled from sequencing data (from Deltaproteobacteria (gwa2 scaffold) and Planctomycetes (gwc2 scaffold)) are set forth as SEQ ID NOS: 51 and 52, respectively. It is important to note that this newly discovered protein (CasX) is short compared to previously identified CRISPR-Cas endonucleases, and thus its use as an alternative offers the advantage of relatively short nucleotide sequences encoding the protein. This is useful, for example, in situations where a nucleic acid encoding a CasX protein is desired, e.g., using a viral vector (e.g., an AAV vector) to deliver the nucleic acid to cells, such as eukaryotic cells (e.g., mammalian cells, human cells, mouse cells, in vitro, ex vivo, or in vivo) for research and / or clinical applications. It is also noted herein that bacteria harboring a CasX CRISPR locus were present in environmental samples collected at low temperatures (e.g., 10-17°C). Therefore, CasX is expected to function well (e.g., better than other Cas endonucleases discovered to date) at low temperatures (e.g., 10-14°C, 10-17°C, 10-20°C).
[0040] In some cases, a CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or greater sequence identity (e.g., 30% or greater, 40% or greater, 50% or greater, 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 1. For example, in some cases, a CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 1. In some cases, the CasX protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 1. In some cases, the CasX protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 1. In some cases, the CasX protein comprises an amino acid sequence having the CasX protein sequence set forth as SEQ ID NO: 1, except that the sequence contains amino acid substitutions (e.g., one, two, or three amino acid substitutions) that reduce the naturally occurring catalytic ability of the protein (e.g., at the amino acid positions described below).
[0041] In some cases, the CasX protein comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 2. In some cases, the CasX protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 2. In some cases, the CasX protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 2. In some cases, the CasX protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 2. In some cases, the CasX protein comprises an amino acid sequence having the CasX protein sequence set forth as SEQ ID NO: 2, except that the sequence contains amino acid substitutions (e.g., one, two, or three amino acid substitutions) that reduce the naturally occurring catalytic ability of the protein (e.g., at the amino acid positions described below).
[0042] In some cases, the CasX protein comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 3. In some cases, the CasX protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 3. In some cases, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 3. In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to the CasX protein sequence set forth as SEQ ID NO: 3. In some cases, the CasX protein comprises an amino acid sequence having the CasX protein sequence set forth as SEQ ID NO: 3, except that the sequence contains amino acid substitutions (e.g., one, two, or three amino acid substitutions) that reduce the naturally occurring catalytic ability of the protein (e.g., at the amino acid positions described below).
[0043] In some cases, the CasX protein comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2. In some cases, the CasX protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2. In some cases, the CasX protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2. In some cases, the CasX protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2. In some cases, the CasX protein comprises an amino acid sequence having the CasX protein sequence set forth in one of SEQ ID NOs: 1 and 2. The CasX protein comprises an amino acid sequence having a CasX protein sequence set forth in any one of SEQ ID NOs: 1 and 2, except that in some cases the sequence contains amino acid substitutions (e.g., 1, 2, or 3 amino acid substitutions) that reduce the naturally occurring catalytic ability of the protein (e.g., at the amino acid positions set forth below).
[0044] In some cases, the CasX protein comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3. In some cases, the CasX protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3. In some cases, the CasX protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3. In some cases, the CasX protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3. In some cases, the CasX protein comprises an amino acid sequence having a CasX protein sequence set forth as one of SEQ ID NOs: 1-3. The CasX protein comprises an amino acid sequence having a CasX protein sequence set forth in any one of SEQ ID NOs: 1-3, except that in some cases the sequence contains amino acid substitutions (e.g., 1, 2, or 3 amino acid substitutions) that reduce the naturally occurring catalytic ability of the protein (e.g., at the amino acid positions set forth below).
[0045] CasX protein domain The domains of the CasX protein are shown in Figure 3. As seen in the schematic diagram in Figure 3 (amino acids are numbered based on the CasX1 protein (SEQ ID NO: 1)), the CasX protein comprises an N-terminal domain that is approximately 650 amino acids long (e.g., 663 for CasX1 and 650 for CasX2) and a C-terminal domain that is not contiguous with respect to the primary amino acid sequence of the CasX protein but contains three partial RuvC domains (also referred to herein as subdomains, RuvC-I, RuvC-II, and RuvC-III) that form the RuvC domain once the protein is produced and folded. Thus, in some cases, a CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having an N-terminal domain (e.g., excluding any fused heterologous sequences, such as an NLS and / or catalytic domain) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). In some cases, the CasX protein (of the subject compositions and / or methods) comprises 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-70 ... and / or 650-670 amino acids).
[0046] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). For example, in some instances, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0047] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO:2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). For example, in some instances, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises the amino acid sequence of SEQ ID NO: 2, which corresponds to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0048] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). For example, in some instances, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises the amino acid sequence of SEQ ID NO: 3, which corresponds to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0049] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). For example, in some cases, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain set forth as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain set forth as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence corresponding to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0050] In some cases, a CasX protein (of the subject compositions and / or methods) comprises a first amino acid sequence having 20% or greater sequence identity (e.g., 30% or greater, 40% or greater, 50% or greater, 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III). For example, in some cases, a CasX protein comprises a first amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III). In some cases, the CasX protein comprises a first amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III).In some cases, the CasX protein comprises a first amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III). In some cases, the CasX protein comprises an amino acid sequence corresponding to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO:1 (e.g., amino acids 1-650 of the CasX protein sequence set forth as SEQ ID NO:2) and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III).
[0051] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). For example, in some cases, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain set forth as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain set forth as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence corresponding to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0052] In some cases, a CasX protein (of the subject compositions and / or methods) comprises a first amino acid sequence having 20% or greater sequence identity (e.g., 30% or greater, 40% or greater, 50% or greater, 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III). For example, in some cases, a CasX protein comprises a first amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III). In some cases, the CasX protein comprises a first amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III).In some cases, the CasX protein comprises a first amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III). In some cases, the CasX protein comprises an amino acid sequence corresponding to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO:1 (e.g., amino acids 1-650 of the CasX protein sequence set forth as SEQ ID NO:2) and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising a split Ruv C domain (e.g., three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III).
[0053] In some embodiments, the split RuvC domain of a CasX protein (of the subject compositions and / or methods) comprises a region between RuvC-II and RuvC-III that is longer than the RuvC-III subdomain. For example, in some cases, the ratio of the length of the region between the RuvC-II and RuvC-III subdomains to the length of the RuvC-III subdomain is 1.1 or greater (e.g., 1.2). In some cases, the ratio of the length of the region between the RuvC-II and RuvC-III subdomains to the length of the RuvC-III subdomain is greater than 1. In some cases, the ratio of the length of the region between the RuvC-II and RuvC-III subdomains to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2).
[0054] In some embodiments (for a CasX protein of the subject compositions and / or methods), the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, or 1.4 or less). For example, in some cases, the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less). In some embodiments, the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is in the range of 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4).
[0055] In some cases (for a CasX protein of the subject compositions and / or methods), the ratio of the length of the region between the RuvC-II and RuvC-III subdomains to the length of the RuvC-III subdomain is greater than 1. In some cases, the ratio of the length of the region between the RuvC-II and RuvC-III subdomains to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2).
[0056] In some cases (for CasX proteins of the subject compositions and / or methods), the region between the RuvC-II and RuvC-III subdomains is at least 73 amino acids in length (e.g., at least 75, 77, 80, 85, or 87 amino acids in length). For example, in some cases, the region between the RuvC-II and RuvC-III subdomains is at least 78 amino acids in length (e.g., at least 80, 85, or 87 amino acids in length). In some cases, the region between the RuvC-II and RuvC-III subdomains is at least 85 amino acids in length. In some cases, the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-90, 83-88, 85-100, 85-95, 85-90, or 85-88 amino acids). In some cases, the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80-95 amino acids (e.g., 80-90, 80-88, 83-95, 83-90, 83-88, 85-95, 85-90, or 85-88 amino acids).
[0057] In some cases, the CasX protein (of the subject compositions and / or methods) shares 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% or more) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III, wherein (i) the ratio of the length of the RuvC-II subdomain and the region between the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), or (ii) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.2 or more (e.g., 1.2). (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and 1 to 1.3 (e.g., 1 to 1.2); or (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less). , 1.5 or less, or 1.4 or less), (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less), or (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4), (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2), (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids long (e.g., at least 75, 77, 80, 85, or 87 amino acids long), (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long). (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids long, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-95, 83-10 ... or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 80 to 88, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0058] For example, in some cases, the CasX protein comprises a first amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1 to 663 of CasX1 in Figure 3, panel A), and a third amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity). and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III, wherein (i) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, etc.). (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less); (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4), (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2), (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids long (e.g., at least 75, 77, 80, 85, or 87 amino acids long), (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long). (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids long, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-95, 83-10 ... or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 80 to 88, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0059] In some cases, the CasX protein comprises a first amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (e.g., the first amino acid sequence having 80% or greater sequence identity ...). (i) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, or (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3. (e.g., 1 to 1.2), (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, or 1.4 or less), (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less), or (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1 (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; or (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids in length (e.g., at least 75, 77, 80, 85, or 87 amino acids in length), or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids in length (e.g., at least 80, 85, or 87 amino acids in length), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids in length, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 75 amino acids in length. or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 75 to 88, 78 to 100, 78 to 95, 78 to 90, 78 to 88, 80 to 100, 80 to 95, 80 to 90, 80 to 88, 83 to 100, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0060] In some cases, the CasX protein comprises a first amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III. (i) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, or (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, or 1.4 or less); (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less); (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; or (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids in length (e.g., at least 75, 77, 80, 85, or 87 amino acids in length), or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids in length (e.g., at least 80, 85, or 87 amino acids in length), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids in length, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 75 amino acids in length. or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 75 to 88, 78 to 100, 78 to 95, 78 to 90, 78 to 88, 80 to 100, 80 to 95, 80 to 90, 80 to 88, 83 to 100, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0061] In some cases, the CasX protein (of the subject compositions and / or methods) shares 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% or more) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A). and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III, wherein (i) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), or (ii) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.2 or more (e.g., 1.2). (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and 1 to 1.3 (e.g., 1 to 1.2); or (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less). , 1.5 or less, or 1.4 or less), (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less), or (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4), (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2), (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids long (e.g., at least 75, 77, 80, 85, or 87 amino acids long), (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long). (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids long, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-95, 83-10 ... or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 80 to 88, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0062] For example, in some cases, the CasX protein comprises a first amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a third amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity). and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III, wherein (i) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, etc.). (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less); (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4), (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2), (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids long (e.g., at least 75, 77, 80, 85, or 87 amino acids long), (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long). (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids long, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-95, 83-10 ... or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 80 to 88, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0063] In some cases, the CasX protein comprises a first amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) that includes three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III. (i) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, or (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3. (e.g., 1 to 1.2), (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, or 1.4 or less), (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less), or (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1 (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; or (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids in length (e.g., at least 75, 77, 80, 85, or 87 amino acids in length), or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids in length (e.g., at least 80, 85, or 87 amino acids in length), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids in length, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 75 amino acids in length. or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 75 to 88, 78 to 100, 78 to 95, 78 to 90, 78 to 88, 80 to 100, 80 to 95, 80 to 90, 80 to 88, 83 to 100, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0064] In some cases, the CasX protein comprises a first amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the N-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 1-663 of CasX1 in Figure 3, panel A), and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III. (i) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, or (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, or 1.4 or less); (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less); (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; or (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids in length (e.g., at least 75, 77, 80, 85, or 87 amino acids in length), or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids in length (e.g., at least 80, 85, or 87 amino acids in length), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids in length, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 75 amino acids in length. or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 75 to 88, 78 to 100, 78 to 95, 78 to 90, 78 to 88, 80 to 100, 80 to 95, 80 to 90, 80 to 88, 83 to 100, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0065] In some instances, the CasX protein comprises an amino acid sequence corresponding to amino acids 1-663 of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., amino acids 1-650 of the CasX protein sequence set forth as SEQ ID NO: 2) and a second amino acid sequence (C-terminal to the first amino acid sequence) comprising three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III, wherein (i) the length of the RuvC-II subdomain and the RuvC-III subdomain are proportional to the length of the RuvC-III subdomain. (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (for example, between 1 and 1.2); or (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (for example, between 1 and 1.2). (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.8 or less, 1.7 or less, 1.6 or less, 1.5 or less, or 1.4 or less); (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; or (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., 1 to 1.2), or (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids in length (e.g., at least 75, 77, 80, 85, or 87 amino acids in length), or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids in length (e.g., at least 80, 85, or 87 amino acids in length), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids in length, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 75 amino acids in length. or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 75 to 88, 78 to 100, 78 to 95, 78 to 90, 78 to 88, 80 to 100, 80 to 95, 80 to 90, 80 to 88, 83 to 100, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0066] In some cases, the CasX protein (of the subject compositions and / or methods) has an N-terminal domain having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). and a second amino acid sequence (C-terminal to the first) having a split RuvC domain with three partial RuvC domains, RuvC-I, RuvC-II, and RuvC-III, wherein (i) the ratio of the length of the region between the RuvC-II and RuvC-III subdomains to the length of the RuvC-III subdomain is 1.1 or greater (e.g., 1.2). (ii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1; (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and 1 to 1.3 (e.g., 1 to 1.2); or (iv) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or less). (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less); or (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4), (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2), (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids long (e.g., at least 75, 77, 80, 85, or 87 amino acids long), (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long). (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids long, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-95, 83-10 ... or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 80 to 88, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0067] In some cases, the CasX protein (of the subject compositions and / or methods) comprises 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650 and (ii) a ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2), or (iii) a ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 1.1 or more (e.g., 1.2). (iii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and 1 to 1.3 (e.g., 1 to 1.2); or (iv) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is 2 or less (e.g., 1.8 or less, 1.7 or more). (v) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1.5 or less (e.g., 1.4 or less); or (vi) the ratio of the length of the RuvC-II subdomain to the length of the RuvC-III subdomain is 1 to 2 (e.g., 1.1 to 2, 1.2 to 2, 1 to 1.8, 1.1 to 1.8, 1.2 to 1.8, 1 to 1.6, 1.1 to 1.6, 1.2 to 1.6, 1 to 14, 1.1 to 1.4, or 1.2 to 1.4), (vii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1, (viii) the ratio of the length of the region between the RuvC-II subdomain and the RuvC-III subdomain to the length of the RuvC-III subdomain is greater than 1 and between 1 and 1.3 (e.g., between 1 and 1.2), (ix) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 73 amino acids long (e.g., at least 75, 77, 80, 85, or 87 amino acids long), (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long), or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 78 amino acids long (e.g., at least 80, 85, or 87 amino acids long). (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is at least 85 amino acids long, or (x) the region between the RuvC-II subdomain and the RuvC-III subdomain is in the range of 75-100 amino acids (e.g., 75-95, 75-90, 75-88, 78-100, 78-95, 78-90, 78-88, 80-100, 80-95, 80-90, 80-88, 83-100, 83-95, 83-95, 83-10 ... or (xi) the region between the RuvC-II subdomain and the RuvC-III subdomain has a length in the range of 80 to 95 amino acids (e.g., 80 to 90, 80 to 88, 83 to 95, 83 to 90, 83 to 88, 85 to 100, 85 to 95, 85 to 90, or 85 to 88 amino acids).
[0068] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). For example, in some instances, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having amino acids 664-986 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0069] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO:2 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). For example, in some instances, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 2 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 2 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO:2 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises the amino acid sequence of SEQ ID NO:2, which corresponds to amino acids 664-986 of the CasX protein sequence set forth as SEQ ID NO:1.
[0070] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). For example, in some instances, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some instances, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of the CasX protein sequence set forth as SEQ ID NO: 3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises the amino acid sequence of SEQ ID NO: 3, which corresponds to amino acids 664-986 of the CasX protein sequence set forth as SEQ ID NO: 1.
[0071] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence (e.g., amino acids 651-978 of the CasX protein sequence set forth as SEQ ID NO: 2) having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). For example, in some cases, the CasX protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence corresponding to amino acids 664-986 of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., amino acids 651-978 of the CasX protein sequence set forth as SEQ ID NO: 2).
[0072] In some cases, the CasX protein (of the subject compositions and / or methods) comprises a first amino acid sequence (N-terminal domain) (e.g., excluding any fused heterologous sequences, such as an NLS and / or catalytic domain) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). ) and a second amino acid sequence positioned C-terminal to the first amino acid sequence (e.g., amino acids 651-978 of the CasX protein sequence set forth as SEQ ID NO: 2) that has 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A).For example, in some cases, the CasX protein comprises a first amino acid sequence (N-terminal domain) (e.g., an NLS and / or a catalytically active domain) having a length in the range of 500 to 750 amino acids (e.g., 550 to 750, 600 to 750, 640 to 750, 650 to 750, 500 to 700, 550 to 700, 600 to 700, 640 to 700, 650 to 700, 500 to 680, 550 to 680, 600 to 680, 640 to 680, 650 to 680, 500 to 670, 550 to 670, 600 to 670, 640 to 670, or 650 to 670 amino acids). and a second amino acid sequence positioned C-terminally to the first amino acid sequence, the second amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 664 to 986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises a first amino acid sequence (N-terminal domain) (e.g., an NLS and / or a nucleotide sequence) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). and a second amino acid sequence positioned C-terminally to the first amino acid sequence, the second amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 664 to 986 of CasX1 in Figure 3, panel A).In some cases, the CasX protein comprises a first amino acid sequence (N-terminal domain) (e.g., an N-terminal domain) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). and a second amino acid sequence positioned C-terminal to the first amino acid sequence, the second amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1 and 2 (e.g., the domain shown as amino acids 664 to 986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein is in the range of 500-750 amino acids in length (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). and a second amino acid sequence positioned C-terminally to the first amino acid sequence, the second amino acid sequence having an amino acid sequence corresponding to amino acids 664 to 986 of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., amino acids 651 to 978 of the CasX protein sequence set forth as SEQ ID NO: 2).
[0073] In some cases, the CasX protein (of the subject compositions and / or methods) comprises an amino acid sequence (amino acids 651-978 of the CasX protein sequence set forth as SEQ ID NO: 2) having 20% or greater sequence identity (e.g., 30% or greater, 40% or greater, 50% or greater, 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). For example, in some cases, the CasX protein comprises an amino acid sequence having 50% or greater sequence identity (e.g., 60% or greater, 70% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 80% or greater sequence identity (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises an amino acid sequence corresponding to amino acids 664-986 of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., amino acids 651-978 of the CasX protein sequence set forth as SEQ ID NO: 2).
[0074] In some cases, the CasX protein (of the subject compositions and / or methods) comprises a first amino acid sequence (N-terminal domain) (e.g., excluding any fused heterologous sequences, such as an NLS and / or catalytic domain) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). ) and a second amino acid sequence positioned C-terminal to the first amino acid sequence (e.g., amino acids 651-978 of the CasX protein sequence set forth as SEQ ID NO: 2) that has 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain set forth as amino acids 664-986 of CasX1 in Figure 3, panel A).For example, in some cases, the CasX protein comprises a first amino acid sequence (N-terminal domain) (e.g., an NLS and / or catalytic domain) having a length in the range of 500 to 750 amino acids (e.g., 550 to 750, 600 to 750, 640 to 750, 650 to 750, 500 to 700, 550 to 700, 600 to 700, 640 to 700, 650 to 700, 500 to 680, 550 to 680, 600 to 680, 640 to 680, 650 to 680, 500 to 670, 550 to 670, 600 to 670, 640 to 670, or 650 to 670 amino acids). and a second amino acid sequence positioned C-terminal to the first amino acid sequence, the second amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein comprises a first amino acid sequence (N-terminal domain) (e.g., an NLS and / or an NLS-like domain) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). or catalytic domain) and a second amino acid sequence positioned C-terminal to the first amino acid sequence, the second amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A).In some cases, the CasX protein comprises a first amino acid sequence (N-terminal domain) (e.g., a nucleotide sequence) having a length in the range of 500-750 amino acids (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). and a second amino acid sequence positioned C-terminal to the first amino acid sequence, the second amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) with the C-terminal domain of any one of the CasX protein sequences set forth as SEQ ID NOs: 1-3 (e.g., the domain shown as amino acids 664-986 of CasX1 in Figure 3, panel A). In some cases, the CasX protein is in the range of 500-750 amino acids in length (e.g., 550-750, 600-750, 640-750, 650-750, 500-700, 550-700, 600-700, 640-700, 650-700, 500-680, 550-680, 600-680, 640-680, 650-680, 500-670, 550-670, 600-670, 640-670, or 650-670 amino acids). and a second amino acid sequence positioned C-terminally to the first amino acid sequence, the second amino acid sequence having an amino acid sequence corresponding to amino acids 664 to 986 of the CasX protein sequence set forth as SEQ ID NO: 1 (e.g., amino acids 651 to 978 of the CasX protein sequence set forth as SEQ ID NO: 2).
[0075] CasX mutants A mutant CasX protein has an amino acid sequence that differs by at least one amino acid (e.g., has a deletion, insertion, substitution, or fusion) when compared with the amino acid sequence of the corresponding wild-type CasX protein. A CasX protein that cleaves one strand of a double-stranded target nucleic acid but not the other is referred to herein as a "nickase" (e.g., "nickase CasX"). A CasX protein that has substantially no nuclease activity is referred to herein as an inactive CasX protein ("dCasX") (although nuclease activity may be provided by a heterologous polypeptide-fusion partner, as described further below, in the case of a chimeric CasX protein). For any of the CasX mutant proteins described herein (e.g., nickase CasX, dCasX, chimeric CasX), the CasX variant can comprise a CasX protein sequence with the same parameters (e.g., domains present, percent identity, etc.) as described above.
[0076] Mutants - Catalytic activity In some cases, the CasX protein is, for example, a mutant CasX protein mutated relative to a naturally occurring catalytically active sequence and exhibits reduced cleavage activity (e.g., 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, or 30% or less) compared to the corresponding naturally occurring sequence. In some cases, such a mutant CasX protein is a catalytically "inactive" protein (has substantially no cleavage activity) and may be referred to as "dCasX." In some cases, the mutant CasX protein is a nickase (cleaves only one strand of a double-stranded target nucleic acid, e.g., double-stranded target DNA). As described in detail herein, in some cases, a CasX protein (in some cases a CasX protein with wild-type cleavage activity, and in some cases a mutant CasX with reduced cleavage activity, e.g., dCasX or nickase CasX) is fused (conjugated) to a heterologous polypeptide having an activity of interest (e.g., catalytic ability of interest) to form a fusion protein (chimeric CasX protein).
[0077] Conserved catalytic residues of CasX, when numbered according to CasX1 (SEQ ID NO: 1), include D672, E769, D935, and when numbered according to CasX2 (SEQ ID NO: 2), include 659D, 756E, and 922D (these residues are underlined in Figure 1). (Note that in the alignment of Figure 2, the numbering does not track either CasX, but instead tracks the alignment itself. In this paragraph, the conserved residues mentioned above are marked in the figure, with CasX2 being the top sequence ("gwc2") and CasX1 being the bottom sequence ("gwa2").
[0078] Thus, in some cases, the CasX protein has reduced activity, and one or more of the above amino acids (or one or more corresponding amino acids in any CasX protein) are mutated (e.g., substituted with alanine). In some cases, the mutant CasX protein is a catalytically "inactive" protein (lacking catalytic activity) and is referred to as "dCasX." The dCasX protein can be fused to a fusion partner that provides activity, and in some cases, dCasX (e.g., one that provides catalytic activity but lacks a fusion partner that may have an NLS when expressed in a eukaryotic cell) can bind to target DNA and block RNA polymerase from translating from the target DNA. In some cases, the mutant CasX protein is a nickase (cleaving only one strand of a double-stranded target nucleic acid, e.g., double-stranded target DNA).
[0079] Mutant forms - chimeric CasX (i.e., fusion proteins) As described above, in some cases, a CasX protein (in some cases a CasX protein with wild-type cleavage activity, and in some cases a mutant CasX with reduced cleavage activity, such as dCasX or nickase CasX) is fused (conjugated) to a heterologous polypeptide having an activity of interest (e.g., catalytic ability of interest) to form a fusion protein (chimeric CasX protein). The heterologous polypeptide to which the CasX protein can be fused is referred to herein as a "fusion partner."
[0080] In some cases, the fusion partner can modulate (e.g., inhibit transcription, increase transcription) transcription of the target DNA. For example, in some cases, the fusion partner is a protein (or a domain from a protein) that inhibits transcription (e.g., a protein that functions via a transcription repressor, recruitment of a transcription inhibitor protein, modification of target DNA such as methylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of histone modifiers such as those that modify histone acetylation and / or methylation, etc.). In some cases, the fusion partner is a protein (or a domain from a protein) that increases transcription (e.g., a transcription activator, recruitment of a transcription activator protein, modification of target DNA such as methylation, recruitment of a DNA modifier, modulation of histones associated with target DNA, recruitment of histone modifiers such as those that modify histone acetylation and / or methylation, etc.).
[0081] In some cases, the chimeric CasX protein comprises a heterologous polypeptide having an enzymatic activity that modifies a target nucleic acid (e.g., nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity).
[0082] In some cases, the chimeric CasX protein comprises a heterologous polypeptide having an enzymatic activity (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity) that modifies a polypeptide (e.g., a histone) associated with a target nucleic acid.
[0083] Examples of proteins (or fragments thereof) that can be used to increase transcription include transcription activators such as VP16, VP64, VP48, VP160, p65 subdomains (e.g., from NFkB), and the activation domain of EDLL and / or TAL activation domains (e.g., for activity in plants); histone lysine methyltransferases (e.g., SET1A, SET1B, MLL1-5, ASH1, SYMD2, NSD1, etc.); histone lysine demethylases (e.g., JHDM2a / b, UTX, JMJD3, etc.); histone acetyltransferases (e.g., GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK, etc.); and DNA demethylases (e.g., Ten-Eleven These include, but are not limited to, TET translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, and the like.
[0084] Examples of proteins (or fragments thereof) that can be used to reduce transcription include transcriptional repressors such as Kruppel-associated box (KRAB or SKD); KOX1 repression domain; Mad mSIN3 interacting domain (SID); ERF repressor domain (ERD), SRDX repression domain (e.g., for repression in plants), etc.; histone lysine methyltransferases (e.g., Pr-SET7 / 8, SUV4-20H1, RIZ1, etc.); histone lysine demethylases (e.g., JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, etc.); histone lysine deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.); DNA methylases (e.g., HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), etc.); and peripheral recruitment elements (e.g., Lamin A, Lamin B, etc.).
[0085] In some cases, the fusion partner has an enzymatic activity that modifies a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities that can be provided by the fusion partner include nuclease activity, such as that provided by a restriction enzyme (e.g., FokI nuclease), methyltransferase activity, such as that provided by a methyltransferase (e.g., HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), etc.); demethylase activity, such as that provided by a Ten-Eleven demethylase activity such as that provided by TET dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.), DNA repair activity, DNA damage activity, deamination activity such as that provided by deaminases (e.g., cytosine deaminase enzymes such as rat APOBEC1), dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase and / or resolvase (e.g., Gin These include, but are not limited to, integrase activity such as that provided by a Gin invertase, such as a hyperactive mutant of invertase (GinH106Y); human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase, etc.), transposase activity, recombinase activity such as that provided by a recombinase (e.g., the catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity.
[0086] In some cases, the fusion partner has an enzymatic activity that modifies a protein (e.g., histone, RNA-binding protein, DNA-binding protein, etc.) associated with a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities (that modify a protein associated with a target nucleic acid) that can be provided by a fusion partner include histone methyltransferases (HMTs) (e.g., suppressor of variegation 3-9 homolog 1 (SUV39H1, also known as KMT1A), euchromatin histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1, etc., SET1A, SET1B, MLL1-5, ASH1, SYMD2, NSD methyltransferase activity, such as that provided by 1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1), histone demethylases (e.g., lysine demethylase 1A (KDM1A, also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, J demethylase activity such as that provided by histone acetylase transferases (e.g., catalytic cores / fragments of human acetyltransferases p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HBO1 / MYST2, HMOF / MYST1, SRC1, ACTR, P160, CLOCK, etc.); histone deacetylases (e.g., These include, but are not limited to, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, and demyristoylating activity, such as those provided by HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.
[0087] Additional examples of suitable fusion partners are dihydrofolate reductase (DHFR) destabilization domains (e.g., to generate chemically controllable chimeric CasX proteins) and chloroplast transit peptides. Suitable chloroplast transit peptides include, but are not limited to:
[0088] MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGRVKCMQVWPPIGKKKFETLSYLPPLTRDSRA (SEQ ID NO: 83), MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGRVKS (SEQ ID NO: 84), MASSMLSSATMVASPAQATMVAPFNGLKSSAAFPATRKANNDITSITSNGGRVNCMQVWPPIE KKKFETLSYLPDLTDSGGRVNC (SEQ ID NO: 85), MAQVSRICNGVQNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIGSELRPLKVMSSVSTAC (SEQ ID NO: 86), MAQVSRICNGVWNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIGSELRPLKVMSSVSTAC (SEQ ID NO: 87), MAQINNMAQGIQTLNPNSNFHK PQVPKSSSFLVFGSKKLKNSANSMLVLKKDSIFMQLFCSFRISASVATAC (SEQ ID NO: 88), MAALVTSQLATSGTVLSVTDRFRRPGFQGLRPRNPADAALGMRTVGASAAPKQSRKPHRFDRRCLSMVV (SEQ ID NO: 89), MAALTTSQLATSATGFGIADRSAPSSLLRHGFQGLKPRSPAGGDATSLSVTTSARATPKQQRSVQRGSRRFPSVVVC (SEQ ID NO: 90), MASSVLSSAAVATRSNVAQANMVAPFTGLKSAASFPVSRKQNLDITSIASNGGRVQC (SEQ ID NO: 91), MESLAATSVFAPSRVAVPAARALVRAGTVVPTRRTSSTSGTSGVKCSAAVTPQASPVISRSAAAA (SEQ ID NO: 92), and MGAAATSMQSLKFSNRLVPPSRRLSPVPNNVTCNNLPKSAAPVRTVKCCASSWNSTINGAAATTNGASAASS (SEQ ID NO: 93).
[0089] In some cases, a CasX fusion polypeptide of the present disclosure comprises a) a CasX polypeptide of the present disclosure and b) a chloroplast transit peptide. Thus, for example, a CRISPR-CasX complex can be targeted to chloroplasts. In some cases, this targeting can be achieved by the presence of an N-terminal extension called a chloroplast transit peptide (CTP) or plastid transit peptide. A chromosomal transgene from a bacterial source must have a sequence encoding a CTP sequence fused to the sequence encoding the expressed polypeptide if the expressed polypeptide is to be compartmentalized within a plant plastid (e.g., a chloroplast). Therefore, localization of an exogenous polypeptide to chloroplasts is often achieved by operably linking a polynucleotide sequence encoding a CTP sequence to the 5' region of the polynucleotide encoding the exogenous polypeptide. The CTP is removed in a processing step during translocation into plastids. However, processing efficiency can be affected by the amino acid sequence of the CTP and nearby sequences at the NH2-terminus of the peptide. Other options for targeting to chloroplasts that have been described are the maize cab-m7 signal sequence (U.S. Pat. No. 7,022,896, WO97 / 41228), the pea glutathione reductase signal sequence (WO97 / 41228), and the CTP described in US2009029861.
[0090] In some instances, a CasX fusion polypeptide of the present disclosure can comprise a) a CasX polypeptide of the present disclosure, and b) an endosomal escape peptide. In some instances, the endosomal escape polypeptide comprises the amino acid sequence GLFXALLXLLXSLWXLLLXA (SEQ ID NO: 94), where each X is independently selected from lysine, histidine, and arginine. In some instances, the endosomal escape polypeptide comprises the amino acid sequence GLFHALLHLLHSLWHLLLHA (SEQ ID NO: 95).
[0091] For examples of some of the above fusion partners (and others) used in the context of fusion with Cas9, zinc finger, and / or TALE proteins (for site-specific targeted nucleic acid modification, modulation of transcription, and / or targeted protein modification, e.g., histone modification), see, e.g., Nomura et al., J Am Chem Soc. 2007 Jul 18;129(28):8676-7; Rivenbark et al., Epigenetics. 2012 Apr;7(4):350-60; Nucleic Acids Res. 2016 Jul 8;44(12):5615-28; Gilbert et al., Cell. 2013 Jul 18;154(2):442-51; Kearns et al., Nat Methods. 2015 May;12(5):401-3; Mendenhall et al., Nat Biotechnol. 2013 Dec;31(12):1133-6, Hilton et.al., Nat Biotechnol.2015 May;33(5):510-7, Gordley et.al., Proc Natl Acad Sci USA.2009 Mar 31;106(13):5053-8, Akopian et.al., Proc Natl Acad Sci USA.2003 Jul 22;100(15):8688-91, Tan et.,al.,J Virol.2006 Feb;80(4):1939-48, Tan et.al.,Proc Natl Acad Sci USA.2003 Oct 14;100(21):11997-2002, Papworth et.al.,Proc Natl Acad Sci USA.2003 Feb 18;100(4):1621-6, Sanjana et.al., Nat Protoc.2012 Jan 5;7(1):171-92, Beerli et.al., Proc Natl Acad Sci USA.1998 Dec 8;95(25):14628-33, Snowden et.al., Curr Biol.2002 Dec 23;12(24):2159-66, Xu et.al.,Xu et.al.,Cell Discov.2016 May 3;2:16009;Komor et al.,Nature.2016 Apr 20;533(7603):420-4;Chaikind et.al.,Nucleic Acids Res.2016 Aug 11;Choudhury et.al.,Oncotarget.2016 Jun 23.Du et al.,Cold Spring Harb Protoc.2016 Jan 4;Pham et al.,Methods Mol Biol.2016;1358:43-57;Balboa et al.,Stem Cell Reports.2015 Sep 8;5(3):448-59;Hara et al.,Sci Rep.2015 Jun 9;5:11221;Piatek et al.,Plant BiotechnolJ.2015 May;13(4):578-89;Hu et al.,Nucleic Acids Res.2014 Apr;42(7):4375-90;Cheng et al.,Cell Res.2013 Oct;23(10):1163-71 and Cheng et al.,Cell Res.2013 Oct;23(10):1163-71;
[0092] Additional suitable heterologous polypeptides include, but are not limited to, polypeptides that directly and / or indirectly provide increased transcription and / or translation of a target nucleic acid (e.g., transcriptional activators or fragments thereof, proteins or fragments thereof that recruit transcriptional activators, small molecule / drug-responsive transcriptional and / or translational regulators, translational regulatory proteins, etc.). Non-limiting examples of heterologous polypeptides that achieve increased or decreased transcription include transcriptional activator and transcriptional repressor domains. In some such cases, chimeric CasX polypeptides are targeted to specific locations (i.e., sequences) in a target nucleic acid by a guide nucleic acid (guide RNA) to provide locus-specific regulation, such as blocking RNA polymerase binding to the promoter (selectively inhibiting transcriptional activator function) and / or modifying the local chromatin state (when a fusion sequence that modifies the target nucleic acid or a polypeptide associated with the target nucleic acid is used). In some cases, the change is transient (e.g., transcriptional repression or activation). In some cases, the change is heritable (eg, when an epigenetic modification is made to the target nucleic acid or to a protein associated with the target nucleic acid, such as a nucleosomal histone).
[0093] Non-limiting examples of heterologous polypeptides for use in targeting ssRNA target nucleic acids include, but are not limited to, splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or release factors, e.g., eIF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases, including, e.g., adenosine deaminases acting on RNA (ADARs), AI and / or CU editing enzymes); helicases; RNA binding proteins, and the like. A heterologous polypeptide can comprise an entire protein, or in some cases, can comprise a fragment of a protein (e.g., a functional domain).
[0094] The heterologous polypeptide of a subject chimeric CasX polypeptide can be any domain (for purposes of this disclosure, intramolecular and / or intermolecular secondary structures, e.g., double-stranded RNA duplexes such as hairpins, stem-loops, etc.) capable of interacting with ssRNA, whether transiently or irreversibly, directly or indirectly, including endonucleases (e.g., RNase III from proteins such as SMG5 and SMG6, CRR22 DYW domain, Dicer, and PIN (PilT N-terminal) domain); proteins and protein domains involved in stimulating RNA cleavage (e.g., CPSF, CstF, CFIm, and CFIIm); exonucleases (e.g., XRN-1 or exonuclease T); deadenylases (e.g., HNT3); proteins and protein domains involved in nonsense-mediated RNA decay (e.g., UPF1, UPF2, UPF3, UPF3b, RNP S1, Y14, DEK, REF2, and SRm160); proteins and protein domains involved in RNA stabilization (e.g., PABP); proteins and protein domains involved in translational repression (e.g., Ago2 and Ago4); proteins and protein domains involved in translational stimulation (e.g., Staufen); proteins and protein domains involved (e.g., capable) in modulating translation (e.g., translation factors such as initiation factors, elongation factors, and release factors, e.g., eIF4G); proteins and protein domains involved in RNA polyadenylation (e.g., PAP1, GLD-2, and Star-PAP); proteins and protein domains involved in RNA polyuridylation (e.g., CI D1 and terminal uridylate transferase); proteins and protein domains involved in RNA localization (e.g., from IMP1, ZBP1, She2p, She3p, and Bicaudal-D); proteins and protein domains involved in the nuclear retention of RNA (e.g., Rrp6); proteins and protein domains involved in the nuclear export of RNA (e.g., TAP, NXF1, THO, TREX, REF, and Aly); proteins and protein domains involved in the repression of RNA splicing (e.g., PTB, Sam68, and hnRNP A1);Proteins and protein domains involved in stimulating RNA splicing (e.g., serine / arginine-rich (SR) domains); proteins and protein domains involved in reducing transcription efficiency (e.g., FUS (TLS)); and proteins and protein domains involved in stimulating transcription (e.g., CDK7 and HIV Tat). Alternatively, the effector domain may be an endonuclease; a protein or protein domain capable of stimulating RNA cleavage; an exonuclease; a deadenylase; a protein or protein domain with nonsense-mediated RNA degradation activity; a protein or protein domain capable of stabilizing RNA; a protein or protein domain capable of repressing translation; a protein or protein domain capable of stimulating translation; a protein or protein domain capable of modulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, e.g., eIF4G); a protein capable of polyadenylation of RNA The heterologous polypeptide may be selected from the group consisting of proteins and protein domains; proteins and protein domains capable of polyuridylation of RNA; proteins and protein domains with RNA localization activity; proteins and protein domains capable of nuclear retention of RNA; proteins and protein domains with RNA nuclear export activity; proteins and protein domains capable of RNA splicing; proteins and protein domains capable of stimulating RNA splicing; proteins and protein domains capable of reducing transcription efficiency; and proteins and protein domains capable of stimulating transcription. Another suitable heterologous polypeptide is the PUF RNA binding domain described in detail in WO2012068627 (incorporated herein by reference in its entirety);
[0095] Some RNA splicing factors that can be used (in whole or as fragments) as heterologous polypeptides in chimeric CasX polypeptides have modular structures, with distinct sequence-specific RNA-binding modules and splicing effector domains. For example, members of the serine / arginine-rich (SR) protein family contain an N-terminal RNA recognition motif (RRM) that binds to exon splicing enhancers (ESEs) in pre-mRNAs and a C-terminal RS domain that promotes exon inclusion. As another example, the hnRNP protein hnRNP A1 binds to exon splicing silencers (ESSs) through its RRM domain and inhibits exon inclusion through its C-terminal glycine-rich domain. Some splicing factors can regulate the use of alternative splice sites (SSs) by binding to regulatory sequences between the two alternative sites. For example, ASF / SF2 can recognize an ESE and promote the use of an intron-proximal site, whereas hnRNP A1 can bind to an ESS and redirect splicing to the use of an intron-distal site. One use of such factors is to generate ESFs that modulate the alternative splicing of endogenous genes, particularly disease-related genes. For example, Bcl-x pre-mRNA produces two splicing isoforms with two alternative 5' splice sites, encoding proteins with opposite functions. The long splicing isoform, Bcl-xL, is a potent apoptosis inhibitor expressed in long-lived, postmitotic cells and is upregulated in many cancer cells, protecting them from apoptotic signals. The short isoform, Bcl-xS, is a pro-apoptotic isoform expressed at high levels in cells with a high turnover rate (e.g., developing lymphocytes). The ratio of the two Bcl-x splicing isoforms is regulated by multiple cω elements located either in the core exon region or in the exon extension region (e.g., between the two alternative 5' splice sites). For further examples, see WO2010075303 (incorporated herein by reference in its entirety).
[0096] Further suitable fusion partners include, but are not limited to, proteins (or fragments thereof) that are boundary elements (e.g., CTCF), proteins and fragments thereof that provide peripheral recruitment (e.g., Lamin A, Lamin B, etc.), protein docking elements (e.g., FKBP / FRB, Pil1 / Aby1, etc.).
[0097] Examples of various additional suitable heterologous polypeptides (or fragments thereof) to the subject chimeric CasX polypeptides include, but are not limited to, those described in the following applications (which publications relate to other CRISPR endonucleases, such as Cas9, although the fusion partners described can alternatively be used with CasX): PCT patent applications: WO2010075303, WO2012068627, and WO2013155555; and, e.g., U.S. Patents and Patent Applications Nos. 8,906,616, 8,895,308, 8,889,418, 8,889,356, 8,871,445, 8,865,406, 8,795,965, 8,771,945, 8,697,No. 359, No. 20140068797, No. 20140170753, No. 20140179006, No. 20140179770, No. 20140186843, No. 20140186919, No. 20140186958, 20140189896, 20140227787, 20140234972, 20140242664, 20140242699, 201402 No. 42700, No. 20140242702, No. 20140248702, No. 20140256046, No. 20140273037, No. 20140273226, No. 20140273230 , No. 20140273231, No. 20140273232, No. 20140273233, No. 20140273234, No. 20140273235, No. 20140287938, No. 2014 No. 0295556, No. 20140295557, No. 20140298547, No. 20140304853, No. 20140309487, No. 20140310828, No. 2014031083 No. 0, No. 20140315985, No. 20140335063, No. 20140335620, No. 20140342456, No. 20140342457, No. 20140342458, No. 20 Nos. 140349400, 20140349405, 20140356867, 20140356956, 20140356958, 20140356959, 20140357523, 20140357530, 20140364333, and 20140377868 (all of which are incorporated herein by reference in their entirety).
[0098] In some cases, the heterologous polypeptide (fusion partner) provides subcellular localization, i.e., the heterologous polypeptide comprises a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence that retains the fusion protein outside the nucleus, e.g., a nuclear export sequence (NES), a sequence that retains the fusion protein in the cytoplasm, a mitochondrial localization signal for targeting to mitochondria, a chloroplast localization signal for targeting to chloroplasts, an ER retention signal, etc.). In some embodiments, the CasX fusion polypeptide does not contain an NLS, and thus the protein is not targeted to the nucleus (which can be advantageous, for example, when the target nucleic acid is RNA present in the cytosol). In some embodiments, the heterologous polypeptide can be provided with a tag to facilitate tracking and / or purification (i.e., the heterologous polypeptide is a detectable label) (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, etc.; a histidine tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag, etc.).
[0099] In some cases, a CasX protein (e.g., a wild-type CasX protein, a mutant CasX protein, a chimeric CasX protein, a dCasX protein, a chimeric CasX protein (wherein the CasX portion has reduced nuclease activity), e.g., a dCasX protein fused to a fusion partner, etc.) comprises (is fused to) a nuclear localization signal (NLS) (e.g., in some cases, two or more, three or more, four or more, or five or more NLSs). Thus, in some cases, a CasX polypeptide comprises one or more NLSs (e.g., two or more, three or more, four or more, or five or more NLSs). In some cases, the one or more NLSs (e.g., two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the N-terminus and / or C-terminus. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the N-terminus. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the C-terminus. In some cases, one or more NLSs (three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) both the N-terminus and the C-terminus. In some cases, an NLS is located at the N-terminus and an NLS is located at the C-terminus.
[0100] In some cases, a CasX protein (e.g., a wild-type CasX protein, a mutant CasX protein, a chimeric CasX protein, a dCasX protein, or a chimeric CasX protein having a CasX portion with reduced nuclease activity, such as a dCasX protein fused to a fusion partner) comprises (is fused to) one to ten NLSs (e.g., one to nine, one to eight, one to seven, one to six, one to five, two to ten, two to nine, two to eight, two to seven, two to six, or two to five NLSs). In some cases, a CasX protein (e.g., a wild-type CasX protein, a mutant CasX protein, a chimeric CasX protein, a dCasX protein, or a chimeric CasX protein having a CasX portion with reduced nuclease activity, such as a dCasX protein fused to a fusion partner) comprises (is fused to) two to five NLSs (e.g., two to four or two to three NLSs).
[0101] Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 96); an NLS derived from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 97); a c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 98) or RQRRNELKRSP (SEQ ID NO: 99); an hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 100); an IBB domain from importin-α having the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 101); the sequences VSRKRPRP (SEQ ID NO: 102) and PPKKARED (SEQ ID NO: 103) of the fibroid T protein; the sequence PQPKKKPL (SEQ ID NO: 104) of human p53; and a mouse c-ab1 The sequences include the NLS sequences derived from the sequence SALIKKKKKMAP (SEQ ID NO: 105) of IV; the sequences DRLRR (SEQ ID NO: 106) and PKQKKRK (SEQ ID NO: 107) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 108) of hepatitis virus Δ antigen; the sequence REKKKFLKRR (SEQ ID NO: 109) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 110) of human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 111) of steroid hormone receptor (human) glucocorticoid. Generally, the NLS (or NLSs) is strong enough to induce the accumulation of detectable amounts of CasX in the nucleus of eukaryotic cells. The detection of nuclear accumulation can be carried out by any suitable technique. For example, a detectable marker can be fused to the CasX protein, thereby visualizing its location within the cell. The cell nucleus can also be isolated from the cell, and its contents can then be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay. The accumulation in the nucleus can also be determined indirectly.
[0102] In some cases, the CasX fusion polypeptide contains a "protein transduction domain" or PTD (also known as a CPP—cell-penetrating peptide), which refers to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates crossing of a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule and / or nanoparticle, which can range from small polar molecules to large macromolecules, facilitates membrane crossing of molecules, for example, moving from the extracellular space to the intracellular space or from the cytosol into an organelle. In some embodiments, the PTD is covalently attached to the amino acid terminus of the polypeptide (e.g., to wild-type CasX to generate a fusion protein or to a mutant CasX protein, e.g., dCasX, nickase CasX, or chimeric CasX protein to generate a fusion protein). In some embodiments, the PTD is covalently attached to the carboxyl terminus of the polypeptide (e.g., linked to wild-type CasX to generate a fusion protein or to a mutant CasX protein, e.g., dCasX, nickase CasX, or chimeric CasX protein to generate a fusion protein). In some cases, the PTD is inserted internally into the CasX fusion polypeptide at a suitable insertion site (i.e., not at the N- or C-terminus of the CasX fusion polypeptide). In some cases, the subject CasX fusion polypeptide comprises (is conjugated to, is fused to) one or more PTDs (e.g., two or more, three or more, four or more PTDs). In some cases, the PTD comprises a nuclear localization signal (NLS) (e.g., in some cases, two or more, three or more, four or more, or five or more NLSs). Thus, in some cases, the CasX fusion polypeptide comprises one or more NLSs (e.g., two or more, three or more, four or more, or five or more NLSs). In some embodiments, the PTD is covalently linked to a nucleic acid (e.g., a CasX guide nucleic acid, a polynucleotide encoding a CasX guide nucleic acid, a polynucleotide encoding a CasX fusion polypeptide, a donor polynucleotide, etc.).Examples of PTDs include a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT, which contains YGRKKRRQRRR (SEQ ID NO: 112)); a polyarginine sequence containing a sufficient number of arginines (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines) for direct entry into a cell; a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); a Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO: 113); transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 114); KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 115); and RQIKIWFQNRRMKWKK (SEQ ID NO: 116). Exemplary PTDs include, but are not limited to, YGRKKRRQRRR (SEQ ID NO: 117), RKKRRQRRR (SEQ ID NO: 118); arginine homopolymers of 3 to 50 arginine residues. Exemplary PTD domain amino acid sequences include, but are not limited to, any of the following: YGRKKRRQRRR (SEQ ID NO: 119), RKKRRQRR (SEQ ID NO: 120), YARAAARQARA (SEQ ID NO: 121), THRLPRRRRRR (SEQ ID NO: 122), and GGRRARRRRRR (SEQ ID NO: 123). In some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June;1(5-6):371-381).ACPPs contain a polycationic CPP (e.g., Arg9 or "R9") connected to a matching polyanion (e.g., Glu9 or "E9") via a cleavable linker, which reduces the net charge to near zero, thereby inhibiting cellular attachment and uptake. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesive properties, thus "activating" the ACPP to cross the membrane.
[0103] Linkers (e.g., for fusion partners) In some embodiments, the CasX protein of interest can be fused to a fusion partner via a linker polypeptide (e.g., one or more linker polypeptides). The linker polypeptide can have any of a variety of amino acid sequences. Proteins can be joined by a generally flexible spacer peptide, although other chemical bonds are not excluded. Suitable linkers include polypeptides between 4 and 40 amino acids in length, or between 4 and 25 amino acids in length. These linkers can be produced by linking proteins using synthetic linker-encoding oligonucleotides or can be encoded by a nucleic acid sequence encoding the fusion protein. Peptide linkers with some degree of flexibility can be used. The linking peptide can have virtually any amino acid sequence, with the caveat that preferred linkers generally have sequences that result in flexible peptides. The use of small amino acids such as glycine and alanine is useful in creating flexible peptides. Creating such sequences is routine for those skilled in the art. A variety of different linkers are commercially available and may be suitable for use.
[0104] Examples of linker polypeptides include glycine polymers (G)n, glycine-serine polymers (e.g., (GS)n, GSGGSn (SEQ ID NO: 124), GGSGGSn (SEQ ID NO: 125), and GGGSn (SEQ ID NO: 126), where n is an integer of at least 1), glycine-alanine polymers, alanine-serine polymers. Exemplary linkers may comprise amino acid sequences including, but not limited to, GGSG (SEQ ID NO: 127), GGSGG (SEQ ID NO: 128), GSGSG (SEQ ID NO: 129), GSGGG (SEQ ID NO: 130), GGGSG (SEQ ID NO: 131), GSSSG (SEQ ID NO: 132), and the like. The skilled artisan will recognize that the design of a peptide conjugated to any desired element may include a linker that is fully or partially flexible, whereby the linker may include a flexible linker as well as one or more moieties that confer a less flexible structure.
[0105] Detectable Label In some cases, the CasX polypeptides of the present disclosure comprise a detectable label. Suitable detectable labels and / or moieties capable of providing a detectable signal can include, but are not limited to, enzymes, radioisotopes, members of specific binding pairs, fluorophores, fluorescent proteins, quantum dots, etc.
[0106] Suitable fluorescent proteins include green fluorescent protein (GFP) or variants thereof, blue fluorescent variants of GFP (BFP), cyan fluorescent variants of GFP (CFP), yellow fluorescent variants of GFP (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilized EGFP (dEGFP), destabilized ECFP (dECFP), destabilized EYFP (dEYFP), mCFPm, Cerulean, T-Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J-Red, dimer2, t-dimer2 (12), mRFP1, pocilloporin, Renilla GFP, Monster Examples of fluorescent proteins include, but are not limited to, GFP, paGFP, Kaede protein and kindling protein, phycobiliproteins and phycobiliprotein complexes including B-phycoerythrin, R-phycoerythrin, and allophycocyanin. Other examples of fluorescent proteins include mHoneydew, mBanana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrape1, mRaspberry, mGrape2, and mPlum (Shaner et al. (2005) Nat. Methods 2:905-909). Suitable fluorescent and chromatic proteins from anthozoan species are described, for example, in Matz et al. (1999) Nature Biotechnol. 17:969-973.
[0107] Suitable enzymes include, but are not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), β-galactosidase (GAL), glucose-6-phosphatase dehydrogenase, β-N-acetylglucosaminidase, β-glucuronidase, invertase, xanthine oxidase, firefly luciferase, glucose oxidase (GO), and the like.
[0108] Protospacer adjacent motif (PAM) The CasX protein binds to target DNA at a target sequence defined by the region of complementarity between the DNA target RNA and the target DNA. As with many CRISPR endonucleases, site-specific binding (and / or cleavage) of double-stranded target DNA occurs at a location determined by both (i) base-pairing complementarity between the guide RNA and the target DNA and (ii) a short motif in the target DNA, termed a protospacer adjacent motif (PAM).
[0109] In some embodiments, the PAM of the CasX protein is immediately 5' to the target sequence on the non-complementary strand of the target DNA (the complementary strand hybridizes to the guide sequence of the guide RNA, while the non-complementary strand does not directly hybridize to the guide RNA and is the reverse complement of the non-complementary strand). In some embodiments (e.g., when CasX1 described herein is used), the PAM sequence on the non-complementary strand is 5'-TCN-3' (and in some cases, TTCN), where N is any DNA nucleotide. For examples, see Figure 6, panel C and Figure 7, where the PAM (TCN) (on the non-complementary strand) is TCA (and the PAM shown in the figure is TTCA), and the PAM is 5' to the target sequence.
[0110] In some cases, different CasX proteins (i.e., CasX proteins from various species) may be advantageous for use in the various provided methods to take advantage of different enzymatic characteristics of different CasX proteins (e.g., for different PAM sequence preferences, increased or decreased enzymatic activity, increased or decreased levels of cytotoxicity, altering the balance between NHEJ, homology-directed repair, single-strand breaks, double-strand breaks, etc., utilizing shorter overall sequences, etc.). CasX proteins from different species may require different PAM sequences in the target DNA. Thus, for certain preferred CasX proteins, the PAM sequence requirements may differ from the 5′-TCN-3′ sequence described above. Various methods (including in silico and / or wet-lab methods) for identifying appropriate PAM sequences are known and routine in the art, and any convenient method may be used. The TCN PAM sequence described herein was identified using a PAM depletion assay (see, e.g., Figure 5 in the Examples below).
[0111] CasX guide RNA A nucleic acid molecule that binds to the CasX protein, forms a ribonucleoprotein complex (RNP), and targets that complex to a specific location within a target nucleic acid (e.g., target DNA) is referred to herein as a "CasX guide RNA" or simply "guide RNA." It should be understood that in some cases, a hybrid DNA / RNA may be generated such that the CasX guide RNA contains DNA bases in addition to RNA bases, but the term "CasX guide RNA" is still used herein to encompass such molecules.
[0112] A CasX guide RNA can be said to contain two segments: a targeting segment and a protein-binding segment. The targeting segment of a CasX guide RNA contains a nucleotide sequence (guide sequence) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid (e.g., a target ssRNA, a target ssDNA, the complementary strand of a double-stranded target DNA, etc.). The protein-binding segment (or "protein-binding sequence") interacts with (binds to) a CasX polypeptide. The protein-binding segment of a subject CasX guide RNA contains two complementary stretches of nucleotides that hybridize to each other to form a double-stranded RNA duplex (dsRNA duplex). Site-specific binding and / or cleavage of a target nucleic acid (e.g., genomic DNA) can occur at a location (e.g., a target sequence at a target locus) determined by base-pairing complementarity between the CasX guide RNA (the guide sequence of the CasX guide RNA) and the target nucleic acid.
[0113] The CasX guide RNA and the CasX protein, e.g., a fusion CasX polypeptide, form a complex (e.g., bind via non-covalent interactions). The CasX guide RNA provides target specificity to the complex by including a targeting segment that includes a guide sequence (a nucleotide sequence complementary to the sequence of the target nucleic acid). The CasX protein of the complex provides site-specific activity (e.g., cleavage activity provided by the CasX protein and / or activity provided by the fusion partner in the case of a chimeric CasX protein). In other words, the CasX protein is guided to the target nucleic acid sequence (e.g., the target sequence) by its association with the CasX guide RNA.
[0114] The "guide sequence," also referred to as the "targeting sequence," of a CasX guide RNA can be modified such that the CasX guide RNA can target a CasX protein (e.g., a naturally occurring CasX protein, a fusion CasX polypeptide (chimeric CasX), etc.) to any desired sequence in any desired target nucleic acid, with the exception that a PAM sequence may be considered (e.g., as described herein). Thus, for example, a CasX guide RNA can have a guide sequence that is complementary to (e.g., hybridizes to) a sequence in a nucleic acid in a eukaryotic cell, e.g., a viral nucleic acid, a eukaryotic nucleic acid (e.g., a eukaryotic chromosome, a chromosomal sequence, a eukaryotic RNA, etc.).
[0115] The subject CasX guide RNAs may also be referred to as including an "activator" and a "targeter" (e.g., an "activator-RNA" and a "targeter-RNA," respectively. When the "activator" and "targeter" are two separate molecules, the guide RNA is referred to herein as a "dual guide RNA," "dgRNA," "dual-molecule guide RNA," or "binomolecular guide RNA" (e.g., a "CasX dual guide RNA"). In some embodiments, the activator and targeter are covalently linked to each other (e.g., via an intervening nucleotide), and the guide RNA is referred to herein as a "single guide RNA," "sgRNA," "single-molecule guide RNA," or "single-molecule guide RNA." The subject CasX single guide RNAs are referred to as "CasX single guide RNAs" (e.g., "CasX single guide RNAs"). Thus, a subject CasX single guide RNA comprises a targeter (e.g., targeter-RNA) and an activator (e.g., activator-RNA) that are linked to each other (e.g., by intervening nucleotides) and hybridize to each other to form a double-stranded RNA duplex (dsRNA duplex) of the protein-binding segment of the guide RNA, thus resulting in a stem-loop structure (Figure 6, panel C). Thus, the targeter and activator each have a duplex-forming segment, and the duplex-forming segment of the targeter and the duplex-forming segment of the activator are complementary to each other and hybridize to each other.
[0116] In some embodiments, the linker of a CasX single guide RNA is a stretch of nucleotides (shown as GAAA in Figure 6, panel C). In some cases, the targeter and activator of a CasX single guide RNA are linked to each other by intervening nucleotides, and the linker can have a length of 3 to 20 nucleotides (nt) (e.g., 3 to 15, 3 to 12, 3 to 10, 3 to 8, 3 to 6, 3 to 5, 3 to 4, 4 to 20, 4 to 15, 4 to 12, 4 to 10, 4 to 8, 4 to 6, or 4 to 5 nt). In some embodiments, the linker of a CasX single guide RNA can have a length of 3 to 100 nucleotides (nt) (e.g., 3 to 80, 3 to 50, 3 to 30, 3 to 25, 3 to 20, 3 to 15, 3 to 12, 3 to 10, 3 to 8, 3 to 6, 3 to 5, 3 to 4, 4 to 100, 4 to 80, 4 to 50, 4 to 30, 4 to 25, 4 to 20, 4 to 15, 4 to 12, 4 to 10, 4 to 8, 4 to 6, or 4 to 5 nt). In some embodiments, the linker of a CasX single guide RNA can have a length of 3 to 10 nucleotides (nt) (e.g., 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 to 4, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, or 4 to 5 nt).
[0117] Guide sequence of CasX guide RNA The targeting segment of a target CasX guide RNA comprises a guide sequence (i.e., a targeting sequence), which is a nucleotide sequence complementary to a sequence (target site) in a target nucleic acid. In other words, the targeting segment of a CasX guide RNA can interact with a target nucleic acid (e.g., double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), single-stranded RNA (ssRNA), or double-stranded RNA (dsRNA)) in a sequence-specific manner by hybridization (i.e., base pairing). The guide sequence of a CasX guide RNA can be modified / designed (e.g., by genetic engineering) to hybridize to any desired target sequence within a target nucleic acid (e.g., a eukaryotic target nucleic acid such as genomic DNA) (e.g., taking PAM into consideration, e.g., when targeting a dsDNA target).
[0118] In some embodiments, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 65% or greater, 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100%.
[0119] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over the seven contiguous 3'-most nucleotides.
[0120] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19 or more (e.g., 20% or greater, 21% or greater, 22% or greater) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19 or more (e.g., 20% or greater, 21% or greater, 22% or greater) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides.
[0121] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19-25 contiguous nucleotides.
[0122] In some cases, the guide sequence has a length in the range of 19-30 nucleotides (nt) (e.g., 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some cases, the guide sequence has a length in the range of 19-25 nucleotides (nt) (e.g., 19-22, 19-20, 20-25, 20-25, or 20-22 nt). In some cases, the guide sequence has a length of 19 or more nt (e.g., 20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 19 nt. In some cases, the guide sequence has a length of 20 nt. In some cases, the guide sequence has a length of 21 nt. In some cases, the guide sequence has a length of 22 nt. In some cases, the guide sequence has a length of 23 nt.
[0123] Protein-binding segment of CasX guide RNA The protein-binding segment of the target CasX guide RNA interacts with the CasX protein. The CasX guide RNA guides the bound CasX protein to a specific nucleotide sequence within the target nucleic acid via the guide sequence described above. The protein-binding segment of the CasX guide RNA contains two stretches of nucleotides (an activator duplex-forming segment and a targeter duplex-forming segment) that are complementary to each other and hybridize to form a double-stranded RNA duplex (dsRNA duplex). Thus, the protein-binding segment comprises a dsRNA duplex.
[0124] In some cases, the dsRNA duplex region formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) (e.g., in a duplex or single guide RNA format) comprises a range of 8-25 base pairs (bp) (e.g., 8-22, 8-18, 8-15, 8-12, 12-25, 12-22, 12-18, 12-15, 13-25, 13-22, 13-18, 13-15, 14-25, 14-22, 14-18, 14-15, 15-25, 15-22, 15-18, 17-25, 17-22, or 17-18 bp, e.g., 15 bp, 16 bp, 17 bp, 18 bp, 19 bp, 20 bp, 21 bp, etc.). In some cases, the duplex region (e.g., in a duplex or single-guide RNA format) comprises 8 or more bp (e.g., 10 or more, 12 or more, 15 or more, or 17 or more bp). In some cases, because not all nucleotides in the duplex region are paired, the duplex-forming region may comprise a bulge (see, e.g., Figure 6, panel C and Figure 7). The term "bulge" is used herein to refer to a stretch of nucleotides (which may be a single nucleotide) that does not contribute to the double-stranded duplex but is surrounded 5' and 3' by contributing nucleotides, and as such, the bulge is considered part of the duplex region. In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) comprises one or more bulges (e.g., two or more, three or more, four or more bulges). In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) contains two or more bulges (e.g., three or more, four or more bulges). In some cases, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) contains one to five bulges (e.g., one to four, one to three, two to five, two to four, or two to three bulges).
[0125] Thus, in some cases, the duplex-forming segments of the activator and targeting factor have 70% to 100% complementarity with each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, the duplex-forming segments of the activator and targeting factor have 70% to 100% complementarity with each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, the duplex-forming segments of the activator and targeting factor have 85% to 100% complementarity with each other (e.g., 90% to 100%, 95% to 100% complementarity). In some cases, the duplex-forming segments of the activator and targeting agent have 70% to 95% complementarity with each other (e.g., 75% to 95%, 80% to 95%, 85% to 95%, 90% to 95% complementarity).
[0126] In other words, in some embodiments, the dsRNA duplex formed between the activator and targeter (i.e., the activator / targeter dsRNA duplex) comprises two stretches of nucleotides that have 70% to 100% complementarity with each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, the activator / targeter dsRNA duplex comprises two stretches of nucleotides that have 85% to 100% complementarity with each other (e.g., 90% to 100%, 95% to 100% complementarity). In some cases, the activator / targeter dsRNA duplex contains two stretches of nucleotides that have 70%-95% complementarity to each other (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity).
[0127] The duplex region of a subject CasX guide RNA (in dual-guide or single-guide RNA format) can include one or more (1, 2, 3, 4, 5, etc.) mutations relative to a naturally occurring duplex region. For example, in some cases, base pairing can be maintained, but the nucleotides contributed to the base pair from each segment (targeter and activator) can be different. In some cases, the duplex region of a subject CasX guide RNA includes more paired bases, fewer paired bases, a smaller bulge, a larger bulge, fewer bulges, more bulges, or any convenient combination thereof, compared to a naturally occurring duplex region (of a naturally occurring CasX guide RNA).
[0128] In some cases, the activator (e.g., activator-RNA) of a subject CasX guide RNA (in duplex or single-guide RNA format) comprises at least two internal RNA duplexes (i.e., two internal hairpins in addition to the activator / targeter dsRNA). The internal RNA duplex (hairpin) of the activator can be positioned 5' of the activator / targeter dsRNA duplex (see, e.g., Figure 6, panel C and Figure 7, both of which include activators with two internal hairpins positioned 5' of the activator / targeter dsRNA duplex). In some cases, the activator comprises one hairpin positioned 5' of the activator / targeter dsRNA duplex. In some cases, the activator comprises two hairpins positioned 5' of the activator / targeter dsRNA duplex. In some cases, the activator comprises three hairpins positioned 5' of the activator / targeter dsRNA duplex. In some cases, the activator comprises two or more hairpins (e.g., three or more or four or more hairpins) positioned 5' of the activator / targeter dsRNA duplex. In some cases, the activator comprises two to five hairpins (e.g., two to four or two to three hairpins) positioned 5' of the activator / targeter dsRNA duplex.
[0129] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises at least two nucleotides (nt) (e.g., at least three or at least four nt) 5' of the 5'-most hairpin stem, e.g., as shown in the tracrRNA in Figures 6 and 7. In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises at least four nt 5' of the 5'-most hairpin stem, e.g., as shown in the tracrRNA in Figures 6 and 7.
[0130] In some cases, the activator-RNA (e.g., in a duplex or single-guide format) has a length of 65 nucleotides (nt) or more (e.g., 66 or more, 67 or more, 68 or more, 69 or more, 70 or more, or 75 or more nt). In some cases, the activator-RNA (e.g., in a duplex or single-guide format) has a length of 66 nt or more (e.g., 67 or more, 68 or more, 69 or more, 70 or more, or 75 or more nt). In some cases, the activator-RNA (e.g., in a duplex or single-guide format) has a length of 67 nt or more (e.g., 68 or more, 69 or more, 70 or more, or 75 or more nt).
[0131] In some cases, the activator-RNA (e.g., in a duplex or single-guide format) comprises 45 or more nucleotides (nt) (e.g., 46 or more, 47 or more, 48 or more, 49 or more, 50 or more, 51 or more, 52 or more, 53 or more, 54 or more, or 55 or more nt) 5' of the dsRNA duplex formed between the activator and targeter (activator / targeter dsRNA duplex). In some cases, the activator is truncated at the 5' end relative to a naturally occurring CasX activator. In some cases, the activator is extended at the 5' end relative to a naturally occurring CasX activator.
[0132] Examples of various Cas9 guide RNAs can be found in the art, and in some cases, similar variations to those introduced into Cas9 guide RNAs can also be introduced into the CasX guide RNAs of the present disclosure. For example, Jinek et al.,Science.2012 Aug 17;337(6096):816-21, Chylinski et al.,RNA Biol.2013 May;10(5):726-37, Ma et al.,Biomed Res Int.2013;2013:270805, Hou et al.,Proc Natl Acad Sci USA.2013 Sep 24;110(39):15644-9, Jinek et al.,Elife.2013;2:e00471, Pattanayak et al.,Nat Biotechnol.2013 Sep;31(9):839-43, Qi et al,Cell.2013 Feb 28;152(5):1173-83, Wang et al. al.,Cell.2013 May 9;153(4):910-8, Auer et.al., Genome Res.2013 Oct 31, Chen et.al., Nucleic Acids Res.2013 Nov 1;41(20):e19, Cheng et.al., Cell Res.2013 Oct;23(10):1163-71, Cho et.al., Genetics.2013 Nov;195(3):1177-80, DiCarlo et.al., Nucleic Acids Res.2013 Apr;41(7):4336-43, Dickinson et.al., Nat Methods.2013 Oct;10(10):1028-34, Ebina et.al., Sci Rep.2013;3:2510, Fujii et.al,Nucleic Acids Res.2013 Nov 1;41(20):e187, Hu et.al., Cell Res.2013 Nov;23(11):1322-5, Jiang et.al., Nucleic Acids Res.2013 Nov 1;41(20):e188, Larson et.al., Nat Protoc.2013 Nov;8(11):2180-96, Mali et.at.,Nat Methods.2013 Oct;10(10):957-63、Nakayama et.al.,Genesis.2013 Dec;51(12):835-43、Ran et.al.,Nat Protoc.2013 Nov;8(11):2281-308、Ran et.al.,Cell.2013 Sep 12;154(6):1380-9、Upadhyay et.al.,G3(Bethesda).2013 Dec 9;3(12):2233-8、Walsh et.al.,Proc Natl Acad Sci USA.2013 Sep 24;110(39):15514-5、Xie et.al.,Mol Plant.2013 Oct 9、Yang et.al.,Cell.2013 Sep 12;154(6):1370-9、Briner et al.,Mol Cell.2014 Oct 23;56(2):333-9, and U.S. Patents and Patent Applications Nos. 8,906,616, 8,895,308, 8,889,418, 8,889,356, 8,871,445, 8,865,406, 8,795,965, 8,771,945, 8,697,359, 20140068797, 20140170753, 20140179006, 20140179770, and 2014018684 No. 3, No. 20140186919, No. 20140186958, No. 20140189896, No. 20140227787, No. 20140234972, No. 20140242664, No. 20140242699, No. 20140242700, 20140242702, 20140248702, 20140256046, 20140273037, 20140273226, 20140273230, 20140 No. 273231, No. 20140273232, No. 20140273233, No. 20140273234, No. 20140273235, No. 20140287938, No. 20140295556, No. 201402955 No. 57, No. 20140298547, No. 20140304853, No. 20140309487, No. 20140310828, No. 20140310830, No. 20140315985, No. 20140335063, See Nos. 20140335620, 20140342456, 20140342457, 20140342458, 20140349400, 20140349405, 20140356867, 20140356956, 20140356958, 20140356959, 20140357523, 20140357530, 20140364333, and 20140377868, all of which are incorporated herein by reference in their entirety.
[0133] The terms "activator" or "activator RNA" are used herein to refer to a tracrRNA-like molecule (tracrRNA: "trans-acting CRISPR RNA") of a CasX dual guide RNA (and thus a CasX single guide RNA when the "activator" and "targeter" are linked together, e.g., by an intervening nucleotide). Thus, for example, a CasX guide RNA (dgRNA or sgRNA) comprises an activator sequence (e.g., a tracrRNA sequence). A tracr molecule (tracrRNA) is a naturally occurring molecule that hybridizes with a CRISPR RNA molecule (crRNA) to form a CasX dual guide RNA. The term "activator" is used herein to encompass not only naturally occurring tracrRNAs, but also tracrRNAs with modifications (e.g., truncations, elongations, sequence mutations, base modifications, backbone modifications, linkage modifications, etc.), where the activator retains at least one function of the tracrRNA (e.g., contributes to the dsRNA duplex to which the CasX protein binds). In some cases, the activator provides one or more stem loops that can interact with the CasX protein. The activator can be referred to as having a tracr sequence (tracrRNA sequence), and in some cases is a tracrRNA, although the term "activator" is not limited to naturally occurring tracrRNA.
[0134] In some cases (e.g., when the guide RNA is in a single-guide format), the activator-RNA is truncated (shorter) relative to the corresponding wild-type tracrRNA. In some cases (e.g., when the guide RNA is in a single-guide format), the activator-RNA is not truncated (shorter) relative to the corresponding wild-type tracrRNA. In some cases (e.g., when the guide RNA is in a single-guide format), the activator-RNA has a length of more than 50 nt (e.g., more than 55 nt, more than 60 nt, more than 65 nt, more than 70 nt, more than 75 nt, more than 80 nt). In some cases (e.g., when the guide RNA is in a single-guide format), the activator-RNA has a length of more than 80 nt. In some cases (e.g., when the guide RNA is in a single-guide format), the activator-RNA has a length in the range of 51-90 nt (e.g., 51-85, 51-84, 55-90, 55-85, 55-84, 60-90, 60-85, 60-84, 65-90, 65-85, 65-84, 70-90, 70-85, 70-84, 75-90, 75-85, 75-84, 80-90, 80-85, or 80-84 nt). In some cases (e.g., when the guide RNA is in a single-guide format), the activator-RNA has a length in the range of 80-90 nt.
[0135] The term "targeting factor" or "targeting factor RNA" is used herein to refer to a crRNA-like molecule (crRNA: "CRISPR RNA") of a CasX dual-guide RNA (and thus of a CasX single-guide RNA when the "activator" and "targeting factor" are linked together, e.g., by an intervening nucleotide). Thus, for example, a CasX guide RNA (dgRNA or sgRNA) comprises a guide sequence and a duplex-forming segment (e.g., a duplex-forming segment of a crRNA, which may also be referred to as a crRNA repeat). Because the sequence of the targeting segment of a targeting factor (the segment that hybridizes with the target sequence of a target nucleic acid) is modified by the user to hybridize with a desired target nucleic acid, the sequence of the targeting factor is often a non-naturally occurring sequence. However, the duplex-forming segment of a targeting factor (described in more detail herein) that hybridizes with the duplex-forming segment of an activator may comprise a naturally occurring sequence (e.g., may comprise the sequence of a duplex-forming segment of a naturally occurring crRNA, which may also be referred to as a crRNA repeat). For this reason, the term targeting agent is used herein to distinguish it from naturally occurring crRNA, despite the fact that portions of targeting agents (e.g., duplex-forming segments) often contain naturally occurring sequences derived from crRNA. However, the term "targeting agent" encompasses naturally occurring crRNA.
[0136] As described above, a targeting factor comprises both the guide sequence of the CasX guide RNA and a stretch of nucleotides (the "duplex-forming segment") that forms one half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA. The corresponding tracrRNA-like molecule (activator) comprises a stretch of nucleotides (the duplex-forming segment) that forms the other half of the dsRNA duplex of the protein-binding segment of the CasX guide RNA. In other words, the stretch of nucleotides of the targeting factor is complementary to the stretch of nucleotides of the activator and hybridizes therewith to form the dsRNA duplex of the protein-binding segment of the CasX guide RNA. As such, each targeting factor can be said to have a corresponding activator (having a region that hybridizes with the targeting factor). The targeting factor molecule additionally provides a guide sequence. Thus, the targeting factor and the activator (as a corresponding pair) hybridize to form the CasX guide RNA. The specific sequence of a given naturally occurring crRNA or tracrRNA molecule may be characteristic of the species in which the RNA molecule is found. Examples of suitable activators and targeting agents are provided herein.
[0137] Exemplary guide RNA sequences The guide RNAs shown in Figure 6 (dual guide format) and Figure 7 (dual guide format) are derived from the native CasX1 locus. For the sequences discussed in the following paragraphs and for the sequences described and tested in the Examples below, the tracrRNA and crRNA sequences were derived from the CasX1 locus. The same parameters and set of possible targeter-RNAs and activator-RNAs are expected and can be derived by comparing the sequence of the CasX1 locus with that of the CasX2 locus. For example, CasX1 tracrRNA sequences: UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 25) and UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 23). a, CasX2 tracrRNA sequence: UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 26) and UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 27).
[0138] For the CasX3 locus, tracr may lie within these 230 nt (region of complementarity is underlined).
[0139] [ka]
[0140] Similarly, the CasX1 crRNA sequence CCGAUAAGUAAAACGCAUCAAAGNNNNNNNNNNNNNNNNNNNN (SEQ ID NO: 11 without N, SEQ ID NO: 61 with N) can be compared to the CasX2 crRNA sequence UCUCCGAUAAAUAAGAAGCAUCAAAGNNNNNNNNNNNNNNNNNNNNNN (SEQ ID NO: 13 without N, SEQ ID NO: 69 with N).
[0141] The crRNA repeats from the CasX3 locus are GTTTACACACTCCCTCTCATAGGGT (SEQ ID NO: 54), GTTTACACACTCCCTCTCATGAGGT (SEQ ID NO: 55), TTTTACATACCCCCTCTCATGGGAT (SEQ ID NO: 56), and GTTTACACACTCCCTCTCATGGGGG (SEQ ID NO: 57). Thus, the crRNA sequence (e.g., from the CasX3 locus) may comprise GUUUACACACUCCCUCUCAUAGGGUNNNNNNNNNNNNNNNNNNNN (SEQ ID NO: 14 without N, SEQ ID NO: 31 with N), GUUUACACACUCCCUCUCAUGAGGUNNNNNNNNNNNNNNNNNNNN (SEQ ID NO: 15 without N, SEQ ID NO: 32 with N), UUUUACAUACCCCCUCUCAUGGGAUNNNNNNNNNNNNNNNNNNNN (SEQ ID NO: 16 without N, SEQ ID NO: 33 with N), and / or GUUUACACACUCCCUCUCAUGGGGGNNNNNNNNNNNNNNNNNNNN (SEQ ID NO: 17 without N, SEQ ID NO: 34 with N).
[0142] Exemplary targeting factor-RNA (e.g., crRNA) sequences In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to the guide sequence) the crRNA sequence CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 11) (see, e.g., the sgRNA in Figure 6, panel C). In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 11).
[0143] In some cases, the targeter-RNA comprises (e.g., in addition to the guide sequence) the crRNA sequence AUUUGAAGGUAUCUCCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 12). In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence AUUUGAAGGUAUCUCCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 12).
[0144] In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to the guide sequence) the crRNA sequence UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 13). In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 13).
[0145] In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to the guide sequence) the crRNA sequence GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 14). In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 14).
[0146] In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to the guide sequence) the crRNA sequence GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 15). In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 15).
[0147] In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to the guide sequence) the crRNA sequence UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 16). In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 16).
[0148] In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to the guide sequence) the crRNA sequence GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 17). In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 17).
[0149] In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises (e.g., in addition to a guide sequence) a crRNA sequence set forth in any one of SEQ ID NOs: 11 and 13. In some cases, the targeter-RNA (e.g., in dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the crRNA sequence set forth in any one of SEQ ID NOs: 11 and 13.
[0150] In some cases, the targeter-RNA (e.g., in addition to a guide sequence) comprises a crRNA sequence set forth in any one of SEQ ID NOs: 11-13. In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a crRNA sequence set forth in any one of SEQ ID NOs: 11-13.
[0151] In some cases, the targeter-RNA comprises (e.g., in addition to a guide sequence) a crRNA sequence set forth in any one of SEQ ID NOs: 14-17. In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a crRNA sequence set forth in any one of SEQ ID NOs: 14-17.
[0152] In some cases, the targeter-RNA comprises (e.g., in addition to a guide sequence) a crRNA sequence set forth in any one of SEQ ID NOs: 11-17. In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a crRNA sequence set forth in any one of SEQ ID NOs: 11-17.
[0153] Exemplary activator-RNA (e.g., tracrRNA) sequences In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 21). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 21).
[0154] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 22). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 22).
[0155] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 23) (see, e.g., the sgRNA in Figure 6). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 23).
[0156] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence AAGUAGUAAAUUACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 24) (see, e.g., sgRNA in Figure 6). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence AAGUAGUAAAUUACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 24).
[0157] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 25) (see, e.g., the sgRNA in Figure 6). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 25).
[0158] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 26). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 26).
[0159] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises the tracrRNA sequence UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 27). In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 27).
[0160] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) includes a tracrRNA sequence from within the following sequences:
[0161] [ka]
[0162] In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to a tracrRNA sequence from within the following sequences:
[0163] [ka]
[0164] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises a tracrRNA sequence set forth in any one of SEQ ID NOs: 21-27. In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to a tracrRNA sequence set forth in any one of SEQ ID NOs: 21-27.
[0165] In some cases, the activator-RNA (e.g., in a dual or single guide RNA format) comprises a tracrRNA sequence set forth in any one of SEQ ID NOs: 21-27. In some cases, the targeter-RNA (e.g., in a dual or single guide RNA format) comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to a tracrRNA sequence set forth in any one of SEQ ID NOs: 21-28.
[0166] In some cases, the CasX single guide RNA comprises the sequence UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGgaaaCCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 41). In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence UUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGgaaaCCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 41).
[0167] In some cases, the CasX single guide RNA comprises the sequence ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGAgaaaCCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 42). In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGAgaaaCCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 42).
[0168] In some cases, the CasX single guide RNA comprises the sequence UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGgaaaUCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 43). In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to the tracrRNA sequence UUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGgaaaUCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 43).
[0169] In some cases, the CasX single guide RNA comprises a sequence set forth in any one of SEQ ID NOs: 41-43. In some cases, the targeter-RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% identity) to a tracrRNA sequence set forth in any one of SEQ ID NOs: 41-43.
[0170] CASX system The present disclosure provides a CasX system. The CasX system of the present disclosure comprises: a) a CasX polypeptide and a CasX guide RNA of the present disclosure; b) a CasX polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; c) a CasX fusion polypeptide and a CasX guide RNA of the present disclosure; d) a CasX fusion polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; e) an mRNA encoding a CasX polypeptide and a CasX guide RNA of the present disclosure; f) an mRNA encoding a CasX polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; g) an mRNA encoding a CasX fusion polypeptide and a CasX guide RNA of the present disclosure; h) an mRNA encoding a CasX fusion polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; i) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide and a nucleotide sequence encoding a CasX guide RNA of the present disclosure; j) a nucleotide sequence encoding a CasX polypeptide, a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid of the present disclosure. k) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure and a nucleotide sequence encoding a CasX guide RNA; l) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, a nucleotide sequence encoding a CasX guide RNA, and a nucleotide sequence encoding a donor template nucleic acid; m) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA; n) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid; o) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA;p) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the disclosure, and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid; q) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the disclosure, a nucleotide sequence encoding a first CasX guide RNA, and a nucleotide sequence encoding a CasX guide RNA; or r) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the disclosure, a nucleotide sequence encoding a first CasX guide RNA, and a nucleotide sequence encoding a second CasX guide RNA; or some variant of one of (a) through (r);
[0171] nucleic acid The present disclosure provides one or more nucleic acids comprising one or more of a donor polynucleotide sequence, a nucleotide sequence encoding a CasX polypeptide (e.g., a wild-type CasX protein, a nickase CasX protein, a dCasX protein, a chimeric CasX protein, etc.), a CasX guide RNA, and a nucleotide sequence encoding a CasX guide RNA (which may comprise two separate nucleotide sequences in a dual-guide RNA format or a single nucleotide sequence in a single-guide RNA format). The present disclosure provides a nucleic acid comprising a nucleotide sequence encoding a CasX fusion polypeptide. The present disclosure provides a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide. The present disclosure provides a recombinant expression vector comprising a) a nucleotide sequence encoding a CasX polypeptide and b) a nucleotide sequence encoding CasX guide RNA(s). The present disclosure provides a recombinant expression vector comprising a) a nucleotide sequence encoding a CasX polypeptide and b) a nucleotide sequence encoding CasX guide RNA(s). In some cases, the nucleotide sequence encoding the CasX protein and / or the nucleotide sequence encoding the CasX guide RNA is operably linked to a promoter that is operable in a cell type of choice (e.g., a prokaryotic cell, a eukaryotic cell, a plant cell, an animal cell, a mammalian cell, a primate cell, a rodent cell, a human cell, etc.).
[0172] In some cases, nucleotide sequences encoding the CasX polypeptides of the present disclosure are codon-optimized. This type of optimization can involve mutating the nucleotide sequence encoding CasX to mimic the codon preferences of the intended host organism or cell while still encoding the same protein. Thus, the codons can be changed, but the encoded protein remains unchanged. For example, if the intended target cell was a human cell, a nucleotide sequence encoding CasX optimized for human codons could be used. As another non-limiting example, if the intended host cell was a mouse cell, a nucleotide sequence encoding CasX optimized for mouse codons could then be generated. As another non-limiting example, if the intended host cell was a plant cell, a nucleotide sequence encoding CasX optimized for plant codons could then be generated. As another non-limiting example, if the intended host cell was an insect cell, a nucleotide sequence encoding CasX optimized for insect codons could then be generated.
[0173] The present disclosure provides one or more recombinant expression vectors that include (in some cases in different recombinant expression vectors, and in some cases in the same recombinant expression vector): (i) a nucleotide sequence of a donor template nucleic acid, where the donor template includes a nucleotide sequence having homology to a target sequence of a target nucleic acid (e.g., a target genome), (ii) a nucleotide sequence encoding a CasX guide RNA (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell) that hybridizes to a target sequence at a target locus of the target genome (e.g., a single or dual guide RNA), and (iii) a nucleotide sequence encoding a CasX protein (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell). The present disclosure provides one or more recombinant expression vectors that include (in some cases in different recombinant expression vectors, and in some cases in the same recombinant expression vector): (i) a nucleotide sequence of a donor template nucleic acid (the donor template comprises a nucleotide sequence having homology to a target sequence of a target nucleic acid (e.g., a target genome)), and (ii) a nucleotide sequence encoding a CasX guide RNA (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell) that hybridizes to a target sequence at a target locus of the target genome (e.g., a single or dual guide RNA). The present disclosure provides one or more recombinant expression vectors that comprise (in some cases in different recombinant expression vectors, and in some cases in the same recombinant expression vector): (i) a nucleotide sequence encoding a CasX guide RNA (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell) that hybridizes to a target sequence at a target locus of the target genome (e.g., a single or dual guide RNA) and (ii) a nucleotide sequence encoding a CasX protein (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell).
[0174] Suitable expression vectors include viral expression vectors (e.g., vaccinia virus-based viral vectors, poliovirus, adenovirus (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:2543 2549, 1994; Borras et al., Gene Ther 6:515 524, 1999; Li and Davidson, PNAS 92:7700 7704, 1995; Sakamoto et al., H Gene Ther 5:1088 1097, 1999; WO94 / 12649; WO93 / 03769; WO93 / 19191; WO94 / 28938; WO95 / 11984; and WO95 / 00655); adeno-associated virus (AAV) (see, e.g., Ali et al., Hum Gene Ther 9:81 86,1998,Flannery et al.,PNAS 94:6916 6921,1997,Bennett et al.,Invest Opthalmol Vis Sci 38:2857 2863,1997,Jomary et al.,Gene Ther 4:683 690,1997,Rolling et al.,Hum Gene Ther 10:641 648,1999, Ali et al., Hum Mol Genet 5:591 594,1996, Srivastava in WO93 / 09239, Samulski et al., J. Vir. (1989) 63:3822-3828, Mendelson et al., Virol. (1988) 166:154-165, and Flotte et al. al., PNAS (1993) 90:10613-10617); SV40; herpes simplex virus; human immunodeficiency virus (see, for example, Miyoshi et al., PNAS 94:10319 23, 1997; Takahashi et al., J Virol 73:7812 7816, 1999); retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus). In some cases, the recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some cases, the recombinant expression vector of the present disclosure is a recombinant lentiviral vector. In some cases, the recombinant expression vector of the present disclosure is a recombinant retroviral vector.
[0175] Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements, including constitutive and inducible promoters, transcriptional enhancer elements, transcription terminators, etc., may be used in the expression vector.
[0176] In some embodiments, the nucleotide sequence encoding the CasX guide RNA is operably linked to a regulatory element, e.g., a transcriptional control element such as a promoter. In some embodiments, the nucleotide sequence encoding the CasX protein or CasX fusion polypeptide is operably linked to a regulatory element, e.g., a transcriptional repression element such as a promoter.
[0177] The transcriptional control element can be a promoter. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell-type-specific promoter. In some cases, the transcriptional control element (e.g., promoter) is functional in a target cell type or target cell population. For example, in some cases, the transcriptional control element can be functional in eukaryotic cells, such as hematopoietic stem cells (e.g., mobilized peripheral blood (mPB) CD34(+) cells, bone marrow (BM) CD34(+) cells, etc.).
[0178] Non-limiting examples of eukaryotic promoters (promoters functional in eukaryotic cells) include EF1α, those derived from immediate early cytomegalovirus (CMV), herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) derived from retroviruses, and mouse metallothionein-I. Selection of an appropriate vector and promoter is well within the level of ordinary skill in the art. Expression vectors can also contain a ribosome binding site for translation initiation and a transcription terminator. Expression vectors can also include appropriate sequences for amplifying expression. Expression vectors can also include a nucleotide sequence encoding a protein tag (e.g., a 6xHis tag, a hemagglutinin tag, a fluorescent protein, etc.) that can be fused to the CasX protein, thus resulting in a chimeric CasX polypeptide.
[0179] In some embodiments, the nucleotide sequence encoding the CasX guide RNA and / or CasX fusion polypeptide is operably linked to an inducible promoter. In some embodiments, the nucleotide sequence encoding the CasX guide RNA and / or CasX fusion protein is operably linked to a constitutive promoter.
[0180] The promoter may be a constitutively active promoter (i.e., a promoter that is constitutively active / "ON"), an inducible promoter (i.e., a promoter whose active / "ON" or inactive / "OFF" state is controlled by an external stimulus, e.g., a particular temperature, compound, or the presence of a protein), a spatially constrained promoter (i.e., a transcriptional control element, enhancer, etc.) (e.g., a tissue-specific promoter, cell-type-specific promoter, etc.), or a temporally constrained promoter (i.e., the promoter is "ON" or "OFF" during a particular stage of embryonic development or during a particular stage of a biological process, e.g., the hair follicle cycle in mice).
[0181] Suitable promoters may be derived from viruses and therefore may be referred to as viral promoters, or they may be derived from any organism, including prokaryotes or eukaryotes. Suitable promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Exemplary promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter; the adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter region (CMVIE), the Rous sarcoma virus (RSV) promoter, the human U6 micronucleus promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497-500 (2002)), the enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1; 31 (17)), the human H1 promoter (H1), and the like.
[0182] In some cases, the nucleotide sequence encoding the CasX guide RNA is operably linked to (under the control of) a promoter operable in eukaryotic cells (e.g., a U6 promoter, an enhanced U6 promoter, an H1 promoter, etc.). As will be understood by those skilled in the art, when expressing an RNA (e.g., a guide RNA) from a nucleic acid (e.g., an expression vector) using a U6 promoter (e.g., in a eukaryotic cell) or another Pol III promoter, the RNA may need to be mutated (encoding Us in the RNA) if several Ts are in a row. This is because a series of Ts in DNA (e.g., five Ts) can act as a terminator for polymerase III (pol III). Therefore, to ensure transcription of the guide RNA (e.g., the activator and / or targeter portions of a dual-guide or single-guide format) in eukaryotic cells, it may sometimes be necessary to modify the sequence encoding the guide RNA to remove the Ts. In some cases, the nucleotide sequence encoding a CasX protein (e.g., a wild-type CasX protein, a nickase CasX protein, a dCasX protein, a chimeric CasX protein, etc.) is operably linked to a promoter operable in eukaryotic cells (e.g., a CMV promoter, an EF1α promoter, an estrogen receptor-regulated promoter, etc.).
[0183] Examples of inducible promoters include, but are not limited to, T7 RNA polymerase promoter, T3 RNA polymerase promoter, isopropyl-β-D-thiogalactopyranoside (IPTG)-regulated promoter, lactose-inducible promoter, heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline; estrogen and / or estrogen analogs; IPTG; etc.
[0184] Inducible promoters suitable for use include any inducible promoter described herein or known to one of skill in the art. Examples of inducible promoters include, without limitation, chemically / biochemically regulated and physically regulated promoters such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)-responsive promoters and other tetracycline-responsive promoter systems (including tetracycline repressor protein (tetR), tetracycline operator sequence (tetO), and tetracycline transactivator fusion protein (tTA))), steroid-regulated promoters (e.g., promoters based on the rat glucocorticoid receptor, human estrogen receptor, and gaecdysone receptor, and promoters from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters derived from metallothionein (a protein that binds and sequesters metal ions) from yeast, mouse, and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene, or benzothiadiazole (BTH)), temperature / heat-inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light-responsive promoters from plant cells).
[0185] In some cases, the promoter is a spatially constrained promoter (i.e., a cell type-specific promoter, a tissue-specific promoter, etc.), such that in a multicellular organism, the promoter is active (i.e., "ON") in a specific subset of cells. A spatially constrained promoter may also be referred to as an enhancer, a transcriptional control element, a regulatory sequence, etc. Any convenient spatially constrained promoter can be used as long as the promoter is functional in the target host cell (e.g., eukaryotic, prokaryotic).
[0186] In some cases, the promoter is a reversible promoter. Suitable reversible promoters, including reversibly inducible promoters, are known in the art. Such reversible promoters can be isolated and derived from many organisms, such as eukaryotes and prokaryotes. The modification of a reversible promoter from a first organism (the first prokaryote and the second eukaryote, or the first eukaryote and the second prokaryote) for use in a second organism is well known in the art. Examples of such reversible promoters and systems based on such reversible promoters but also containing additional regulatory proteins include, but are not limited to, alcohol-regulated promoters (e.g., alcohol dehydrogenase I (alcA) gene promoter, promoters responsive to alcohol transactivator protein (AlcR), etc.), tetracycline-regulated promoters (e.g., promoter systems including TetActivators, TetON, TetOFF, etc.), steroid-regulated promoters (e.g., rat glucocorticoid receptor promoter system, human estrogen receptor promoter system, retinoid promoter system, thyroid promoter system, ecdysone promoter system, mifepristone promoter system, etc.), metal-regulated promoters (e.g., metallothionein promoter system, etc.), pathogen-associated regulated promoters (e.g., salicylic acid-regulated promoter, ethylene-regulated promoter, benzothiadiazole-regulated promoter, etc.), temperature-regulated promoters (e.g., heat shock-inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter, etc.), light-regulated promoters, synthetic inducible promoters, etc.
[0187] Methods for introducing nucleic acids (e.g., nucleic acids comprising donor polynucleotide sequences, one or more nucleic acids encoding a CasX protein and / or a CasX guide RNA, etc.) into host cells are known in the art, and any convenient method can be used to introduce a nucleic acid (e.g., an expression construct) into a cell. Suitable methods include, for example, viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc.
[0188] Introduction of the recombinant expression vector into cells can occur in any culture medium and under any culture conditions that promote cell survival. Introduction of the recombinant expression vector into target cells can be performed in vivo or ex vivo. Introduction of the recombinant expression vector into target cells can be performed in vitro.
[0189] In some embodiments, the CasX protein can be provided as RNA. The RNA can be provided by direct chemical synthesis or can be transcribed in vitro from DNA (e.g., encoding the CasX protein). Once synthesized, the RNA can be introduced into cells by any of the well-known techniques for introducing nucleic acids into cells (e.g., microinjection, electroporation, transfection, etc.).
[0190] Nucleic acids can be provided to cells using well-developed transfection techniques (see, e.g., Angel and Yanik (2010) PLoS ONE 5(7):e11756), and commercially available TransMessenger® reagents from Qiagen, Stemfect™ RNA transfection kits from Stemgent, and TransIT®-mRNA transfection kits from Mirus Bio LLC. See also Beumer et al. (2008) PNAS 105(50):19821-19826.
[0191] The vector can be directly provided to the target host cell. In other words, the cell is contacted with a vector containing the nucleic acid of interest (e.g., a recombinant expression vector having a donor template sequence and encoding a CasX guide RNA, a recombinant expression vector encoding a CasX protein, etc.), thereby the vector is taken up by the cell. Methods for contacting cells with a nucleic acid vector that is a plasmid include electroporation, calcium chloride transfection, microinjection, and lipofection, which are well known in the art. In the case of viral vector delivery, the cell can be contacted with a viral particle containing the viral expression vector of interest.
[0192] Retroviruses, such as lentiviruses, are suitable for use in the methods of the present disclosure. Commonly used retroviral vectors are "defective," i.e., unable to produce viral proteins necessary for productive infection. Rather, vector replication requires growth in a packaging cell line. To generate viral particles containing a nucleic acid of interest, the retroviral nucleic acid containing the nucleic acid is packaged into a viral capsid by a packaging cell line. Different packaging cell lines provide different envelope proteins (ecotropic, amphotropic, or xenotropic) that are incorporated into the capsid, and this envelope protein determines the specificity of the viral particle to the cell (ecotropic for mouse and rat, amphotropic for most mammalian cell types, including human, dog, and mouse, and xenotropic for most mammalian cells except mouse cells). Using an appropriate packaging cell line can ensure that cells are targeted by the packaged viral particles. Methods for introducing a target vector expression vector into a packaging cell line and harvesting viral particles produced by the packaging cell line are well known in the art. Nucleic acids can also be introduced by direct microinjection (eg, injection of RNA).
[0193] The vector used to provide the nucleic acid encoding the CasX guide RNA and / or CasX polypeptide to the target host cell can include a suitable promoter for driving expression of the nucleic acid of interest, i.e., transcriptional activation. In other words, in some cases, the nucleic acid of interest is operably linked to a promoter. This can include a ubiquitously acting promoter, such as the CMV-β-actin promoter, or an inducible promoter, a promoter that is active in specific cell populations or responds to the presence of a drug such as tetracycline. By transcriptional activation, it is intended to increase transcription by 10-fold, 100-fold, or more usually 1000-fold over basal levels in the target cell. In addition, the vector used to provide the nucleic acid encoding the CasX guide RNA and / or CasX protein to the cell can include a nucleic acid sequence encoding a selectable marker in the target cell to identify cells that have taken up the CasX guide RNA and / or CasX protein.
[0194] A nucleic acid containing a nucleotide sequence encoding a CasX polypeptide or a CasX fusion polypeptide may, in some cases, be RNA. Thus, a CasX fusion protein may be introduced into a cell as RNA. Methods for introducing RNA into a cell are known in the art and may include, for example, direct injection, transfection, or any other method used for DNA introduction. Alternatively, a CasX protein may be provided to a cell as a polypeptide. Such a polypeptide may optionally be fused to a polypeptide domain that increases the solubility of the product. This domain may be linked to the polypeptide through a defined protease cleavage site, e.g., a TEV sequence, that is cleaved by TEV protease. The linker may also contain one or more flexible sequences, e.g., 1 to 10 glycine residues. In some embodiments, cleavage of the fusion protein is performed in a buffer that maintains the solubility of the product, e.g., in the presence of 0.5 to 2 M urea, in the presence of a solubility-enhancing polypeptide and / or polynucleotide, etc. Domains of interest include endosomal destabilizing domains, such as influenza HA domains, and other polypeptides that aid in production, such as IF2 domains, GST domains, GRPE domains, etc. Polypeptides can be formulated for improved stability. For example, peptides can be PEGylated, where the polyethyleneoxy groups provide for improved longevity in the bloodstream.
[0195] Additionally or alternatively, the CasX polypeptides of the present disclosure can be fused to a polypeptide permeation domain to facilitate cellular uptake. Multiple permeation domains are known in the art and can be used in the non-integrated polypeptides of the present disclosure, including peptides, peptidomimetics, and non-peptide carriers. For example, a permeation peptide can be derived from the third alpha helix of the Drosophila melanogaster transcription factor Antennapedia, designated penetratin, which contains the amino acid sequence RQIKIWFQNRRMKWKK (SEQ ID NO: 133). As another example, a permeation peptide can contain the HIV-1 tat basic region amino acid sequence, which can include, for example, amino acids 49-57 of the naturally occurring tat protein. Other permeation domains include poly-arginine motifs, such as the region of amino acids 34-56 of the HIV-1 rev protein, nona-arginine, octa-arginine, etc. (See, e.g., Futaki et al. (2003) Curr Protein Pept Sci. 2003 Apr; 4(2):87-9 and 446, and Wender et al. (2000) Proc. Natl. Acad. Sci. USA 2000 Nov. 21; 97(24):13003-8, published U.S. patent applications 20030220334, 20030083256, 20030032593, and 20030022831, which are specifically incorporated by reference herein for their teachings of translocating peptides and peptoids.) The nona-arginine (R9) sequence is one of the more efficient PTDs characterized (Wender et al. 2000, Uemura et al. 2002). The site at which the fusion is made may be selected in order to optimize the biological activity, secretion, or binding characteristics of the polypeptide. Optimal sites can be determined by routine experimentation.
[0196] The CasX polypeptides of the present disclosure can be produced in vitro, or by eukaryotic cells, or by prokaryotic cells, and can be further treated by unfolding, e.g., heat denaturation, dithiothreitol reduction, etc., and further refolded using methods known in the art.
[0197] Modifications of interest that do not alter the primary sequence include chemical derivatization of the polypeptide, such as acylation, acetylation, carboxylation, amidation, etc. Also included are glycosylation modifications, e.g., those made by modifying the glycosylation pattern of a polypeptide during its synthesis and processing or during further processing steps, e.g., by exposing the polypeptide to enzymes that affect glycosylation, such as mammalian glycosylation or deglycosylation enzymes. Also encompassed are sequences having phosphorylated amino acid residues, e.g., phosphotyrosine, phosphoserine, or phosphothreonine.
[0198] Also suitable for inclusion in embodiments of the present disclosure are nucleic acids (e.g., encoding CasX guide RNAs, encoding CasX fusion proteins, etc.) and proteins (e.g., encoding CasX fusion proteins derived from wild-type or mutant proteins) modified using standard molecular biology techniques and synthetic chemistry to improve their resistance to proteolysis, alter target sequence specificity, optimize solubility properties, modify protein activity (e.g., transcriptional regulatory activity, enzymatic activity, etc.), or make them more suitable. Analogs of such polypeptides include those containing residues other than naturally occurring L-amino acids, such as D-amino acids or non-naturally occurring synthetic amino acids. D-amino acids can be substituted for some or all of the amino acid residues.
[0199] The CasX polypeptides of the present disclosure can be prepared by in vitro synthesis using convenient methods known in the art. Various commercial synthesizers are available, such as automated synthesizers from Applied Biosystems, Inc., Beckman, etc. By using a synthesizer, naturally occurring amino acids can be substituted with unnatural amino acids. The particular sequence and preparation method will be determined by convenience, economy, required purity, etc.
[0200] If desired, various groups can be introduced into the peptide during synthesis or expression, which allow for linkage to other molecules or surfaces: thus, cysteine can be used to create thioethers, histidine for linking to metal ion complexes, carboxyl groups to form amides or esters, amino groups to form amides, etc.
[0201] The CasX polypeptides of the present disclosure may also be isolated and purified according to conventional methods of recombinant synthesis. A lysate of the expression host may be prepared, and the lysate may be purified using liquid chromatography (HPLC), exclusion chromatography, gel electroporation, affinity chromatography, or other purification techniques. Typically, the composition used will represent 20% or more by weight of the desired product, more usually 75% or more by weight, preferably 95% or more by weight, and for therapeutic purposes, usually 99.5% by weight, with respect to contaminants associated with the method of preparation and purification of the product. Typically, percentages are based on total protein. Thus, in some cases, the CasX polypeptides or CasX fusion polypeptides of the present disclosure are at least 80% pure, at least 85% pure, at least 90% pure, at least 95% pure, at least 98% pure, or at least 99% pure (e.g., free of contaminants, non-CasX proteins, other macromolecules, etc.).
[0202] To induce cleavage or any desired modification in a target nucleic acid (e.g., genomic DNA) or to induce any desired modification in a polypeptide associated with a target nucleic acid, the CasX guide RNA and / or CasX polypeptide and / or donor template sequence of the present disclosure, whether introduced as a nucleic acid or as a polypeptide, is provided to a cell for about 30 minutes to about 24 hours, e.g., 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 12 hours, 16 hours, 18 hours, 20 hours, or any other period from about 30 minutes to about 24 hours, which may be repeated at a frequency of about every day to about every 4 days, e.g., every 1.5 days, every 2 days, every 3 days, or any other frequency from about every day to about every 4 days. The agent(s) may be provided to the cells of interest one or more times, e.g., one, two, three, or four or more times, and the cells are incubated with the agent(s) for some period of time following each contact event, e.g., 16-24 hours, after which the medium is replaced with fresh medium and the cells are further cultured.
[0203] In cases where two or more different targeting complexes (e.g., two different CasX guide RNAs that are complementary to different sequences within the same or different target nucleic acid) are provided to a cell, the complexes can be provided simultaneously (e.g., as two polypeptides and / or nucleic acids) or delivered simultaneously. Alternatively, they can be provided sequentially, e.g., a targeting complex is provided first, followed by a second targeting complex, or vice versa.
[0204] To improve delivery of DNA vectors to target cells, DNA can be protected from damage and its entry into cells, facilitated by, for example, the use of lipoplexes and polyplexes. For this reason, in some cases, nucleic acids of the present disclosure (e.g., recombinant expression vectors of the present disclosure) can be coated with lipids in organized structures such as micelles or liposomes. When the organized structures are complexed with DNA, they are called lipoplexes. Three types of lipids exist: anionic (negatively charged), neutral, and cationic (positively charged). Lipoplexes using cationic lipids have proven useful for gene transfer. Due to their positive charge, cationic lipids naturally complex with negatively charged DNA. Furthermore, they interact with cell membranes as a result of their charge. Lipoplex endocytosis then occurs, releasing the DNA into the cytoplasm. Cationic lipids also protect against cellular denaturation of the DNA.
[0205] Polymer-DNA complexes are called polyplexes. Most polyplexes are composed of cationic polymers, and their production is regulated by ionic interactions. One major difference between the way polyplexes and lipoplexes work is that polyplexes cannot release their DNA payload into the cytoplasm; therefore, cotransfection with an endosomolytic agent, such as an inactivated adenovirus, which dissolves endosomes created during endocytosis, must occur. However, this is not always the case. Polymers such as polyethyleneimine, as well as chitosan and trimethylchitosan, have their own methods of disrupting endosomes.
[0206] Dendrimers, highly branched macromolecules with spherical shapes, can also be used to genetically modify stem cells. The surface of dendrimer particles can be functionalized to alter their properties. Specifically, cationic dendrimers (i.e., those with a positive surface charge) can be constructed. In the presence of genetic material such as a DNA plasmid, charge complementarity leads to a transient association of the nucleic acid with the cationic dendrimer. Upon reaching its destination, the dendrimer-nucleic acid complex can be taken up by the cell by endocytosis.
[0207] In some cases, the disclosed nucleic acids (e.g., expression vectors) include an insertion site for a guide sequence of interest. For example, a nucleic acid can include an insertion site for a guide sequence of interest, where the insertion site is immediately adjacent to a nucleotide sequence encoding a portion of the CasX guide RNA that remains unchanged when the guide sequence is changed to hybridize to a desired target sequence (e.g., a sequence that contributes to the CasX-binding aspects of the guide RNA, e.g., a sequence that contributes to the dsRNA duplex(es) of the CasX guide RNA. This portion of the guide RNA can also be referred to as the "scaffold" or "constant region" of the guide RNA). Thus, in some cases, a nucleic acid of interest (e.g., an expression vector) includes a nucleotide sequence encoding a CasX guide RNA, except that the portion encoding the guide sequence portion of the guide RNA is the insertion sequence (insertion site). An insertion site is any nucleotide sequence used for insertion of a desired sequence. "Insertion sites" for use in various techniques are known to those of skill in the art, and any convenient insertion site can be used. The insertion site can be for any method for inserting a nucleic acid sequence. For example, in some cases, the insertion site is a multiple cloning site (MCS) (e.g., a site containing one or more restriction enzyme recognition sequences), a site for ligation-independent cloning, a site for recombination-based cloning (e.g., att site-based recombination), a nucleotide sequence recognized by CRISPR / Cas (e.g., Cas9)-based technologies, etc.
[0208] The insertion site can be any desired length and can depend on the type of insertion site (e.g., whether (and how many) the site includes one or more restriction enzyme recognition sequences, whether the site includes a target site for a CRISPR / Cas protein, etc.). In some cases, the insertion site of a nucleic acid of interest is 3 or more nucleotides (nt) in length (e.g., 5 or more, 8 or more, 10 or more, 15 or more, 17 or more, 18 or more, 19 or more, 20 or more, or 25 or more, or 30 or more nt in length). In some cases, the length of the insertion site of the nucleic acid of interest is in the range of 2 to 50 nucleotides (nt) (e.g., 2 to 40 nt, 2 to 30 nt, 2 to 25 nt, 2 to 20 nt, 5 to 50 nt, 5 to 40 nt, 5 to 30 nt, 5 to 25 nt, 5 to 20 nt, 10 to 50 nt, 10 to 40 nt, 10 to 30 nt, 10 to 25 nt, 10 to 20 nt, 17 to 50 nt, 17 to 40 nt, 17 to 30 nt, 17 to 25 nt). In some cases, the length of the insertion site of the nucleic acid of interest is in the range of 5 to 40 nt.
[0209] Nucleic acid modification In some embodiments, a nucleic acid of interest (e.g., a CasX guide RNA) has one or more modifications, such as base modifications, backbone modifications, etc., to provide the nucleic acid with new or enhanced characteristics (e.g., improved stability). A nucleoside is a base-sugar combination. The base portion of a nucleoside is typically a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide is a nucleoside that further includes a phosphate group covalently linked to the sugar portion of the nucleoside. In the case of nucleosides containing a pentofuranosyl sugar, the phosphate group can be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming an oligonucleotide, the phosphate group covalently links adjacent nucleosides to one another to form a linear polymeric compound. In turn, the respective ends of this linear polymeric compound can be further joined to form a circular compound, although linear compounds are preferred. In addition, linear compounds may have internal nucleotide base complementarity and therefore may fold in such a way as to produce fully or partially double-stranded compounds. Within oligonucleotides, the phosphate groups are generally considered to form the internucleoside backbone of the oligonucleotide. The usual linkage or backbone of RNA and DNA is the 3' to 5' phosphodiester linkage.
[0210] Suitable nucleic acid modifications include, but are not limited to, 2'O-methyl modified nucleotides, 2'fluoro modified nucleotides, locked nucleic acid (LNA) modified nucleotides, peptide nucleic acid (PNA) modified nucleotides, nucleotides with phosphorothioate linkages, and 5' caps (e.g., 7-methylguanylate caps (m7G)). Additional details and modifications are described below.
[0211] 2'-O-methyl modified nucleotides (also called 2'-O-methyl RNA) are naturally occurring modifications of RNA found in tRNA and other small RNAs that arise as post-transcriptional modifications. Oligonucleotides containing 2'-O-methyl RNA can be directly synthesized. This modification increases the Tm of RNA:RNA duplexes but results in only minor changes in RNA:DNA stability. They are stable against attack by single-stranded ribonucleases and are typically 5-10 times less susceptible to DNases than DNA. They are commonly used in antisense oligos as a means of increasing stability and binding affinity for target messages.
[0212] 2'-fluoro-modified nucleotides (e.g., 2'-fluoro bases) have a fluorine-modified ribose that increases binding affinity (Tm) and confers some relative nuclease resistance when compared to natural RNA. These modifications are commonly used in ribozymes and siRNAs to improve stability in serum or other biological fluids.
[0213] LNA bases have a modification to the ribose backbone that locks the base at the C3'-terminal position, favoring an RNA A-form helical duplex geometry. This modification significantly increases Tm and is highly nuclease-resistant. Multiple LNA insertions can be placed in oligos at any position except the 3'-terminus. Applications ranging from antisense oligos to hybridization probes for SNP detection and allele-specific PCR have been described. Due to the significant increase in Tm conferred by LNAs, they can also increase primer-dimer formation and self-hairpin formation. In some cases, the number of LNAs incorporated into a single oligo is 10 or fewer bases.
[0214] Phosphorothioate (PS) bonds (i.e., phosphorothioate linkages) substitute sulfur atoms for non-bridging oxygen atoms in the phosphate backbone of nucleic acids (e.g., oligos). This modification renders internucleotide linkages resistant to nuclease degradation. Phosphorothioate linkages can be introduced between the last 3-5 nucleotides at the 5' or 3' end of an oligo to inhibit exonuclease degradation. Inclusion of phosphorothioate linkages within an oligo (e.g., throughout the oligo) can also help reduce attack by endonucleases.
[0215] In some embodiments, the nucleic acid of interest has one or more nucleotides that are 2'-O-methyl modified nucleotides. In some embodiments, the nucleic acid of interest (e.g., dsRNA, siNA, etc.) has one or more 2' fluoro-modified nucleotides. In some embodiments, the nucleic acid of interest (e.g., dsRNA, siNA, etc.) has one or more LNA bases. In some embodiments, the nucleic acid of interest (e.g., dsRNA, siNA, etc.) has one or more nucleotides linked by phosphorothioate linkages (i.e., the nucleic acid of interest has one or more phosphorothioate linkages). In some embodiments, a nucleic acid of interest (e.g., dsRNA, siNA, etc.) has a 5' cap (e.g., a 7-methylguanylate cap (m7G)). In some embodiments, a nucleic acid of interest (e.g., dsRNA, siNA, etc.) has a combination of modified nucleotides. For example, a nucleic acid of interest (e.g., dsRNA, siNA, etc.) can have a 5' cap (e.g., a 7-methylguanylate cap (m7G)) in addition to having one or more nucleotides with other modifications (e.g., 2'-O-methyl nucleotides and / or 2' fluoro-modified nucleotides and / or LNA bases and / or phosphorothioate linkages).
[0216] Modified backbones and modified internucleoside linkages Examples of suitable nucleic acids (e.g., CasX guide RNAs) containing modifications include nucleic acids containing modified backbones or non-natural internucleoside linkages. Nucleic acids with modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.
[0217] Suitable modified oligonucleotide backbones containing phosphate atoms include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates (3'-alkylene phosphonates, 5'-alkylene phosphonates, and chiral phosphonates), phosphinates, phosphoramidates (including 3'-amino phosphoramidate and aminoalkyl phosphoramidate), phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates, which typically have 3'-5' linkages, their 2'-5' linked analogs, and those with inverted polarity in which one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages. Preferred oligonucleotides with reversed polarity contain a single 3'-3' linkage at the 3'-most internucleotide linkage, i.e., a single reversed nucleoside residue that may be basic (either lacking a nucleobase or having a hydroxyl group instead).Various salts (e.g., potassium or sodium), mixed salts, and free acid forms are also included.
[0218] In some embodiments, the subject nucleic acids contain one or more phosphorothioate and / or heteroatom internucleoside linkages, particularly -CH-NH-O-CH-, -CH-N(CH)-O-CH- (methylene (also known as methylimino or MMI backbone) or MMI backbone), -CH-ON(CH)-CH-, -CH-N(CH)-N(CH)-CH-, and -ON(CH)-CH-CH- (natural phosphodiester internucleotide linkages are represented as -OP(=O)(OH)-O-CH-). MMI-type internucleoside linkages are disclosed in the above-referenced U.S. Pat. No. 5,489,677, the disclosure of which is incorporated herein by reference in its entirety. Suitable amide internucleoside linkages are disclosed in U.S. Pat. No. 5,602,240, the disclosure of which is incorporated herein by reference in its entirety.
[0219] Also suitable are nucleic acids with morpholino backbone structures, as described in, for example, U.S. Patent No. 5,034,506.For example, in some embodiments, the nucleic acid of interest comprises a 6-membered morpholino ring instead of a ribose ring.In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages replace phosphodiester linkages.
[0220] Suitable modified polynucleotide backbones that do not contain phosphorus atoms have backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including those with morpholino linkages, siloxane backbones, sulfide, sulfoxide and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonic acid and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH moieties.
[0221] Mimics The nucleic acid of interest may be a nucleic acid mimic. When applied to polynucleotides, the term "mimetic" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring is also referred to in the art as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety is maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid, a polynucleotide mimic that has been shown to have excellent hybridization properties, is called a peptide nucleic acid (PNA). In PNAs, the sugar backbone of a polynucleotide is replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. Nucleotides are retained and are linked directly or indirectly to the aza nitrogen atoms of the amide portion of the backbone.
[0222] One polynucleotide mimic that has been reported to have excellent hybridization properties is peptide nucleic acid (PNA). The backbone in PNA compounds is two or more linked aminoethylglycine units, giving PNA an amide-containing backbone. Heterocyclic base moieties are directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone. Representative U.S. patents that describe the preparation of PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082, 5,714,331, and 5,719,262 (the disclosures of which are incorporated herein by reference in their entireties).
[0223] Another class of polynucleotide mimics that has been studied is based on linked morpholino units (morpholino nucleic acids) with a heterocyclic base attached to the morpholino ring. Several linking groups have been reported to link the morpholino monomer units in morpholino nucleic acids. One class of linking group was selected to obtain nonionic oligomeric compounds. Nonionic morpholino-based oligomeric compounds are less likely to have undesired interactions with cellular proteins. Morpholino-based polynucleotides are nonionic mimics of oligonucleotides that are less likely to form undesired interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholino-based polynucleotides are disclosed in U.S. Patent No. 5,034,506, the disclosure of which is incorporated herein by reference in its entirety. Various compounds within the morpholino class of polynucleotides have been prepared with a variety of different linking groups joining the monomer subunits.
[0224] Another class of polynucleotide mimics is called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in DNA / RNA molecules is replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers have been prepared and used for oligomeric compound synthesis using classical phosphoramidite chemistry. Fully modified CeNA oligomeric compounds and oligonucleotides with specific CeNA modifications have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, 122, 8595-8602, the disclosure of which is incorporated herein by reference in its entirety). In general, the incorporation of CeNA monomers into DNA strands increases the stability of DNA / RNA hybrids. CeNA oligoadenylates formed complexes with RNA and DNA complements with similar stability to the native complexes. Studies incorporating the CeNA structure into native nucleic acid structures have demonstrated facile conformational adaptation, as demonstrated by NMR and circular dichroism.
[0225] A further modification includes locked nucleic acids (LNAs), in which the 2'-hydroxyl group is attached to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2 (Singh et al., Chem. Commun., 1998, 4, 455-456, the disclosure of which is incorporated herein by reference in its entirety). LNAs and LNA analogs exhibit very high duplex thermal stability with complementary DNA and RNA (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility properties. Potent and non-toxic antisense oligonucleotides containing LNAs have been described (see, e.g., Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 2000, 97, 5633-5638, the disclosure of which is incorporated herein by reference in its entirety).
[0226] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine, and uracil, along with their oligomerization and nucleic acid recognition properties, have been described (e.g., Koshkin et al., Tetrahedron, 1998, 54, 3607-3630, the disclosure of which is incorporated herein by reference in its entirety). LNAs and their preparation are also described in WO98 / 39352 and WO99 / 14226, and U.S. Application Nos. 20120165514, 20100216983, 20090041809, 20060117410, 20040014959, 20020094555, and 20020086998, the disclosures of which are incorporated herein by reference in their entireties.
[0227] Modified sugar moiety The subject nucleic acids may also contain one or more substituted sugar moieties. Suitable polynucleotides contain sugar substituents selected from OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl, and alkynyl are substituted or unsubstituted C to C. 10 Alkyl or C2-C 10 It may be alkenyl or alkynyl. O((CH2) n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON(CH2) n CH3)2 is particularly preferred, where n and m are from 1 to about 10. Other preferred polynucleotides include C1 to C 10Sugar substituents include lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH, OCN, Cl, Br, CN, CF, OCF, SOCH, SOCH, ONO, NO, N, NH, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, groups for improving the pharmacokinetic properties of oligonucleotides, or groups for improving the pharmacological properties of oligonucleotides, and other substituents with similar properties. Suitable modifications include 2'-methoxyethoxy (2'-O-CHCHOCH, also known as 2'-O-(2-methoxyethyl) or 2'-MOE) (Martin et al., Helv. Chim. Acta, 1995, 78, 486-504, the disclosure of which is incorporated herein by reference in its entirety), i.e., alkoxyalkoxy groups. Further suitable modifications include the 2'-dimethylaminooxyethoxy, i.e., O(CH2)2ON(CH3)2 group (also known as 2'-DMAOE), described in the Examples below, and 2'-dimethylaminoethoxyethoxy (known in the art as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2.
[0228] Other suitable sugar substituents include methoxy (-O-CH), aminopropoxy (-OCHCHCHNH), allyl (-CH-CH=CH), -O-allyl (-O-CH-CH=CH), and fluoro (F). The 2'-sugar substituent may be at the arabino (up) or ribo (down) position. A preferred 2'-arabino modification is 2'-F. Similar modifications may be made at other positions on the oligomeric compound, particularly the 3' position of the sugar on the 3'-terminal nucleoside or in 2'-5'-linked oligonucleotides, and the 5' position of the 5'-terminal nucleotide. Oligomeric compounds may also have sugar mimetics, such as cyclobutyl moieties, in place of the pentofuranosyl sugar.
[0229] Base modifications and substitutions The subject nucleic acids may also contain modifications or substitutions of nucleobases (often simply referred to in the art as "bases"). As used herein, "unmodified" or "natural" nucleobases include the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include 5-hydroxymethylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine and These include thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine. Further modified nucleobases include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).
[0230] Heterocyclic base moieties can also include those in which the purine or pyrimidine base is replaced with other heterocycles, such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone. Additional nucleobases include those disclosed in United States Patent No. 3,687,808, those disclosed in The Concise Encyclopedia of Polymer Science and Engineering, pages 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, those disclosed by Englisch et al., Angewandte Chemie, International Edition, 1991, 30,613, and those disclosed by Sanghvi, YS, Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, ST and Lebleu, B., ed., CRC Press, 1993 (the disclosures of which are incorporated herein by reference in their entirety).Some of these nucleobases are useful for increasing the binding affinity of oligomeric compounds. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, such as 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-Methylcytosine substitutions have been shown to increase nucleic acid duplex stability by 0.6 to 1.2°C (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278, the disclosure of which is incorporated herein by reference in its entirety), and are preferred base substitutions, for example, when combined with 2'-O-methoxyethyl sugar modifications.
[0231] conjugate Another possible modification of a nucleic acid of interest involves chemically linking one or more moieties or conjugates to the polynucleotide that improve the activity, cellular distribution, or cellular uptake of the oligonucleotide. Such moieties or conjugates can include conjugate groups covalently attached to functional groups such as primary or secondary hydroxyl groups. Conjugate groups include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that improve the pharmacodynamic properties of oligomers, and groups that improve the pharmacokinetic properties of oligomers. Suitable conjugate groups include cholesterol, lipids, phospholipids, biotin, phenazine, folate, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that improve pharmacodynamic properties include groups that improve uptake, improve resistance to degradation, and / or enhance sequence-specific hybridization with the target nucleic acid. Groups that improve pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of the nucleic acid of interest.
[0232] Conjugate moieties include cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, 6553-6556), cholic acid (Manoharan et al., Bioorg. Med. Chem. Let., 1994, 4, 1053-1060), thioethers such as hexyl-S-tritylthiol (Manoharan et al., Ann. NY Acad. Sci., 1992, 660, 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, 2765-2770), and thiocholesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, 533-538), aliphatic chains such as dodecanediol or undecyl residues (Saison-Behmoaras et al., EMBO J., 1991, 10, 1111-1118; Kabanov et al., FEBS Lett., 1990, 259, 327-330; Svinarchuk et al., Biochimie, 1993, 75, 49-54), phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654; Shea et al., Nucl. Acids Res., 1990, 18, 3777-3783), polyamine or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, 969-973), or adamantane acetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654), palmityl moiety (Mishra et al., Biochim. Biophys. Acta, 1995, 1264, 229-237), or octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al., J. Pharmacol. Exp. Ther.Lipid moieties include, but are not limited to, lipid moieties such as those listed above (Beckmann, 1996, 277, 923-937).
[0233] The conjugate may include a "protein transduction domain" or PTD (also known as a CPP - cell penetrating peptide), which may refer to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates crossing of a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule, which may range from a small polar molecule to a large macromolecule, facilitates membrane crossing of the molecule, for example, moving from the extracellular space to the cytosol or from the cytosol into an organelle (e.g., the nucleus). In some embodiments, the PTD is covalently attached to the 3' end of the exogenous polynucleotide. In some embodiments, the PTD is covalently attached to the 5' end of the exogenous polynucleotide. Exemplary PTDs include a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT, containing YGRKKRRQRRR (SEQ ID NO: 112)); a polyarginine sequence containing multiple arginines sufficient for direct entry into cells (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines); a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); a Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO: 113); transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 114); KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 115); and RQIKIWFQNRRMKWKK (SEQ ID NO: 116). Exemplary PTDs include, but are not limited to, YGRKKRRQRRR (SEQ ID NO: 117), RKKRRQRRR (SEQ ID NO: 118); arginine homopolymers of 3 to 50 arginine residues.Exemplary PTD domain amino acid sequences include, but are not limited to, any of the following: YGRKKRRQRRR (SEQ ID NO: 119), RKKRRQRR (SEQ ID NO: 120), YARAAARQARA (SEQ ID NO: 121), THRLPRRRRRR (SEQ ID NO: 122), and GGRRARRRRRR (SEQ ID NO: 123). In some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June;1(5-6):371-381). ACPPs contain a polycationic CPP (e.g., Arg9 or "R9") connected to a matching polyanion (e.g., Glu9 or "E9") via a cleavable linker, which reduces the net charge to near zero, thereby inhibiting cellular attachment and uptake. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesive properties, thus "activating" the ACPP to cross the membrane.
[0234] Introducing components into target cells A CasX guide RNA (or a nucleic acid comprising a nucleotide sequence encoding same) and / or a CasX polypeptide (or a nucleic acid comprising a nucleotide sequence encoding same) and / or a CasX fusion polypeptide (or a nucleic acid comprising a nucleotide sequence encoding a CasX fusion polypeptide) and / or a donor polynucleotide (donor template) of the present disclosure can be introduced into a host cell by any of a variety of well-known methods.
[0235] Any of a variety of compounds and methods can be used to deliver the CasX system of the present disclosure to a target cell (e.g., a CasX system can be delivered using a recombinant expression vector comprising: a) a CasX polypeptide and a CasX guide RNA of the present disclosure; b) a CasX polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; c) a CasX fusion polypeptide and a CasX guide RNA of the present disclosure; d) a CasX fusion polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; e) an mRNA encoding a CasX polypeptide of the present disclosure and a CasX guide RNA; f) an mRNA encoding a CasX polypeptide of the present disclosure, a CasX guide RNA, and a donor template nucleic acid; g) an mRNA encoding a CasX fusion polypeptide of the present disclosure and a CasX guide RNA; h) an mRNA encoding a CasX fusion polypeptide of the present disclosure, a CasX guide RNA, and a donor template nucleic acid; i) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a nucleotide sequence encoding a CasX guide RNA). j) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, a nucleotide sequence encoding a CasX guide RNA, and a nucleotide sequence encoding a donor template nucleic acid; k) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure and a nucleotide sequence encoding a CasX guide RNA; l) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, a nucleotide sequence encoding a CasX guide RNA, and a nucleotide sequence encoding a donor template nucleic acid; m) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA; n) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid;o) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA; p) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid; q) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, a nucleotide sequence encoding a first CasX guide RNA, and a nucleotide sequence encoding a CasX guide RNA; or r) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, a nucleotide sequence encoding a first CasX guide RNA, and a nucleotide sequence encoding a second CasX guide RNA; or some variant of one of (a) through (r). As a non-limiting example, the CasX system of the present disclosure can be combined with a lipid. As another non-limiting example, the CasX system of the present disclosure can be combined with or formulated into a particle.
[0236] Methods for introducing nucleic acids into host cells are known in the art, and any convenient method can be used to introduce a nucleic acid of interest (e.g., an expression construct / vector) into target cells (e.g., prokaryotic cells, eukaryotic cells, plant cells, animal cells, mammalian cells, human cells, etc.). Suitable methods include, for example, viral infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (e.g., Panyam et., al Adv Drug Deliv Rev. 2012 Sep 13. pii:S0169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023), and the like.
[0237] In some cases, a CasX polypeptide of the present disclosure is provided as a nucleic acid (e.g., mRNA, DNA, plasmid, expression vector, viral vector, etc.) encoding the CasX polypeptide. In some cases, a CasX polypeptide of the present disclosure is provided directly as a protein (e.g., without or with an associated guide RNA, i.e., as a ribonucleoprotein complex). A CasX polypeptide of the present disclosure can be introduced into (provided to) a cell by any convenient method, and such methods are known to those of skill in the art. As an illustrative example, a CasX polypeptide of the present disclosure can be directly injected into a cell (e.g., with or without a CasX guide RNA or a nucleic acid encoding a CasX guide RNA, and with or without a donor polynucleotide). As another example, a preformed complex of a CasX polypeptide of the present disclosure and a CasX guide RNA (RNP) can be introduced into a cell (e.g., a eukaryotic cell) (e.g., via injection, via nucleofection, conjugated to one or more components, e.g., conjugated to a CasX protein, conjugated to a guide RNA, via a protein transduction domain (PTD) conjugated to a CasX polypeptide of the present disclosure and a guide RNA, etc.).
[0238] In some cases, a CasX fusion polypeptide of the present disclosure (e.g., dCasX fused to a fusion partner, nickase-CasX fused to a fusion partner, etc.) is provided as a nucleic acid (e.g., mRNA, DNA, plasmid, expression vector, viral vector, etc.) encoding the CasX fusion polypeptide. In some cases, a CasX fusion polypeptide of the present disclosure is provided directly as a protein (e.g., without or with an associated guide RNA, i.e., as a ribonucleoprotein complex). A CasX fusion polypeptide of the present disclosure can be introduced into (provided to) a cell by any convenient method, and such methods are known to those of skill in the art. As an illustrative example, a CasX fusion polypeptide of the present disclosure can be directly injected into a cell (e.g., with or without a nucleic acid encoding a CasX guide RNA and with or without a donor polynucleotide). As another example, a preformed complex of a CasX fusion polypeptide of the present disclosure and a CasX guide RNA (RNP) can be introduced into a cell (e.g., via injection, via nucleofection, conjugated to one or more components, e.g., conjugated to a CasX protein, conjugated to a guide RNA, via a protein transduction domain (PTD) conjugated to a CasX polypeptide of the present disclosure and a guide RNA, etc.).
[0239] In some cases, nucleic acids (e.g., CasX guide RNAs, nucleic acids comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, etc.) are delivered to cells (e.g., target host cells) and / or polypeptides (e.g., CasX polypeptides, CasX fusion polypeptides) in particles or associated with particles. In some cases, CasX systems of the present disclosure are delivered to cells in particles or associated with particles. The terms "particle" and "nanoparticle" may be used interchangeably, as appropriate. A recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide and / or CasX guide RNA of the present disclosure, an mRNA comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, and a guide RNA may be delivered simultaneously using a particle or lipid envelope. For example, the CasX polypeptide and CasX guide RNA, e.g., as a complex (e.g., as a ribonucleoprotein (RNP) complex), can be delivered via a particle, e.g., a delivery particle comprising a lipid or lipidoid and a hydrophilic polymer, e.g., a cationic lipid and a hydrophilic polymer, e.g., the cationic lipid comprises 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP) or 1,2-ditetradecanoyl-sn-glycero-3-phosphocholine (DMPC), and / or the hydrophilic polymer comprises ethylene glycol or polyethylene glycol (PEG), and / or the particle further comprises cholesterol (e.g., particles from Formulation 1 = DOTAP 100, DMPC 0, PEG 0, cholesterol 0; Formulation No. 2 = DOTAP 90, DMPC 0, PEG 10, cholesterol 0; Formulation No. 3 = DOTAP 90, DMPC 0, PEG 5, cholesterol 5).For example, particles can be formed using a multi-step process in which the CasX polypeptide and CasX guide RNA are mixed together, e.g., in a 1:1 molar ratio, e.g., at room temperature for, e.g., 30 minutes, in, e.g., sterile, nuclease-free 1x phosphate buffered saline (PBS); separately, DOTAP, DMPC, PEG, and cholesterol, if required for the formulation, are dissolved in alcohol, e.g., 100% ethanol, and the two solutions are mixed together to form particles containing the complexes).
[0240] The CasX polypeptide of the present disclosure (or mRNA comprising a nucleotide sequence encoding the CasX polypeptide of the present disclosure, or a recombinant expression vector comprising a nucleotide sequence encoding the CasX polypeptide of the present disclosure) and / or CasX guide RNA (or nucleic acids, such as one or more expression vectors encoding CasX guide RNAs) can be co-delivered using particles or lipid envelopes. For example, biodegradable core-shell nanoparticles having a poly(β-amino ester) (PBAE) core encapsulated by a phospholipid bilayer shell can be used. In some cases, particles / nanoparticles based on self-assembling bioadhesive polymers are used, and such particles / nanoparticles can be applied, for example, to oral delivery of peptides to the brain, intravenous delivery of peptides, and intranasal delivery of peptides. Other embodiments, such as oral absorption and intraocular delivery of hydrophobic drugs, are also contemplated. Molecular envelope technology, which requires an engineered polymer envelope to be protected and delivered to the site of disease, can be used. A dose of approximately 5 mg / kg can be used in single or multiple doses, depending on various factors, such as the target tissue.
[0241] Lipidoid compounds (e.g., those described in U.S. Patent Application No. 20110293703) are also useful in administering polynucleotides and can be used to deliver a CasX polypeptide of the present disclosure, a CasX fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure, or a CasX system of the present disclosure (e.g., a CasX system comprising: a) a CasX polypeptide and a CasX guide RNA of the present disclosure; b) a CasX polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; c) a CasX fusion polypeptide and a CasX guide RNA of the present disclosure; d) a CasX fusion polypeptide, a CasX guide RNA, and a donor template nucleic acid of the present disclosure; e) an mRNA encoding a CasX polypeptide of the present disclosure and a CasX guide RNA; f) an mRNA, a CasX guide RNA, and a donor template nucleic acid encoding a CasX polypeptide of the present disclosure; g) an mRNA encoding a CasX fusion polypeptide of the present disclosure and a CasX guide RNA; h) an RNP encoding a CasX fusion polypeptide of the present disclosure. i) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a nucleotide sequence encoding a CasX guide RNA; j) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, a nucleotide sequence encoding a CasX guide RNA, and a nucleotide sequence encoding a donor template nucleic acid; k) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure and a nucleotide sequence encoding a CasX guide RNA; l) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, a nucleotide sequence encoding a CasX guide RNA, and a nucleotide sequence encoding a donor template nucleic acid; m) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA;n) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid; o) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA; p) a first recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, and a second recombinant expression vector comprising a nucleotide sequence encoding a CasX guide RNA, and a donor template nucleic acid; q) a recombinant expression vector comprising a nucleotide sequence encoding a CasX polypeptide of the present disclosure, a nucleotide sequence encoding a first CasX guide RNA, and a nucleotide sequence encoding a CasX guide RNA; or r) a recombinant expression vector comprising a nucleotide sequence encoding a CasX fusion polypeptide of the present disclosure, a nucleotide sequence encoding a first CasX guide RNA, and a nucleotide sequence encoding a second CasX guide RNA; or some variant of one of (a) through (r). In one embodiment, the amino alcohol lipidoid compound is combined with a drug to be delivered to a cell or subject to form a microparticle, nanoparticle, liposome, or micelle. The amino alcohol lipidoid compound can be combined with other amino alcohol lipidoid compounds, polymers (synthetic or natural), surfactants, cholesterol, carbohydrates, lipids, etc. to form particles. These particles can then be optionally combined with pharmaceutical excipients to form pharmaceutical compositions.
[0242] Poly(β-amino alcohols) (PBAAs) can be used to deliver a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell. U.S. Patent Publication No. 20130302401 relates to a class of poly(β-amino alcohols) (PBAAs) prepared using combinatorial polymerization.
[0243] Sugar-based particles, e.g., GalNAc, can be used as described with reference to WO2014118272 (incorporated herein by reference) and Nair, JK et al., 2014, Journal of the American Chemical Society 136(49), 16958-16961), to deliver a CasX polypeptide of the present disclosure, a CasX fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure, or a CasX system of the present disclosure to a target cell.
[0244] In some cases, lipid nanoparticles (LNPs) are used to deliver the disclosed CasX polypeptides, disclosed CasX fusion polypeptides, disclosed RNPs, disclosed nucleic acids, or disclosed CasX systems to target cells. Negatively charged polymers such as RNA can be loaded into LNPs at low pH (e.g., pH 4), and ionized lipids exhibit positive charges. However, at physiological pH, LNPs exhibit low surface charge, which corresponds to longer circulation times. The focus is on four ionizable cationic lipids: 1,2-dilineoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxy-keto-N,N-dimethyl-3-aminopropane (DLinKDMA), and 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA). Preparation of LNPs is described, for example, in Rosin et al. (2011) Molecular Therapy 19:1286-2200. The cationic lipids 1,2-dilineoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA), (3-o-[2″-(methoxypolyethylene glycol 2000)sucrineoyl]-1,2-dimyristoyl-sn-glycol (PEG-S-DMG), and R-3- [(.omega.-methoxy-poly(ethylene glycol)2000)carbamoyl]-1,2-dimyristoyloxylpropyl-3-amine (PEG-C-DOMG) can be used. Nucleic acids (e.g., CasX guide RNAs, nucleic acids of the present disclosure, etc.) can be encapsulated in LNPs containing DLinDAP, DLinDMA, DLinK-DMA, and DLinKC2-DMA (cationic lipid:DSPC:CHOL:PEGS-DMG or PEG-C-DOMG in a 40:10:40:10 molar ratio). In some cases, 0.2% SP-DiOC18 is incorporated.
[0245] Spherical nucleic acid (SNA™) constructs and other nanoparticles (particularly gold nanoparticles) can be used to deliver a CasX polypeptide of the present disclosure, a CasX fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure, or a CasX system of the present disclosure to a target cell. See, for example, Cutler et al., J. Am. Chem. Soc. 2011 133:9254-9257; Hao et al., Small. 2011 7:3158-3162; Zhang et al., ACS Nano. 2011 5:6962-6970; Cutler et al., J. Am. Chem. Soc. 2012 134:1376-1391; Young et al., Nano Lett. 2012 12:3867-71; Zheng et al. al.,Proc.Natl.Acad.Sci.USA.2012 109:11975-80, Mirkin, Nanomedicine 20127:635-638, Zhang et al.,J.Am.Chem.Soc.2012 134:16488-1691, Weintraub,Nature 2013 495:S14-S16, Choi et al., Proc. Natl. Acad. Sci. USA.2013 110(19):7625-7630, Jensen et al., Sci.
[0246] Self-assembled nanoparticles bearing RNA can be constructed from polyethyleneimine (PEI) that is PEGylated with Arg-Gly-Asp (RGD) attached to the distal end of the polyethylene glycol (PEG).
[0247] Generally, "nanoparticle" refers to any particle having a diameter of less than 1000 nm. In some cases, nanoparticles suitable for use in delivering a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell have a diameter of 500 nm or less, e.g., 25 nm to 35 nm, 35 nm to 50 nm, 50 nm to 75 nm, 75 nm to 100 nm, 100 nm to 150 nm, 150 nm to 200 nm, 200 nm to 300 nm, 300 nm to 400 nm, or 400 nm to 500 nm. In some cases, nanoparticles suitable for use in delivering a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell have a diameter of 25 nm to 200 nm. In some cases, nanoparticles suitable for use in delivering a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell have a diameter of 100 nm or less. In some cases, nanoparticles suitable for use in delivering a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell have a diameter of 35 nm to 60 nm.
[0248] Nanoparticles suitable for use in delivering a CasX polypeptide of the present disclosure, a CasX fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure, or a CasX system of the present disclosure to target cells can be provided in different forms, such as solid nanoparticles (e.g., metals such as silver, gold, iron, and titanium), non-metallic, lipid-based solids, polymers, suspensions of nanoparticles, or combinations thereof. Metallic, insulating, and semiconductor nanoparticles, as well as hybrid structures (e.g., core-shell nanoparticles), can be prepared. Nanoparticles made of semiconductor materials can also be labeled quantum dots if they are small enough (typically 10 nm or less) for quantization of electronic energy levels to occur. Such nanoscale particles are used as drug carriers or imaging agents in biomedical applications and can be adapted for similar purposes in the present disclosure.
[0249] Semi-solid and soft nanoparticles are also suitable for use in delivering a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell. A prototypical nanoparticle of semi-solid nature is the liposome.
[0250] In some cases, exosomes are used to deliver a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to target cells. Exosomes are endogenous nanovesicle particles that can transport RNA and proteins and deliver RNA to the brain and other target organs.
[0251] In some cases, liposomes are used to deliver the disclosed CasX polypeptides, the disclosed CasX fusion polypeptides, the disclosed RNPs, the disclosed nucleic acids, or the disclosed CasX systems to target cells. Liposomes are spherical vesicular structures composed of a unilamellar or multilamellar lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes can be made from several different types of lipids, but phospholipids are most commonly used to generate liposomes. Liposome formation occurs spontaneously when a lipid membrane is mixed with an aqueous solution, but it can also be promoted by applying force in the form of shaking using a homogenizer, sonicator, or extrusion device. Several other additives may be added to liposomes to modify their structure and properties. For example, either cholesterol or sphingomyelin may be added to the liposome mixture to help stabilize the liposome structure and prevent leakage of the liposome's internal cargo. Liposomal formulations may be composed primarily of lipids such as natural phospholipids and 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, and monosialogangliosides.
[0252] Stable nucleic acid-lipid particles (SNALPs) can be used to deliver the disclosed CasX polypeptides, disclosed CasX fusion polypeptides, disclosed RNPs, disclosed nucleic acids, or disclosed CasX systems to target cells. The SNALP formulation can contain the lipids 3-N-[(methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristoyloxy-propylamine (PEG-C-DMA), 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), and cholesterol in a 2:40:10:48 molar ratio. SNALP liposomes can be prepared by combining D-Lin-DMA and PEG-C-DMA with distearoylphosphatidylcholine (DSPC), cholesterol, and siRNA using a lipid / siRNA ratio of 25:1 and a 48 / 40 / 10 / 2 molar ratio of cholesterol / D-Lin-DMA / DSPC / PEG-C-DMA. The resulting SNALP liposomes can be approximately 80-100 nm in size. SNALP can contain synthetic cholesterol (Sigma-Aldrich, St. Louis, MO, USA), dipalmitoylphosphatidylcholine (Avanti Polar Lipids, Alabaster, Ala., USA), 3-N-[(w-methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristoyloxypropylamine, and cationic 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane. SNALPs can include synthetic cholesterol (Sigma-Aldrich), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC, Avanti Polar Lipids Inc.), PEG-cDMA, and 1,2-dilinoleyloxy-3-(N;N-dimethyl)aminopropane (DLinDMA).
[0253] Other cationic lipids, such as the amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA), can be used to deliver the disclosed CasX polypeptides, the disclosed CasX fusion polypeptides, the disclosed RNPs, the disclosed nucleic acids, or the disclosed CasX systems to target cells. Preformed vesicles with the following lipid composition can be contemplated: amino lipid, distearoylphosphatidylcholine (DSPC), cholesterol, and (R)-2,3-bis(octadecyloxy)propyl-1-(methoxypoly(ethylene glycol)2000)propylcarbamate (PEG-lipid) in a 40 / 10 / 40 / 10 molar ratio, respectively, and a FVII siRNA / total lipid ratio of approximately 0.05 (w / w). To ensure a narrow particle size distribution in the 70-90 nm range and a low polydispersity index of 0.11 ± 0.04 (n = 56), particles can be extruded up to three times through an 80 nm membrane before adding guide RNA. Particles containing highly potent amino lipids 16 may also be used, where the molar ratio of the four lipid components 16, DSPC, cholesterol, and PEG-lipid (50 / 10 / 38.5 / 1.5), can be further optimized to improve in vivo activity.
[0254] Lipids can be combined with the disclosed CasX system or its component(s) or nucleic acids encoding same to form lipid nanoparticles (LNPs). Suitable lipids include, but are not limited to, DLin-KC2-DMA4, C12-200, and the co-lipids disteroylphosphatidylcholine, cholesterol, and PEG-DMG can be combined with the disclosed CasX system or its components using a spontaneous vesicle formation procedure. The molar ratio of the components can be approximately 50 / 10 / 38.5 / 1.5 (DLin-KC2-DMA or C12-200 / disteroylphosphatidylcholine / cholesterol / PEG-DMG).
[0255] The CasX system of the present disclosure, or components thereof, can be delivered encapsulated in PLGA microspheres, as further described in U.S. Published Application Nos. 20130252281, 20130245107, and 20130244279.
[0256] Supercharged proteins can be used to deliver the disclosed CasX polypeptides, disclosed CasX fusion polypeptides, disclosed RNPs, disclosed nucleic acids, or disclosed CasX systems to target cells. Supercharged proteins are a class of engineered or naturally occurring proteins that have an unusually high net positive or negative theoretical charge. Both supernegatively and superpositively charged proteins exhibit the ability to resist thermally or chemically induced aggregation. Superpositively charged proteins can also penetrate mammalian cells. Associating cargo with these proteins, such as plasmid DNA, RNA, or other proteins, can facilitate the functional delivery of these macromolecules to mammalian cells both in vitro and in vivo.
[0257] A cell-penetrating peptide (CPP) can be used to deliver a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell. CPPs typically have an amino acid composition that either contains a high relative abundance of positively charged amino acids, such as lysine or arginine, or has a sequence containing an alternating pattern of polar / charged amino acids and nonpolar, hydrophobic amino acids.
[0258] An implantable device can be used to deliver a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure (e.g., a CasX guide RNA, a nucleic acid encoding a CasX guide RNA, a nucleic acid encoding a CasX polypeptide, a donor template, etc.), or a CasX system of this disclosure to a target cell (e.g., a target cell in vivo, where the target cell is a target cell in the circulation, a target cell in a tissue, a target cell in an organ, etc.). An implantable device suitable for use in delivering a CasX polypeptide of this disclosure, a CasX fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasX system of this disclosure to a target cell (e.g., a target cell in vivo, where the target cell is a target cell in the circulation, a target cell in a tissue, a target cell in an organ, etc.) can include a container (e.g., a reservoir, a matrix, etc.) that contains the CasX polypeptide, CasX fusion polypeptide, RNP, or CasX system (or a component thereof, e.g., a nucleic acid of this disclosure).
[0259] Suitable implantable devices may include, for example, a polymer substrate, such as a matrix, used as the device body, and in some cases, additional scaffolding materials, such as metals, additional polymers, and materials to improve visibility and imaging. Implantable delivery devices may be advantageous in providing localized and prolonged release, with the delivered polypeptide and / or nucleic acid being released directly to the target site (e.g., extracellular matrix (ECM), the vasculature surrounding a tumor, diseased tissue, etc.). Suitable implantable devices include devices suitable for use in delivery to cavities such as the peritoneal cavity and / or any other type of administration in which the drug delivery system is not tethered or attached, and include a biostable and / or biodegradable and / or bioabsorbable polymer substrate (e.g., which may optionally be a matrix). In some cases, suitable implantable drug delivery devices include degradable polymers, and the primary release mechanism is bulk erosion. In some cases, suitable implantable drug delivery devices include non-degradable or slowly degrading polymers, and the primary release mechanism is diffusion rather than bulk erosion, whereby the outer portion functions as a membrane and the inner portion functions as a drug reservoir that is virtually unaffected by the environment for extended periods (e.g., from about one week to about several months). Combinations of different polymers with different release mechanisms may optionally be used. The concentration gradient may be maintained effectively constant for a substantial period of the total release period, so that the diffusion rate is effectively constant (referred to as "zero-mode" diffusion). The term "constant" refers to a diffusion rate that is maintained above a low threshold of therapeutic efficacy but may still optionally be characterized by an initial burst and / or may fluctuate (e.g., increase and decrease to a certain extent). The diffusion rate may be maintained in this manner for extended periods and may be considered constant at a certain level to optimize the therapeutic effective period (e.g., effective silencing period).
[0260] In some cases, implantable delivery systems are designed to shield the nucleotide-based therapeutic agent from degradation, whether chemical in nature or due to attack from enzymes and other factors within the subject's body.
[0261] The site for device implantation, or target site, can be selected for maximum therapeutic effect. For example, the delivery device can be implanted within or proximal to the tumor environment or tumor-associated blood supply. Target locations can include, for example, 1) the brain at degenerative sites such as Parkinson's disease or Alzheimer's disease in the basal ganglia, white matter, and gray matter; 2) the spine, as in amyotrophic lateral sclerosis (ALS); 3) the cervix; 4) active and chronically inflamed joints; 5) the dermis, as in psoriasis; 6) exchange and sensory nerve sites for analgesic effects; 7) bone; 8) sites of acute or chronic infection; 9) the vagina; 10) the inner ear - auditory system, labyrinth of the cochlear ear, vestibular system, 11) in the trachea, 12) in the heart, annulus, epicardium, 13) in the urinary tract or bladder, 14) in the biliary system, 15) in parenchymal tissues including, but not limited to, the kidney, liver, and spleen, 16) in lymph nodes, 17) in salivary glands, 18) in the gums, 19) in articular tissues (in joints), 20) in the eye, 21) in brain tissue, 22) in the ventricles, 23) in cavities including the peritoneal cavity (for example, but not limited to, in the case of ovarian cancer), 24) in the esophagus, 25) in the rectum, and 26) in the vasculature.
[0262] Methods of insertion, such as implantation, may optionally be used for other types of tissue implantation and / or insertion and / or tissue harvesting, optionally without modification or optionally with only minor modifications to such methods, optionally including, but not limited to, brachytherapy, biopsy, endoscopy with and / or without ultrasound, e.g., stereotactic radioscopy of brain tissue, laparoscopy, including implantation of ...
Claims
1. 1. A composition comprising: a) a CasX polypeptide or a nucleic acid molecule encoding said CasX polypeptide; b) a CasX guide RNA or one or more DNA molecules encoding said CasX guide RNA; It contains the CasX polypeptide comprises a RuvC domain comprising a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain; The amino acid sequence of the CasX polypeptide has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, the CasX guide RNA comprises a nucleotide sequence complementary to a nucleotide sequence within a target nucleic acid; The CasX polypeptide forms a complex with the CasX guide RNA and binds to the target nucleic acid. composition.
2. The composition of claim 1, wherein the amino acid sequence of the CasX polypeptide has 95% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO:
2.
3. 3. The composition of claim 1 or claim 2, wherein the CasX guide RNA is a single guide RNA or a dual guide RNA.
4. The CasX polypeptide is a) a nickase capable of cleaving only one strand of a double-stranded target nucleic acid molecule, or b) a catalytically inactive CasX polypeptide (dCasX); Optionally, the CasX polypeptide comprises one or more mutations at positions corresponding to those selected from D672, E769, and D935 of SEQ ID NO:
1. The composition according to any one of claims 1 to 3.
5. The composition of any one of claims 1 to 4, further comprising a DNA donor template, optionally wherein the DNA donor template comprises a nucleotide sequence having homology to a target nucleic acid.
6. 1. A CasX single guide RNA molecule, comprising: a) (i) a guide sequence that hybridizes to a target nucleic acid, and (ii) duplex-forming segment and a targeting factor sequence comprising: b) an activator sequence that hybridizes with the duplex-forming segment of the targeter sequence to form a double-stranded RNA (dsRNA) duplex that can bind to a CasX polypeptide, binds to a target nucleic acid when incorporated into a single guide RNA, and comprises a nucleotide sequence having 80% or more identity to any one of SEQ ID NOs: 21 to 27; It contains the sequence of the targeting factor sequence is a non-naturally occurring sequence; the CasX polypeptide comprises a RuvC domain comprising a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain; The amino acid sequence of the CasX polypeptide has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO:
2. CasX single guide RNA molecule.
7. A DNA molecule comprising a nucleotide sequence encoding the CasX single guide RNA molecule of claim 6.
8. A CasX fusion polypeptide, a) a CasX polypeptide comprising a RuvC domain comprising a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain; b) a heterologous polypeptide fused to the N-terminus and / or C-terminus of the CasX polypeptide; It contains The amino acid sequence of the CasX polypeptide has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, the CasX polypeptide, when complexed with a CasX guide RNA, binds to a target nucleic acid; CasX fusion polypeptides.
9. The CasX fusion polypeptide of claim 8, wherein the amino acid sequence of the CasX polypeptide has 95% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO:
2.
10. The CasX polypeptide is a) a nickase capable of cleaving only one strand of a double-stranded target nucleic acid molecule, or b) a catalytically inactive CasX polypeptide (dCasX); Optionally, the CasX polypeptide comprises one or more mutations at positions corresponding to those selected from D672, E769, and D935 of SEQ ID NO:
1. A CasX fusion polypeptide according to claim 8 or claim 9.
11. the heterologous polypeptide is a) a targeting polypeptide that provides binding to a cell surface moiety on a target cell or target cell type; b) exhibits an enzymatic activity that modifies target DNA, optionally exhibiting one or more enzymatic activities selected from nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity; c) exhibits an enzymatic activity that modifies a target polypeptide associated with a target nucleic acid, optionally exhibiting one or more enzymatic activities selected from methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, demyristoylating activity, glycosylation activity (e.g., from an O-GlcNAc transferase), and deglycosylating activity; d) is an endosomal escape polypeptide; e) a chloroplast transit peptide, optionally wherein the chloroplast transit peptide is selected from the amino acid sequences set forth in SEQ ID NOs: 83-93; or f) a protein that increases or decreases transcription, optionally selected from a transcription repressor domain or a transcription activator domain; g) a protein-binding domain; or h) a nuclear localization signal; A CasX fusion polypeptide according to any one of claims 8 to 10.
12. A nucleic acid comprising a nucleotide sequence encoding the CasX fusion polypeptide of any one of claims 8 to 11, Optionally, the nucleotide sequence is operably linked to a promoter. Nucleic acid.
13. a) a CasX polypeptide comprising a RuvC domain including a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain, the amino acid sequence of which has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, and which binds to a target nucleic acid when complexed with a CasX guide RNA, or a nucleic acid molecule encoding the CasX polypeptide; or b) A CasX fusion polypeptide comprising a RuvC domain including a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain, and having an amino acid sequence that is 90% or more identical to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, and a fusion partner that binds the CasX fusion polypeptide to a target nucleic acid when complexed with a CasX guide RNA, or a nucleic acid molecule encoding the CasX fusion polypeptide. and are not human germline or human embryonic cells, eukaryotic cells.
14. 14. The eukaryotic cell of claim 13, comprising the nucleic acid molecule encoding the CasX polypeptide, wherein the nucleic acid molecule is integrated into the genomic DNA of the eukaryotic cell.
15. 15. The eukaryotic cell of claim 13 or claim 14, which is a plant cell, a mammalian cell, an insect cell, an arachnid cell, a fungal cell, an avian cell, a reptilian cell, an amphibian cell, an invertebrate cell, a mouse cell, a rat cell, a primate cell, a non-human primate cell, or a human cell.
16. A cell comprising a CasX fusion polypeptide according to any one of claims 8 to 11, or a nucleic acid comprising a nucleotide sequence encoding said CasX fusion polypeptide, wherein the cell is not a human germline cell or a human embryonic cell.
17. 1. An in vitro method for modifying a target nucleic acid, comprising: The target nucleic acid a) a CasX polypeptide comprising a RuvC domain comprising a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain; and b) contacting the target nucleic acid with a CasX guide RNA comprising a guide sequence that hybridizes to a target sequence of the target nucleic acid; The amino acid sequence of the CasX polypeptide has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, said contacting results in modification of said target nucleic acid by said CasX polypeptide; Optionally, said contacting comprises: (i) the CasX polypeptide or a nucleic acid molecule encoding the CasX polypeptide; (ii) the CasX guide RNA or a nucleic acid molecule encoding the CasX guide RNA, and optionally (iii) a DNA donor template; and into a cell, Optionally, the DNA donor template comprises a nucleotide sequence having homology to the target nucleic acid; the modification is integration of the DNA donor template into the target nucleic acid; the cell is a eukaryotic cell, optionally the eukaryotic cell is selected from a fungal cell, a mammalian cell, a reptilian cell, an insect cell, an avian cell, a fish cell, a parasite cell, an arthropod cell, an invertebrate cell, a vertebrate cell, a rodent cell, a mouse cell, a rat cell, a primate cell, a non-human primate cell, and a human cell; the cell is not a human germline cell or a human embryonic cell; method.
18. 1. An in vitro or ex vivo method for modulating transcription from a target DNA, modifying a target nucleic acid, or modifying a protein associated with a target nucleic acid, comprising: The target nucleic acid a) a CasX fusion polypeptide comprising a CasX polypeptide fused to a heterologous polypeptide; and b) contacting the target nucleic acid with a CasX guide RNA comprising a guide sequence that hybridizes to a target sequence of the target nucleic acid; the CasX polypeptide comprises a RuvC domain comprising a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain; The amino acid sequence of the CasX polypeptide has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, Optionally, said contacting comprises contacting the cell with: (i) the CasX fusion polypeptide, or a nucleic acid molecule encoding the CasX fusion polypeptide; and (ii) introducing the CasX guide RNA or a nucleic acid molecule encoding the CasX guide RNA; the cell is not a human germline cell or a human embryonic cell; method.
19. 1. A transgenic multicellular non-human organism comprising: Those genomes are a) a CasX polypeptide comprising a RuvC domain including a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain, the amino acid sequence of which has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, and which binds to a target nucleic acid when complexed with a CasX guide RNA; b) a CasX fusion polypeptide comprising a RuvC domain including a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain, and a CasX polypeptide having an amino acid sequence that is 90% or more identical to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO: 2, and a fusion partner, wherein the CasX fusion polypeptide binds to a target nucleic acid when complexed with a CasX guide RNA; and c) I) a CasX guide RNA comprising: i) a guide sequence that hybridizes to a target nucleic acid; and (ii) a targeting factor sequence comprising a duplex-forming segment; and II) an activator sequence that hybridizes to the duplex-forming segment of the targeting factor sequence to form a double-stranded RNA (dsRNA) duplex that can bind to a CasX polypeptide, and that binds to the target nucleic acid when incorporated into the single guide RNA, and that comprises a nucleotide sequence having 80% or more identity to any one of SEQ ID NOs: 21 to 27, wherein the CasX polypeptide comprises a RuvC domain that includes a RuvC-I subdomain, a RuvC-II subdomain, and a RuvC-III subdomain, and the amino acid sequence of the CasX polypeptide has 90% or more identity to the amino acid sequence set forth in SEQ ID NO: 1 or SEQ ID NO:
2. a transgene comprising a nucleotide sequence encoding one or more of: Transgenic multicellular non-human organisms.