CasZ Compositions and Methods of Use
The modified CRISPR-Cas system, incorporating a CasZ protein, guide RNA, and trancRNA, addresses the challenges of precise sequence targeting and modification, achieving efficient and controlled editing in nucleic acids.
Patent Information
- Application Number
- JP2023059813
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-11-01
- Filing Date
- 2023-04-03
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2038-10-31
AI Technical Summary
Current CRISPR-Cas systems face challenges in efficiently targeting and modifying specific sequences in nucleic acids, particularly in achieving precise and controlled editing in various biological contexts.
The development of a modified CRISPR-Cas system utilizing a CasZ protein, a CasZ guide RNA, and a CasZ transactivated non-coding RNA (trancRNA), which form a ribonucleoprotein complex that provides sequence specificity and site-specific activity for targeted nucleic acid modification.
This modified CRISPR-Cas system enables efficient and precise targeting of specific sequences, facilitating controlled editing and modification of nucleic acids, thereby overcoming the limitations of existing systems.
Smart Images

Figure 0007672445000037 
Figure 0007672445000038 
Figure 0007672445000039
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 580,395, filed November 1, 2017, which is incorporated herein by reference in its entirety.
[0002] Incorporation by reference of sequence listings provided as text files A sequence listing is provided herewith as a text file "BERK-374WO_SEQ_LISTING_ST25.txt," created on October 30, 2018, and having a size of 536 KB. The contents of the text file are incorporated herein by reference in their entirety. [Background technology]
[0003] preface The CRISPR-Cas system, an example of a pathway unknown to science before the era of DNA sequencing, is now understood to confer adaptive immunity to phages and viruses to bacteria and archaea. Extensive research has clarified the biochemistry of this system. CRISPR-Cas systems consist of a CRISPR array, which contains direct repeats of Cas proteins, responsible for acquiring, targeting, and cleaving foreign DNA or RNA, and short, flanking spacer sequences that guide the Cas proteins to their targets. Class 2 CRISPR-Cas is a streamlined version in which a single Cas protein, bound to RNA, is responsible for binding to and cleaving the targeted sequence. The programmable nature of these minimal systems facilitates their use as a versatile technology that is revolutionizing the field of genome engineering. Summary of the Invention
[0004] The present disclosure provides compositions and methods, including one or more of: (1) a "CasZ" protein (also referred to as a CasZ polypeptide), a nucleic acid encoding a CasZ protein, and / or a modified host cell comprising a CasZ protein (and / or a nucleic acid encoding same); (2) a modified host cell comprising a CasZ guide RNA that binds to and provides sequence specificity to a CasZ protein, a nucleic acid encoding a CasZ guide RNA, and / or a CasZ guide RNA (and / or a nucleic acid encoding same); and (3) a modified host cell comprising a CasZ trans-activating non-coding RNA (trancRNA) (referred to herein as a "CasZ trancRNA"), a nucleic acid encoding a CasZ trancRNA, and / or a CasZ trancRNA (and / or a nucleic acid encoding same). [Brief explanation of the drawings]
[0005] [Figure 1-1] An example of a native CasZ protein sequence is shown below. [Figure 1-2] See description of Figure 1-1. [Figure 1-3] See description of Figure 1-1. [Figure 1-4] See description of Figure 1-1. [Figure 1-5] See description of Figure 1-1. [Figure 1-6] See description of Figure 1-1. [Figure 1-7] See description of Figure 1-1. [Figure 1-8] See description of Figure 1-1. [Figure 1-9] See description of Figure 1-1. [Figure 1-10] See description of Figure 1-1. [Figure 1-11] See description of Figure 1-1. [Figure 1-12] See description of Figure 1-1. [Figure 1-13] See description of Figure 1-1. [Figure 1-14] See description of Figure 1-1. [Figure 1-15] See description of Figure 1-1. [Figure 1-16] See description of Figure 1-1. [Figure 1-17] See description of Figure 1-1. [Figure 2] A schematic diagram of the CasZ locus is shown, which contains the Cas1 protein in addition to the CasZ protein. [Figure 3] Phylogenetic tree of CasZ sequences in relation to other class 2 CRISPR / Cas effector protein sequences. [Figure 4] Phylogenetic tree of Cas1 sequences from the CasZ locus in relation to Cas1 sequences from other class 2 CRISPR / Cas loci. [Figure 5-1] Figure 5 shows transcriptome RNA mapping data representing expression of transcribed RNAs from the CasZ locus. The transcribed RNAs are adjacent to the CasZ repeat array but do not contain or are complementary to the repeat sequences. RNA mapping data is shown for the following loci: CasZa3, CasZb4, CasZc5, CasZd1, and CasZe3. Small, repetitive, aligned arrows represent repeats of the CRISPR array (indicating the presence of guide RNA coding sequences). Peaks outside and adjacent to the repeat array represent highly transcribed transcribed RNAs. Figure 5 (continued 1) Nucleotide sequence (top to bottom: SEQ ID NOs: 312-331). Figure 5 (continued 3) Nucleotide sequence (top to bottom: SEQ ID NOs: 161-177). [Figure 5-2] See description of Figure 5-1. [Figure 5-3] See description of Figure 5-1. [Figure 5-4] See description of Figure 5-1. [Figure 5-5] See description of Figure 5-1. [Figure 5-6] See description of Figure 5-1. [Figure 5-7] See description of Figure 5-1. [Figure 6]The results of PAM preference are shown for CasZc (top) and CasZb (bottom) when assayed using a PAM depletion assay. [Figure 7-1] 1 shows the sequence of the Cas14 protein described herein. [Figure 7-2] See description of Figure 7-1. [Figure 7-3] See description of Figure 7-1. [Figure 7-4] See description of Figure 7-1. [Figure 7-5] See description of Figure 7-1. [Figure 7-6] See description of Figure 7-1. [Figure 7-7] See description of Figure 7-1. [Figure 7-8] See description of Figure 7-1. [Figure 7-9] See description of Figure 7-1. [Figure 7-10] See description of Figure 7-1. [Figure 7-11] See description of Figure 7-1. [Figure 7-12] See description of Figure 7-1. [Figure 7-13] See description of Figure 7-1. [Figure 7-14] See description of Figure 7-1. [Figure 8-1] Panels A-D show the structure and phylogeny of the CRISPR-Cas14 genomic locus. [Figure 8-2] See description of Figure 8-1. [Figure 9] Phylogenetic analysis of Cas14 orthologs. [Figure 10] A maximum likelihood tree of CAS1 from known CRISPR systems is shown. [Figure 11] Panels A and B show the acquisition of a new spacer using the CRISPR-Cas14 system. [Figure 12] Panels A-D show CRISPR-Cas14a actively encoding tracrRNA. [Figure 13-1]Panels A-B show metatranscriptomics of the CRISPR-Cas14 locus. [Figure 13-2] See description of Figure 13-1. [Figure 13-3] See description of Figure 13-1. [Figure 13-4] See description of Figure 13-1. [Figure 13-5] See description of Figure 13-1. [Figure 14] Panels A-B show RNA processing and heterologous expression by CRISPR-Cas14. [Figure 15-1] Panels A to D show plasmid depletion by Cas14a1 and SpCas9. [Figure 15-2] See description of Figure 15-1. [Figure 15-3] See description of Figure 15-1. [Figure 15-4] See description of Figure 15-1. [Figure 16-1] Panels A to D show that CRISPR-Cas14a is an RNA-guided DNA endonuclease. [Figure 16-2] See description of Figure 16-1. [Figure 16-3] See description of Figure 16-1. [Figure 17-1] Panels A to E show degradation of ssDNA by Cas14a1. [Figure 17-2] See description of Figure 17-1. [Figure 17-3] See description of Figure 17-1. [Figure 17-4] See description of Figure 17-1. [Figure 18] Figure 1 shows the kinetics of Cas14a1 cleavage of ssDNA with various guide RNA components. [Figure 19-1] Panels A-F show optimization of the Cas14a1 guide RNA component. [Figure 19-2] See description of Figure 19-1. [Figure 19-3] See description of Figure 19-1. [Figure 20-1] Panels A-E show high-fidelity ssDNA DNP detection by CRISPR-Cas14a. Panel C provides the nucleotide sequences (top to bottom: SEQ ID NOs: 367-370). [Figure 20-2] See description of Figure 20-1. [Figure 20-3] See description of Figure 20-1. [Figure 20-4] See description of Figure 20-1. [Figure 21-1] Panels A-F show the effect of various activators on Cas14a1 cleavage kinetics. [Figure 21-2] See description of Figure 21-1. [Figure 21-3] See description of Figure 21-1. [Figure 22-1] Panels A and B show the diversity of the CRISPR-Cas14 system. [Figure 22-2] See description of Figure 22-1. [Figure 23-1] Panels A-C show testing of Cas14a1-mediated interference in a heterologous host. Diagram of Cas14a1 and LbCas12a constructs for testing interference in E. coli. [Figure 23-2] See description of Figure 23-1. [Figure 24-1] 1 shows the Cas14 nucleotide sequence of the plasmid used in the present invention. [Figure 24-2] See description of Figure 24-1. [Figure 24-3] See description of Figure 24-1. [Figure 24-4] See description of Figure 24-1. [Figure 24-5] See description of Figure 24-1. [Figure 24-6] See description of Figure 24-1. [Figure 24-7] See description of Figure 24-1. [Figure 24-8] See description of Figure 24-1. [Figure 24-9] See description of Figure 24-1. [Figure 24-10] See description of Figure 24-1. [Figure 24-11] See description of Figure 24-1. [Figure 24-12] See description of Figure 24-1. [Figure 24-13] See description of Figure 24-1. [Figure 24-14] See description of Figure 24-1. [Figure 24-15] See description of Figure 24-1. [Figure 24-16] See description of Figure 24-1. [Figure 24-17] See description of Figure 24-1. [Figure 25-1] Panels A through E show the sequence maps of each of the plasmids disclosed in FIG. [Figure 25-2] See description of Figure 25-1. [Figure 25-3] See description of Figure 25-1. [Figure 25-4] See description of Figure 25-1. [Figure 25-5] See description of Figure 25-1. DETAILED DESCRIPTION OF THE INVENTION
[0006] definition As used herein, "heterologous" refers to a nucleotide or polypeptide sequence that is not found in a naturally occurring nucleic acid or protein, respectively. For example, with respect to a CasZ polypeptide, a heterologous polypeptide includes an amino acid sequence from a protein other than the CasZ polypeptide. In some cases, a portion of a CasZ protein from one species is fused with a portion of a CasZ protein from a different species. Thus, CasZ sequences from each species can be considered heterologous to each other. As another example, a CasZ protein (e.g., a dCasZ protein) can be fused with an active domain from a non-CasZ protein (e.g., a histone deacetylase), and the sequence of the active domain can be considered a heterologous polypeptide (heterologous to the CasZ protein).
[0007] The terms "polynucleotide" and "nucleic acid," used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxynucleotides. Thus, the terms include, but are not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The terms "polynucleotide" and "nucleic acid" should be understood to include single-stranded (such as sense or antisense) and double-stranded polynucleotides, as applicable to the described embodiments.
[0008] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymeric forms of amino acids of any length, including genetically encoded and non-genetically encoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones. This term encompasses fusion proteins, including, but not limited to, fusion proteins with heterologous amino acid sequences, fusions with heterologous and homologous leader sequences, with or without an N-terminal methionine residue; immunologically tagged proteins; and the like.
[0009] As used herein, the term "natively occurring," when applied to a nucleic acid, protein, cell, or organism, refers to a nucleic acid, cell, protein, or organism found in nature.
[0010] As used herein, the term "isolated" is meant to describe a polynucleotide, polypeptide, or cell that is in an environment that is different from the environment that the polynucleotide, polypeptide, or cell naturally occurs in. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.
[0011] As used herein, the term "exogenous nucleic acid" refers to a nucleic acid that is not normally or naturally found in and / or produced by a given bacterium, organism, or cell in nature. As used herein, the term "endogenous nucleic acid" refers to a nucleic acid that is normally found in and / or produced by a given bacterium, organism, or cell in nature. "Endogenous nucleic acid" is also referred to as a "naturally occurring nucleic acid," or a nucleic acid that is "native" to a given bacterium, organism, or cell.
[0012] As used herein, "recombinant" means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps that result in a construct having structural coding or non-coding sequences distinguishable from endogenous nucleic acids found in natural systems. Generally, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers or from a series of synthetic oligonucleotides to provide a synthetic nucleic acid capable of expression from a recombinant transcription unit contained in a cell or cell-free transcription and translation system. Such sequences can be provided in the form of an open reading frame uninterrupted by internal non-translated sequences, or introns, typically present in eukaryotic genes. Genomic DNA containing relevant sequences can also be used in forming recombinant genes or transcription units. Sequences of non-translated DNA can be present 5' or 3' from the open reading frame, where such sequences do not interfere with the manipulation or expression of the coding region, or may actually act to regulate the production of a desired product by various mechanisms (see "DNA Regulatory Sequences," below).
[0013] Thus, for example, the terms "recombinant" polynucleotide or "recombinant" nucleic acid refer to something that does not occur in nature, e.g., created through human intervention, by the artificial combination of two or isolated segments of sequence. This artificial combination is often achieved either by chemical synthesis means or by the artificial manipulation of isolated segments of nucleic acid, e.g., genetic engineering techniques. This is usually done to replace a codon with a redundant codon or a conserved amino acid that encodes it, typically while introducing or removing sequence recognition sites. Alternatively, it is performed by joining nucleic acid segments of desired functions together to produce a desired combination of functions. This artificial combination is often achieved either by chemical synthesis means or by the artificial manipulation of isolated segments of nucleic acid, e.g., genetic engineering techniques.
[0014] Similarly, the term "recombinant" polypeptide refers to a polypeptide that does not occur in nature, e.g., that is made by the artificial combination of two or separate segments of amino acid sequence through human intervention. Thus, for example, a polypeptide that includes a heterologous amino acid sequence is recombinant.
[0015] "Construct" or "vector" means a recombinant nucleic acid, generally a recombinant DNA that has been generated for the purpose of expressing and / or propagating a particular nucleotide sequence(s) or for use in the construction of other recombinant nucleotide sequences.
[0016] The terms "DNA regulatory sequence," "control element," and "regulatory element," used interchangeably herein, refer to transcriptional and translational control sequences, e.g., promoters, enhancers, polyadenylation signals, terminators, proteolysis signals, and the like, that provide and / or regulate the expression of a coding sequence and / or the production of the encoded polypeptide in a host cell.
[0017] The term "transformation" is used interchangeably herein with "genetic modification" and refers to a permanent or transient genetic change induced in a cell after introducing new nucleic acid (e.g., DNA exogenous to the cell) into the cell. The genetic change ("modification") can be achieved either by integration of the new nucleic acid into the host cell's genome or by transient or stable maintenance of the new nucleic acid as an episomal element. If the cell is eukaryotic, a permanent genetic change is generally achieved by introducing the new DNA into the cell's genome. In prokaryotic cells, permanent changes can be introduced into the chromosome or via extrachromosomal elements such as plasmids and expression vectors, which may contain one or more selectable markers to aid in their maintenance in the recombinant host cell. Suitable methods of genetic modification include viral infection, transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, and the like. The choice of method generally depends on the type of cell being transformed and the context in which the transformation is occurring (i.e., in vitro, ex vivo, or in vivo). A general discussion of these methods can be found in Ausubel, et al., Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995.
[0018] "Operably linked" refers to a juxtaposition, where the components so described are in a relationship permitting them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression. As used herein, the terms "heterologous promoter" and "heterologous control region" refer to promoters and other control regions that are not normally associated with a particular nucleic acid in nature. For example, a "transcriptional control region heterologous to a coding region" is a transcriptional control region that is not normally associated with that coding region in nature.
[0019] As used herein, "host cell" refers to a eukaryotic cell, a prokaryotic cell, or a cell from a multicellular organism cultured as a unicellular entity (e.g., a cell line) in vivo or in vitro; a eukaryotic or prokaryotic cell can be or has been used as a recipient of nucleic acid (e.g., an expression vector), including the progeny of the original cell that has been genetically modified with the nucleic acid. It is understood that the progeny of a unicellular cell may not necessarily be completely identical in morphology or in genomic or total DNA complement to the original parent due to natural, inadvertent, or deliberate mutation. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, e.g., an expression vector, has been introduced. For example, a prokaryotic host cell of interest is a prokaryotic host cell (e.g., a bacterium) that has been genetically modified by the introduction into a suitable prokaryotic host cell of heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the prokaryotic host cell (not normally naturally found in the prokaryotic host cell), or a recombinant nucleic acid that is not normally found in the prokaryotic host cell; and a eukaryotic host cell of interest is a eukaryotic host cell that has been genetically modified by the introduction into a suitable eukaryotic host cell of heterologous nucleic acid, e.g., an exogenous nucleic acid that is foreign to the eukaryotic host cell, or a recombinant nucleic acid that is not normally found in the eukaryotic host cell.
[0020] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues in proteins that have similar side chains. For example, the group of amino acids with aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic-hydroxyl side chains consists of serine and threonine; the group of amino acids with amide-containing side chains consists of asparagine and glutamine; the group of amino acids with aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains consists of lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0021] A polynucleotide or polypeptide has a certain percentage of "sequence identity" with another polynucleotide or polypeptide. This means that, when aligned, that percentage of bases or amino acids are identical and in the same relative positions when comparing the two sequences. Sequence similarity can be determined in several different ways. To determine sequence identity, sequences can be aligned using methods and computer programs, including BLAST, available at www.ncbi.nlm.nih.gov / BLAST / . See, e.g., Altschul et al. (1990), J. Mol. Biol. 215:403-10. Another alignment algorithm is FASTA, available in the Genetics Computing Group (GCG) package (Madison, Wisconsin, USA), a wholly owned subsidiary of Oxford Molecular Group, Inc. Other techniques for alignment are described in Methods in Enzymology, vol. 266: Computer Methods for Macromolecular Sequence Analysis (1996), ed. Doolittle, Academic Press, Inc. (a division of Harcourt Brace & Co., San Diego, California, USA). Of particular interest are alignment programs that allow gaps between sequences. Smith-Waterman is one type of algorithm that allows gaps in sequence alignment. See Meth. Mol. Biol. 70:173-187 (1997). Sequences can also be aligned using the GAP program, which uses the Needleman-Wunsch alignment method. See J. Mol. Biol. 48:443-453 (1970).
[0022] As used herein, the terms "treatment," "treating," and the like refer to obtaining a pharmacological and / or physiological effect. The effect can be prophylactic, in terms of completely or partially preventing the disease or condition, and / or therapeutic, in terms of partially or completely curing the disease and / or side effects caused by the disease. As used herein, "treatment" encompasses any treatment of disease in a mammal, e.g., a human, including (a) preventing the onset of the disease in a subject who is susceptible to the disease but has not yet been diagnosed as having it, (b) inhibiting the disease, i.e., preventing its development, and (c) relieving the disease, i.e., causing regression of the disease.
[0023] The terms "individual," "subject," "host," and "patient," used interchangeably herein, refer to individual organisms, e.g., mammals, including, but not limited to, mice, monkeys, non-human primates, humans, mammalian livestock, mammalian sport animals, and mammalian pets.
[0024] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0025] Where a range of values is provided, it is understood that each intervening value between the upper and lower limit of that range, to the tenth of the unit of the lower limit, and any other stated or intervening value in that stated range, is encompassed within the invention, unless the context clearly dictates otherwise. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specific excluded limits in the stated range. Where the stated range includes one or both of its limits, ranges excluding either or both of those included limits are also encompassed within the invention.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0027] It should be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a CasZ polypeptide" includes a plurality of such polypeptides; a reference to "the guide RNA" includes a reference to one or more guide RNAs and equivalents thereof known to those skilled in the art; and so forth. It should be further noted that the claims may be drafted to exclude any and all optional elements. Accordingly, this statement is intended to serve as a guide against the use of exclusive terminology such as "solely," "only," and the like, or the use of "negative" limitations, with regard to the recitation of claim elements.
[0028] It is understood that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of the embodiments related to the present invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination were individually and explicitly disclosed herein. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such subcombination were individually and explicitly disclosed herein.
[0029] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein should be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0030] Detailed Description The present disclosure provides compositions and methods, including one or more of: (1) a "CasZ" protein (also referred to as a CasZ polypeptide), a nucleic acid encoding a CasZ protein, and / or a modified host cell comprising a CasZ protein (and / or a nucleic acid encoding same); (2) a modified host cell comprising a CasZ guide RNA that binds to and provides sequence specificity to a CasZ protein, a nucleic acid encoding a CasZ guide RNA, and / or a CasZ guide RNA (and / or a nucleic acid encoding same); and (3) a modified host cell comprising a CasZ trans-activating non-coding RNA (trancRNA) (referred to herein as a "CasZ trancRNA"), a nucleic acid encoding a CasZ trancRNA, and / or a CasZ trancRNA (and / or a nucleic acid encoding same).
[0031] composition CRISPR / CasZ proteins, guide RNAs, and trancRNAs Class 2 CRISPR-Cas systems are characterized by an effector module containing a single multidomain protein. In the CasZ system, a CRISPR / Cas endonuclease (e.g., a CasZ protein) interacts with (binds to) a corresponding guide RNA (e.g., a CasZ guide RNA) to form a ribonucleoprotein (RNP) complex that is targeted to a specific site in the target nucleic acid through base pairing between the guide RNA and the target sequence within the target nucleic acid molecule. The guide RNA contains a nucleotide sequence (guide sequence) that is complementary to the sequence of the target nucleic acid (target site). Thus, the CasZ protein forms a complex with the CasZ guide RNA, and the guide RNA provides sequence specificity to the RNP complex via the guide sequence. The CasZ protein in the complex provides site-specific activity. In other words, the CasZ protein is guided to (e.g., stabilized at) a target site within a target nucleic acid (e.g., a target nucleotide sequence within a target chromosomal nucleic acid; or a target nucleotide sequence within a target extrachromosomal nucleic acid, such as, for example, an episomal nucleic acid, a minicircle nucleic acid, a mitochondrial nucleic acid, or a chloroplast nucleic acid) by its association with the guide RNA.
[0032] The present disclosure provides compositions comprising a CasZ polypeptide (and / or a nucleic acid encoding a CasZ polypeptide) (e.g., the CasZ polypeptide can be a naturally occurring CasZ protein, a nickase CasZ protein, a dCasZ protein, a chimeric CasZ protein, etc.) (CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein). The present disclosure provides compositions comprising a CasZ guide RNA (and / or a nucleic acid encoding a CasZ guide RNA). For example, the present disclosure provides a composition comprising (a) a CasZ polypeptide (and / or a nucleic acid encoding a CasZ polypeptide), and (b) a CasZ guide RNA (and / or a nucleic acid encoding a CasZ guide RNA). The present disclosure provides a nucleic acid / protein complex (RNP complex) comprising (a) a CasZ polypeptide, and (b) a CasZ guide RNA. The present disclosure provides compositions comprising a CasZ trancRNA. The present disclosure provides compositions comprising a CasZ trancRNA and one or more of (a) a CasZ protein and (b) a CasZ guide RNA (e.g., a CasZ trancRNA and a CasZ protein, a CasZ trancRNA and a CasZ guide RNA, or a CasZ trancRNA and a CasZ protein and a CasZ guide RNA). The present disclosure provides nucleic acid / protein complexes (RNP complexes) comprising (a) a CasZ polypeptide, (b) a CasZ guide RNA, and (c) a CasZ trancRNA. The present disclosure provides compositions comprising a CasZ protein and one or more of (a) a CasZ trancRNA, and (b) a CasZ guide RNA.
[0033] CasZ protein A CasZ polypeptide (this term is used interchangeably with the terms "CasZ protein," "Cas14," "Cas14 polypeptide," or "Cas14 protein") can bind to and / or modify (e.g., cleave, nick, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with the target nucleic acid (e.g., methylation or acetylation of histone tails) (e.g., in some cases, the CasZ protein includes an active fusion partner, and in some cases, the CasZ protein provides nuclease activity). In some cases, the CasZ protein is a naturally occurring protein (e.g., occurring naturally in a prokaryotic cell). In other cases, the CasZ protein is not a naturally occurring polypeptide (e.g., the CasZ protein is a variant CasZ protein, a chimeric protein, etc.). The CasZ protein contains three partial RuvC domains (RuvC-I, RuvC-II, and RuvC-III, also referred to herein as subdomains), which are not contiguous with the primary amino acid sequence of the CasZ protein but form the RuvC domain when the protein is produced and folded. Native CasZ proteins function as endonucleases, catalyzing cleavage at specific sequences in targeted nucleic acids (e.g., double-stranded DNA (dsDNA)). Sequence specificity is provided by an associated guide RNA that hybridizes to a target sequence within the target DNA. The native CasZ guide RNA is a crRNA, which contains (i) a guide sequence that hybridizes to a target sequence in the target DNA and (ii) a protein-binding segment that binds to the CasZ protein.
[0034] In some embodiments, the CasZ protein of the subject methods and / or compositions is (or is derived from) a naturally occurring (wild-type) protein. Examples of naturally occurring CasZ proteins (e.g., CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, and CasZl) are shown in FIG. 1. In some cases, the CasZ protein of interest is a CasZa protein. In some cases, the CasZ protein of interest is a CasZb protein. In some cases, the CasZ protein of interest is a CasZc protein. In some cases, the CasZ protein of interest is a CasZd protein. In some cases, the CasZ protein of interest is a CasZe protein. In some cases, the CasZ protein of interest is a CasZf protein. In some cases, the CasZ protein of interest is a CasZg protein. In some cases, the CasZ protein of interest is a CasZh protein. In some cases, the CasZ protein of interest is a CasZi protein. In some cases, the CasZ protein of interest is a CasZj protein. In some cases, the CasZ protein of interest is a CasZk protein. In some cases, the CasZ protein of interest is a CasZl protein. In some cases, the CasZ protein of interest is a CasZe, CasZf, CasZg, or CasZh protein. In some cases, the CasZ protein of interest is a CasZj, CasZk, or CasZl protein.
[0035] It is important to note that this newly discovered protein (CasZ) is short compared to previously identified CRISPR-Cas endonucleases, and therefore, its use as a replacement offers the advantage of a relatively short nucleotide sequence encoding the protein. This is useful for research and / or clinical applications, for example, when nucleic acids encoding the CasZ protein are desired for delivery into cells, such as eukaryotic cells (e.g., mammalian cells, human cells, mouse cells, in vitro, ex vivo, or in vivo), for example, in situations using viral vectors (e.g., AAV vectors). Additionally, in their natural context, CasZ-encoding DNA sequences occur in a locus that also harbors the CAS1 protein.
[0036] In some cases, the subject CasZ protein has a length of 900 amino acids or less (e.g., 850 amino acids or less, 800 amino acids or less, 750 amino acids or less, or 700 amino acids or less). In some cases, the subject CasZ protein has a length of 850 amino acids or less (e.g., 850 amino acids or less). In some cases, the subject CasZ protein has a length of 800 amino acids or less (e.g., 750 amino acids or less). In some cases, the subject CasZ protein has a length of 700 amino acids or less. In some cases, the subject CasZ protein has a length of 650 amino acids or less.
[0037] In some cases, the subject CasZ protein has a length in the range of 350-900 amino acids (e.g., 350-850, 350-800, 350-750, 350-700, 400-900, 400-850, 400-800, 400-750, or 400-700 amino acids).
[0038] In some cases, the CasZ protein of interest (e.g., CasZa) has a length in the range of 350-750 amino acids (e.g., 350-700, 350-550, 450-550, 450-750, 450-650, or 450-550 amino acids). In some cases, the CasZ protein of interest (e.g., CasZa) has a length in the range of 450-750 amino acids (e.g., 500-700 amino acids). In some cases, the CasZ protein of interest (e.g., CasZa) has a length in the range of 350-700 amino acids (e.g., 350-650, 350-600, or 350-550 amino acids). In some cases, the CasZ protein of interest (e.g., CasZa) has a length in the range of 500-700 amino acids. In some cases, the CasZ protein of interest (e.g., CasZa) has a length in the range of 450-550 amino acids. In some cases, the CasZ protein of interest (e.g., CasZa) has a length in the range of 350 to 550 amino acids.
[0039] In some cases, the CasZ protein of interest (e.g., CasZb) has a length in the range of 350-700 amino acids (e.g., 350-650, or 350-620 amino acids). In some cases, the CasZ protein of interest (e.g., CasZb) has a length in the range of 450-700 amino acids (e.g., 450-650, 500-650, or 500-620 amino acids). In some cases, the CasZ protein of interest (e.g., CasZb) has a length in the range of 500-650 amino acids (e.g., 500-620 amino acids). In some cases, the CasZ protein of interest (e.g., CasZb) has a length in the range of 500-620 amino acids.
[0040] In some cases, the subject CasZ protein (e.g., CasZc) has a length in the range of 600 to 800 amino acids (e.g., 600 to 650 or 700 to 800 amino acids). In some cases, the subject CasZ protein (e.g., CasZc) has a length in the range of 600 to 650 amino acids. In some cases, the subject CasZ protein (e.g., CasZc) has a length in the range of 700 to 800 amino acids.
[0041] In some cases, the subject CasZ protein (e.g., CasZd) has a length in the range of 400 to 650 amino acids (e.g., 400 to 600, 400 to 550, 500 to 650, 500 to 600, or 500 to 550 amino acids). In some cases, the subject CasZ protein (e.g., CasZd) has a length in the range of 500 to 600 amino acids. In some cases, the subject CasZ protein (e.g., CasZd) has a length in the range of 500 to 550 amino acids. In some cases, the subject CasZ protein (e.g., CasZd) has a length in the range of 400 to 550 amino acids.
[0042] In some cases, the subject CasZ protein (e.g., CasZe) has a length in the range of 450-700 amino acids (e.g., 450-650, 450-615, 475-700, 475-650, or 475-615 amino acids). In some cases, the subject CasZ protein (e.g., CasZe) has a length in the range of 450-675 amino acids. In some cases, the subject CasZ protein (e.g., CasZe) has a length in the range of 475-675 amino acids.
[0043] In some cases, the subject CasZ protein (e.g., CasZf) has a length in the range of 400-550 amino acids (e.g., 400-520, 400-500, 400-475, 415-550, 415-520, 415-500, or 415-475 amino acids). In some cases, the subject CasZ protein (e.g., CasZf) has a length in the range of 400-475 amino acids (e.g., 400-450 amino acids).
[0044] In some cases, the CasZ protein of interest (e.g., CasZg) has a length in the range of 500-750 amino acids (e.g., 550-750 or 500-700 amino acids). In some cases, the CasZ protein of interest (e.g., CasZg) has a length in the range of 700-750 amino acids. In some cases, the CasZ protein of interest (e.g., CasZg) has a length in the range of 550-600 amino acids.
[0045] In some cases, the subject CasZ protein (e.g., CasZh) has a length in the range of 380-450 amino acids (e.g., 380-420, 400-450, 400-420 amino acids). In some cases, the subject CasZ protein (e.g., CasZh) has a length in the range of 400-420 amino acids.
[0046] In some cases, the subject CasZ protein (e.g., CasZi) has a length in the range of 700-800 amino acids (e.g., 700-750, 720-800, or 720-750 amino acids). In some cases, the subject CasZ protein (e.g., CasZi) has a length in the range of 720-780 amino acids.
[0047] In some cases, the CasZ protein of interest (e.g., CasZj) has a length in the range of 600-750 amino acids (e.g., 600-700 or 650-700 amino acids), hi some cases, the CasZ protein of interest (e.g., CasZj) has a length in the range of 400-420 amino acids.
[0048] In some cases, the CasZ protein of interest (e.g., CasZk) has a length in the range of 450-600 amino acids (e.g., 450-580, 480-600, 480-580, or 500-600 amino acids). In some cases, the CasZ protein of interest (e.g., CasZk) has a length in the range of 480-580 amino acids.
[0049] In some cases, the CasZ protein of interest (e.g., CasZl) has a length in the range of 350-500 amino acids (e.g., 350-450, 380-450, 350-420, or 380-420 amino acids). In some cases, the CasZ protein of interest (e.g., CasZl) has a length in the range of 380-420 amino acids.
[0050] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZ protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZ protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZa protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZa protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZa amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZa amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for a sequence that includes one or more amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZa protein of FIG. 1 or FIG. 7 and has a length in the range of 350-800 amino acids (e.g., 350-800, 350-750, 350-700, 350-550, 450-550, 450-750, 450-650, or 450-550 amino acids).
[0051] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZb protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZb protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZb protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZb protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZb amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZb amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZb protein of FIG. 1 or FIG. 7, and has a length in the range of 350-700 amino acids (e.g., 350-650, or 350-620 amino acids).
[0052] In some cases, a subject CasZ protein (of a subject composition and / or method) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZc protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZc protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZc protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZc protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZc amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZc amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZc protein of FIG. 1 or FIG. 7 and has a length in the range of 600-800 amino acids (e.g., 600-650, or 700-800 amino acids).
[0053] In some cases, a subject CasZ protein (of a subject composition and / or method) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZd protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZd protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZd protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZd protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZd amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZd amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZd protein of FIG. 1 or FIG. 7 and has a length in the range of 400-650 amino acids (e.g., 400-600, 400-550, 500-650, 500-600, or 500-550 amino acids).
[0054] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZe amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZe amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe protein of FIG. 1 or FIG. 7 and has a length in the range of 450-700 amino acids (e.g., 450-650, 450-615, 475-700, 475-650, or 475-615 amino acids).
[0055] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZf protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZf protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZf protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZf protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZf amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZf amino acid sequence of Figure 1 or Figure 7 (e.g., when the CasZ protein is dCasZ), except for a sequence that includes amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to the CasZf protein of FIG. 1 or FIG. 7 and has a length in the range of 400-750 amino acids (e.g., 400-700, 700-650, 400-620, 400-600, 400-550, 400-520, 400-500, 400-475, 415-550, 415-520, 415-500, or 415-475 amino acids).
[0056] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZg protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZg protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZg protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZg protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZg amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZg amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences that include amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZg protein of FIG. 1 or FIG. 7 and has a length in the range of 500-750 amino acids (e.g., 500-750 amino acids (e.g., 550-750 amino acids)).
[0057] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZh protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZh protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZh protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZh protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZh amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZh amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to the CasZh protein of FIG. 1 or FIG. 7 and has a length in the range of 380-450 amino acids (e.g., 380-420, 400-450, or 400-420 amino acids).
[0058] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZi protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZi protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZi protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZi protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZi amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZi amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZi protein of FIG. 1 or FIG. 7, and has a length in the range of 700-800 amino acids (e.g., 700-750, 720-800, or 720-750 amino acids).
[0059] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZj protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZj protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZj protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZj protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZj amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZj amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences that include amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZj protein of FIG. 1 or FIG. 7, and has a length in the range of 600-750 amino acids (e.g., 600-700, or 650-700 amino acids).
[0060] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZk protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZk protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZk protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZk protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZk amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZk amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZk protein of FIG. 1 or FIG. 7 and has a length in the range of 450-600 amino acids (e.g., 450-580, 480-600, 480-580, or 500-600 amino acids).
[0061] In some cases, a subject CasZ protein (of a subject composition and / or method) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZl protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZl protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZl protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZl protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZl amino acid sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises the CasZl amino acid sequence of FIG. 1 or FIG. 7 (e.g., when the CasZ protein is dCasZ), except for sequences containing amino acid substitutions (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduce the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZl protein of FIG. 1 or FIG. 7 and has a length in the range of 450-600 amino acids (e.g., 450-580, 480-600, 480-580, or 500-600 amino acids).
[0062] In some cases, a subject CasZ protein (of the subject compositions and / or methods) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZe, CasZf, CasZg, or CasZh protein of Figure 1 or Figure 7. For example, in some cases, a subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZe, CasZf, CasZg, or CasZh protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe, CasZf, CasZg, or CasZh protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZe, CasZf, CasZg, or CasZh protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having the CasZe, CasZf, CasZg, or CasZh protein sequence of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having the CasZe, CasZf, CasZg, or CasZh protein sequence of Figure 1 or Figure 7, except for the sequence containing an amino acid substitution (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduces the native catalytic activity of the protein.In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or greater sequence identity (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100% sequence identity) to a CasZe, CasZf, CasZg, or CasZh protein of Figure 1 or Figure 7, and has a length in the range of 350 to 900 amino acids (e.g., 350 to 850, 350 to 800, 400 to 900, 400 to 850, or 400 to 800 amino acids).
[0063] In some cases, a subject CasZ protein (of a subject composition and / or method) comprises an amino acid sequence having 20% or more sequence identity (e.g., 30% or more, 40% or more, 50% or more, 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein of FIG. 1 or FIG. 7. For example, in some cases, the subject CasZ protein comprises an amino acid sequence having 50% or more sequence identity (e.g., 60% or more, 70% or more, 80% or more, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 80% or more sequence identity (e.g., 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to the CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein of Figure 1 or Figure 7. In some cases, the subject CasZ protein comprises an amino acid sequence having the CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein sequence of Figure 1 or Figure 7.In some cases, the subject CasZ protein comprises an amino acid sequence having the CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein sequence of Figure 1 or Figure 7, except for the sequence containing an amino acid substitution (e.g., one, two, or three amino acid substitutions) (e.g., at one or more catalytic amino acid positions) that reduces the native catalytic activity of the protein. In some cases, the subject CasZ protein comprises an amino acid sequence having 90% or more sequence identity (e.g., 95% or more, 97% or more, 98% or more, 99% or more, or 100% sequence identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, CasZi, CasZj, CasZk, or CasZl protein of Figure 1 or Figure 7, and has a length in the range of 350 to 900 amino acids (e.g., 350 to 850, 350 to 800, 400 to 900, 400 to 850, or 400 to 800 amino acids).
[0064] CasZ variants A variant CasZ protein has an amino acid sequence that differs by at least one amino acid (e.g., has a deletion, insertion, substitution, or fusion) when compared with the amino acid sequence of the corresponding wild-type CasZ protein. A CasZ protein that cleaves one strand of a double-stranded target nucleic acid but not the other is referred to herein as a "nickase" (e.g., "nickase CasZ"). A CasZ protein that has substantially no nuclease activity is referred to herein as a non-viable CasZ protein ("dCasZ") (although nuclease activity can be provided by a heterologous polypeptide, i.e., a fusion partner, in the case of a chimeric CasZ protein, as described in more detail below). For any of the CasZ variant proteins described herein (e.g., nickase CasZ, dCasZ, chimeric CasZ), the CasZ variant can include a CasZ protein sequence with the same parameters (e.g., domains present, percent identity, length, etc.) as described above.
[0065] Variant-catalytic activity In some cases, the CasZ protein is, for example, a variant CasZ protein mutated relative to a native catalytically active sequence and exhibits reduced cleavage activity (e.g., 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, or 30% or less) when compared to the corresponding native sequence. In some cases, such a variant CasZ protein is a catalytically "dead" protein (having substantially no cleavage activity) and may be referred to as "dCasZ." In some cases, the variant CasZ protein is a nickase (which cleaves only one strand of a double-stranded target nucleic acid, e.g., double-stranded target DNA). As described in more detail herein, in some cases, a CasZ protein (in some cases, a CasZ protein having wild-type cleavage activity, in some cases, a variant CasZ having reduced cleavage activity, e.g., dCasZ or nickase CasZ) is fused (conjugated) to a heterologous polypeptide having an activity of interest (e.g., a catalytic activity of interest) to form a fusion protein (chimeric CasZ protein).
[0066] The catalytic residues of CasZ, numbered according to CasZi.1, include D405, E586, and D684 (see, e.g., Figure 1). Thus, in some cases, a CasZ protein has reduced activity, and one or more of the above amino acids (or one or more corresponding amino acids in any CasZ protein) are mutated (e.g., substituted with alanine). In some cases, the variant CasZ protein is catalytically "dead" (catalytically inactive) and is referred to as "dCasZ." The dCasZ protein can be fused to a fusion partner that provides activity, and in some cases, dCasZ (e.g., one that does not have a fusion partner that provides catalytic activity but may have an NLS when expressed in eukaryotic cells) can bind to target DNA, be used for imaging (e.g., the protein can be tagged / labeled), and / or block RNA polymerase from transcribing from target DNA. In some cases, the variant CasZ protein is a nickase (which cleaves only one strand of a double-stranded target nucleic acid, e.g., double-stranded target DNA).
[0067] Variant-chimeric CasZ (i.e., fusion proteins) As described above, in some cases, a CasZ protein (in some cases, a CasZ protein having wild-type cleavage activity, in some cases, a variant CasZ having reduced cleavage activity, e.g., dCasZ or nickase CasZ) is fused (conjugated) to a heterologous polypeptide having an activity of interest (e.g., a catalytic activity of interest) to form a fusion protein (chimeric CasZ protein). The heterologous polypeptide to which the CasZ protein can be fused is referred to herein as a "fusion partner."
[0068] In some cases, the fusion partner can modulate (e.g., inhibit transcription, increase transcription) transcription of the target DNA. For example, in some cases, the fusion partner is a protein (or domain from a protein) that inhibits transcription (e.g., a transcription repressor, a protein that functions via recruitment of a transcription inhibitor protein, modification of target DNA such as methylation, recruitment of a DNA modifier, regulation of histones associated with the target DNA, recruitment of histone modifiers such as those that modify histone acetylation and / or methylation, etc.). In some cases, the fusion partner is a protein (or domain from a protein) that increases transcription (e.g., a transcription activator, a protein that functions via recruitment of a transcription activator protein, modification of target DNA such as demethylation, recruitment of a DNA modifier, regulation of histones associated with the target DNA, recruitment of histone modifiers such as those that modify histone acetylation and / or methylation, etc.).
[0069] In some cases, the chimeric CasZ protein comprises a heterologous polypeptide having an enzymatic activity that modifies a target nucleic acid (e.g., a nuclease activity such as FokI nuclease activity, a methyltransferase activity, a demethylase activity, a DNA repair activity, a DNA damage activity, a deamination activity, a dismutase activity, an alkylation activity, a depurination activity, an oxidation activity, a pyrimidine dimer formation activity, an integrase activity, a transposase activity, a recombinase activity, a polymerase activity, a ligase activity, a helicase activity, a photolyase activity, or a glycosylase activity).
[0070] In some cases, the chimeric CasZ protein comprises a heterologous polypeptide having an enzymatic activity (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, or demyristoylating activity) that modifies a polypeptide (e.g., a histone) associated with a target nucleic acid.
[0071] Examples of proteins (or fragments thereof) that can be used to increase transcription include transcription activators such as VP16, VP64, VP48, VP160, the p65 subdomain (e.g., from NFkB), and the activation domain of EDLL and / or the TAL activation domain (e.g., for activity in plants); histone lysine methyltransferases such as SET1A, SET1B, MLL1-5, ASH1, SYMD2, NSD1; histone lysine demethylases such as JHDM2a / b, UTX, JMJD3; histone acetyltransferases such as GCN5, PCAF, CBP, p300, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, SRC1, ACTR, P160, CLOCK; and Ten-Eleven These include, but are not limited to, DNA demethylases such as Translocation (TET) dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, and ROS1.
[0072] Examples of proteins (or fragments thereof) that can be used to reduce transcription include transcription repressors such as Kruppel-associated box (KRAB or SKD); KOX1 repression domain; Mad mSIN3 interacting domain (SID); ERF repressor domain (ERD), SRDX repression domain (e.g., for repression in plants), etc.; histone lysine methyltransferases such as Pr-SET7 / 8, SUV4-20H1, and RIZ1; histone lysine demethylases such as JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, and JARID1D / SMCY; histone lysine deacetylases such as HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, and HDAC11; HhaI DNA These include, but are not limited to, DNA methylases such as m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants); and peripheral mobilization elements such as Lamin A and Lamin B.
[0073] In some cases, the fusion partner has an enzymatic activity that modifies a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities that can be provided by the fusion partner include nuclease activity, such as that provided by a restriction enzyme (e.g., FokI nuclease), methyltransferase activity, such as that provided by a methyltransferase (e.g., HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plants), ZMET2, CMT1, CMT2 (plants), etc.); demethylase activity, such as that provided by a Ten-Eleven These include, but are not limited to, demethylase activity such as that provided by TET dioxygenase 1 (TET1CD), TET1, DME, DML1, DML2, ROS1, etc.), DNA repair activity, DNA damage activity, deamination activity such as that provided by deaminases (e.g., cytosine deaminase enzymes such as rat APOBEC1), dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity such as that provided by integrases and / or resolvases (e.g., Gin invertase such as a hyperactive mutant of Gin invertase, GinH106Y; human immunodeficiency virus type 1 integrase (IN); Tn3 resolvase, etc.), transposase activity, recombinase activity such as that provided by recombinases (e.g., the catalytic domain of Gin recombinase), polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity.
[0074] In some cases, the fusion partner has an enzymatic activity that modifies a protein (e.g., histone, RNA-binding protein, DNA-binding protein, etc.) associated with a target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activities (that modify a protein associated with a target nucleic acid) that can be provided by a fusion partner include histone methyltransferases (HMTs) (e.g., suppressor of variegation 3-9 homolog 1 (SUV39H1, also known as KMT1A), euchromatin histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1, etc., SET1A, SET1B, MLL1-5, ASH1, SYMD2, NS methyltransferase activity, such as that provided by D1, DOT1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1), histone demethylases (e.g., lysine demethylase 1A (KDM1A, also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, demethylase activity such as that provided by histone acetylase transferases (e.g., catalytic cores / fragments of human acetyltransferases p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HBO1 / MYST2, HMOF / MYST1, SRC1, ACTR, P160, CLOCK, etc.), histone deacetylases (e.g., These include, but are not limited to, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitinating activity, adenylating activity, deadenylating activity, sumoylating activity, desumoylating activity, ribosylation activity, deribosylation activity, myristoylating activity, and demyristoylating activity, such as those provided by HDAC1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.
[0075] Additional examples of suitable fusion partners are dihydrofolate reductase (DHFR) destabilization domains (e.g., to generate chemically controllable chimeric CasZ proteins) and chloroplast transit peptides. Suitable chloroplast transit peptides include: These include, but are not limited to, TIFF0007672445000001.tif145164.
[0076] In some cases, the CasZ fusion polypeptide of the present disclosure comprises a) a CasZ polypeptide of the present disclosure and b) a chloroplast transit peptide. Thus, for example, a CRISPR-CasZ complex can be targeted to chloroplasts. In some cases, this targeting can be achieved by the presence of an N-terminal extension called a chloroplast transit peptide (CTP) or plastid transit peptide. If an expressed polypeptide is to be compartmentalized within a plant plastid (e.g., a chloroplast), a chromosomal transgene from a bacterial source must have a sequence encoding a CTP sequence fused to the sequence encoding the expressed polypeptide. Therefore, localization of an exogenous polypeptide to chloroplasts is often achieved by operably linking a polynucleotide sequence encoding the CTP sequence to the 5' region of a polynucleotide encoding the exogenous polypeptide. The CTP is removed in a processing step during translocation into plastids. However, processing efficiency can be affected by the amino acid sequence of the CTP and nearby sequences at the NH2-terminus of the peptide. Other options for targeting to chloroplasts that have been described are the maize cab-m7 signal sequence (U.S. Pat. No. 7,022,896, WO97 / 41228), the pea glutathione reductase signal sequence (WO97 / 41228), and the CTP described in US2009029861.
[0077] In some cases, a CasZ fusion polypeptide of the disclosure can include a) a CasZ polypeptide of the disclosure, and b) an endosomal escape peptide. In some cases, the endosomal escape polypeptide can include the amino acid sequence TIFF0007672445000002.tif4128, wherein each X is independently selected from lysine, histidine, and arginine. In some cases, the endosomal escape polypeptide has the amino acid sequence Includes TIFF0007672445000003.tif4128.
[0078] For examples of some of the above fusion partners (and others) used in the context of fusion with Cas9, zinc finger, and / or TALE proteins (for site-specific targeted nucleic acid modification, regulation of transcription, and / or targeted protein modification, e.g., histone modification), see, e.g., Nomura et al., J Am Chem Soc. 2007 Jul 18;129(28):8676-7; Rivenbark et al., Epigenetics. 2012 Apr;7(4):350-60; Nucleic Acids Res. 2016 Jul 8;44(12):5615-28; Gilbert et al., Cell. 2013 Jul 18;154(2):442-51; Kearns et al., Nat Methods. 2015 May;12(5):401-3; Mendenhall et al., Nat Biotechnol. 2013 Dec;31(12):1133-6, Hilton et al.,Nat Biotechnol.2015 May;33(5):510-7, Gordley et al.,Proc Natl Acad Sci USA.2009 Mar 31;106(13):5053-8, Akopian et al.,Proc Natl Acad Sci USA.2003 Jul 22;100(15):8688-91, Tan et.,al.,J Virol.2006 Feb;80(4):1939-48, Tan et al.,Proc Natl Acad Sci USA.2003 Oct 14;100(21):11997-2002, Papworth et al.,Proc Natl Acad Sci USA.2003 Feb 18;100(4):1621-6, Sanjana et al., Nat Protoc.2012 Jan 5;7(1):171-92, Beerli et al., Proc Natl Acad Sci USA.1998 Dec 8;95(25):14628-33, Snowden et al., Curr Biol.2002 Dec 23;12(24):2159-66, Xu et.al.,Xu et.al.,Cell Discov.See 2016 May 3;2:16009, Komor et al., Nature. 2016 Apr 20;533(7603):420-4, Chaikind et al., Nucleic Acids Res. 2016 Aug 11, Choudhury at.al., Oncotarget. 2016 Jun 23, Du et al., Cold Spring Harb Protoc. 2016 Jan 4, Pham et al., Methods Mol Biol. 2016;1358:43-57, Balboa et al., Stem Cell Reports. 2015 Sep 8;5(3):448-59, Hara et al., Sci Rep. 2015 Jun 9;5:11221, Piatek et al., Plant Biotechnol J. 2015 May;13(4):578-89, Hu et al., Nucleic Acids Res. 2014 Apr;42(7):4375-90, Cheng et al., Cell Res. 2013 Oct;23(10):1163-71, and Maeder et.al., Nat Methods. 2013 Oct;10(10):977-9.
[0079] Additional suitable heterologous polypeptides include, but are not limited to, polypeptides that directly and / or indirectly provide increased transcription and / or translation of a target nucleic acid (e.g., transcriptional activators or fragments thereof, proteins or fragments thereof that recruit transcriptional activators, small molecule / drug-responsive transcriptional and / or translational regulators, translational regulatory proteins, etc.). Non-limiting examples of heterologous polypeptides that achieve increased or decreased transcription include transcriptional activator and transcriptional repressor domains. In some such cases, a chimeric CasZ polypeptide is targeted to a specific location (i.e., sequence) in a target nucleic acid by a guide nucleic acid (guide RNA), resulting in locus-specific regulation, such as blocking RNA polymerase binding to the promoter (selectively inhibiting transcriptional activator function) and / or modifying the local chromatin state (e.g., when a fusion sequence that modifies the target nucleic acid or a polypeptide associated with the target nucleic acid is used). In some cases, the change is transient (e.g., transcriptional repression or activation). In some cases, the change is heritable (eg, when an epigenetic modification is made to the target nucleic acid or to a protein associated with the target nucleic acid, such as a nucleosomal histone).
[0080] Non-limiting examples of heterologous polypeptides for use in targeting ssRNA target nucleic acids include, but are not limited to, splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or release factors, e.g., eIF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases, e.g., adenosine deaminase acting on RNA (ADAR)), including A to I and / or C to U editing enzymes); helicases; RNA binding proteins, and the like. It is understood that a heterologous polypeptide can comprise an entire protein or, in some cases, can comprise a fragment of a protein (e.g., a functional domain).
[0081] The heterologous polypeptide of a subject chimeric CasZ polypeptide can be any domain (for purposes of this disclosure, this includes intramolecular and / or intermolecular secondary structures, e.g., double-stranded RNA duplexes such as hairpins, stem-loops, etc.) that can interact with ssRNA, whether transiently or irreversibly, directly or indirectly, including endonucleases (e.g., RNase III from proteins such as SMG5 and SMG6, CRR22 DYW domain, Dicer, and PIN (PilT N-terminal) domain); proteins and protein domains involved in stimulation of RNA cleavage (e.g., CPSF, CstF, CFIm, and CFIIm); exonucleases (e.g., XRN-1 or exonuclease T); deadenylases (e.g., HNT3); proteins and protein domains involved in nonsense-mediated RNA decay (e.g., UPF1, UPF2, UPF3, UPF3b, RNP S1, Y14, DEK, REF2, and SRm160); proteins and protein domains involved in RNA stabilization (e.g., PABP); proteins and protein domains involved in translational repression (e.g., Ago2 and Ago4); proteins and protein domains involved in translational stimulation (e.g., Staufen); proteins and protein domains involved in (e.g., capable of regulating) translation (e.g., translation factors such as initiation factors, elongation factors, and release factors, e.g., eIF4G); proteins and protein domains involved in RNA polyadenylation (e.g., PAP1, GLD-2, and Star-PAP); proteins and protein domains involved in RNA polyuridylation (e.g., CI D1 and terminal uridylate transferase); proteins and protein domains involved in RNA localization (e.g., from IMP1, ZBP1, She2p, She3p, and Bicaudal-D); proteins and protein domains involved in the nuclear retention of RNA (e.g., Rrp6); proteins and protein domains involved in the nuclear export of RNA (e.g., TAP, NXF1, THO, TREX, REF, and Aly); proteins and protein domains involved in the repression of RNA splicing (e.g., PTB, Sam68, and hnRNP A1);These include, but are not limited to, effector domains selected from the group including proteins and protein domains involved in stimulating RNA splicing (e.g., serine / arginine-rich (SR) domains); proteins and protein domains involved in reducing transcription efficiency (e.g., FUS (TLS)); and proteins and protein domains involved in stimulating transcription (e.g., CDK7 and HIV Tat). Alternatively, the effector domain may be selected from the group comprising endonucleases; proteins and protein domains capable of stimulating RNA cleavage; exonucleases; deadenylases; proteins and protein domains with nonsense-mediated RNA degradation activity; proteins and protein domains capable of stabilizing RNA; proteins and protein domains capable of repressing translation; proteins and protein domains capable of stimulating translation; proteins and protein domains capable of regulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, e.g., eIF4G); proteins and protein domains capable of polyadenylation of RNA; proteins and protein domains capable of polyuridylation of RNA; proteins and protein domains with RNA localization activity; proteins and protein domains capable of nuclear retention of RNA; proteins and protein domains with RNA nuclear export activity; proteins and protein domains capable of suppressing RNA splicing; proteins and protein domains capable of stimulating RNA splicing; proteins and protein domains capable of reducing transcription efficiency; and proteins and protein domains capable of stimulating transcription. Another suitable heterologous polypeptide is a PUF RNA-binding domain, as described in more detail in WO2012068627, which is incorporated herein by reference in its entirety.
[0082] Some RNA splicing factors that can be used (in whole or as fragments) as heterologous polypeptides in chimeric CasZ polypeptides have a modular organization, possessing distinct sequence-specific RNA-binding modules and splicing effector domains. For example, members of the serine / arginine-rich (SR) protein family contain an N-terminal RNA recognition motif (RRM) that binds to exon splicing enhancers (ESEs) in pre-mRNAs and a C-terminal RS domain that promotes exon inclusion. As another example, the hnRNP protein hnRNP A1 binds to exon splicing silencers (ESSs) through its RRM domain and inhibits exon inclusion through its C-terminal glycine-rich domain. Some splicing factors can regulate the use of alternative splice sites (SSs) by binding to regulatory sequences between the two alternative sites. For example, ASF / SF2 can recognize an ESE and promote the use of an intron-proximal site, whereas hnRNP A1 can bind to an ESS and redirect splicing to the use of an intron-distal site. One application of such factors is to generate ESFs that regulate alternative splicing of endogenous genes, particularly disease-related genes. For example, Bcl-x pre-mRNA produces two splicing isoforms with two alternative 5' splice sites, encoding proteins with opposing functions. The long splicing isoform, Bcl-xL, is a potent apoptosis inhibitor expressed in long-lived, postmitotic cells and is upregulated in many cancer cells, protecting them from apoptotic signals. The short isoform, Bcl-xS, is a pro-apoptotic isoform expressed at high levels in cells with a high turnover rate (e.g., developing lymphocytes). The ratio of the two Bcl-x splicing isoforms is regulated by multiple c-elements located either in the core exon region or in the exon extension region (i.e., between the two alternative 5' splice sites). For further examples, see WO2010075303 (incorporated herein by reference in its entirety).
[0083] Further suitable fusion partners include, but are not limited to, proteins (or fragments thereof) that are boundary elements (e.g., CTCF), proteins and fragments thereof that provide peripheral recruitment (e.g., Lamin A, Lamin B, etc.), protein docking elements (e.g., FKBP / FRB, Pil1 / Aby1, etc.).
[0084] Examples of various additional suitable heterologous polypeptides (or fragments thereof) to the subject chimeric CasZ polypeptides include, but are not limited to, those described in the following applications (which publications relate to other CRISPR endonucleases, such as Cas9, although the fusion partners described can alternatively be used with CasZ): PCT patent applications: WO2010075303, WO2012068627, and WO2013155555; and U.S. patents and patent applications, e.g., U.S. patents and applications Ser. Nos. 8,906,616, 8,895,308, 8,889,418, 8,889,356, 8,871,445, 8,865,406, 8,795,965, 8,771,945, 8,697,No. 359, No. 20140068797, No. 20140170753, No. 20140179006, No. 20140179770, No. 20140186843, No. 20140186919, No. 20140186958, 20140189896, 20140227787, 20140234972, 20140242664, 20140242699, 201402 No. 42700, No. 20140242702, No. 20140248702, No. 20140256046, No. 20140273037, No. 20140273226, No. 20140273230 , No. 20140273231, No. 20140273232, No. 20140273233, No. 20140273234, No. 20140273235, No. 20140287938, No. 2014 No. 0295556, No. 20140295557, No. 20140298547, No. 20140304853, No. 20140309487, No. 20140310828, No. 2014031083 No. 0, No. 20140315985, No. 20140335063, No. 20140335620, No. 20140342456, No. 20140342457, No. 20140342458, No. 20 Nos. 140349400, 20140349405, 20140356867, 20140356956, 20140356958, 20140356959, 20140357523, 20140357530, 20140364333, and 20140377868 (all of which are incorporated herein by reference in their entirety).
[0085] In some cases, the heterologous polypeptide (fusion partner) provides subcellular localization. That is, the heterologous polypeptide contains a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence that retains the fusion protein outside the nucleus, e.g., a nuclear export sequence (NES), a sequence that retains the fusion protein in the cytoplasm, a mitochondrial localization signal for targeting to mitochondria, a chloroplast localization signal for targeting to chloroplasts, an ER retention signal, etc.). In some embodiments, the CasZ fusion polypeptide does not contain an NLS, thereby preventing the protein from being targeted to the nucleus (which can be advantageous, for example, when the target nucleic acid is RNA present in the cytosol). In some embodiments, the heterologous polypeptide can be provided with a tag to facilitate tracking and / or purification (i.e., the heterologous polypeptide is detectably labeled) (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, etc.; a histidine tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag, etc.).
[0086] In some cases, a CasZ protein (e.g., a wild-type CasZ protein, a variant CasZ protein, a chimeric CasZ protein, a dCasZ protein, a chimeric CasZ protein in which a CasZ portion has reduced nuclease activity, e.g., a dCasZ protein fused to a fusion partner, etc.) includes (is fused to) a nuclear localization signal (NLS) (e.g., in some cases, two or more, three or more, four or more, or five or more NLSs). Thus, in some cases, a CasZ polypeptide includes one or more NLSs (e.g., two or more, three or more, four or more, or five or more NLSs). In some cases, the one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the N-terminus and / or C-terminus. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the N-terminus. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the C-terminus. In some cases, one or more NLSs (three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) both the N-terminus and the C-terminus. In some cases, an NLS is located at the N-terminus and an NLS is located at the C-terminus.
[0087] In some cases, a CasZ protein (e.g., a wild-type CasZ protein, a variant CasZ protein, a chimeric CasZ protein, a dCasZ protein, a chimeric CasZ protein having a CasZ portion with reduced nuclease activity, e.g., a dCasZ protein fused to a fusion partner, etc.) comprises (is fused to) one to ten NLSs (e.g., one to nine, one to eight, one to seven, one to six, one to five, two to ten, two to nine, two to eight, two to seven, two to six, or two to five NLSs). In some cases, a CasZ protein (e.g., a wild-type CasZ protein, a variant CasZ protein, a chimeric CasZ protein, a dCasZ protein, a chimeric CasZ protein having a CasZ portion with reduced nuclease activity, e.g., a dCasZ protein fused to a fusion partner, etc.) comprises (is fused to) two to five NLSs (e.g., two to four or two to three NLSs).
[0088] Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen, having the amino acid sequence PKKKRKV (SEQ ID NO:114); an NLS from nucleoplasmin (e.g., the sequence Nucleoplasmin double NLS with TIFF0007672445000004.tif4128); amino acid sequence c-myc NLS with TIFF0007672445000005.tif4128; sequence hRNPA1 M9 NLS with TIFF0007672445000006.tif4131; sequence of the IBB domain from importin-alpha TIFF0007672445000007.tif4139; the sequences of VSRKRPRP (SEQ ID NO: 120) and PPKKARED (SEQ ID NO: 121) of fibroid T protein; the sequence of PQPKKKPL (SEQ ID NO: 122) of human p53; the sequence of SALIKKKKKMAP (SEQ ID NO: 123) of mouse c-abl IV; the sequences of DRLRR (SEQ ID NO: 124) and PKQKKRK (SEQ ID NO: 125) of influenza virus NS1; the sequence of RKLKKKIKKL (SEQ ID NO: 126) of hepatitis virus delta antigen; the sequence of REKKKFLKRR (SEQ ID NO: 127) of mouse Mx1 protein; and the sequence of human poly(ADP-ribose) polymerase. TIFF0007672445000008.tif4128; and the sequence of the steroid hormone receptor (human) glucocorticoid Included are NLS sequences derived from TIFF0007672445000009.tif4128. Generally, NLS (or multiple NLSs) are strong enough to drive the accumulation of detectable amounts of CasZ in the nucleus of eukaryotic cells. The detection of accumulation in the nucleus can be carried out by any suitable technique. For example, a detectable marker can be fused with CasZ protein, thereby making its location within the cell visible. Cell nuclei can also be isolated from cells, and then their contents can be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus can also be determined indirectly.
[0089] In some cases, the CasZ fusion polypeptide comprises a "protein transduction domain" or PTD (also known as a CPP—cell-penetrating peptide), which refers to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates crossing of a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule and / or nanoparticle, which can range from small polar molecules to large macromolecules, facilitates membrane crossing of molecules, for example, moving from the extracellular space to the intracellular space or from the cytosol into an organelle. In some embodiments, the PTD is covalently linked to the amino acid terminus of the polypeptide (e.g., linked to wild-type CasZ to generate a fusion protein or linked to a variant CasZ protein, e.g., dCasZ, nickase CasZ, or chimeric CasZ protein to generate a fusion protein). In some embodiments, the PTD is covalently linked to the carboxyl terminus of the polypeptide (e.g., linked to wild-type CasZ to generate a fusion protein, or linked to a variant CasZ protein, e.g., dCasZ, nickase CasZ, or chimeric CasZ protein to generate a fusion protein). In some cases, the PTD is inserted internally into the CasZ fusion polypeptide at a suitable insertion site (i.e., not at the N- or C-terminus of the CasZ fusion polypeptide). In some cases, the subject CasZ fusion polypeptide comprises (conjugated to, fused to) one or more PTDs (e.g., two or more, three or more, four or more PTDs). In some cases, the PTD comprises a nuclear localization signal (NLS) (e.g., in some cases, two or more, three or more, four or more, or five or more NLSs). Thus, in some cases, the CasZ fusion polypeptide comprises one or more NLSs (e.g., two or more, three or more, four or more, or five or more NLSs). In some cases, the PTD is covalently linked to a nucleic acid (e.g., a CasZ guide nucleic acid, a polynucleotide encoding a CasZ guide nucleic acid, a polynucleotide encoding a CasZ fusion polypeptide, a donor polynucleotide, etc.).Examples of PTDs include a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT, including YGRKKRRQRRR (SEQ ID NO:130)); a polyarginine sequence containing a sufficient number of arginines (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines) for direct entry into a cell; a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); a Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008);RRQRRTSKLMKR (SEQ ID NO: 131);Transportan. Exemplary PTDs include, but are not limited to, TIFF0007672445000010.tif17132. TIFF0007672445000011.tif4128; arginine homopolymers of 3 to 50 arginine residues. Exemplary PTD domain amino acid sequences include, but are not limited to: TIFF0007672445000012.tif17156. In some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June;1(5-6):371-381). ACPPs contain a polycationic CPP (e.g., Arg9 or "R9") connected to a matching polyanion (e.g., Glu9 or "E9") via a cleavable linker, which reduces the net charge to near zero, thereby inhibiting cellular attachment and uptake. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesive properties, thus "activating" the ACPP to cross membranes.
[0090] Linkers (e.g., for fusion partners) In some cases, the CasZ protein of interest is fused to the fusion partner via a linker polypeptide (e.g., one or more linker polypeptides). The linker polypeptide can have any of a variety of amino acid sequences. Proteins can be joined by a generally flexible spacer peptide, although other chemical bonds are not excluded. Suitable linkers include polypeptides between 4 and 40 amino acids in length, or between 4 and 25 amino acids in length. These linkers can be produced by linking proteins using synthetic linker-encoding oligonucleotides, or can be encoded by a nucleic acid sequence encoding the fusion protein. Peptide linkers with some degree of flexibility can be used. The linking peptide can have virtually any amino acid sequence, keeping in mind that preferred linkers generally have sequences that result in flexible peptides. The use of small amino acids, such as glycine and alanine, is useful in creating flexible peptides. Creating such sequences is routine for those of skill in the art. A variety of different linkers are commercially available and may be suitable for use.
[0091] Examples of linker polypeptides include glycine polymers (G) n , glycine-serine polymers (e.g., (GS) n , G.S.G.S.G.S. n (SEQ ID NO:140), GGSGGS n (SEQ ID NO: 141), and GGGS n (SEQ ID NO:142), where n is an integer of at least 1), glycine-alanine polymers, and alanine-serine polymers. Exemplary linkers include: TIFF0007672445000013.tif17158, etc. Those skilled in the art will recognize that the design of a peptide conjugated with any desired element can include a linker that is fully or partially flexible, whereby the linker can include a flexible linker as well as one or more moieties that confer a less flexible structure.
[0092] Detectable Label In some cases, the CasZ polypeptides of the present disclosure comprise (can be linked to / fused to) a detectable label. Suitable detectable labels and / or moieties capable of providing a detectable signal can include, but are not limited to, enzymes, radioisotopes, members of specific binding pairs, fluorophores, fluorescent proteins, quantum dots, etc.
[0093] Suitable fluorescent proteins include green fluorescent protein (GFP) or variants thereof, blue fluorescent variants of GFP (BFP), cyan fluorescent variants of GFP (CFP), yellow fluorescent variants of GFP (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilized EGFP (dEGFP), destabilized ECFP (dECFP), destabilized EYFP (dEYFP), mCFPm, Cerulean, T-Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J-Red, dimer2, t-dimer2(12), mRFP1, pocilloporin, Renilla GFP, and Monster. Examples of fluorescent proteins include, but are not limited to, GFP, paGFP, Kaede protein and kindling protein, phycobiliproteins, and phycobiliprotein conjugates including B-phycoerythrin, R-phycoerythrin, and allophycocyanin. Other examples of fluorescent proteins include mHoneydew, mBanana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrape1, mRaspberry, mGrape2, and mPlum (Shaner et al. (2005) Nat. Methods 2:905-909). Suitable for use are any of the various fluorescent and chromatic proteins from anthozoan species, such as those described in Matz et al. (1999) Nature Biotechnol. 17:969-973.
[0094] Suitable enzymes include, but are not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), beta-galactosidase (GAL), glucose-6-phosphatase dehydrogenase, beta-N-acetylglucosaminidase, β-glucuronidase, invertase, xanthine oxidase, firefly luciferase, glucose oxidase (GO), and the like.
[0095] Protospacer adjacent motif (PAM) Naturally occurring CasZ proteins bind to target DNA at target sequences defined by the region of complementarity between the DNA-targeting RNA and the target DNA. As with many CRISPR endonucleases, site-specific binding (and / or cleavage) of double-stranded target DNA occurs at a location determined by both (i) base-pairing complementarity between the guide RNA and the target DNA and (ii) a short motif in the target DNA, termed a protospacer adjacent motif (PAM).
[0096] In some embodiments, the PAM of a CasZ protein is immediately 5' to the target sequence of a non-complementary strand of a target DNA (also referred to as the non-target strand; a complementary strand hybridizes to the guide sequence of a guide RNA, while a non-complementary strand does not directly hybridize to the guide RNA and is the reverse complement of the non-complementary strand). In some cases (e.g., in the case of CasZc), the PAM sequence of the non-complementary strand is 5'-TTA-3'. In some cases (e.g., in the case of CasZb), the PAM sequence of the non-complementary strand is 5'-TTTN-3'. In some cases (e.g., in the case of CasZb), the PAM sequence of the non-complementary strand is 5'-TTTA-3'.
[0097] In some cases, different CasZ proteins (i.e., CasZ proteins from various species) may be advantageous for use in the various provided methods to take advantage of different enzymatic characteristics of different CasZ proteins (e.g., because of different PAM sequence preferences; because of increased or decreased enzymatic activity; because of increased or decreased levels of cytotoxicity; to alter the balance between NHEJ, homology-directed repair, single-strand breaks, double-strand breaks, etc.; to utilize shorter overall sequences, etc.). CasZ proteins from different species may require different PAM sequences in the target DNA. Thus, for a particular CasZ protein selected, the PAM sequence preference may differ from the sequence(s) described above. Various methods (including in silico and / or wet-lab methods) for identifying appropriate PAM sequences are known and routine in the art, and any convenient method can be used.
[0098] CasZ guide RNA A nucleic acid molecule that binds to a CasZ protein, forms a ribonucleoprotein complex (RNP), and targets that complex to a specific location within a target nucleic acid (e.g., target DNA) is referred to herein as a "CasZ guide RNA" or simply a "guide RNA." It should be understood that in some cases, a hybrid DNA / RNA can be made such that the CasZ guide RNA contains DNA bases in addition to RNA bases, but the term "CasZ guide RNA" is still used herein to encompass such molecules.
[0099] A CasZ guide RNA can be said to contain two segments: a targeting segment and a protein-binding segment. The targeting segment of a CasZ guide RNA contains a nucleotide sequence (guide sequence) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid (e.g., a target ssRNA, a target ssDNA, the complementary strand of a double-stranded target DNA, etc.). The protein-binding segment (or "protein-binding sequence") interacts with (binds to) a CasZ polypeptide. The protein-binding segment of a subject CasZ guide RNA contains two complementary stretches of nucleotides that hybridize to each other to form a double-stranded RNA duplex (dsRNA duplex). Site-specific binding and / or cleavage of a target nucleic acid (e.g., genomic DNA) can occur at a location (e.g., a target sequence at a target locus) determined by base-pairing complementarity between the CasZ guide RNA (the guide sequence of the CasZ guide RNA) and the target nucleic acid.
[0100] The CasZ guide RNA and the CasZ protein, e.g., a fusion CasZ polypeptide, form a complex (e.g., bind via non-covalent interactions). The CasZ guide RNA provides target specificity to the complex by including a targeting segment that includes a guide sequence (a nucleotide sequence complementary to the sequence of a target nucleic acid). The CasZ protein of the complex provides site-specific activity (e.g., cleavage activity provided by the CasZ protein and / or activity provided by a fusion partner in the case of a chimeric CasZ protein). In other words, the CasZ protein is guided to a target nucleic acid sequence (e.g., a target sequence) by its association with the CasZ guide RNA.
[0101] The "guide sequence," also referred to as the "targeting sequence" of a CasZ guide RNA, can be modified such that the CasZ guide RNA can target a CasZ protein (e.g., a native CasZ protein, a fusion CasZ polypeptide (chimeric CasZ), etc.) to any desired sequence in any desired target nucleic acid, with the exception that a PAM sequence can be taken into account (e.g., as described herein). Thus, for example, a CasZ guide RNA can have a guide sequence that is complementary to (e.g., can hybridize to) a sequence in a nucleic acid in a eukaryotic cell, e.g., a viral nucleic acid, a eukaryotic nucleic acid (e.g., a eukaryotic chromosome, a chromosomal sequence, a eukaryotic RNA, etc.).
[0102] In some cases, the CasZ guide RNA has a length of 30 nucleotides (nt) or more (e.g., 35 nt or more, 40 nt or more, 45 nt or more, 50 nt or more, 55 nt or more, 60 nt or more). In some embodiments, the CasZ guide RNA has a length of 40 nucleotides (nt) or more (e.g., 45 nt or more, 50 nt or more, 55 nt or more, 60 nt or more). In some cases, the CasZ guide RNA has a length of 30 nucleotides (nt) to 100 nt (e.g., 30-90, 30-80, 30-75, 30-70, 30-65, 40-100, 40-90, 40-80, 40-75, 40-70, or 40-65 nt). In some cases, the CasZ guide RNA has a length of 40 nucleotides (nt) to 100 nt (e.g., 40-90, 40-80, 40-75, 40-70, or 40-65 nt).
[0103] Guide sequence of CasZ guide RNA The subject CasZ guide RNA comprises a guide sequence (i.e., a targeting sequence), which is a nucleotide sequence complementary to a sequence (target site) in the target nucleic acid. In other words, the guide sequence of the CasZ guide RNA can interact with the target nucleic acid (e.g., double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), single-stranded RNA (ssRNA), or double-stranded RNA (dsRNA)) in a sequence-specific manner by hybridization (i.e., base pairing). The guide sequence of the CasZ guide RNA can be modified (e.g., by genetic engineering) / designed (e.g., taking into account PAM, e.g., when targeting a dsDNA target) to hybridize to any desired target sequence within the target nucleic acid (e.g., a eukaryotic target nucleic acid such as genomic DNA).
[0104] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 65% or greater, 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%). In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100%.
[0105] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over the seven contiguous 3'-most nucleotides of the target site of the target nucleic acid.
[0106] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 17 or more (e.g., 18 or greater, 19 or greater, 20 or greater, 21 or greater, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 17 or more (e.g., 18 or greater, 19 or greater, 20 or greater, 21 or greater, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 17 or more (e.g., 18 or greater, 19 or greater, 20 or greater, 21 or greater, 22 or greater) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 17 or more (e.g., 18 or greater, 19 or greater, 20 or greater, 21 or greater, 22 or greater) contiguous nucleotides.
[0107] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19 or more (e.g., 20% or greater, 21% or greater, 22% or greater) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19 or more (e.g., 20% or greater, 21% or greater, 22% or greater) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19 or more (e.g., 20 or more, 21 or more, 22 or more) contiguous nucleotides.
[0108] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 17-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 17-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 17-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 17-25 contiguous nucleotides.
[0109] In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 60% or greater (e.g., 70% or greater, 75% or greater, 80% or greater, 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 80% or greater (e.g., 85% or greater, 90% or greater, 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 90% or greater (e.g., 95% or greater, 97% or greater, 98% or greater, 99% or greater, or 100%) over 19-25 contiguous nucleotides. In some cases, the percent complementarity between the guide sequence and the target site of the target nucleic acid is 100% over 19-25 contiguous nucleotides.
[0110] In some cases, the guide sequence has a length in the range of 17-30 nucleotides (nt) (e.g., 17-25, 17-22, 17-20, 19-30, 19-25, 19-22, 19-20, 20-30, 20-25, or 20-22 nt). In some cases, the guide sequence has a length in the range of 17-25 nucleotides (nt) (e.g., 17-22, 17-20, 19-25, 19-22, 19-20, 20-25, or 20-22 nt). In some cases, the guide sequence has a length of 17 nt or more (e.g., 18 or more, 19 or more, 20 or more, 21 or more, or 22 or more nt; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 19 nt or more (e.g., 20 nt or more, 21 nt or more, or 22 nt or more; 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, etc.). In some cases, the guide sequence has a length of 17 nt. In some cases, the guide sequence has a length of 18 nt. In some cases, the guide sequence has a length of 19 nt. In some cases, the guide sequence has a length of 20 nt. In some cases, the guide sequence has a length of 21 nt. In some cases, the guide sequence has a length of 22 nt. In some cases, the guide sequence has a length of 23 nt.
[0111] Protein-binding segment of CasZ guide RNA The protein-binding segment of the target CasZ guide RNA interacts with CasZ protein. CasZ guide RNA guides the bound CasZ protein to a specific nucleotide sequence in the target nucleic acid via the above-mentioned guide sequence. The protein-binding segment of CasZ guide RNA is complementary to each other and comprises two stretches of nucleotides that hybridize to form a double-stranded RNA duplex (dsRNA duplex). Thus, the protein-binding segment comprises a dsRNA duplex.
[0112] In some cases, the dsRNA duplex region comprises a range of 5 to 25 base pairs (bp) (e.g., 5 to 22, 5 to 20, 5 to 18, 5 to 15, 5 to 12, 5 to 10, 5 to 8, 8 to 25, 8 to 22, 8 to 18, 8 to 15, 8 to 12, 12 to 25, 12 to 22, 12 to 18, 12 to 15, 13 to 25, 13 to 22, 13 to 18, 13 to 15, 14 to 25, 14 to 22, 14 to 18, 14 to 15, 15 to 25, 15 to 22, 15 to 18, 17 to 25, 17 to 22, or 17 to 18 bp, e.g., 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the dsRNA duplex region comprises a range of 6 to 15 base pairs (bp) (e.g., 6 to 12, 6 to 10, or 6 to 8 bp, e.g., 6 bp, 7 bp, 8 bp, 9 bp, 10 bp, etc.). In some cases, the duplex region comprises 5 or more bp (e.g., 6 or more, 7 or more, or 8 or more bp). In some cases, the duplex region comprises 6 or more bp (e.g., 7 or more, or 8 or more bp). In some cases, not all nucleotides in the duplex region are paired, and thus the duplex-forming region can comprise a bulge. As used herein, the term "bulge" refers to a stretch of nucleotides (which may be one nucleotide or multiple nucleotides) that do not contribute to the double-stranded duplex and are surrounded 5' and 3' by contributing nucleotides, and as such, the bulge is considered part of the duplex region. In some cases, the dsRNA comprises one or more bulges (e.g., two or more, three or more, four or more bulges). In some cases, the dsRNA duplex contains two or more bulges (e.g., three or more, four or more bulges), and in some cases, the dsRNA duplex contains one to five bulges (e.g., one to four, one to three, two to five, two to four, or two to three bulges).
[0113] Thus, in some cases, stretches of nucleotides that hybridize to each other to form a dsRNA duplex have 70% to 100% complementarity to each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, stretches of nucleotides that hybridize to each other to form a dsRNA duplex have 70% to 100% complementarity to each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, stretches of nucleotides that hybridize to each other to form a dsRNA duplex have 85% to 100% complementarity to each other (e.g., 90% to 100%, 95% to 100% complementarity). In some cases, the stretches of nucleotides that hybridize to each other to form the dsRNA duplex have 70%-95% complementarity to each other (e.g., 75%-95%, 80%-95%, 85%-95%, 90%-95% complementarity).
[0114] In other words, in some embodiments, the dsRNA duplex comprises two stretches of nucleotides that have 70% to 100% complementarity to each other (e.g., 75% to 100%, 80% to 10%, 85% to 100%, 90% to 100%, 95% to 100% complementarity). In some cases, the dsRNA duplex comprises two stretches of nucleotides that have 85% to 100% complementarity to each other (e.g., 90% to 100%, 95% to 100% complementarity). In some cases, the dsRNA duplex comprises two stretches of nucleotides that have 70% to 95% complementarity to each other (e.g., 75% to 95%, 80% to 95%, 85% to 95%, 90% to 95% complementarity).
[0115] The duplex region of a subject CasZ guide RNA can include one or more (1, 2, 3, 4, 5, etc.) mutations relative to the native duplex region. For example, in some cases, base pairs can be maintained, while the nucleotides contributing to the base pair from each segment can be different. In some cases, the duplex region of a subject CasZ guide RNA includes more base pairs, fewer base pairs, smaller bulges, larger bulges, fewer bulges, more bulges, or any convenient combination thereof, compared to the native duplex region (of the native CasZ guide RNA).
[0116] Examples of various Cas9 and cpf1 guide RNAs can be found in the art, and in some cases, modifications similar to those introduced into Cas9 guide RNAs can also be introduced into the CasZ guide RNAs of the present disclosure (e.g., mutations to the dsRNA duplex region, extensions at the 5' or 3' ends to add stabilization to provide for interaction with another protein, etc.). For example, Jinek et al.,Science.2012 Aug 17;337(6096):816-21, Chylinski et al.,RNA Biol.2013 May;10(5):726-37, Ma et al.,Biomed Res Int.2013;2013:270805, Hou et al.,Proc Natl Acad Sci USA.2013 Sep 24;110(39):15644-9, Jinek et al.,Elife.2013;2:e00471, Pattanayak et al.,Nat Biotechnol.2013 Sep;31(9):839-43, Qi et al,Cell.2013 Feb 28;152(5):1173-83, Wang et al. al.,Cell.2013 May 9;153(4):910-8, Auer et al.,Genome Res.2013 Oct 31, Chen et al.,Nucleic Acids Res.2013 Nov 1;41(20):e19, Cheng et al.,Cell Res.2013 Oct;23(10):1163-71, Cho et al.,Genetics.2013 Nov;195(3):1177-80, DiCarlo et al., Nucleic Acids Res.2013 Apr;41(7):4336-43, Dickinson et al., Nat Methods.2013 Oct;10(10):1028-34, Ebina et al., Sci Rep.2013;3:2510, Fujii et.al,Nucleic Acids Res.2013 Nov 1;41(20):e187, Hu et al., Cell Res.2013 Nov;23(11):1322-5, Jiang et al., Nucleic Acids Res.2013 Nov 1;41(20):e188、Larson et al.,Nat Protoc.2013 Nov;8(11):2180-96、Mali et.at.,Nat Methods.2013 Oct;10(10):957-63、Nakayama et al.,Genesis.2013 Dec;51(12):835-43、Ran et al.,Nat Protoc.2013 Nov;8(11):2281-308、Ran et al.,Cell.2013 Sep 12;154(6):1380-9、Upadhyay et al.,G3(Bethesda).2013 Dec 9;3(12):2233-8、Walsh et al.,Proc Natl Acad Sci USA.2013 Sep 24;110(39):15514-5、Xie et al.,Mol Plant.2013 Oct 9、Yang et al.,Cell.2013 Sep 12;154(6):1370-9、Briner et al.,Mol Cell.2014 Oct 23;56(2):333-9, and U.S. Patents and Patent Applications Nos. 8,906,616, 8,895,308, 8,889,418, 8,889,356, 8,871,445, 8,865,406, 8,795,965, 8,771,945, 8,697,359, 20140068797, 20140170753, 20140179006, 20140179770, 20140186843, 2014 No. 0186919, No. 20140186958, No. 20140189896, No. 20140227787, No. 20140234972, No. 20140242664, No. 20140242699, No. 20140242700 , No. 20140242702, No. 20140248702, No. 20140256046, No. 20140273037, No. 20140273226, No. 20140273230, No. 20140273231, No. 2014027 No. 3232, No. 20140273233, No. 20140273234, No. 20140273235, No. 20140287938, No. 20140295556, No. 20140295557, No. 20140298547, No. 2 No. 0140304853, No. 20140309487, No. 20140310828, No. 20140310830, No. 20140315985, No. 20140335063, No. 20140335620, No. 201403424 56, 20140342457, 20140342458, 20140349400, 20140349405, 20140356867, 20140356956, 20140356958, 20140356959, 20140357523, 20140357530, 20140364333, and 20140377868, all of which are incorporated by reference herein in their entirety.
[0117] A CasZ guide RNA guide comprises both a guide sequence and two stretches of nucleotides ("duplex-forming segments") that hybridize to form a dsRNA duplex of the protein-binding segment. The specific sequence of a given CasZ guide RNA can be characteristic of the species in which the crRNA is found. Examples of suitable CasZ guide RNAs are provided herein.
[0118] Exemplary guide RNA sequences The repeat sequences (non-guide sequence portions of CasZ guide RNAs) of crRNAs of naturally occurring CasZ proteins (see, e.g., Figures 1 and 7) are shown in Tables 1 and 3.
[0119] Table 1. CasZ protein crRNA repeat sequences TIFF0007672445000014.tif222165TIFF0007672445000015.tif127165
[0120] In some cases, the subject CasZ guide RNA comprises a crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a crRNA sequence in Table 1 or Table 3.
[0121] In some cases, a subject CasZ guide RNA comprises (e.g., in addition to a guide sequence, e.g., as part of a protein binding region) a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, or CasZi crRNA sequence of Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or greater identity (e.g., 75% or greater, 80% or greater, 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, or CasZi crRNA sequence of Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, or CasZi crRNA sequence in Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZa, CasZb, CasZc, CasZd, CasZe, CasZf, CasZg, CasZh, or CasZi crRNA sequence in Table 1 or Table 3.
[0122] In some cases, a subject CasZ guide RNA comprises a CasZa crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZa crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZa crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZa crRNA sequence in Table 1 or Table 3.
[0123] In some cases, a subject CasZ guide RNA comprises a CasZb crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZb crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZb crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZb crRNA sequence in Table 1 or Table 3.
[0124] In some cases, a subject CasZ guide RNA comprises a CasZc crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZc crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZc crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZc crRNA sequence in Table 1 or Table 3.
[0125] In some cases, a subject CasZ guide RNA comprises a CasZd crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZd crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZd crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZd crRNA sequence in Table 1 or Table 3.
[0126] In some cases, a subject CasZ guide RNA comprises a CasZe crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZe crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZe crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZe crRNA sequence in Table 1 or Table 3.
[0127] In some cases, a subject CasZ guide RNA comprises a CasZf crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZf crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZf crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZf crRNA sequence in Table 1 or Table 3.
[0128] In some cases, a subject CasZ guide RNA comprises a CasZg crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZg crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZg crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZg crRNA sequence in Table 1 or Table 3.
[0129] In some cases, a subject CasZ guide RNA comprises a CasZh crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZh crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZh crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZh crRNA sequence in Table 1 or Table 3.
[0130] In some cases, a subject CasZ guide RNA comprises a CasZi crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZi crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZi crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZi crRNA sequence in Table 1 or Table 3.
[0131] In some cases, a subject CasZ guide RNA comprises a CasZj crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZj crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZj crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZj crRNA sequence in Table 1 or Table 3.
[0132] In some cases, a subject CasZ guide RNA comprises a CasZk crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZk crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZk crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZk crRNA sequence in Table 1 or Table 3.
[0133] In some cases, a subject CasZ guide RNA comprises a CasZl crRNA sequence from Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZl crRNA sequence from Table 1 or Table 3. In some cases, a subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZl crRNA sequence from Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZl crRNA sequence in Table 1 or Table 3.
[0134] In some cases, the subject CasZ guide RNA comprises a CasZj, CasZl, or CasZk crRNA sequence of Table 1 or Table 3 (e.g., in addition to the guide sequence, e.g., as part of a protein-binding region). In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZj, CasZl, or CasZk crRNA sequence of Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to a CasZj, CasZl, or CasZk crRNA sequence of Table 1 or Table 3. In some cases, the subject CasZ guide RNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to a CasZj, CasZl, or CasZk crRNA sequence in Table 1 or Table 3.
[0135] CasZ trans-activating non-coding RNA (trancRNA) The compositions and methods of the present disclosure include a CasZ trans-activating non-coding RNA (trancRNA) ("trancRNA"; also referred to herein as "CasZ trancRNA"). In some cases, the trancRNA forms a complex with a CasZ polypeptide and a CasZ guide RNA of the present disclosure. A trancRNA can be identified as a highly transcribed RNA encoded by a nucleotide sequence present in the CasZ locus. The sequence encoding the trancRNA is usually located between the cas gene and the array (repeat) of the CasZ locus (e.g., it can be located adjacent to the repeat sequence). The following examples demonstrate detection of CasZ trancRNA. In some cases, the CasZ trancRNA co-immunoprecipitates (forms a complex) with a CasZ polypeptide. In some cases, the presence of the CasZ trancRNA is required for system function. Data related to trancRNAs (e.g., their expression and their location on the native array) is presented in the Examples section below.
[0136] In some cases, the CasZ trancRNA has a length of 60 nucleotides (nt) to 270 nt (e.g., 60 to 260, 70 to 270, 70 to 260, or 75 to 255 nt). In some cases, the CasZ trancRNA (e.g., CasZa trancRNA) has a length of 60 to 150 nt (e.g., 60 to 140, 60 to 130, 65 to 150, 65 to 140, 65 to 130, 70 to 150, 70 to 140, or 70 to 130 nt). In some cases, the CasZ trancRNA (e.g., CasZa trancRNA) has a length of 70 to 130 nt. In some cases, the CasZ trancRNA (e.g., CasZa trancRNA) has a length of about 80 nt. In some cases, the CasZ trancRNA (e.g., CasZa trancRNA) has a length of about 90 nt. In some cases, the CasZ trancRNA (e.g., CasZa trancRNA) has a length of about 120 nt.
[0137] In some cases, the CasZ trancRNA (e.g., CasZb trancRNA) has a length of 85 to 240 nt (e.g., 85 to 230, 85 to 220, 85 to 150, 85 to 130, 95 to 240, 95 to 230, 95 to 220, 95 to 150, or 95 to 130 nt). In some cases, the CasZ trancRNA (e.g., CasZb trancRNA) has a length of 95 to 120 nt. In some cases, the CasZ trancRNA (e.g., CasZb trancRNA) has a length of about 105 nt. In some cases, the CasZ trancRNA (e.g., CasZb trancRNA) has a length of about 115 nt. In some cases, the CasZ trancRNA (e.g., CasZb trancRNA) has a length of about 215 nt.
[0138] In some cases, the CasZ trancRNA (e.g., CasZc trancRNA) has a length of 80 to 275 nt (e.g., 85 to 260 nt). In some cases, the CasZ trancRNA (e.g., CasZc trancRNA) has a length of 80 to 110 nt (e.g., 85 to 105 nt). In some cases, the CasZ trancRNA (e.g., CasZc trancRNA) has a length of 235 to 270 nt (e.g., 240 to 260 nt). In some cases, the CasZ trancRNA (e.g., CasZc trancRNA) has a length of about 95 nt. In some cases, the CasZ trancRNA (e.g., CasZc trancRNA) has a length of about 250 nt.
[0139] Exemplary trancRNA sequences Examples of naturally occurring trancRNA sequences for CasZ proteins are shown in Table 2.
[0140] (Table 2) CasZ trancRNA sequence TIFF0007672445000016.tif145147
[0141] In some cases, the target CasZ trancRNA comprises the above-mentioned CasZa trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZa trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZa trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 90% or more identity (e.g., 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the above-mentioned CasZ a trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the above-mentioned CasZ a trancRNA sequence and has a length of 60 to 150 nt (e.g., 60 to 140, 60 to 130, 65 to 150, 65 to 140, 65 to 130, 70 to 150, 70 to 140, or 70 to 130 nt).
[0142] In some cases, the target CasZ trancRNA comprises the above-mentioned CasZb trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZb trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZb trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 90% or more identity (e.g., 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the above-mentioned CasZb trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the above-mentioned CasZb trancRNA sequence and has a length of 85 to 240 nt (e.g., 85 to 230, 85 to 220, 85 to 150, 85 to 130, 95 to 240, 95 to 230, 95 to 220, 95 to 150, or 95 to 130 nt).
[0143] In some cases, the target CasZ trancRNA comprises the above-mentioned CasZc trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZc trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZc trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 90% or more identity (e.g., 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the above-mentioned CasZc trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) to the above-mentioned CasZc trancRNA sequence, and has a length of 80 to 110 nt (e.g., 85 to 105 nt) or 235 to 270 nt (e.g., 240 to 260 nt).
[0144] In some cases, the target CasZ trancRNA comprises the above-mentioned CasZa, CasZb, or CasZc trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 70% or more identity (e.g., 75% or more, 80% or more, 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZa, CasZb, or CasZc trancRNA sequence. In some cases, the target CasZ trancRNA comprises a nucleotide sequence having 80% or more identity (e.g., 85% or more, 90% or more, 93% or more, 95% or more, 97% or more, 98% or more, or 100% identity) with the above-mentioned CasZa, CasZb, or CasZc trancRNA sequence. In some cases, the subject CasZ trancRNA comprises a nucleotide sequence having 90% or greater identity (e.g., 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the above-described CasZa, CasZb, or CasZc trancRNA sequence. In some cases, the subject CasZ trancRNA comprises a nucleotide sequence having 80% or greater identity (e.g., 85% or greater, 90% or greater, 93% or greater, 95% or greater, 97% or greater, 98% or greater, or 100% identity) to the above-described CasZa, CasZb, or CasZc trancRNA sequence and has a length of 60 nucleotides (nt) to 270 nt (e.g., 60-260, 70-270, 70-260, or 75-255 nt).
[0145] In some cases, the CasZ trancRNA contains modified nucleotides (e.g., methylated). In some cases, the CasZ trancRNA contains one or more of: i) base modifications or substitutions, ii) backbone modifications, iii) modified internucleoside linkages, and iv) modified sugar moieties. Possible nucleic acid modifications are described below.
[0146] CasZ System The present disclosure provides a CasZ system. The CasZ system of the present disclosure can include one or more of: (1) a CasZ trans-activating non-coding RNA (trancRNA) (referred to herein as "CasZ trancRNA") or a nucleic acid (e.g., an expression vector) encoding the CasZ trancRNA; (2) a CasZ protein (e.g., a wild-type protein, a variant, a catalytically impaired variant, a CasZ fusion protein, etc.) or a nucleic acid (e.g., an RNA, an expression vector, etc.) encoding the CasZ protein; and (3) a CasZ guide RNA (a guide RNA that binds to the CasZ protein to provide sequence specificity, e.g., capable of binding to a target sequence in a eukaryotic genome) or a nucleic acid (e.g., an expression vector) encoding the CasZ guide RNA. The CasZ system can include a host cell (e.g., a eukaryotic cell, a plant cell, a mammalian cell, a human cell) comprising one or more of (1), (2), and (3) (in any combination); for example, in some cases, the host cell comprises the trancRNA and / or a nucleic acid encoding the trancRNA. In some cases, the CasZ system includes a donor template nucleic acid (e.g., in addition to the above). In some cases, the CasZ system is a system of one or more nucleic acids (e.g., one or more expression vectors encoding any combination of the above).
[0147] nucleic acid The present disclosure provides one or more nucleic acids comprising one or more of a CasZ trancRNA sequence, a nucleotide sequence encoding a CasZ trancRNA, a nucleotide sequence encoding a CasZ polypeptide (e.g., a wild-type CasZ protein, a nickase CasZ protein, a dCasZ protein, a chimeric CasZ protein / CasZ fusion protein, etc.), a CasZ guide RNA sequence, a nucleotide sequence encoding a CasZ guide RNA, and a donor polynucleotide (donor template, donor DNA) sequence. In some cases, the nucleic acid of interest (e.g., one or more nucleic acids) is a recombinant expression vector (e.g., a plasmid, a viral vector, a minicircle DNA, etc.). In some cases, the nucleotide sequence encoding the CasZ trancRNA, the nucleotide sequence encoding the CasZ protein, and / or the nucleotide sequence encoding the CasZ guide RNA is operably linked to a promoter (e.g., an inducible promoter), e.g., one that is operable in a selected cell type (e.g., a prokaryotic cell, a eukaryotic cell, a plant cell, an animal cell, a mammalian cell, a primate cell, a rodent cell, a human cell, etc.).
[0148] In some cases, the nucleotide sequence encoding the CasZ polypeptide of the present disclosure is codon-optimized. This type of optimization can involve altering the CasZ-encoding nucleotide sequence to mimic the codon preferences of the intended host organism or cell while still encoding the same protein. Thus, the codons can be altered, but the encoded protein remains unchanged. For example, if the intended target cell is a human cell, a human-codon-optimized CasZ-encoding nucleotide sequence can be used. As another non-limiting example, if the intended host cell is a mouse cell, a mouse-codon-optimized CasZ-encoding nucleotide sequence can be generated. As another non-limiting example, if the intended host cell is a plant cell, a plant-codon-optimized CasZ-encoding nucleotide sequence can be generated. As another non-limiting example, if the intended host cell is an insect cell, an insect-codon-optimized CasZ-encoding nucleotide sequence can be generated.
[0149] The present disclosure provides one or more recombinant expression vectors comprising (in some cases in different recombinant expression vectors, and in some cases in the same recombinant expression vector) a CasZ trancRNA sequence, a nucleotide sequence encoding the CasZ trancRNA, a nucleotide sequence encoding a CasZ polypeptide (e.g., a wild-type CasZ protein, a nickase CasZ protein, a dCasZ protein, a chimeric CasZ protein / CasZ fusion protein, etc.), a CasZ guide RNA sequence, a nucleotide sequence encoding the CasZ guide RNA, and a donor polynucleotide (donor template, donor DNA) sequence. In some cases, the nucleic acid of interest (e.g., one or more nucleic acids) is a recombinant expression vector (e.g., a plasmid, a viral vector, a minicircle DNA, etc.). In some cases, the nucleotide sequence encoding the CasZ transcript RNA, the nucleotide sequence encoding the CasZ protein, and / or the nucleotide sequence encoding the CasZ guide RNA is operably linked to a promoter (e.g., an inducible promoter), e.g., one that is operable in a selected cell type (e.g., a prokaryotic cell, a eukaryotic cell, a plant cell, an animal cell, a mammalian cell, a primate cell, a rodent cell, a human cell, etc.).
[0150] Suitable expression vectors include viral expression vectors (e.g., vaccinia virus, poliovirus, adenovirus (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:2543 2549, 1994; Borras et al., Gene Ther 6:515 524, 1999; Li and Davidson, PNAS 92:7700 7704, 1995; Sakamoto et al., H Gene Ther 5:1088 1097, 1999; WO94 / 12649; WO93 / 03769; WO93 / 19191; WO94 / 28938; WO95 / 11984; and WO95 / 00655), adeno-associated virus (AAV) (see, e.g., Ali et al., Hum Gene Ther 9:81 86,1998, Flannery et al.,PNAS 94:6916 6921,1997,Bennett et al.,Invest Opthalmol Vis Sci 38:2857 2863,1997,Jomary et al.,Gene Ther 4:683 690,1997,Rolling et al.,Hum Gene Ther 10:641 648,1999, Ali et al., Hum Mol Genet 5:591 594,1996, Srivastava in WO93 / 09239, Samulski et al., J. Vir. (1989) 63:3822-3828, Mendelson et al., Virol. (1988) 166:154-165, and Flotte et al. al., PNAS (1993) 90:10613-10617), SV40, herpes simplex virus, human immunodeficiency virus (e.g., Miyoshi et al., PNAS 94:10319 23, 1997, Takahashi et al., J Virol 73:7812 7816, 1999), retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus). In some cases, the recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some cases, the recombinant expression vector of the present disclosure is a recombinant lentiviral vector. In some cases, the recombinant expression vector of the present disclosure is a recombinant retroviral vector.
[0151] Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc., may be used in the expression vector.
[0152] In some embodiments, the nucleotide sequence encoding the CasZ guide RNA is operably linked to a regulatory element, e.g., a transcriptional control element such as a promoter. In some embodiments, the nucleotide sequence encoding the CasZ protein or CasZ fusion polypeptide is operably linked to a regulatory element, e.g., a transcriptional repression element such as a promoter.
[0153] The transcriptional control element can be a promoter. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell-type-specific promoter. In some cases, the transcriptional control element (e.g., promoter) is functional in a targeted cell type or a targeted cell population. For example, in some cases, the transcriptional control element can be functional in eukaryotic cells, such as hematopoietic stem cells (e.g., mobilized peripheral blood (mPB) CD34(+) cells, bone marrow (BM) CD34(+) cells, etc.).
[0154] Non-limiting examples of eukaryotic promoters (promoters functional in eukaryotic cells) include EF1α, those derived from immediate early cytomegalovirus (CMV), herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) derived from retroviruses, and mouse metallothionein-I. Selection of an appropriate vector and promoter is well within the level of ordinary skill in the art. Expression vectors can also contain a ribosome binding site for translation initiation and a transcription terminator. Expression vectors can also include appropriate sequences for amplifying expression. Expression vectors can also include a nucleotide sequence encoding a protein tag (e.g., a 6xHis tag, a hemagglutinin tag, a fluorescent protein, etc.) that can be fused to the CasZ protein, thus resulting in a chimeric CasZ polypeptide.
[0155] In some cases, the nucleotide sequence encoding the CasZ guide RNA and / or the CasZ fusion polypeptide is operably linked to an inducible promoter. In some cases, the nucleotide sequence encoding the CasZ guide RNA and / or the CasZ fusion protein is operably linked to a constitutive promoter.
[0156] The promoter may be a constitutively active promoter (i.e., a promoter that is constitutively active / "ON" state), an inducible promoter (i.e., a promoter whose active / "ON" or inactive / "OFF" state is controlled by an external stimulus, e.g., a particular temperature, compound, or the presence of a protein), a spatially constrained promoter (i.e., a transcriptional control element, enhancer, etc.) (e.g., a tissue-specific promoter, cell-type-specific promoter, etc.), or a temporally constrained promoter (i.e., the promoter is "ON" or "OFF" state during a particular stage of embryonic development or during a particular stage of a biological process, e.g., the hair follicle cycle in mice).
[0157] Suitable promoters may be derived from viruses and therefore may be referred to as viral promoters, or they may be derived from any organism, including prokaryotes or eukaryotes. Suitable promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Exemplary promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter; the adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter region (CMVIE), the Rous sarcoma virus (RSV) promoter, the human U6 micronucleus promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497-500 (2002)), the enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1; 31 (17)), the human H1 promoter (H1), and the like.
[0158] In some cases, the nucleotide sequence encoding the CasZ guide RNA is operably linked to (under the control of) a promoter operable in eukaryotic cells (e.g., a U6 promoter, an enhanced U6 promoter, an H1 promoter, etc.). As will be understood by those skilled in the art, when expressing an RNA (e.g., a guide RNA) from a nucleic acid (e.g., an expression vector) using a U6 promoter or another Pol III promoter (e.g., in a eukaryotic cell), if several Ts are in a row (encoding Us in the RNA), the RNA may need to be mutated. This is because consecutive Ts (e.g., five Ts) in DNA can act as a terminator for polymerase III (Pol III). Therefore, to ensure transcription of the guide RNA in eukaryotic cells, it may sometimes be necessary to modify the sequence encoding the guide RNA to remove consecutive Ts. In some cases, the nucleotide sequence encoding a CasZ protein (e.g., a wild-type CasZ protein, a nickase CasZ protein, a dCasZ protein, a chimeric CasZ protein, etc.) is operably linked to a promoter operable in a eukaryotic cell (e.g., a CMV promoter, an EF1α promoter, an estrogen receptor-regulated promoter, etc.).
[0159] Examples of inducible promoters include, but are not limited to, T7 RNA polymerase promoter, T3 RNA polymerase promoter, isopropyl-beta-D-thiogalactopyranoside (IPTG)-regulated promoter, lactose-inducible promoter, heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline, estrogen and / or estrogen analogs, IPTG, etc.
[0160] Inducible promoters suitable for use include any inducible promoter described herein or known to those of skill in the art. Examples of inducible promoters include alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)-responsive promoters and other tetracycline-responsive promoter systems including tetracycline repressor protein (tetR), tetracycline operator sequence (tetO), and tetracycline transactivator fusion protein (tTA)), steroid-regulated promoters (e.g., promoters based on rat glucocorticoid receptor, human estrogen receptor, gaecdysone receptor, and steroid / retinoid receptors). These promoters include, but are not limited to, chemically / biochemically regulated and physically regulated promoters such as promoters from the thyroid receptor superfamily, metal-regulated promoters (e.g., promoters derived from metallothioneins (proteins that bind and sequester metal ions) from yeast, mouse, and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene, or benzothiadiazole (BTH)), temperature / heat-inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light-responsive promoters from plant cells).
[0161] In some cases, the promoter is a spatially constrained promoter (i.e., a cell-type specific promoter, a tissue-specific promoter, etc.), such that in a multicellular organism, the promoter is active (i.e., "ON") in a specific subset of cells. A spatially constrained promoter may also be referred to as an enhancer, a transcriptional control element, a regulatory sequence, etc. Any convenient spatially constrained promoter can be used, so long as the promoter is functional in the target host cell (e.g., eukaryotic, prokaryotic).
[0162] In some cases, the promoter is a reversible promoter. Suitable reversible promoters, including reversibly inducible promoters, are known in the art. Such reversible promoters can be isolated from and derived from many organisms, for example, eukaryotes and prokaryotes. The modification of a reversible promoter from a first organism (e.g., a first organism is a prokaryote and a second organism is a eukaryote, a first organism is a eukaryote and a second organism is a prokaryote, etc.) for use in a second organism is well known in the art. Examples of such reversible promoters and systems based on such reversible promoters but also containing additional regulatory proteins include, but are not limited to, alcohol-regulated promoters (e.g., alcohol dehydrogenase I (alcA) gene promoter, promoters responsive to alcohol transactivator protein (AlcR), etc.), tetracycline-regulated promoters (e.g., promoter systems including TetActivator, TetON, TetOFF, etc.), steroid-regulated promoters (e.g., rat glucocorticoid receptor promoter system, human estrogen receptor promoter system, retinoid promoter system, thyroid promoter system, ecdysone promoter system, mifepristone promoter system, etc.), metal-regulated promoters (e.g., metallothionein promoter system, etc.), pathogen-associated regulated promoters (e.g., salicylic acid-regulated promoter, ethylene-regulated promoter, benzothiadiazole-regulated promoter, etc.), temperature-regulated promoters (e.g., heat shock-inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter, etc.), light-regulated promoters, synthetic inducible promoters, etc.
[0163] Methods for introducing nucleic acids (e.g., DNA or RNA) (e.g., nucleic acids comprising donor polynucleotide sequences, one or more nucleic acids encoding a CasZ protein and / or a CasZ guide RNA and / or a CasZ transcript RNA, etc.) into host cells are known in the art, and any convenient method can be used to introduce a nucleic acid (e.g., an expression construct) into a cell. Suitable methods include, for example, viral infection, transfection, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc.
[0164] Introduction of the recombinant expression vector into cells can occur in any culture medium and under any culture conditions that promote cell survival. Introduction of the recombinant expression vector into target cells can be performed in vivo or ex vivo. Introduction of the recombinant expression vector into target cells can be performed in vitro.
[0165] In some cases, the CasZ protein can be provided as RNA. The RNA can be provided by direct chemical synthesis or can be transcribed in vitro from DNA (e.g., encoding the CasZ protein). Once synthesized, the RNA can be introduced into cells by any of the well-known techniques for introducing nucleic acids into cells (e.g., microinjection, electroporation, transfection, etc.).
[0166] Nucleic acids can be provided to cells using well-developed transfection techniques (see, e.g., Angel and Yanik (2010) PLoS ONE 5(7):e11756), as well as commercially available TransMessenger® reagents from Qiagen, Stemfect™ RNA transfection kits from Stemgent, and TransIT®-mRNA transfection kits from Mirus Bio LLC. See also Beumer et al. (2008) PNAS 105(50):19821-19826.
[0167] The vector can be directly provided to the target host cell. In other words, the cell is contacted with a vector containing the target nucleic acid (for example, a recombinant expression vector having a donor template sequence and encoding a CasZ guide RNA, a recombinant expression vector encoding a CasZ protein, etc.), and the vector is thereby taken up by the cell. Methods for contacting cells with a nucleic acid vector that is a plasmid include electroporation, calcium chloride transfection, microinjection, and lipofection, which are well known in the art. In the case of viral vector delivery, the cell can be contacted with a viral particle containing the target viral expression vector.
[0168] Retroviruses, such as lentiviruses, are suitable for use in the methods of the present disclosure. Commonly used retroviral vectors are "defective," i.e., unable to produce viral proteins necessary for productive infection. Rather, vector replication requires propagation in a packaging cell line. To produce viral particles containing a nucleic acid of interest, the retroviral nucleic acid containing the nucleic acid is packaged into a viral capsid by the packaging cell line. Different packaging cell lines provide different envelope proteins (ecotropic, amphotropic, or xenotropic) that are incorporated into the capsid, and this envelope protein determines the specificity of the viral particle to the cell (ecotropic for mouse and rat, amphotropic for most mammalian cell types, including human, dog, and mouse, and xenotropic for most mammalian cells except mouse cells). Using an appropriate packaging cell line can ensure that the cell is targeted by the packaged viral particle. Methods for introducing a subject expression vector into a packaging cell line and harvesting viral particles produced by the packaging line are well known in the art. Nucleic acids can also be introduced by direct microinjection (e.g., injection of RNA).
[0169] The vector used to provide the nucleic acid encoding the CasZ guide RNA and / or CasZ polypeptide to the target host cell can include a suitable promoter to drive expression of the nucleic acid of interest, i.e., transcriptional activation. In other words, in some cases, the nucleic acid of interest is operably linked to a promoter. This can include a ubiquitously acting promoter, such as the CMV-β-actin promoter, or an inducible promoter, such as a promoter that is active in a specific cell population or responds to the presence of a drug such as tetracycline. By transcriptional activation, it is intended that transcription be increased 10-fold, 100-fold, or more usually 1000-fold over basal levels in the target cell. In addition, the vector used to provide the nucleic acid encoding the CasZ guide RNA and / or CasZ protein to the cell can include a nucleic acid sequence encoding a selectable marker in the target cell to identify cells that have taken up the CasZ guide RNA and / or CasZ protein.
[0170] A nucleic acid comprising a nucleotide sequence encoding a CasZ polypeptide or a CasZ fusion polypeptide may be RNA in some cases. Thus, a CasZ fusion protein can be introduced into a cell as RNA. Methods for introducing RNA into a cell are known in the art and may include, for example, direct injection, transfection, or any other method used for DNA introduction. Alternatively, a CasZ protein may be provided to a cell as a polypeptide. Such a polypeptide may optionally be fused to a polypeptide domain that increases the solubility of the product. This domain may be linked to the polypeptide through a defined protease cleavage site, e.g., a TEV sequence, which is cleaved by TEV protease. The linker may also contain one or more flexible sequences, e.g., 1 to 10 glycine residues. In some embodiments, cleavage of the fusion protein is performed in a buffer that maintains the solubility of the product, e.g., in the presence of 0.5 to 2 M urea, in the presence of a solubility-increasing polypeptide and / or polynucleotide, etc. Domains of interest include endosomal destabilizing domains, such as influenza HA domains, and other polypeptides that aid in production, such as IF2 domains, GST domains, GRPE domains, etc. Polypeptides can be formulated to improve stability. For example, peptides can be PEGylated, where polyethyleneoxy groups provide improved longevity in the bloodstream.
[0171] Additionally or alternatively, the CasZ polypeptides of the present disclosure can be fused to a polypeptide permeation domain to facilitate cellular uptake. Several permeation domains are known in the art and can be used in non-integrated polypeptides of the present disclosure, including peptides, peptidomimetics, and non-peptide carriers. For example, a permeation peptide can be derived from the third alpha helix of the Drosophila melanogaster transcription factor Antennapedia, designated penetratin, which contains the amino acid sequence RQIKIWFQNRRMKWKK (SEQ ID NO: 134). As another example, a permeation peptide can contain the HIV-1 tat basic region amino acid sequence, which can include, for example, amino acids 49-57 of the native tat protein. Other permeation domains include poly-arginine motifs, such as the region of amino acids 34-56 of the HIV-1 rev protein, nona-arginine, octa-arginine, etc. (See, e.g., Futaki et al. (2003) Curr Protein Pept Sci. 2003 Apr;4(2):87-9 and 446, and Wender et al. (2000) Proc. Natl. Acad. Sci. USA 2000 Nov. 21;97(24):13003-8, U.S. Patent Application Publication Nos. 20030220334, 20030083256, 20030032593, and 20030022831, which are specifically incorporated by reference herein for their teachings of translocating peptides and peptoids.) The nona-arginine (R9) sequence is one of the more efficient PTDs characterized (Wender et al. 2000, Uemura et al. 2002). The site at which the fusion is made may be selected in order to optimize the biological activity, secretion, or binding characteristics of the polypeptide. Optimal sites will be determined by routine experimentation.
[0172] The CasZ polypeptides of the present disclosure can be produced in vitro, or by eukaryotic cells, or by prokaryotic cells, and can be further treated by unfolding, e.g., heat denaturation, dithiothreitol reduction, etc., and further refolded using methods known in the art.
[0173] Modifications of interest that do not alter the primary sequence include chemical derivatization of the polypeptide, such as acylation, acetylation, carboxylation, amidation, etc. Also included are glycosylation modifications, e.g., those made by modifying the glycosylation pattern of a polypeptide during its synthesis and processing or during further processing steps, e.g., by exposing the polypeptide to enzymes that affect glycosylation, such as mammalian glycosylation or deglycosylation enzymes. Also encompassed are sequences having phosphorylated amino acid residues, e.g., phosphotyrosine, phosphoserine, or phosphothreonine.
[0174] Also suitable for inclusion in embodiments of the present disclosure are nucleic acids (e.g., encoding CasZ guide RNAs, encoding CasZ fusion proteins, etc.) and proteins (e.g., CasZ fusion proteins derived from wild-type or variant proteins) that have been modified using standard molecular biology techniques and synthetic chemistry to improve resistance to proteolysis, alter target sequence specificity, optimize solubility properties, or alter or enhance protein activity (e.g., transcriptional regulatory activity, enzymatic activity, etc.). Analogs of such polypeptides include those that contain residues other than naturally occurring L-amino acids, such as D-amino acids or non-naturally occurring synthetic amino acids. D-amino acids can be substituted for some or all of the amino acid residues.
[0175] The CasZ polypeptides of the present disclosure can be prepared by in vitro synthesis using convenient methods known in the art. Various commercially available synthesizers are available, such as automated synthesizers from Applied Biosystems, Inc., Beckman, etc. By using a synthesizer, natural amino acids may be substituted with unnatural amino acids. The particular sequence and preparation method will be determined by convenience, economy, required purity, etc.
[0176] If desired, various groups may be introduced into the peptide during synthesis or expression, allowing for linkage to other molecules or surfaces. Thus, cysteine can be used to create thioethers, histidine for linking to metal ion complexes, carboxyl groups for forming amides or esters, amino groups for forming amides, etc.
[0177] The CasZ polypeptides of the present disclosure may also be isolated and purified according to conventional methods of recombinant synthesis. A lysate of the expression host may be prepared, and the lysate may be purified using high-performance liquid chromatography (HPLC), exclusion chromatography, gel electroporation, affinity chromatography, or other purification techniques. Typically, the compositions used contain 20% or more by weight of the desired product, more usually 75% or more, preferably 95% or more, and for therapeutic purposes usually 99.5% by weight, relative to contaminants associated with the method of preparation and purification of the product. Percentages are typically based on total protein. Thus, in some cases, the CasZ polypeptides or CasZ fusion polypeptides of the present disclosure are at least 80% pure, at least 85% pure, at least 90% pure, at least 95% pure, at least 98% pure, or at least 99% pure (e.g., free of contaminants, non-CasZ proteins, other macromolecules, etc.).
[0178] To induce cleavage or any desired modification in a target nucleic acid (e.g., genomic DNA) or to induce any desired modification in a polypeptide associated with a target nucleic acid, the CasZ guide RNA and / or CasZ polypeptide and / or CasZ trancRNA, and / or donor template sequence, whether introduced as a nucleic acid or as a polypeptide, can be provided to the cell for a period of about 30 minutes to about 24 hours, e.g., 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 12 hours, 16 hours, 18 hours, 20 hours, or any other period of about 30 minutes to about 24 hours, which can be repeated at a frequency of about every day to about every 4 days, e.g., every 1.5 days, every 2 days, every 3 days, or any other frequency of about every day to about every 4 days. The agent(s) may be provided to the cells of interest one or more times, e.g., one, two, three, or four or more times, and the cells are incubated with the agent(s) for some period of time following each contact event, e.g., 16-24 hours, after which time the medium is replaced with fresh medium and the cells are further cultured.
[0179] When two or more different targeting complexes (e.g., two different CasZ guide RNAs that are complementary to different sequences within the same or different target nucleic acid) are provided to a cell, the complexes can be provided simultaneously (e.g., as two polypeptides and / or nucleic acids) or delivered simultaneously, or they can be provided sequentially, e.g., a targeting complex is provided first, followed by a second targeting complex, or vice versa.
[0180] To improve delivery of DNA vectors to target cells, DNA can be protected from damage and its entry into cells, for example, by using lipoplexes and polyplexes. Thus, in some cases, nucleic acids of the present disclosure (e.g., recombinant expression vectors of the present disclosure) can be coated with lipids in organized structures such as micelles or liposomes. When the organized structures are complexed with DNA, they are called lipoplexes. Three types of lipids exist: anionic (negatively charged), neutral, or cationic (positively charged). Lipoplexes using cationic lipids have proven useful for gene transfer. Due to their positive charge, cationic lipids naturally complex with negatively charged DNA. Also, as a result of their charge, they interact with cell membranes. Lipoplex endocytosis then occurs, releasing the DNA into the cytoplasm. Cationic lipids also protect against cellular denaturation of the DNA.
[0181] Polymer-DNA complexes are called polyplexes. Most polyplexes are composed of cationic polymers, and their production is regulated by ionic interactions. One major difference between the way polyplexes and lipoplexes work is that polyplexes cannot release their DNA payload into the cytoplasm; therefore, cotransfection with an endosomolytic agent, such as an inactivated adenovirus, which dissolves endosomes created during endocytosis, is essential. However, this is not always the case; polymers such as polyethyleneimine have their own endosome-disrupting methods, as do chitosan and trimethylchitosan.
[0182] Dendrimers, highly branched macromolecules with spherical shapes, can also be used to genetically modify stem cells. The surface of dendrimer particles can be functionalized to alter their properties. In particular, cationic dendrimers (i.e., those with a positive surface charge) can be constructed. When in the presence of genetic material such as a DNA plasmid, charge complementarity results in a transient association of the nucleic acid with the cationic dendrimer. Upon reaching the destination, the dendrimer-nucleic acid complex can be taken up by the cell by endocytosis.
[0183] In some cases, a nucleic acid of the present disclosure (e.g., an expression vector) includes an insertion site for a guide sequence of interest. For example, a nucleic acid can include an insertion site for a guide sequence of interest, where the insertion site is immediately adjacent to a nucleotide sequence encoding a portion of the CasZ guide RNA that remains unchanged when the guide sequence is changed to hybridize to a desired target sequence (e.g., a sequence that contributes to the CasZ binding aspects of the guide RNA, e.g., a sequence that contributes to the dsRNA duplex(es) of the CasZ guide RNA; this portion of the guide RNA may also be referred to as the "scaffold" or "constant region" of the guide RNA). Thus, in some cases, a nucleic acid of interest (e.g., an expression vector) includes a nucleotide sequence encoding a CasZ guide RNA, except that the portion encoding the guide sequence portion of the guide RNA is the insertion sequence (insertion site). An insertion site is any nucleotide sequence used for insertion of a desired sequence. "Insertion sites" for use in various techniques are known to those of skill in the art, and any convenient insertion site can be used. The insertion site can be for any method for utilizing a nucleic acid sequence. For example, in some cases, the insertion site is a multiple cloning site (MCS) (e.g., a site containing one or more restriction enzyme recognition sequences), a site for ligation-independent cloning, a site for recombination-based cloning (e.g., att site-based recombination), a nucleotide sequence recognized by CRISPR / Cas (e.g., Cas9)-based technologies, etc.
[0184] The insertion site can be any desired length and can depend on the type of insertion site (e.g., whether (and how many) the site includes one or more restriction enzyme recognition sequences, whether the site includes a target site for a CRISPR / Cas protein, etc.). In some cases, the insertion site for a nucleic acid of interest is 3 or more nucleotides (nt) in length (e.g., 5 or more, 8 or more, 10 or more, 15 or more, 17 or more, 18 or more, 19 or more, 20 or more, or 25 or more, or 30 or more nt in length). In some cases, the length of the insertion site of the nucleic acid of interest is in the range of 2 to 50 nucleotides (nt) (e.g., 2 to 40 nt, 2 to 30 nt, 2 to 25 nt, 2 to 20 nt, 5 to 50 nt, 5 to 40 nt, 5 to 30 nt, 5 to 25 nt, 5 to 20 nt, 10 to 50 nt, 10 to 40 nt, 10 to 30 nt, 10 to 25 nt, 10 to 20 nt, 17 to 50 nt, 17 to 40 nt, 17 to 30 nt, 17 to 25 nt). In some cases, the length of the insertion site of the nucleic acid of interest is in the range of 5 to 40 nt.
[0185] Nucleic acid modification In some embodiments, the nucleic acid of interest (e.g., a CasZ guide RNA or a truncRNA) has one or more modifications, such as base modifications, backbone modifications, etc., to provide the nucleic acid with new or enhanced characteristics (e.g., improved stability). A nucleoside is a base-sugar combination. The base portion of a nucleoside is typically a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide is a nucleoside that further includes a phosphate group covalently attached to the sugar portion of the nucleoside. In the case of nucleosides containing a pentofuranosyl sugar, the phosphate group can be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming oligonucleotides, the phosphate groups covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound can then be further joined to form a circular compound, although linear compounds are preferred. In addition, linear compounds may have internal nucleotide base complementarity and thus may fold in such a way as to produce fully or partially double-stranded compounds. Within oligonucleotides, the phosphate groups are commonly referred to as forming the internucleoside backbone of the oligonucleotide. The usual linkage or backbone of RNA and DNA is the 3' to 5' phosphodiester linkage.
[0186] Suitable nucleic acid modifications include, but are not limited to, 2'O-methyl modified nucleotides, 2'fluoro modified nucleotides, locked nucleic acid (LNA) modified nucleotides, peptide nucleic acid (PNA) modified nucleotides, nucleotides with phosphorothioate linkages, and 5' caps (e.g., 7-methylguanylate caps (m7G)). Further details and additional modifications are described below.
[0187] 2'-O-methyl modified nucleotides (also called 2'-O-methyl RNA) are a naturally occurring post-transcriptional modification of RNA found in tRNA and other small RNAs. Oligonucleotides containing 2'-O-methyl RNA can be directly synthesized. This modification increases the Tm of the RNA:RNA duplex but results in only minor changes in RNA:DNA stability. It is stable to attack by single-stranded ribonucleases and is typically 5-10 times less susceptible to DNases than DNA. It is commonly used in antisense oligos as a means of increasing stability and binding affinity for target messages.
[0188] 2'-fluoro-modified nucleotides (e.g., 2'-fluoro bases) have a fluorine-modified ribose that increases binding affinity (Tm) and confers some relative nuclease resistance when compared to natural RNA. These modifications are commonly used in ribozymes and siRNAs to improve stability in serum or other biological fluids.
[0189] LNA bases have a modification to the ribose backbone that locks the base at the C3'-terminal position, favoring an RNA A-form helical duplex geometry. This modification significantly increases Tm and is highly nuclease-resistant. Multiple LNA insertions can be placed in oligos at any position except the 3'-terminus. Applications ranging from antisense oligos to hybridization probes for SNP detection and allele-specific PCR have been described. Due to the significant increase in Tm conferred by LNAs, they can also increase primer-dimer formation and self-hairpin formation. In some cases, the number of LNAs incorporated into a single oligo is 10 or fewer bases.
[0190] Phosphorothioate (PS) bonds (i.e., phosphorothioate linkages) replace non-bridging oxygens in the phosphate backbone of nucleic acids (e.g., oligos) with sulfur atoms. This modification confers resistance to nuclease degradation to internucleotide linkages. Phosphorothioate linkages can be introduced between the last 3-5 nucleotides at the 5' or 3' end of an oligo to inhibit exonuclease degradation. Inclusion of phosphorothioate linkages within an oligo (e.g., throughout the oligo) can also help reduce endonuclease attack.
[0191] In some embodiments, the nucleic acid of interest has one or more nucleotides that are 2'-O-methyl modified nucleotides. In some embodiments, the nucleic acid of interest (e.g., guide RNA, transcription RNA, etc.) has one or more 2' fluoro-modified nucleotides. In some embodiments, the nucleic acid of interest (e.g., dsRNA, siNA, etc.) has one or more LNA bases. In some embodiments, the nucleic acid of interest (e.g., dsRNA, siNA, etc.) has one or more nucleotides linked by phosphorothioate linkages (i.e., the nucleic acid of interest has one or more phosphorothioate linkages). In some embodiments, a nucleic acid of interest (e.g., dsRNA, siNA, etc.) has a 5' cap (e.g., a 7-methylguanylate cap (m7G)). In some embodiments, a nucleic acid of interest (e.g., a guide RNA, a transcript RNA, etc.) has a combination of modified nucleotides. For example, a nucleic acid of interest (e.g., a guide RNA, a transcript RNA, etc.) can have a 5' cap (e.g., a 7-methylguanylate cap (m7G)) in addition to having one or more nucleotides with other modifications (e.g., 2'-O-methyl nucleotides and / or 2' fluoro-modified nucleotides and / or LNA bases and / or phosphorothioate linkages).
[0192] Modified backbones and modified internucleoside linkages Examples of suitable nucleic acids (e.g., CasZ guide RNAs and / or CasZ trancRNAs) containing modifications include nucleic acids containing modified backbones or non-natural internucleoside linkages. Nucleic acids with modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.
[0193] Suitable modified oligonucleotide backbones containing phosphate atoms therein include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates (including 3'-alkylene phosphonates, 5'-alkylene phosphonates, and chiral phosphonates), phosphinates, phosphoramidates (including 3'-amino phosphoramidate and aminoalkyl phosphoramidate), phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates, which typically have 3'-5' linkages, their 2'-5' linked analogs, and those with inverted polarity in which one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages. Preferred oligonucleotides with reversed polarity contain a single 3'-3' linkage at the 3'-most internucleotide linkage, i.e., a single reversed nucleoside residue that may be basic (either lacking a nucleobase or having a hydroxyl group instead).Various salts (e.g., potassium or sodium), mixed salts, and free acid forms are also included.
[0194] In some embodiments, the nucleic acids of interest contain one or more phosphorothioate and / or heteroatom internucleoside linkages, particularly -CH-NH-O-CH-, -CH-N(CH)-O-CH- (known as a methylene(methylimino) or MMI backbone), -CH-ON(CH)-CH-, -CH-N(CH)-N(CH)-CH-, and -ON(CH)-CH-CH- (natural phosphodiester internucleotide linkages are represented as -OP(=O)(OH)-O-CH-). MMI-type internucleoside linkages are disclosed in the above-referenced U.S. Pat. No. 5,489,677, the disclosure of which is incorporated herein by reference in its entirety. Suitable amide internucleoside linkages are disclosed in U.S. Pat. No. 5,602,240, the disclosure of which is incorporated herein by reference in its entirety.
[0195] Also suitable is the nucleic acid having morpholino backbone structure, as described in U.S. Patent No. 5,034,506.For example, in some embodiments, the nucleic acid of interest comprises a 6-membered morpholino ring instead of a ribose ring.In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages replace phosphodiester linkages.
[0196] Suitable modified polynucleotide backbones that do not contain phosphorus atoms have backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonic acid and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH constituent moieties.
[0197] Mimics The nucleic acid of interest may be a nucleic acid mimic. When applied to polynucleotides, the term "mimetic" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring is also referred to in the art as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety is maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid, a polynucleotide mimic that has been shown to have excellent hybridization properties, is called a peptide nucleic acid (PNA). In PNA, the sugar backbone of a polynucleotide is replaced with an amide backbone, particularly an aminoethylglycine backbone. The nucleotides are retained and are bound directly or indirectly to the aza nitrogen atoms of the amide portion of the backbone.
[0198] One polynucleotide mimic that has been reported to have excellent hybridization properties is peptide nucleic acid (PNA). The backbone in PNA compounds is two or more linked aminoethylglycine units, giving PNA an amide-containing backbone. The heterocyclic base moiety is directly or indirectly linked to the aza nitrogen atom of the amide portion of the backbone. Representative U.S. patents describing the preparation of PNA compounds include, but are not limited to, U.S. Pat. Nos. 5,539,082, 5,714,331, and 5,719,262 (the disclosures of which are incorporated herein by reference in their entirety).
[0199] Another class of polynucleotide mimics being studied is based on linked morpholino units (morpholino nucleic acids) with heterocyclic bases attached to the morpholino ring. Several linking groups have been reported to link the morpholino monomer units in morpholino nucleic acids. One class of linking group was selected to provide nonionic oligomeric compounds. Nonionic morpholino-based oligomeric compounds are less likely to have undesired interactions with cellular proteins. Morpholino-based polynucleotides are nonionic mimics of oligonucleotides that are less likely to form undesired interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholino-based polynucleotides are disclosed in U.S. Patent No. 5,034,506, the disclosure of which is incorporated herein by reference in its entirety. Various compounds within the morpholino class of polynucleotides have been prepared with a variety of different linking groups joining the monomer subunits.
[0200] Another class of polynucleotide mimics is called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in DNA / RNA molecules is replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers have been prepared and used for the synthesis of oligomeric compounds following classical phosphoramidite chemistry. Fully modified CeNA oligomeric compounds and oligonucleotides with specific CeNA modifications have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, 122, 8595-8602, the disclosure of which is incorporated herein by reference in its entirety). In general, the incorporation of CeNA monomers into DNA strands increases the stability of DNA / RNA hybrids. CeNA oligoadenylates formed complexes with RNA and DNA complements with similar stability to the native complexes. Studies of incorporating CeNA structures into native nucleic acid structures have been carried out, with facile conformational adaptation demonstrated by NMR and circular dichroism.
[0201] A further modification includes locked nucleic acids (LNAs), in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2 (Singh et al., Chem. Commun. 1998, 4, 455-456, the disclosure of which is incorporated herein by reference in its entirety). LNAs and LNA analogs exhibit very high duplex thermal stability with complementary DNA and RNA (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility properties. Potent and non-toxic antisense oligonucleotides containing LNAs have been described (see, e.g., Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 2000, 97, 5633-5638, the disclosure of which is incorporated herein by reference in its entirety).
[0202] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine, and uracil have been described, along with their oligomerization and nucleic acid recognition properties (e.g., Koshkin et al., Tetrahedron, 1998, 54, 3607-3630, the disclosure of which is incorporated herein by reference in its entirety). LNAs and their preparation are also described in WO98 / 39352 and WO99 / 14226, and U.S. Application Nos. 20120165514, 20100216983, 20090041809, 20060117410, 20040014959, 20020094555, and 20020086998, the disclosures of which are incorporated herein by reference in their entireties.
[0203] modified sugar moiety The subject nucleic acids can also contain one or more substituted sugar moieties. Suitable polynucleotides contain sugar substituents selected from OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, where the alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C6. 10 Alkyl or C2-C 10 It may be alkenyl or alkynyl. O((CH2) n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2 is particularly preferred, where n and m are from 1 to about 10. Other preferred polynucleotides include C1 to C 10Sugar substituents include those selected from lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH, OCN, Cl, Br, CN, CF, OCF, SOCH, SOCH, ONO, NO, N, NH, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving groups, reporter groups, intercalators, groups for improving the pharmacokinetic properties of oligonucleotides, or groups for improving the pharmacological properties of oligonucleotides, and other substituents with similar properties. Suitable modifications include 2'-methoxyethoxy (2'-O-CHCHOCH, also known as 2'-O-(2-methoxyethyl) or 2'-MOE) (Martin et al., Helv. Chim. ACTA, 1995, 78, 486-504, the disclosure of which is incorporated herein by reference in its entirety), i.e., alkoxyalkoxy groups. Further suitable modifications include the 2'-dimethylaminooxyethoxy, i.e., O(CH2)2ON(CH3)2 group (also known as 2'-DMAOE), as described in the Examples herein below, and 2'-dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2.
[0204] Other suitable sugar substituents include methoxy (-O-CH), aminopropoxy (-OCHCHCHNH), allyl (-CH-CH=CH), -O-allyl (-O-CH-CH=CH), and fluoro (F). The 2'-sugar substituent may be at the arabino (up) or ribo (down) position. A preferred 2'-arabino modification is 2'-F. Similar modifications may be made at other positions on the oligomeric compound, particularly the 3' position of the sugar on the 3'-terminal nucleoside or in 2'-5'-linked oligonucleotides, and the 5' position of the 5'-terminal nucleotide. Oligomeric compounds may also have sugar mimetics, such as cyclobutyl moieties, in place of the pentofuranosyl sugar.
[0205] Base modifications and substitutions Nucleic acids of interest may also contain modifications or substitutions of nucleobases (often simply referred to in the art as "bases"). As used herein, "unmodified" or "natural" nucleobases include the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include 5-hydroxymethylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine and thymine. cytosine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine. Further modified nucleobases include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).
[0206] Heterocyclic base moieties can also include those in which the purine or pyrimidine base is replaced with other heterocycles, such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone. Additional nucleobases include those disclosed in U.S. Patent No. 3,687,808, those disclosed in The Concise Encyclopedia of Polymer Science and Engineering, pages 858-859, Kroschwitz, JI, ed. John Wiley & Sons, 1990, those disclosed by Englisch et al., Angewandte Chemie, International Edition, 1991, 30,613, and those disclosed by Sanghvi, YS, Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, ST and Lebleu, B., ed., CRC Press, 1993, the disclosures of which are incorporated herein by reference in their entirety.Some of these nucleobases are useful for increasing the binding affinity of oligomeric compounds. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, such as 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-Methylcytosine substitutions have been shown to increase nucleic acid duplex stability by 0.6 to 1.2°C (Sanghvi et al., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278, the disclosure of which is incorporated herein by reference in its entirety), and are preferred base substitutions, for example, when combined with 2'-O-methoxyethyl sugar modifications.
[0207] Conjugates Another possible modification of a nucleic acid of interest involves chemically linking one or more moieties or conjugates to the polynucleotide that improve the activity, cellular distribution, or cellular uptake of the oligonucleotide. These moieties or conjugates can include conjugate groups covalently attached to functional groups such as primary or secondary hydroxyl groups. Conjugate groups include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that improve the pharmacodynamic properties of oligomers, and groups that improve the pharmacokinetic properties of oligomers. Suitable conjugate groups include cholesterol, lipids, phospholipids, biotin, phenazine, folate, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that improve pharmacodynamic properties include groups that improve uptake, improve resistance to degradation, and / or enhance sequence-specific hybridization with the target nucleic acid. Groups that improve pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of the nucleic acid of interest.
[0208] Conjugate moieties include cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, 6553-6556), cholic acid (Manoharan et al., Bioorg. Med. Chem. Let., 1994, 4, 1053-1060), thioethers such as hexyl-S-tritylthiol (Manoharan et al., Ann. NY Acad. Sci., 1992, 660, 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, 2765-2770), thiocholesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, 533-538), aliphatic chains such as dodecanediol or undecyl residues (Saison-Behmoaras et al., EMBO J., 1991, 10, 1111-1118; Kabanov et al., FEBS Lett., 1990, 259, 327-330; Svinarchuk et al., Biochimie, 1993, 75, 49-54), phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654; Shea et al., Nucl. Acids Res., 1990, 18, 3777-3783), polyamine or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, 969-973), or adamantane acetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654), palmityl moiety (Mishra et al., Biochim. Biophys. Acta, 1995, 1264, 229-237), or octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al., J. Pharmacol. Exp. Ther.Lipid moieties include, but are not limited to, lipid moieties such as those listed above (Beckmann, 1996, 277, 923-937).
[0209] The conjugate may include a "protein transduction domain" or PTD (also known as a CPP - cell penetrating peptide), which may refer to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates crossing of a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule and / or nanoparticle, which may range from a small polar molecule to a large macromolecule, facilitates membrane crossing of the molecule, for example, moving from the extracellular space to the intracellular space or from the cytosol into an organelle (e.g., the nucleus). In some cases, the PTD is covalently attached to the 3' end of the exogenous polynucleotide. In some cases, the PTD is covalently attached to the 5' end of the exogenous polynucleotide. Exemplary PTDs include a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT, including YGRKKRRQRRR (SEQ ID NO:130)), a polyarginine sequence containing multiple arginines sufficient for direct entry into cells (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines), a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96), a Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737), a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256), polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008), TIFF0007672445000017.tif4128, Transportan TIFF0007672445000018.tif4128, TIFF0007672445000019.tif4128, and Exemplary PTDs include, but are not limited to, TIFF0007672445000020.tif4128. TIFF0007672445000021.tif4128, and arginine homopolymers of 3 to 50 arginine residues. Exemplary PTD domain amino acid sequences include the following: TIFF0007672445000022.tif17147. In some cases, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June;1(5-6):371-381). ACPPs contain a polycationic CPP (e.g., Arg9 or "R9") connected to a matching polyanion (e.g., Glu9 or "E9") via a cleavable linker, which reduces the net charge to near zero, thereby inhibiting cellular attachment and uptake. Upon cleavage of the linker, the polyanion is released, locally unmasking the polyarginine and its inherent adhesive properties, thus "activating" the ACPP to cross the membrane.
[0210] Introduction of components into target cells The CasZ guide RNA (or a nucleic acid comprising a nucleotide sequence encoding it) and / or the CasZ polypeptide (or a nucleic acid comprising a nucleotide sequence encoding it) and / or the CasZ trancRNA (or a nucleic acid comprising a nucleotide sequence encoding it) and / or the donor polynucleotide (donor template) can be introduced into a host cell by any of a variety of well-known methods.
[0211] The CasZ system of the present disclosure can be delivered to target cells using any of a variety of compounds and methods. As a non-limiting example, the CasZ system of the present disclosure can be combined with lipids. As another non-limiting example, the CasZ system of the present disclosure can be combined with particles or incorporated into particles.
[0212] Methods for introducing nucleic acids into host cells are known in the art, and any convenient method can be used to introduce a nucleic acid of interest (e.g., an expression construct / vector) into a target cell (e.g., a prokaryotic cell, a eukaryotic cell, a plant cell, an animal cell, a mammalian cell, a human cell, etc.). Suitable methods include, for example, viral infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et al., Adv Drug Deliv Rev. 2012 Sep 13. pii:S0169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023), and the like.
[0213] In some cases, a CasZ polypeptide of the present disclosure (e.g., a wild-type protein, a variant protein, a chimeric / fusion protein, dCasZ, etc.) is provided as a nucleic acid (e.g., mRNA, DNA, a plasmid, an expression vector, a viral vector, etc.) encoding the CasZ polypeptide. In some cases, a CasZ polypeptide of the present disclosure is provided directly as a protein (e.g., without an associated guide RNA or with an associated guide RNA, i.e., as a ribonucleoprotein complex). A CasZ polypeptide of the present disclosure can be introduced into (provided to) a cell by any convenient method, and such methods are known to those of skill in the art. As an illustrative example, a CasZ polypeptide of the present disclosure can be directly injected into a cell (e.g., with or without a CasZ guide RNA or a nucleic acid encoding a CasZ guide RNA, with or without a donor polynucleotide, and with or without a CasZ transcript RNA). As another example, a preformed complex of a CasZ polypeptide of the present disclosure and a CasZ guide RNA (RNP) can be introduced into a cell (e.g., a eukaryotic cell) (e.g., via injection, via nucleofection, conjugated to one or more components, e.g., conjugated to a CasZ protein, conjugated to a guide RNA, conjugated to a CasZ trancRNA, conjugated to a CasZ polypeptide and a guide RNA of the present disclosure, via a protein transduction domain (PTD), etc.).
[0214] In some cases, nucleic acids (e.g., CasZ guide RNAs and / or nucleic acids encoding them, nucleic acids encoding CasZ proteins, CasZ transcripts and / or nucleic acids encoding them, etc.) and / or polypeptides (e.g., CasZ polypeptides, CasZ fusion polypeptides) are delivered to cells (e.g., target host cells) in particles or associated with particles. In some cases, a CasZ system of the present disclosure is delivered to cells in particles or associated with particles. The terms "particle" and "nanoparticle" can be used interchangeably where appropriate. For example, a recombinant expression vector comprising a nucleotide sequence encoding a CasZ polypeptide and / or CasZ guide RNA of the present disclosure, an mRNA comprising a nucleotide sequence encoding a CasZ polypeptide of the present disclosure, and a guide RNA can be delivered simultaneously using a particle or lipid envelope. For example, CasZ polypeptides and / or CasZ guide RNAs and / or transRNAs can be delivered, e.g., as a complex (e.g., a ribonucleoprotein (RNP) complex), via a particle, e.g., a delivery particle comprising a lipid or lipidoid and a hydrophilic polymer, e.g., a cationic lipid and a hydrophilic polymer, e.g., where the cationic lipid comprises 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP) or 1,2-ditetradecanoyl-sn-glycero-3-phosphocholine (DMPC), and / or the hydrophilic polymer comprises ethylene glycol or polyethylene glycol (PEG), and / or the particle further comprises cholesterol (e.g., particles from Formulation 1 = DOTAP 100, DMPC 0, PEG 0, cholesterol 0; Formulation No. 2 = DOTAP 90, DMPC 0, PEG 10, cholesterol 0; Formulation No. 3 = DOTAP 90, DMPC 0, PEG 5, cholesterol 5).For example, particles can be formed using a multi-step process in which the CasZ polypeptide and CasZ guide RNA are mixed together, e.g., in a 1:1 molar ratio, e.g., at room temperature, for e.g., 30 minutes, in, e.g., sterile, nuclease-free 1x phosphate buffered saline (PBS); separately, if applicable to the formulation, DOTAP, DMPC, PEG, and cholesterol are dissolved in alcohol, e.g., 100% ethanol, and the two solutions are mixed together to form particles containing the complexes.
[0215] The CasZ polypeptides of the present disclosure (or mRNAs comprising a nucleotide sequence encoding the CasZ polypeptides of the present disclosure, or recombinant expression vectors comprising a nucleotide sequence encoding the CasZ polypeptides of the present disclosure) and / or CasZ guide RNAs (or nucleic acids, such as one or more expression vectors encoding CasZ guide RNAs) can be co-delivered using particles or lipid envelopes. For example, biodegradable core-shell nanoparticles having a poly(β-amino ester) (PBAE) core encapsulated by a phospholipid bilayer shell can be used. In some cases, particles / nanoparticles based on self-assembling bioadhesive polymers are used, and such particles / nanoparticles can be applied, for example, to oral delivery of peptides to the brain, intravenous delivery of peptides, and intranasal delivery of peptides. Other embodiments, such as oral absorption and intraocular delivery of hydrophobic drugs, are also contemplated. Molecular enveloping technology, which requires an engineered polymer envelope to be protected and delivered to the site of disease, can be used. A dose of approximately 5 mg / kg can be used in single or multiple doses, depending on various factors, such as the target tissue.
[0216] Lipidoid compounds are also useful in the administration of polynucleotides (e.g., as described in U.S. Patent Application Publication No. 20110293703) and can be used to deliver the disclosed CasZ polypeptides, disclosed CasZ fusion polypeptides, disclosed RNPs, disclosed nucleic acids, or disclosed CasZ systems. In one embodiment, amino alcohol lipidoid compounds are combined with agents to be delivered to cells or subjects to form microparticles, nanoparticles, liposomes, or micelles. Amino alcohol lipidoid compounds can be combined with other amino alcohol lipidoid compounds, polymers (synthetic or natural), surfactants, cholesterol, carbohydrates, lipids, etc. to form particles. These particles can then be optionally combined with pharmaceutical excipients to form pharmaceutical compositions.
[0217] Poly(beta-amino alcohols) (PBAAs) can be used to deliver a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA), or a CasZ system of the present disclosure to a target cell. U.S. Patent Publication No. 20130302401 relates to a class of poly(beta-amino alcohols) (PBAAs) that have been prepared using combinatorial polymerization.
[0218] Sugar-based particles, e.g., GalNAc, can be used as described in WO2014118272 (incorporated herein by reference) and Nair, JK et al., 2014, Journal of the American Chemical Society 136(49), 16958-16961) to deliver a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure, or a CasZ system of the present disclosure to a target cell.
[0219] In some cases, lipid nanoparticles (LNPs) are used to deliver the disclosed CasZ polypeptides, disclosed CasZ fusion polypeptides, disclosed RNPs, disclosed nucleic acids (e.g., CasZ guide RNAs and / or CasZ trancRNAs), or disclosed CasZ systems to target cells. Negatively charged polymers such as RNA can be loaded into LNPs at low pH values (e.g., pH 4), where ionized lipids exhibit a positive charge. However, at physiological pH values, LNPs exhibit a low surface charge, corresponding to longer circulation times. The focus is on four ionizable cationic lipids: 1,2-dilineoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxy-keto-N,N-dimethyl-3-aminopropane (DLinKDMA), and 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA). Preparation of LNPs is described, for example, in Rosin et al. (2011) Molecular Therapy 19:1286-2200. The cationic lipids 1,2-dilineoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA), (3-o-[2''-(methoxypolyethylene glycol 2000)sucrineoyl]-1,2-dimyristoyl-sn-glycol (PEG-S-DMG), and R-3-[(omega-methoxy-poly(ethylene glycol) 2000)carbamoyl]-1,2-dimyristoyloxylpropyl-3-amine (PEG-C-DOMG) may also be used.Nucleic acids (e.g., CasZ guide RNAs, nucleic acids of the present disclosure, etc.) can be encapsulated in LNPs containing DLinDAP, DLinDMA, DLinK-DMA, and DLinKC2-DMA (cationic lipid:DSPC:CHOL:PEGS-DMG or PEG-C-DOMG in a 40:10:40:10 molar ratio). In some cases, 0.2% SP-DiOC18 is incorporated.
[0220] Spherical nucleic acid (SNA™) constructs and other nanoparticles (particularly gold nanoparticles) can be used to deliver a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA), or a CasZ system of the present disclosure to a target cell. For example, Cutler et al.,J.Am.Chem.Soc.2011 133:9254-9257, Hao et al.,Small.2011 7:3158-3162, Zhang et al.,ACS Nano.2011 5:6962-6970, Cutler et al.,J.Am.Chem.Soc.2012 134:1376-1391, Young et al., Nano Lett.2012 12:3867-71, Zheng et al., Proc. Natl. Acad. Sci. USA.2012 109:11975-80, Mirkin, Nanomedicine 2012 7:635-638, Zhang et al. al.,J.Am.Chem.Soc.2012 134:16488-1691; Weintraub, Nature 2013 495:S14-S16; Choi et al., Proc. Natl. Acad. Sci. USA. 2013 110(19):7625-7630; Jensen et al., Sci. Transl. Med. 5, 209ra152 (2013); and Mirkin, et al., Small, 10:186-192.
[0221] Self-assembled nanoparticles bearing RNA can be constructed from polyethyleneimine (PEI) that is PEGylated with an Arg-Gly-Asp (RGD) peptide ligand attached to the distal end of the polyethylene glycol (PEG).
[0222] Generally, a "nanoparticle" refers to any particle having a diameter of less than 1000 nm. In some cases, nanoparticles suitable for use in delivering a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ transcript RNA), or a CasZ system of the present disclosure to a target cell have a diameter of 500 nm or less, e.g., 25 nm to 35 nm, 35 nm to 50 nm, 50 nm to 75 nm, 75 nm to 100 nm, 100 nm to 150 nm, 150 nm to 200 nm, 200 nm to 300 nm, 300 nm to 400 nm, or 400 nm to 500 nm. In some cases, nanoparticles suitable for use in delivering a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure, or a CasZ system of the present disclosure to a target cell have a diameter of 25 nm to 200 nm. In some cases, nanoparticles suitable for use in delivering a CasZ polypeptide of this disclosure, a CasZ fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasZ system of this disclosure to a target cell have a diameter of 100 nm or less. In some cases, nanoparticles suitable for use in delivering a CasZ polypeptide of this disclosure, a CasZ fusion polypeptide of this disclosure, an RNP of this disclosure, a nucleic acid of this disclosure, or a CasZ system of this disclosure to a target cell have a diameter of 35 nm to 60 nm.
[0223] Nanoparticles suitable for use in delivering a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA), or a CasZ system of the present disclosure to a target cell can be provided in different forms, such as solid nanoparticles (e.g., metallic nanoparticles such as silver, gold, iron, or titanium; non-metallic nanoparticles; lipid-based solids; polymeric nanoparticles), suspensions of nanoparticles, or combinations thereof. Metallic, insulating, and semiconductor nanoparticles, as well as hybrid structures (e.g., core-shell nanoparticles), can be prepared. Nanoparticles made of semiconductor materials can also be referred to as labeled quantum dots if they are small enough (typically 10 nm or less) for quantization of electronic energy levels to occur. Such nanoscale particles are used as drug carriers or imaging agents in biomedical applications and can be adapted for similar purposes in the present disclosure.
[0224] Semi-solid and soft nanoparticles are also suitable for use in delivering a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA), or a CasZ system of the present disclosure to a target cell. A prototypical nanoparticle of semi-solid nature is the liposome.
[0225] In some cases, exosomes are used to deliver a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA), or a CasZ system of the present disclosure to a target cell. Exosomes are endogenous nanovesicle particles that can transport RNA and proteins and deliver RNA to the brain and other target organs.
[0226] In some cases, liposomes are used to deliver the disclosed CasZ polypeptides, the disclosed CasZ fusion polypeptides, the disclosed RNPs, the disclosed nucleic acids (e.g., CasZ guide RNAs and / or CasZ trancRNAs), or the disclosed CasZ system to target cells. Liposomes are spherical vesicular structures composed of a unilamellar or multilamellar lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes can be made from several different types of lipids, but phospholipids are most commonly used to generate liposomes. Liposome formation occurs spontaneously when a lipid membrane is mixed with an aqueous solution, but it can also be promoted by applying force in the form of shaking using a homogenizer, sonicator, or extrusion device. Several other additives may be added to liposomes to modify their structure and properties. For example, either cholesterol or sphingomyelin may be added to the liposome mixture to help stabilize the liposome structure and prevent leakage of the liposome's internal cargo. Liposomal formulations may be composed primarily of lipids such as natural phospholipids and 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, and monosialogangliosides.
[0227] Stable nucleic acid-lipid particles (SNALPs) can be used to deliver a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ transcript RNA), or a CasZ system of the present disclosure to target cells. The SNALP formulation can contain the lipids 3-N-[(methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristoyloxy-propylamine (PEG-C-DMA), 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), and cholesterol in a molar ratio of 2:40:10:48. SNALP liposomes can be prepared by combining D-Lin-DMA and PEG-C-DMA with distearoylphosphatidylcholine (DSPC), cholesterol, and siRNA using a lipid / siRNA ratio of 25:1 and a 48 / 40 / 10 / 2 molar ratio of cholesterol / D-Lin-DMA / DSPC / PEG-C-DMA. The resulting SNALP liposomes can be approximately 80-100 nm in size. SNALP can contain synthetic cholesterol (Sigma-Aldrich, St. Louis, MO, USA), dipalmitoylphosphatidylcholine (Avanti Polar Lipids, Alabaster, Ala., USA), 3-N-[(w-methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristoyloxypropylamine, and cationic 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane. SNALPs can include synthetic cholesterol (Sigma-Aldrich), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC, Avanti Polar Lipids Inc.), PEG-cDMA, and 1,2-dilinoleyloxy-3-(N;N-dimethyl)aminopropane (DLinDMA).
[0228] Other cationic lipids, such as the amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA), can be used to deliver the disclosed CasZ polypeptides, the disclosed CasZ fusion polypeptides, the disclosed RNPs, the disclosed nucleic acids (e.g., CasZ guide RNAs and / or CasZ trancRNAs), or the disclosed CasZ systems to target cells. Preformed vesicles having the following lipid composition can be contemplated: amino lipid, distearoylphosphatidylcholine (DSPC), cholesterol, and (R)-2,3-bis(octadecyloxy)propyl-1-(methoxypoly(ethylene glycol)2000)propylcarbamate (PEG-lipid) in a 40 / 10 / 40 / 10 molar ratio, respectively, and a FVII siRNA / total lipid ratio of approximately 0.05 (w / w). To ensure a narrow particle size distribution in the 70-90 nm range and a low polydispersity index of 0.11 ± 0.04 (n = 56), particles can be extruded up to three times through an 80 nm membrane before adding guide RNA. Particles containing highly potent amino lipids 16 may also be used, where the molar ratio of the four lipid components 16, DSPC, cholesterol, and PEG-lipid (50 / 10 / 38.5 / 1.5) can be further optimized to improve in vivo activity.
[0229] Lipids can be combined with the disclosed CasZ system or its component(s) or nucleic acids encoding same to form lipid nanoparticles (LNPs). Suitable lipids include, but are not limited to, DLin-KC2-DMA4, C12-200, and the co-lipids disteroylphosphatidylcholine, cholesterol, and PEG-DMG, which can be combined with the disclosed CasZ system or its components using a spontaneous vesicle formation procedure. The molar ratio of the components can be approximately 50 / 10 / 38.5 / 1.5 (DLin-KC2-DMA or C12-200 / disteroylphosphatidylcholine / cholesterol / PEG-DMG).
[0230] The CasZ system of the present disclosure, or components thereof, can be delivered encapsulated in PLGA microspheres, as further described in U.S. Patent Publication Nos. 20130252281, 20130245107, and 20130244279.
[0231] Supercharged proteins can be used to deliver the disclosed CasZ polypeptides, the disclosed CasZ fusion polypeptides, the disclosed RNPs, the disclosed nucleic acids (e.g., CasZ guide RNAs and / or CasZ trancRNAs), or the disclosed CasZ systems to target cells. Supercharged proteins are a class of engineered or naturally occurring proteins that have an unusually high net positive or negative theoretical charge. Both supernegatively and superpositively charged proteins exhibit the ability to resist thermally or chemically induced aggregation. Superpositively charged proteins can also penetrate mammalian cells. Associating cargo with these proteins, such as plasmid DNA, RNA, or other proteins, can enable the functional delivery of these macromolecules into mammalian cells both in vitro and in vivo.
[0232] Cell-penetrating peptides (CPPs) can be used to deliver the disclosed CasZ polypeptides, disclosed CasZ fusion polypeptides, disclosed RNPs, disclosed nucleic acids (e.g., CasZ guide RNAs and / or CasZ trancRNAs), or disclosed CasZ systems to target cells. CPPs typically have an amino acid composition that either contains a high relative abundance of positively charged amino acids, such as lysine or arginine, or has a sequence containing an alternating pattern of polar / charged and nonpolar, hydrophobic amino acids.
[0233] An implantable device can be used to deliver a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA) (e.g., a CasZ guide RNA, a nucleic acid encoding a CasZ guide RNA, a nucleic acid encoding a CasZ polypeptide, a donor template, etc.), or a CasZ system of the present disclosure to a target cell (e.g., a target cell in vivo, where the target cell is a target cell in the circulation, a target cell in a tissue, a target cell within an organ, etc.). Implantable devices suitable for use in delivering a CasZ polypeptide of the present disclosure, a CasZ fusion polypeptide of the present disclosure, an RNP of the present disclosure, a nucleic acid of the present disclosure (e.g., a CasZ guide RNA and / or a CasZ trancRNA), or a CasZ system of the present disclosure to a target cell (e.g., a target cell in vivo, where the target cell is a target cell in the circulation, a target cell in a tissue, a target cell within an organ, etc.) can include a container (e.g., a reservoir, a matrix, etc.) containing a CasZ polypeptide, a CasZ fusion polypeptide, an RNP, or a CasZ system (or a component thereof, e.g., a nucleic acid of the present disclosure).
[0234] Suitable implantable devices can include, for example, a polymer substrate, such as a matrix, used as the device body, and in some cases, additional scaffold materials, such as metals or additional polymers, as well as materials to improve visibility and imaging. Implantable delivery devices can be advantageous in providing localized and prolonged release, with the delivered polypeptide and / or nucleic acid being released directly into the target site, e.g., the extracellular matrix (ECM), the vasculature surrounding a tumor, diseased tissue, etc. Suitable implantable devices include devices suitable for delivery to cavities such as the peritoneal cavity and / or for use in any other type of administration where the drug delivery system is not tethered or attached, and include a biostable and / or biodegradable and / or bioabsorbable polymer substrate (e.g., which may optionally be a matrix). In some cases, suitable implantable drug delivery devices include degradable polymers, and the primary release mechanism is bulk erosion. In some cases, suitable implantable drug delivery devices include non-degradable or slowly degrading polymers, and the primary release mechanism is diffusion rather than bulk erosion, whereby the outer portion functions as a membrane and the inner portion functions as a drug reservoir that is virtually unaffected by the environment for extended periods (e.g., from about one week to about several months). Combinations of different polymers with different release mechanisms may optionally be used. The concentration gradient may be maintained effectively constant for a significant portion of the total release period, so that the diffusion rate is effectively constant (referred to as "zero-mode" diffusion). The term "constant" refers to a diffusion rate that is maintained above a low threshold of therapeutic efficacy but may still optionally be characterized by an initial burst and / or may fluctuate (e.g., increase and decrease to a certain extent). The diffusion rate may be maintained in this manner for extended periods and may be considered constant at a certain level to optimize the therapeutic effective period, e.g., the effective silencing period.
[0235] In some cases, implantable delivery systems are designed to shield the nucleotide-based therapeutic agent from degradation, either due to its chemical nature or due to attack from enzymes and other factors within the subject's body.
[0236] The site for device implantation, or target site, can be selected for maximum therapeutic effect. For example, the delivery device can be implanted within or proximal to the tumor environment or tumor-associated blood supply. Target locations can include, for example, 1) the brain at degenerative sites such as Parkinson's disease or Alzheimer's disease in the basal ganglia, white matter, and gray matter; 2) the spine, as in amyotrophic lateral sclerosis (ALS); 3) the cervix; 4) active and chronically inflamed joints; 5) the dermis, as in psoriasis; 6) sympathetic and sensory nerve sites for analgesic effects; 7) bone; 8) sites of acute or chronic infection; 9) the vagina; 10) the cochlear-auditory system, the membranous labyrinth of the inner ear, the vestibular system, 11) in the trachea, 12) in the heart, annulus, epicardium, 13) in the urinary tract or bladder, 14) in the biliary system, 15) in parenchymal tissues including, but not limited to, the kidney, liver, and spleen, 16) in lymph nodes, 17) in salivary glands, 18) in the gums, 19) in articular tissues (in joints), 20) in the eye, 21) in brain tissue, 22) in the ventricles, 23) in cavities including the peritoneal cavity (for example, but not limited to, in the case of ovarian cancer), 24) in the esophagus, 25) in the rectum, and 26) in the vasculature.
[0237] The method of insertion, such as implantation, may optionally already be used for other types of tissue implantation and / or insertion and / or tissue harvesting, optionally without modification or optionally with only minor modifications in such methods, optionally including, but not limited to, brachytherapy, biopsy, endoscopy with and / or without ultrasound, e.g., stereotactic radioscopy of brain tissue, laparoscopy, including implantation of a laparoscope into joints, abdominal organs, bladder walls, and body cavities.
[0238] Modified host cells The present disclosure provides modified cells comprising a CasZ polypeptide of the present disclosure and / or a nucleic acid comprising a nucleotide sequence encoding a CasZ polypeptide of the present disclosure. The present disclosure provides modified cells comprising a CasZ polypeptide of the present disclosure, where the modified cell is a cell that does not normally comprise a CasZ polypeptide of the present disclosure. The present disclosure provides modified cells (e.g., genetically modified cells) comprising a nucleic acid comprising a nucleotide sequence encoding a CasZ polypeptide of the present disclosure. The present disclosure provides genetically modified cells genetically modified with an mRNA comprising a nucleotide sequence encoding a CasZ polypeptide of the present disclosure. The present disclosure provides genetically modified cells genetically modified with a recombinant expression vector comprising a nucleotide sequence encoding a CasZ polypeptide of the present disclosure. The present disclosure provides genetically modified cells genetically modified with a recombinant expression vector comprising a) a nucleotide sequence encoding a CasZ polypeptide of the present disclosure, and b) a nucleotide sequence encoding a CasZ guide RNA of the present disclosure. The present disclosure provides a genetically modified cell that has been genetically modified with a recombinant expression vector comprising: a) a nucleotide sequence encoding a CasZ polypeptide of the disclosure; b) a nucleotide sequence encoding a CasZ guide RNA of the disclosure; and c) a nucleotide sequence encoding a donor template.
[0239] Cells that serve as recipients of the CasZ guide RNAs, CasZ polypeptides of the present disclosure, and / or nucleic acids comprising nucleotide sequences encoding CasZ polypeptides of the present disclosure, and / or CasZ guide RNAs (or nucleic acids encoding same), and / or CasZ trancRNAs (or nucleic acids encoding same) can be any of a variety of cells, including, for example, in vitro cells, in vivo cells, ex vivo cells, primary cells, cancer cells, animal cells, plant cells, algal cells, fungal cells, etc. Cells that serve as recipients of the CasZ polypeptides of the present disclosure, and / or nucleic acids comprising nucleotide sequences encoding CasZ polypeptides of the present disclosure, and / or CasZ guide RNAs of the present disclosure are referred to as "host cells" or "target cells." A host cell or target cell can be a recipient of a CasZ system of the present disclosure. A host cell or target cell can be a recipient of a CasZ RNP of the present disclosure. A host cell or target cell can be a recipient of a single component of a CasZ system of the present disclosure.
[0240] Non-limiting examples of cells (target cells) include prokaryotic cells, fungal cells, bacterial cells, archaeal cells, cells of unicellular eukaryotes, protist cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, angiosperms, ferns, club mosses, hornworts, liverworts, bryophytes, liverworts, dicotyledons, monocotyledons, etc.), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, etc.), marine algae (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., Drosophila, worms, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., ungulates (e.g., pigs, cows, goats, sheep), rodents (e.g., rats, mice), non-human primates, humans, felines (e.g., cats), canines (e.g., dogs), etc.). In some cases, the cell is not derived from a natural organism (e.g., the cell can be synthetically produced, also referred to as an artificial cell).
[0241] The cell can be an in vitro cell (e.g., a cell in culture, e.g., an established cultured cell line). The cell can be an ex vivo cell (a cultured cell from an individual). The cell can be an in vivo cell (e.g., a cell within an individual). The cell can be an isolated cell. The cell can be a cell inside an organism. The cell can be an organism. The cell can be a cell in cell culture (e.g., an in vitro cell culture). The cell can be one of a population of cells. The cell can be a prokaryotic cell or derived from a prokaryotic cell. The cell can be a bacterial cell or derived from a bacterial cell. The cell can be an archaeal cell or derived from an archaeal cell. The cell can be a eukaryotic cell or derived from a eukaryotic cell. The cell can be a plant cell or derived from a plant cell. The cell can be an animal cell or derived from an animal cell. The cell can be an invertebrate cell or derived from an invertebrate cell. The cell may be a vertebrate cell or may be derived from a vertebrate. The cell may be a mammalian cell or may be derived from a mammalian cell. The cell may be a rodent cell or may be derived from a rodent cell. The cell may be a human cell or may be derived from a human cell. The cell may be a microbial cell or may be derived from a microbial cell. The cell may be a fungal cell or may be derived from a fungal cell. The cell may be an insect cell. The cell may be an arthropod cell. The cell may be a protozoan cell. The cell may be a helminth cell.
[0242] Suitable cells include stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells (e.g., oocytes, sperm, oogonia, spermatogonia, etc.), somatic cells such as fibroblasts, oligodendrocytes, glial cells, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells, etc.
[0243] Suitable cells include human embryonic stem cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, autologous expanded cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, blood stem cells, myoblasts, adult hepatocytes, bone marrow cells, mesenchymal cells, embryonic stem cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, osteoblasts, chondrocytes, exogenous cells, endogenous cells, stem cells, hematopoietic stem cells, bone marrow derived progenitor cells, cardiomyocytes, skeletal cells, fetal cells, undifferentiated cells, multipotent progenitor cells, unipotent progenitor cells, monocytes, cardiac myoblasts, skeletal myoblasts, macrophages, capillary endothelial cells, xenogeneic cells, allogeneic cells, and postnatal stem cells.
[0244] In some cases, the cell is an immune cell, a neuron, an epithelial cell, an endothelial cell, or a stem cell. In some cases, the immune cell is a T cell, a B cell, a monocyte, a natural killer cell, a dendritic cell, or a macrophage. In some cases, the immune cell is a cytotoxic T cell. In some cases, the immune cell is a helper T cell. In some cases, the immune cell is a regulatory T cell (Treg).
[0245] In some cases, the cells are stem cells. Stem cells include adult stem cells, which are also referred to as somatic stem cells.
[0246] Adult stem cells reside in differentiated tissues but retain the property of self-renewal and the ability to give rise to multiple cell types, usually cell types typical of the tissue they are found in. Many examples of somatic stem cells are known to those skilled in the art, including muscle stem cells, hematopoietic stem cells, epithelial stem cells, neural stem cells, mesenchymal stem cells, mammalian stem cells, intestinal stem cells, mesodermal stem cells, endothelial stem cells, olfactory stem cells, neural crest stem cells, etc.
[0247] Stem cells of interest include mammalian stem cells, where the term "mammal" refers to any animal classified as a mammal, including humans, non-human primates, domestic and farm animals, and zoo, laboratory, sport, or pet animals, such as dogs, horses, cats, cows, mice, rats, rabbits, etc. In some cases, the stem cells are human stem cells. In some cases, the stem cells are rodent (e.g., mouse, rat) stem cells. In some cases, the stem cells are non-human primate stem cells.
[0248] The stem cells may express one or more stem cell markers, such as SOX9, KRT19, KRT7, LGR5, CA9, FXYD2, CDH6, CLDN18, TSPAN8, BPIFB1, OLFM4, CDH17, and PPARGC1A.
[0249] In some cases, the stem cells are hematopoietic stem cells (HSCs). HSCs are mesodermally derived cells that can be isolated from bone marrow, blood, umbilical cord blood, fetal liver, and yolk sac. HSCs express CD34 + and CD3 - HSCs are characterized as being capable of regenerating in vivo into erythroid, neutrophil-macrophage, megakaryocyte, and lymphoid hematopoietic cell lineages. In vitro, HSCs can be induced to undergo at least some self-renewal cell division and can be induced to differentiate into the same lineages as those found in vivo. Thus, HSCs can be induced to differentiate into one or more of erythroid cells, megakaryocytes, neutrophils, macrophages, and lymphoid cells.
[0250] In other cases, the stem cell is a neural stem cell (NSC). Neural stem cells (NSC) can differentiate into neurons and glia (including oligodendrocytes and astrocytes). Neural stem cells are pluripotent stem cells that have the ability to divide multiple times and, under certain conditions, can produce daughter cells that are neural stem cells, or neural progenitor cells that can be neuroblasts or glioblasts, for example, cells that are committed to becoming one or more types of neurons and glial cells, respectively. Methods for obtaining NSCs are known in the art.
[0251] In other cases, the stem cells are mesenchymal stem cells (MSCs). MSCs, originally derived from embryonic mesoderm and isolated from adult bone marrow, can differentiate to form muscle, bone, cartilage, fat, bone marrow matrix, and tendon. Methods for isolating MSCs are known in the art, and any known method can be used to obtain MSCs. See, for example, U.S. Patent No. 5,736,396, which describes the isolation of human MSCs.
[0252] The cell is, in some cases, a plant cell. The plant cell can be a monocotyledonous plant cell. The cell can be a dicotyledonous plant cell.
[0253] In some cases, the cell is a plant cell. For example, the cell can be a cell of a major agricultural plant, such as barley, bean (dry edible), canola, corn, cotton (pima), cotton (apland), flaxseed, hay (alfalfa), hay (non-alfalfa), oats, peanuts, rice, sorghum, soybean, sugar beet, sugarcane, sunflower (oil), sunflower (non-oil), sweet potato, tobacco (burley), tobacco (flue-cured), tomato, wheat (durham), wheat (spring), wheat (winter), etc. In another example, the cells may be derived from, for example, alfalfa sprouts, aloe leaves, arrowroot, arrowheads, artichokes, asparagus, bamboo shoots, banana flowers, bean sprouts, beans, beet stalks, beets, bitter melon, bok choy, broccoli, broccoli rabe (rapini), Brussels sprouts, cabbage, cabbage sprouts, cactus leaf (nopal), pumpkin, cardoon, carrot, cauliflower, celery, chayote, Chinese artichoke (kronu), Chinese cabbage, Chinese cabbage, Chinese chives, Chinese leek, Chinese spinach, Chinese chrysanthemum (tung)ho), collard greens, corn stalks, sweet corn, cucumber, radish, edible dandelion greens, taro, pea sprouts (pea leaves), touki (winter melon), eggplant, endive, esculenta, osmanthus, shepherd's purse, frisee, mustard greens (Chinese mustard), kailan, galangal (Siamese ginger), garlic, ginger, burdock, greens, Hanover salad greens, ouau sotto le, Jerusalem artichoke, jicama, kale greens, kohlrabi, and Kaza (Querite), Bibb lettuce, Boston lettuce, Boston red lettuce, Green leaf lettuce, Iceberg lettuce, Sunny lettuce, Oakleaf green lettuce, Oakleaf red lettuce, Processed lettuce, Red leaf lettuce, Romaine lettuce, Ruby romaine lettuce, Russian red mustard lettuce, Linkok, Loboc, Cowpea, Lotus root, Marsh, Agave leaf lettuce, Yam, Mesclun Mix, mizuna, moap (luffa), moo, mokua (fuzzy squash), mushroom, mustard, yam, okra, water spinach, spring onion, opo (long squash), decorative corn, decorative gourd, parsley, parsnip, peas, capsicum (bell type), capsicum, pumpkin, radicchio, radish sprouts, radish, rape greens, rape greens, rhubarb, romaine (baby red), rutabaga, hamcho (sea bean), loofah (toca The cells of vegetable crops include, but are not limited to, taro, taro leaves, taro sprouts, tatsoi, tepeguaje (leucaena), tindora, tomatillo, tomato, tomato (cherry), tomato (grape type), tomato (plum type), turmeric, turnip greens, turnip, water chestnut, yampi, yams, rapeseed, and yuca (cassava).
[0254] The cell, in some cases, is an arthropod cell. For example, the cell may be an arthropod cell from any of the following families: Chelicerata, Myriapodia, Hexipodia, Arachnida, Insecta, Archaeognatha, Thysanura, Palaeoptera, Ephemeroptera, Odonata, Anisoptera, Zygoptera, Neoptera, Exopterygota, Plecoptera, Embiooptera, Orthoptera, Zoraptera, Dermaptera, Dictyoptera, Notoptera, Grylloblattidae, Mantophasmatidae, Phasmatodea, etc. The cell may be of a suborder, family, subfamily, group, subgroup, or species of the order Parapneuroptera, Blattaria, Isoptera, Mantodea, Parapneuroptera, Psocoptera, Thysanoptera, Phthiraptera, Hemiptera, Endopterygota, or Holometabola, Hymenoptera, Coleoptera, Strepsiptera, Raphidioptera, Megaloptera, Neuroptera, Mecoptera, Siphonaptera, Diptera, Trichoptera, or Lepidoptera.
[0255] The cell, in some cases, is an insect cell, for example, in some cases, the cell is a mosquito, grasshopper, hemipteran, fly, flea, bee, wasp, ant, louse, moth, or beetle cell.
[0256] kit The present disclosure provides kits that include a CasZ system of the present disclosure or components of a CasZ system of the present disclosure.
[0257] Kits of the present disclosure can include any combination, as listed for the CasZ system (see, e.g., above). Kits of the present disclosure can include a) the above-described components of the CasZ system of the present disclosure, or components that may comprise the CasZ system of the present disclosure, and b) one or more additional reagents, such as i) buffers, ii) protease inhibitors, iii) nuclease inhibitors, iv) reagents necessary for developing or visualizing detectable labels, v) positive and / or negative control target DNA, vi) positive and / or negative control CasZ guide RNAs, vii) CasZ trancRNA, etc. Kits of the present disclosure can include a) the above-described components of the CasZ system of the present disclosure, or components that may comprise the CasZ system of the present disclosure, and b) a therapeutic agent.
[0258] The kits of the present disclosure can include a recombinant expression vector comprising: a) an insertion site for inserting a nucleic acid comprising a nucleotide sequence encoding a portion of a CasZ guide RNA that hybridizes to a target nucleotide sequence in a target nucleic acid, and b) a nucleotide sequence encoding a CasZ-binding portion of the CasZ guide RNA. The kits of the present disclosure can include a recombinant expression vector comprising: a) an insertion site for inserting a nucleic acid comprising a nucleotide sequence encoding a portion of a CasZ guide RNA that hybridizes to a target nucleotide sequence in a target nucleic acid, b) a nucleotide sequence encoding a CasZ-binding portion of the CasZ guide RNA, and c) a nucleotide sequence encoding a CasZ polypeptide of the present disclosure. The kits of the present disclosure can include a recombinant expression vector comprising a nucleotide sequence encoding a CasZ trancRNA.
[0259] Detection of ssDNA The CasZ (Cas14) polypeptide of the present disclosure, when activated by detection of target DNA (double-stranded or single-stranded), can indiscriminately cleave non-targeted single-stranded DNA (ssDNA). When CasZ (Cas14) is activated by a guide RNA (which occurs when the guide RNA hybridizes to a target sequence in the target DNA (i.e., when the sample contains target DNA, e.g., target ssDNA)), the protein becomes a nuclease that indiscriminately cleaves ssDNA (i.e., the nuclease cleaves non-target ssDNA (i.e., ssDNA to which the guide sequence of the guide RNA does not hybridize)). Thus, when target DNA is present in a sample (e.g., above a threshold amount in some cases), the result is cleavage of ssDNA in the sample, which can be detected using any convenient method (e.g., using labeled single-stranded detection DNA). In some cases, the CasZ polypeptide requires a transcript RNA for activation in addition to the CasZ guide RNA.
[0260] Compositions and methods are provided for detecting target DNA (double-stranded or single-stranded) in a sample. In some cases, a detection DNA is used that is single-stranded (ssDNA) and does not hybridize to the guide sequence of the guide RNA (i.e., the detection ssDNA is non-target ssDNA). Such methods can include (a) contacting a sample with (i) a CasZ polypeptide, (ii) a guide RNA comprising a region that binds to the CasZ polypeptide and a guide sequence that hybridizes to the target DNA, and (iii) a detection DNA that is single-stranded and does not hybridize to the guide sequence of the guide RNA; and (b) detecting the target DNA by measuring a detectable signal generated by cleavage of the single-stranded detection DNA by the CasZ polypeptide. In some cases, the method can include (a) contacting a sample with (i) a CasZ polypeptide, (ii) a guide RNA comprising a region that binds to the CasZ polypeptide and a guide sequence that hybridizes to the target DNA, (iii) a CasZ transcript RNA, and (iv) a detection DNA that is single-stranded and does not hybridize to the guide sequence of the guide RNA, and (b) detecting the target DNA by measuring a detectable signal generated by cleavage of the single-stranded detection DNA by the CasZ polypeptide. As described above, when a sample is activated by the guide RNA (which occurs when the sample contains target DNA to which the guide RNA hybridizes (i.e., the sample contains targeted target DNA)), the CasZ polypeptide is activated and functions as an endoribonuclease that nonspecifically cleaves ssDNA present in the sample (including non-target ssDNA). Thus, when targeted target DNA is present in the sample (e.g., above a threshold amount), the result is cleavage of ssDNA (including non-target ssDNA) in the sample, which can be detected using any convenient method (e.g., using labeled detection ssDNA).
[0261] Also provided are compositions and methods for cleaving single-stranded DNA (ssDNA) (e.g., non-target ssDNA). Such methods can include contacting a population of nucleic acids comprising a target DNA and a plurality of non-target ssDNAs with (i) a CasZ polypeptide and (ii) a guide RNA comprising a region that binds to the CasZ polypeptide and a guide sequence that hybridizes to the target DNA, where the CasZ polypeptide cleaves the plurality of non-target ssDNAs. Such methods can include contacting a population of nucleic acids comprising a target DNA and a plurality of non-target ssDNAs with (i) a CasZ polypeptide, (ii) a guide RNA comprising a region that binds to the CasZ polypeptide and a guide sequence that hybridizes to the target DNA, and (iii) a CasZ transcript RNA, where the CasZ polypeptide cleaves the plurality of non-target ssDNAs. Such methods can be used, for example, to cleave foreign ssDNA (e.g., viral DNA) in a cell.
[0262] The contacting step of the subject method can be performed in a composition comprising a divalent metal ion. The contacting step can be performed in a cell-free environment, e.g., outside a cell. The contacting step can be performed inside a cell. The contacting step can be performed on an in vitro cell. The contacting step can be performed on an ex vivo cell. The contacting step can be performed on an in vivo cell.
[0263] The guide RNA may be provided as RNA or as a nucleic acid encoding the guide RNA (e.g., DNA, such as a recombinant expression vector). The tranc RNA may be provided as RNA or as a nucleic acid encoding the guide RNA (e.g., DNA, such as a recombinant expression vector). The CasZ polypeptide may be provided as the protein itself or as a nucleic acid encoding the protein (e.g., DNA, such as mRNA, a recombinant expression vector). In some cases, two or more (e.g., three or more, four or more, five or more, or six or more) guide RNAs may be provided. In some cases, a single-molecule RNA (or a nucleic acid comprising a nucleotide sequence encoding a single-molecule RNA) is used, comprising i) a CasZ guide RNA and ii) a tranc RNA.
[0264] In some cases (e.g., when contacting a sample with a guide RNA and a CasZ polypeptide, or when contacting a sample with a guide RNA, a CasZ polypeptide, and a transcript RNA), the sample is contacted for 2 hours or less (e.g., 1.5 hours or less, 1 hour or less, 40 minutes or less, 30 minutes or less, 20 minutes or less, 10 minutes or less, 5 minutes or less, or 1 minute or less) prior to the measuring step. For example, in some cases, the sample is contacted for 40 minutes or less prior to the measuring step. In some cases, the sample is contacted for 20 minutes or less prior to the measuring step. In some cases, the sample is contacted for 10 minutes or less prior to the measuring step. In some cases, the sample is contacted for 5 minutes or less prior to the measuring step. In some cases, the sample is contacted for 50 to 60 seconds prior to the measuring step. In some cases, the sample is contacted for 40 to 50 seconds prior to the measuring step. In some cases, the sample is contacted for 30 to 40 seconds prior to the measuring step. In some cases, the sample is contacted for 20 to 30 seconds prior to the measuring step. In some cases, the sample is allowed to contact for 10 to 20 seconds before the measurement step.
[0265] In some cases, the disclosed method for detecting target DNA includes: a) contacting a sample with a guide RNA, a CasZ polypeptide, and a detection DNA, where the sample is contacted under conditions that provide for trans cleavage of the detection DNA for 2 hours or less (e.g., 1.5 hours or less, 1 hour or less, 40 minutes or less, 30 minutes or less, 20 minutes or less, 10 minutes or less, 5 minutes or less, or 1 minute or less); b) maintaining the sample from step (a) for a period of time under conditions that do not provide for trans cleavage of the detection DNA; and c) measuring a detectable signal generated by cleavage of the single-stranded detection DNA by the CasZ polypeptide after the period of step (b), thereby detecting the target DNA. Conditions that provide for trans cleavage of the detection DNA include temperature conditions such as 17°C to about 39°C (e.g., about 37°C). Conditions that do not provide for trans cleavage of the detection DNA include temperatures of 10°C or less, 5°C or less, 4°C or less, or 0°C.
[0266] In some cases, a method of detecting a target DNA of the present disclosure includes: a) contacting a sample with a guide RNA, a trans RNA, a CasZ polypeptide, and a detection DNA (or contacting a sample with i) a single-molecule RNA comprising a guide RNA and a trans RNA, i) a CasZ polypeptide, and iii) a detection DNA, wherein the sample is contacted for 2 hours or less (e.g., 1.5 hours or less, 1 hour or less, 40 minutes or less, 30 minutes or less, 20 minutes or less, 10 minutes or less, 5 minutes or less, or 1 minute or less) under conditions that provide for trans-cleavage of the detection DNA; b) maintaining the sample from step (a) for a period of time under conditions that do not provide for trans-cleavage of the detection DNA; and c) measuring a detectable signal generated by cleavage of the single-stranded detection DNA by the CasZ polypeptide after the period of step (b), thereby detecting the target DNA. Conditions that provide for trans-cleavage of the detection DNA include temperature conditions, such as 17°C to about 39°C (e.g., about 37°C). Conditions that do not provide for trans cleavage of the detected DNA include temperatures below 10°C, below 5°C, below 4°C, or below 0°C.
[0267] In some cases, the detectable signal generated by cleavage of the single-stranded detection DNA is generated for 60 minutes or less. For example, in some cases, the detectable signal generated by cleavage of the single-stranded detection DNA is generated for 60 minutes or less, 45 minutes or less, 30 minutes or less, 15 minutes or less, 10 minutes or less, or 5 minutes or less. For example, in some cases, the detectable signal generated by cleavage of the single-stranded detection DNA is generated over a period of 1 minute to 60 minutes, e.g., 1 minute to 5 minutes, 5 minutes to 10 minutes, 10 minutes to 15 minutes, 15 minutes to 30 minutes, 30 minutes to 45 minutes, or 45 minutes to 60 minutes. In some cases, after a detectable signal is generated (e.g., generated for 60 minutes or less), the generation of the detectable signal can be stopped by, for example, lowering the temperature of the sample (e.g., lowering the temperature to 10°C or less, 5°C or less, 4°C or less, or 0°C or less), adding an inhibitor to the sample, lyophilizing the sample, heating the sample to above 40°C, etc. The measuring step can be performed at any time after the generation of the detectable signal has stopped. For example, the measuring step can be performed 5 minutes to 48 hours after the generation of the detectable signal has stopped. For example, the measuring step can be performed 5 to 15 minutes, 15 to 30 minutes, 30 to 60 minutes, 1 to 4 hours, 4 to 8 hours, 8 to 12 hours, 12 to 24 hours, 24 to 36 hours, or 36 to 48 hours after the generation of the detectable signal has stopped. The measuring step can be performed 48 hours or more after the generation of the detectable signal has stopped.
[0268] The disclosed methods for detecting target DNA (single-stranded or double-stranded) in a sample can detect target DNA with high sensitivity. In some cases, the disclosed methods can be used to detect target DNA present in a sample containing multiple DNAs (including target DNA and multiple non-target DNAs), and the target DNA can be more than 10 7 One or more copies (e.g., 10) of non-target DNA 6 1 or more copies per 10 non-target DNA 51 or more copies per 10 non-target DNA 4 1 or more copies per 10 non-target DNA 3 1 or more copies per 10 non-target DNA 2 In some cases, the methods of the present disclosure can be used to detect target DNA present in a sample containing multiple DNAs (including target DNA and multiple non-target DNAs), and the target DNA is present at 1 or more copies per 10 non-target DNAs, 1 or more copies per 50 non-target DNAs, 1 or more copies per 20 non-target DNAs, 1 or more copies per 10 non-target DNAs, or 1 or more copies per 5 non-target DNAs. 18 One or more copies (e.g., 10) of non-target DNA 15 1 or more copies per 10 non-target DNA 12 1 or more copies per 10 non-target DNA 9 1 or more copies per 10 non-target DNA 6 1 or more copies per 10 non-target DNA 5 1 or more copies per 10 non-target DNA 4 1 or more copies per 10 non-target DNA 3 1 or more copies per 10 non-target DNA 2 Present at 1 or more copies per 10 non-target DNAs, 1 or more copies per 50 non-target DNAs, 1 or more copies per 20 non-target DNAs, 1 or more copies per 10 non-target DNAs, or 1 or more copies per 5 non-target DNAs).
[0269] In some cases, the methods of the present disclosure can be used to detect target DNA present in a sample, and the target DNA can be in the range of 10 7 From 1 replicate per 10 non-target DNA to 1 replicate per 10 non-target DNA (e.g., 10 7 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 71 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 6 1 copy per 10 non-target DNA 6 From 1 copy per 10 non-target DNA to 1 copy per 10 non-target DNA, 10 6 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 5 From 1 copy per 10 non-target DNA to 1 copy per 10 non-target DNA, 10 5 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 5 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 5 1 to 10 copies per non-target DNA 4 It is present at 1 copy per 100 copies of non-target DNA.
[0270] In some cases, the methods of the present disclosure can be used to detect target DNA present in a sample, and the target DNA can be in the range of 10 18 From 1 replicate per 10 non-target DNA to 1 replicate per 10 non-target DNA (e.g., 10 18 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 15 1 to 10 copies per non-target DNA2 1 copy per 10 non-target DNA 12 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 9 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 6 1 copy per 10 non-target DNA 6 From 1 copy per 10 non-target DNA to 1 copy per 10 non-target DNA, 10 6 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 5 From 1 copy per 10 non-target DNA to 1 copy per 10 non-target DNA, 10 5 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 5 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 51 to 10 copies per non-target DNA 4 It is present at 1 copy per 100 copies of non-target DNA.
[0271] In some cases, the methods of the present disclosure can be used to detect target DNA present in a sample, and the target DNA can be in the range of 10 7 From 1 replicate per 10 non-target DNA to 1 replicate per 100 non-target DNA (e.g., 10 7 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 7 1 to 10 copies per non-target DNA 6 1 copy per 10 non-target DNA 6 From 1 copy per 10 non-target DNA to 1 copy per 10 non-target DNA, 10 6 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 3 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 5 From 1 copy per 100 non-target DNA to 1 copy per 100 non-target DNA 5 1 to 10 copies per non-target DNA 2 1 copy per 10 non-target DNA 5 1 to 10 copies per non-target DNA3 1 copy per 10 non-target DNA 5 1 to 10 copies per non-target DNA 4 It is present at 1 copy per 100 copies of non-target DNA.
[0272] In some cases, in the subject methods for detecting target DNA in a sample, the detection threshold is 10 nM or less. Thus, for example, the target DNA may be present in the sample at a concentration of 10 nM or less. The term "detection threshold" is used herein to describe the minimum amount of target DNA that must be present in a sample for detection to occur. Thus, as an illustrative example, if the detection threshold is 10 nM, a signal can be detected when the target DNA is present in the sample at a concentration of 10 nM or more. In some cases, the disclosed methods have a detection threshold of 5 nM or less. In some cases, the disclosed methods have a detection threshold of 1 nM or less. In some cases, the disclosed methods have a detection threshold of 0.5 nM or less. In some cases, the disclosed methods have a detection threshold of 0.1 nM or less. In some cases, the disclosed methods have a detection threshold of 0.05 nM or less. In some cases, the disclosed methods have a detection threshold of 0.01 nM or less. In some cases, the disclosed methods have a detection threshold of 0.005 nM or less. In some cases, the disclosed methods have a detection threshold of 0.001 nM or less. In some cases, the methods of the present disclosure have a detection threshold of 0.0005 nM or less. In some cases, the methods of the present disclosure have a detection threshold of 0.0001 nM or less. In some cases, the methods of the present disclosure have a detection threshold of 0.00005 nM or less. In some cases, the methods of the present disclosure have a detection threshold of 0.00001 nM or less. In some cases, the methods of the present disclosure have a detection threshold of 10 pM or less. In some cases, the methods of the present disclosure have a detection threshold of 1 pM or less. In some cases, the methods of the present disclosure have a detection threshold of 500 fM or less. In some cases, the methods of the present disclosure have a detection threshold of 250 fM or less. In some cases, the methods of the present disclosure have a detection threshold of 100 fM or less. In some cases, the methods of the present disclosure have a detection threshold of 50 fM or less. In some cases, the methods of the present disclosure have a detection threshold of 500 aM (attomoles) or less. In some cases, the methods of the present disclosure have a detection threshold of 250 aM or less. In some cases, the methods of the present disclosure have a detection threshold of 100 aM or less.In some cases, the disclosed methods have a detection threshold of 50 aM or less. In some cases, the disclosed methods have a detection threshold of 10 aM or less. In some cases, the disclosed methods have a detection threshold of 1 aM or less.
[0273] In some cases, the detection threshold (for detecting target DNA in the subject methods) ranges from 500 fM to 1 nM (e.g., 500 fM to 500 pM, 500 fM to 200 pM, 500 fM to 100 pM, 500 fM to 10 pM, 500 fM to 1 pM, 800 fM to 1 nM, 800 fM to 500 pM, 800 fM to 200 pM, 800 fM to 100 pM, 800 fM to 1 pM, 1 pM to 1 nM, 1 pM to 500 pM, 1 pM to 200 pM, 1 pM to 100 pM, 1 pM to 10 pM) (concentration refers to the threshold concentration of target DNA at which target DNA can be detected). In some cases, the methods of the present disclosure have a detection threshold in the range of 800 fM to 100 pM. In some cases, the disclosed methods have a detection threshold in the range of 1 pM to 10 pM. In some cases, the disclosed methods have a detection threshold in the range of 10 fM to 500 fM, e.g., 10 fM to 50 fM, 50 fM to 100 fM, 100 fM to 250 fM, or 250 fM to 500 fM.
[0274] In some cases, the minimum concentration at which target DNA can be detected in a sample ranges from 500 fM to 1 nM (e.g., 500 fM to 500 pM, 500 fM to 200 pM, 500 fM to 100 pM, 500 fM to 10 pM, 500 fM to 1 pM, 800 fM to 1 nM, 800 fM to 500 pM, 800 fM to 200 pM, 800 fM to 100 pM, 800 fM to 1 pM, 1 pM to 1 nM, 1 pM to 500 pM, 1 pM to 200 pM, 1 pM to 100 pM, 1 pM to 10 pM). In some cases, the minimum concentration at which target DNA can be detected in a sample ranges from 800 fM to 100 pM. In some cases, the minimum concentration at which target DNA can be detected in a sample ranges from 1 pM to 10 pM.
[0275] In some cases, the detection threshold (for detecting target DNA in a subject method) is between 1 aM and 1 nM (e.g., between 1 aM and 500 pM, between 1 aM and 200 pM, between 1 aM and 100 pM, between 1 aM and 10 pM, between 1 aM and 1 pM, between 100 aM and 1 nM, between 100 aM and 500 pM, between 100 aM and 200 pM, between 100 aM and 1 ... aM~1pM, 250aM~1nM, 250aM~500pM, 250aM~200pM, 250aM~100pM, 250aM~10pM, 250aM~1pM, 500aM ~1nM, 500aM~500pM, 500aM~200pM, 500aM~100pM, 500aM~10pM, 500aM~1pM, 750aM~1nM, 750aM~50 0pM, 750aM~200pM, 750aM~100pM, 750aM~10pM, 750aM~1pM, 1fM~1nM, 1fM~500pM, 1fM~200pM, 1f M~100pM, 1fM~10pM, 1fM~1pM, 500fM~500pM, 500fM~200pM, 500fM~100pM, 500fM~10pM, 500fM~1p The detection threshold may range from 1 aM to 800 aM. In some cases, the method of the present disclosure has a detection threshold in the range of 50 aM to 1 pM. In some cases, the method of the present disclosure has a detection threshold in the range of 50 aM to 500 fM. In some cases, the method of the present disclosure has a detection threshold in the range of 50 aM to 500 fM.
[0276] In some cases, the target DNA may be at a concentration of 1 aM to 1 nM (e.g., 1 aM to 500 pM, 1 aM to 200 pM, 1 aM to 100 pM, 1 aM to 10 pM, 1 aM to 1 pM, 100 aM to 1 nM, 100 aM to 500 pM, 100 aM to 200 pM, 100 aM to 100 pM, 100 aM to 10 pM, 100 aM to 1 pM, 25 0aM~1nM, 250aM~500pM, 250aM~200pM, 250aM~100pM, 250aM~10pM, 250aM~1pM, 500aM~1n M, 500aM~500pM, 500aM~200pM, 500aM~100pM, 500aM~10pM, 500aM~1pM, 750aM~1nM, 750a M~500pM, 750aM~200pM, 750aM~100pM, 750aM~10pM, 750aM~1pM, 1fM~1nM, 1fM~500pM, 1f M~200pM, 1fM~100pM, 1fM~10pM, 1fM~1pM, 500fM~500pM, 500fM~200pM, 500fM~100pM, 50 In some cases, the target DNA is present in the sample in a range of 0 fM to 10 pM, 500 fM to 1 pM, 800 fM to 1 nM, 800 fM to 500 pM, 800 fM to 200 pM, 800 fM to 100 pM, 800 fM to 10 pM, 800 fM to 1 pM, 1 pM to 1 nM, 1 pM to 500 pM, 1 pM to 200 pM, 1 pM to 100 pM, 1 pM to 10 pM). In some cases, the target DNA is present in the sample in a range of 1 aM to 800 aM. In some cases, the target DNA is present in the sample in a range of 50 aM to 1 pM. In some cases, the target DNA is present in the sample in a range of 50 aM to 500 fM.
[0277] In some cases, the minimum concentration at which target DNA can be detected in a sample is between 1 aM and 1 nM (e.g., 1 aM to 500 pM, 1 aM to 200 pM, 1 aM to 100 pM, 1 aM to 10 pM, 1 aM to 1 pM, 100 aM to 1 nM, 100 aM to 500 pM, 100 aM to 200 pM, 100 aM to 100 pM, 100 aM to 10 pM, 100aM~1pM, 250aM~1nM, 250aM~500pM, 250aM~200pM, 250aM~100pM, 250aM~10pM, 250a M~1pM, 500aM~1nM, 500aM~500pM, 500aM~200pM, 500aM~100pM, 500aM~10pM, 500aM~1pM, 75 0aM~1nM, 750aM~500pM, 750aM~200pM, 750aM~100pM, 750aM~10pM, 750aM~1pM, 1fM~1nM, 1 fM~500pM, 1fM~200pM, 1fM~100pM, 1fM~10pM, 1fM~1pM, 500fM~500pM, 500fM~200pM, 500fM The minimum concentration at which target DNA can be detected in a sample ranges from 1 aM to 500 pM. In some cases, the minimum concentration at which target DNA can be detected in a sample ranges from 1 aM to 500 pM. In some cases, the minimum concentration at which target DNA can be detected in a sample ranges from 100 aM to 500 pM.
[0278] In some cases, the target DNA may be at a concentration of 1 aM to 1 nM (e.g., 1 aM to 500 pM, 1 aM to 200 pM, 1 aM to 100 pM, 1 aM to 10 pM, 1 aM to 1 pM, 100 aM to 1 nM, 100 aM to 500 pM, 100 aM to 200 pM, 100 aM to 100 pM, 100 aM to 10 pM, 100 aM to 1 pM, 25 0aM~1nM, 250aM~500pM, 250aM~200pM, 250aM~100pM, 250aM~10pM, 250aM~1pM, 500aM~1n M, 500aM~500pM, 500aM~200pM, 500aM~100pM, 500aM~10pM, 500aM~1pM, 750aM~1nM, 750a M~500pM, 750aM~200pM, 750aM~100pM, 750aM~10pM, 750aM~1pM, 1fM~1nM, 1fM~500pM, 1f M~200pM, 1fM~100pM, 1fM~10pM, 1fM~1pM, 500fM~500pM, 500fM~200pM, 500fM~100pM, 50 In some cases, the target DNA is present in the sample in a range of 0 fM to 10 pM, 500 fM to 1 pM, 800 fM to 1 nM, 800 fM to 500 pM, 800 fM to 200 pM, 800 fM to 100 pM, 800 fM to 10 pM, 800 fM to 1 pM, 1 pM to 1 nM, 1 pM to 500 pM, 1 pM to 200 pM, 1 pM to 100 pM, 1 pM to 10 pM). In some cases, the target DNA is present in the sample in a range of 1 aM to 500 pM. In some cases, the target DNA is present in the sample in a range of 100 aM to 500 pM.
[0279] In some cases, the subject compositions or methods exhibit attomolar (aM) detection sensitivity. In some cases, the subject compositions or methods exhibit femtomolar (fM) detection sensitivity. In some cases, the subject compositions or methods exhibit picomolar (pM) detection sensitivity. In some cases, the subject compositions or methods exhibit nanomolar (nM) detection sensitivity.
[0280] target DNA Target DNA can be single-stranded (ssDNA) or double-stranded (dsDNA).When target DNA is single-stranded, there is no preference or requirement for PAM sequence in target DNA.However, when target DNA is dsDNA, PAM is usually located adjacent to the target sequence of target DNA (see, for example, the discussion of PAM elsewhere in this specification).The source of target DNA can be the same as the source of sample, for example, as described below.
[0281] The source of the target DNA can be any source. In some cases, the target DNA is viral DNA (e.g., genomic DNA of a DNA virus). Thus, the subject method can be used to detect the presence of viral DNA in a population of nucleic acids (e.g., in a sample). The subject method can also be used to cleave non-target ssDNA in the presence of target DNA. For example, when the method is performed intracellularly, the subject method can be used to indiscriminately cleave non-target ssDNA (ssDNA that does not hybridize with the guide sequence of the guide RNA) in the cell if a specific target DNA is present in the cell (e.g., when the cell is infected with a virus and viral target DNA is detected).
[0282] Examples of potential target DNA include, for example, papovaviruses (e.g., human papillomavirus (HPV), polyomaviruses), hepadnaviruses (e.g., hepatitis B virus (HBV)), herpesviruses (e.g., herpes simplex virus (HSV), varicella-zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), herpesvirus lymphocytic, pityriasis rosea, Kaposi's sarcoma-associated herpesvirus), adenoviruses ( Examples of target DNA include, but are not limited to, viral DNA from viruses such as atadenovirus, aviadenovirus, ictadenovirus, mastadenovirus, and siadenovirus), poxviruses (e.g., smallpox, vaccinia virus, cowpox virus, monkeypox virus, orf virus, pseudocowpox, bovine papular stomatitis virus, variola virus, yabasa tumor virus, and molluscum contagiosum virus (MCV)), parvoviruses (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, and human parv4 G1), geminiviruses, nanoviruses, and phycodnaviruses. In some cases, the target DNA is parasitic DNA. In some cases, the target DNA is bacterial DNA, such as the DNA of a pathogenic bacterium.
[0283] sample A subject sample contains nucleic acids (e.g., multiple nucleic acids). The term "multiple" is used herein to mean two or more. Thus, in some cases, a sample contains two or more (e.g., three or more, five or more, ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, one thousand or more, or five thousand or more) nucleic acids (e.g., DNA). The subject method can be used as a highly sensitive method for detecting the presence of target DNA in a sample (e.g., in a complex mixture of nucleic acids such as DNA). In some cases, a sample contains five or more DNAs (e.g., ten or more, twenty or more, fifty or more, one hundred or more, five hundred or more, one thousand or more, or five thousand or more DNAs) that differ from each other in sequence. In some cases, a sample contains 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 103 More than 5×10 3 More than 10 pieces 4 More than 5×10 4 More than 10 pieces 5 More than 5×10 5 More than 10 pieces 6 More than 5×10 6 More than 10 7 In some cases, the sample contains 10-20, 20-50, 50-100, 100-500, 500-10 3 pieces, 10 3 ~5×10 3 pieces, 5×10 3 ~10 4 pieces, 10 4 ~5×10 4 pieces, 5×10 4 ~10 5 pieces, 10 5 ~5×10 5 pieces, 5×10 5 ~10 6 pieces, 10 6 ~5×10 6 pieces, or 5 x 10 6 ~10 7 pieces or 10 7 In some cases, a sample may contain 5-10 DNA fragments (e.g., fragments that differ from each other in sequence). 7 DNA (e.g., 5-10 6 pieces, 5~10 5 pieces, 5~50,000 pieces, 5~30,000 pieces, 10~10 6 pieces, 10~10 5 pieces, 10~50,000 pieces, 10~30,000 pieces, 20~10 6 pieces, 20~10 5 In some cases, the sample contains 20 or more DNAs that differ in sequence from one another. In some cases, the sample contains DNA from a cell lysate (e.g., a eukaryotic cell lysate, a mammalian cell lysate, a human cell lysate, a prokaryotic cell lysate, a plant cell lysate, etc.). For example, in some cases, the sample contains DNA from a cell, such as a eukaryotic cell, e.g., a mammalian cell, such as a human cell.
[0284] The term "sample" is used herein to mean any sample containing DNA (e.g., for determining whether a target DNA is present in a population of DNA). A sample can be derived from any source. For example, a sample can be a synthetic combination of purified DNA, a cell lysate, a DNA-enriched cell lysate, or DNA isolated and / or purified from a cell lysate. A sample can be derived from a patient (e.g., for diagnostic purposes). A sample can be derived from permeabilized cells. A sample can be derived from crosslinked cells. A sample can be in the form of a tissue section. A sample can be derived from tissue prepared by crosslinking, followed by delipidation and adjustment to a uniform refractive index. Examples of tissue preparations obtained by crosslinking, followed by delipidation and adjustment to a uniform refractive index are described, for example, in Shah et al., Development (2016) 143, 2862-2867 doi:10.1242 / dev.138560.
[0285] A "sample" can include target DNA and multiple non-target DNAs. In some cases, the target DNA is 1 copy per 10 non-target DNAs, 1 copy per 20 non-target DNAs, 1 copy per 25 non-target DNAs, 1 copy per 50 non-target DNAs, 1 copy per 100 non-target DNAs, 1 copy per 500 non-target DNAs, 1 copy per 10 3 1 copy per 5 x 10 non-target DNA 3 1 copy per 10 non-target DNA 4 1 copy per 5 x 10 non-target DNA 4 1 copy per 10 non-target DNA 5 1 copy per 5 x 10 non-target DNA 5 1 copy per 10 non-target DNA 6 1 copy per 10 non-target DNA, or 10 6In some cases, the target DNA is present in the sample at less than 1 copy per 10 non-target DNAs, 1 copy per 20 non-target DNAs, 1 copy per 20 non-target DNAs, 1 copy per 50 non-target DNAs, 1 copy per 50 non-target DNAs, 1 copy per 100 non-target DNAs, 1 copy per 100 non-target DNAs, 1 copy per 500 non-target DNAs, 1 copy per 500 non-target DNAs, 1 copy per 100 non-target DNAs, 3 1 copy per 10 non-target DNA 3 1 to 5 x 10 copies of non-target DNA 3 1 copy per 5 x 10 non-target DNA 3 1 to 10 copies per non-target DNA 4 1 copy per 10 non-target DNA 4 1 to 10 copies per non-target DNA 5 1 copy per 10 non-target DNA 5 1 to 10 copies per non-target DNA 6 1 copy per 10 non-target DNA 6 1 to 10 copies per non-target DNA 7 Non-target DNA is present in the sample at one copy per 1000 copies.
[0286] Suitable samples include, but are not limited to, saliva, blood, serum, plasma, urine, aspirates, and biopsy samples. Thus, the term "sample" with respect to a patient encompasses blood and other liquid samples of biological origin, solid tissue samples, e.g., biopsy specimens or tissue cultures or cells derived therefrom, and their progeny. The definition also includes samples that have been manipulated in any way after procurement, e.g., by treatment with reagents, washing, or enrichment for certain cell populations, e.g., cancer cells. The definition also includes samples enriched for specific types of molecules, e.g., DNA. The term "sample" encompasses biological samples, e.g., clinical samples such as blood, plasma, serum, aspirates, cerebrospinal fluid (CSF), etc., and also includes tissue obtained by surgical resection, tissue obtained by biopsy, cells in culture, cell supernatants, cell lysates, tissue samples, organs, bone marrow, etc. A "biological sample" also includes biological fluids derived therefrom (e.g., cancer cells, infected cells, etc.), e.g., samples containing DNA obtained from such cells (e.g., cell lysates or other cell extracts containing DNA).
[0287] Samples can include or be obtained from any variety of cells, tissues, organs, or cell-free fluids. Suitable sample sources include eukaryotic cells, bacterial cells, and archaeal cells. Suitable sample sources include unicellular and multicellular organisms. Suitable sample sources include unicellular eukaryotes; plants or plant cells; algae cells, such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Sargassum patens, C. agardh, etc.; fungal cells (e.g., yeast cells); animal cells, tissues, or organs; cells, tissues, or organs from invertebrates (e.g., Drosophila, Cnidaria, Echinoderms, Nematodes, Insecta, Arachnids, etc.); cells, tissues, fluids, or organs from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals); cells, tissues, fluids, or organs from mammals (e.g., humans, non-human primates, ungulates, cats, cows, sheep, goats, etc.). Suitable sample sources include nematodes, protozoans, etc. Suitable sample sources include, for example, helminths, parasites such as malaria parasites.
[0288] Suitable sample sources include cells, tissues, or organisms from any of six kingdoms, such as bacteria (e.g., eubacteria), archaea, protista, fungi, plantae, and animalia. Suitable sample sources include plant-like members of the kingdom protista, including but not limited to algae (e.g., green algae, red algae, glaucophytes, cyanobacteria); fungal members of the kingdom protista, such as slime molds, aquatic fungi, etc.; animal-like members of the kingdom protista, such as flagellates (e.g., Euglena), amoeboids (e.g., Amoeba), sporozoans (e.g., Apicomplexa, Myxozoa, Microsporidia), and ciliates (e.g., Paramecium). Suitable sample sources include members of the kingdom Fungi, including, but not limited to, members of any of the following phyla: Basidiomycota (members of the phylum Basidiomycota, e.g., Agaricus, Amanita, Boletus, Cantherellus, etc.), Ascomycota (e.g., the phylum Ascomycota, which includes Saccharomyces), Mycophycophyta (lichens), Zygomycota (conjugated fungi), and Deuteromycota. Suitable sample sources include members of the kingdom Plantae, including, but not limited to, members of any of the following classes: Bryophyta (e.g., mosses), Anthocerotophyta (e.g., hornworts), Hepaticophyta (e.g., liverworts), Lycophyta (e.g., club mosses), Sphenophyta (e.g., horsetails), Psilophyta (e.g., spiraea), Ophioglossophyta, Pterophyta (e.g., ferns), Cycadophyta, Gingkophyta, Pinophyta, Gnetophyta, and Magnoliophyta (e.g., flowering plants).Suitable sample sources include the following phyla: Porifera (Sponges), Placozoa, Orthonectida (parasites of marine invertebrates), Rhombozoa, Cnidaria (corals, sea anemones, jellyfish, sea pansies, and Chironex), Ctenophora (Ctenophores), Platyhelminthes (Flatworms), Nemertina (Nemertina), Ngathostomulida (Gnathostoma), Gastrotricha, Rotifera, Priapulida, Kinorhyncha, Loricifera, Acanthocephala, Entoprocta, Nemotoda, Nematomorpha, Cyclophora, Mollusca (Molluscs), Sipuncula (Stereostomes), Annelida (Cestodes), Tardigrada (Tardigrads), Onychophora (Onychophores), Arthropods, and the like. and members of the kingdom Animalia, including, but not limited to, any member of the following subphyla: Porionida (including the following subphyla: Chelicerata, Myriapoda, Hexapoda, Crustacea, where Chelicerata includes, for example, Arachnids, Merostomata, and Pycnogonida, Myriapoda includes, for example, Chilopoda (centipedes), Diplopoda (millipedes), Paropoda, and Symphyla, where Hexapoda includes insects, and Crustacea includes shrimp, krill, barnacles, etc.), Phoronida, Ectoprocta (External Proctida), Brachiopoda, Echinodermata (e.g., starfish, sea pansies, sea lilies, sea urchins, sea cucumbers, brittle stars, starfish, etc.), Chaetognatha (Chaetognatha), Hemichordata (Hemichordata), and Chordata.Suitable members of the phylum Chordata include any member of the following subphylum: Urochordata (sea squirts, including Ascidiacea, Thaliacea, and Larvacea); Cephalochordata (lampoxes); Myxini (hagfish); and the subphylum Vertebrata, including, for example, Petromyzontida (lampreys), Chondrichthyces (cartilaginous fish), Actinopterygii (bony fish), Actinista (coelacanths), Dipnoi (lungfish), Reptilia (reptiles, such as snakes, crocodiles, alligators, lizards, etc.), Aves (birds), and Mammalian (mammals). Suitable plants include any monocotyledonous plant and any dicotyledonous plant.
[0289] Suitable sample sources include cells, fluids, tissues, or organs collected from an organism, or specific cells or cell populations isolated from an organism. For example, when the organism is a plant, suitable sources include xylem, phloem, cambium, leaves, roots, etc. When the organism is an animal, suitable sources include specific tissues (e.g., lung, liver, heart, kidney, brain, spleen, skin, fetal tissue, etc.) or specific cell types (e.g., neuronal cells, epithelial cells, endothelial cells, astrocytes, macrophages, glial cells, pancreatic islet cells, T lymphocytes, B lymphocytes, etc.).
[0290] In some cases, the source of the sample is a diseased (or suspected of being diseased) cell, fluid, tissue, or organ. In some cases, the source of the sample is a normal (non-diseased) cell, fluid, tissue, or organ. In some cases, the source of the sample is a cell, tissue, or organ infected (or suspected of being infected) with a pathogen. For example, the source of the sample can be an infected or non-infected individual, and the sample can be any biological sample collected from the individual (e.g., blood, saliva, biopsy, plasma, serum, bronchoalveolar lavage fluid, sputum, stool sample, cerebrospinal fluid, fine needle aspirate, swab sample (e.g., buccal swab, cervical swab, nasal swab), interstitial fluid, synovial fluid, nasal secretion, tears, buffy coat, mucosal sample, epithelial cell sample (e.g., epithelial cell scraping), etc.). In some cases, the sample is a liquid sample that does not contain cells. In some cases, the sample is a liquid sample that may contain cells. Pathogens include viruses, fungi, helminths, protozoans, malaria parasites, Plasmodium parasites, Toxoplasma parasites, Schistosoma parasites, and the like. "Helminths" include roundworms, heartworms, and plant-eating nematodes (Nematoda), trematodes (Tematoda), Acanthocephala, and tapeworms (Cestoda). Protozoan infections include infections from Giardia species, Trichomonas species, African trypanosomiasis, amebic dysentery, babesiosis, balantidial dysentery, Chaga disease, coccidiosis, malaria, and toxoplasmosis. Examples of pathogens, such as parasitic / protozoan pathogens, include, but are not limited to, Plasmodium falciparum, Plasmodium vivax, Trypanosoma cruzi, and Toxoplasma gondii. Fungal pathogens include Cryptococcus neoformans, Histoplasma capsulatum, Coccidioides immitis, Blastomyces dermatitidis, Chlamydia trachomatis, and CandidaExamples of pathogenic viruses include, but are not limited to, immunodeficiency viruses (e.g., HIV), influenza virus, dengue virus, West Nile virus, herpes virus, yellow fever virus, hepatitis C virus, hepatitis A virus, hepatitis B virus, and papillomavirus. Examples of pathogenic viruses include papovaviruses (e.g., human papillomavirus (HPV) and polyomavirus), hepadnaviruses (e.g., hepatitis B virus (HBV)), herpesviruses (e.g., herpes simplex virus (HSV), varicella-zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), herpesvirus lymphocytic, pityriasis rosea, and Kaposi's sarcoma-associated herpesvirus), and adenoviruses (e.g., A. Tadennoviruses, Aviadenoviruses, Ictadenoviruses, Mastadenoviruses, and Siadenoviruses), poxviruses (e.g., smallpox, vaccinia virus, cowpox virus, monkeypox virus, orf virus, pseudocowpox, bovine papular stomatitis virus, variola virus, yaba monkey tumor virus, and molluscum contagiosum virus (MCV)), parvoviruses (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, and human parv4)Examples of pathogens include DNA viruses such as papovaviruses (e.g., human papillomavirus (HPV) and polyomaviruses), hepadnaviruses (e.g., hepatitis B virus (HBV)), herpesviruses (e.g., herpes simplex virus (HSV), varicella-zoster virus (VZV), Epstein-Barr virus (EBV), cytomegalovirus (CMV), lymphocytic herpesvirus, pityriasis rosea, and Kaposi's sarcoma-associated herpesvirus), and adenoviruses. (e.g., atadenovirus, aviadenovirus, ictadenovirus, mastadenovirus, siadenovirus), poxvirus (e.g., smallpox, vaccinia virus, cowpox virus, monkeypox virus, orf virus, pseudocowpox, bovine papular stomatitis virus, variola virus, yaba monkey tumor virus, molluscum contagiosum virus (MCV)), parvovirus (e.g., adeno-associated virus (AAV), parvovirus B19, human bocavirus, bufavirus, human parv4) G1), geminiviruses, nanoviruses, phycodnaviruses, etc.], Mycobacterium tuberculosis, Streptococcus agalactiae, methicillin-resistant Staphylococcus aureus, Legionella pneumophila, Streptococcus pyogenes, Escherichia coli, Neisseria gonorrhoeae, Neisseria meningitidis, Pneumococcus, Cryptococcus neoformans, Histoplasma capsulatum, Haemophilus influenzae type B, Treponema pallidum, Lyme disease spirochete, Pseudomonas aeruginosa, Mycobacterium leprae, Brucellaabortus, rabies virus, influenza virus, cytomegalovirus, herpes simplex virus type I, herpes simplex virus type II, human serum parvo-like virus, respiratory syncytial virus, varicella-zoster virus, hepatitis B virus, hepatitis C virus, measles virus, adenovirus, human T-cell leukemia virus, Epstein-Barr virus, murine leukemia virus, mumps virus, vesicular stomatitis virus, Sindbis virus, lymphocytic choriomeningitis virus, wart virus, bluetongue virus, Sendai virus, feline leukemia virus, reovirus, poliovirus, simian virus 40, mouse mammary tumor virus, dengue virus, rubella virus, West Nile virus, Plasmodium falciparum, Plasmodium vivax, Toxoplasma gondii, Trypanosoma rangeli, Trypanosoma cruzi, Trypanosoma rhodesiense, Trypanosoma brucei, Schistosoma mansoni, Schistosoma japonicum, Babesia bovis, Eimeria tenella, Onchocerca volvulus, Leishmania tropica, Mycobacterium tuberculosis, Trichinella spiralis, Theileria parva, Taenia hydatigena, Taenia ovis, Taenia saginata, Echinococcus granulosus, Mesocestoides corti, Mycoplasma arthritidis, M.hyorhinis, M.orale, M.arginini, Acholeplasma laidlawii, M.salivarium and M.pneumoniae.
[0291] Measuring detectable signals In some cases, the subject methods include a step of measuring (e.g., measuring a detectable signal generated by CasZ-mediated ssDNA cleavage). Because a CasZ polypeptide cleaves non-targeted ssDNA when activated (which occurs when a guide RNA hybridizes to target DNA in the presence of a CasZ polypeptide (and, in some cases, also a trancRNA)), the detectable signal can be any signal generated when ssDNA is cleaved. For example, in some cases, the measurement step may involve gold nanoparticle-based detection (see, e.g., Xu et al., Angew Chem Int Ed Engl. 2007;46(19):3468-70, and Xia et al., Proc Natl Acad Sci USA. 2010 Jun 15;107(24):10837-41), fluorescence polarization, colloidal phase transition / dispersion (see, e.g., Baksh et al., Nature. 2004 Jan 8;427(6970):139-41), electrochemical detection, semiconductor-based sensing (see, e.g., Rothberg et al., Nature. 2011 Jul 15;107(24):10837-41), or other techniques. 20;475(7356):348-52, for example, a phosphatase can be used to generate a pH change after the ssDNA cleavage reaction by opening the 2'-3' cyclic phosphate and releasing inorganic phosphate into the solution), and detection of labeled ssDNA (see elsewhere in this specification for details). The readout for such a detection method can be any convenient readout. Examples of possible readouts include, but are not limited to, a measured amount of a detectable fluorescent signal, visual analysis of bands on a gel (e.g., bands representing cleavage products relative to uncleaved substrate), visual or sensor-based detection of the presence or absence of color (i.e., color detection methods), and the presence or absence of an electrical signal (or a specific amount thereof).
[0292] In some cases, the measurement can be quantitative, for example, in the sense that the amount of signal detected can be used to determine the amount of target DNA present in a sample. In some cases, the measurement can be qualitative, for example, in the sense that the presence or absence of a detectable signal can indicate the presence or absence of target DNA (e.g., a virus, SNP, etc.). In some cases, a detectable signal will not be present (e.g., above a predetermined threshold level) unless the target DNA(s) (e.g., a virus, SNP, etc.) are present above a certain threshold concentration. In some cases, the detection threshold can be titrated by varying the amount of CasZ polypeptide, guide RNA, sample, and / or detection ssDNA (if used). Thus, for example, as will be understood by those skilled in the art, several controls can be used as needed to set up one or more reactions, each set to detect a different threshold level of target DNA, and such a series of reactions can then be used to determine the amount of target DNA present in a sample (e.g., such a series of reactions can be used to determine target DNA present in a sample "at a concentration of at least X"). Non-limiting examples of compositions and applications / uses of the disclosed compositions include detection of single nucleotide polymorphisms (SNPs), cancer screening, detection of bacterial infections, detection of antibiotic resistance, detection of viral infections, etc. The disclosed compositions and methods can be used to detect any DNA target. For example, since a subject's sample can contain cellular genomic DNA, any virus that integrates nucleic acid material into the genome can be detected, and a guide RNA can be designed to detect the integrated nucleotide sequence. In some cases, the disclosed methods do not include an amplification step. In some cases, the disclosed methods include an amplification step.
[0293] In some cases, the methods of the present disclosure can be used to determine the amount of target DNA in a sample (e.g., a sample containing target DNA and multiple non-target DNAs). Determining the amount of target DNA in a sample can include comparing the amount of detectable signal generated from a test sample to the amount of detectable signal generated from a reference sample. Determining the amount of target DNA in a sample includes measuring the detectable signal to generate a test measurement, measuring the detectable signal generated by the reference sample to generate a reference measurement, and comparing the test measurement to the reference measurement to determine the amount of target DNA present in the sample.
[0294] For example, in some cases, a method of the present disclosure for determining the amount of target DNA in a sample includes: a) contacting a sample (e.g., a sample containing target DNA and multiple non-target DNAs) with (i) a guide RNA that hybridizes to the target DNA, (ii) a CasZ polypeptide that cleaves DNA present in the sample, and (iii) a detection ssDNA; b) measuring a detectable signal generated by the CasZ polypeptide-mediated ssDNA cleavage (e.g., cleavage of the detection ssDNA) to generate a test measurement; c) measuring a detectable signal generated by a reference sample to generate a reference measurement; and d) comparing the test measurement to the reference measurement to determine the amount of target DNA present in the sample.
[0295] As another example, in some cases, a method of the present disclosure for determining the amount of target DNA in a sample includes: a) contacting a sample (e.g., a sample containing target DNA and multiple non-target DNAs) with (i) a guide RNA that hybridizes to the target DNA, (ii) a CasZ polypeptide that cleaves DNA present in the sample, (iii) a transcription RNA, and (iv) a detection ssDNA; b) measuring a detectable signal generated by the CasZ polypeptide-mediated ssDNA cleavage (e.g., cleavage of the detection ssDNA) to generate a test measurement; c) measuring a detectable signal generated by a reference sample to generate a reference measurement; and d) comparing the test measurement to the reference measurement to determine the amount of target DNA present in the sample.
[0296] Amplification of nucleic acids in a sample In some embodiments, the sensitivity of a subject composition and / or method (e.g., for detecting the presence of target DNA, such as viral DNA or SNPs, in cellular genomic DNA) can be increased by combining detection with nucleic acid amplification. In some cases, nucleic acids in a sample are amplified before contact with a CasZ polypeptide that cleaves ssDNA (e.g., amplification of nucleic acids in a sample can be initiated before contact with a CasZ polypeptide). In some cases, nucleic acids in a sample are amplified simultaneously with contact with a CasZ polypeptide. For example, in some cases, a subject method involves amplifying nucleic acids in a sample (e.g., by contacting the sample with amplification components) and then contacting the amplified sample with a CasZ polypeptide. In some cases, a subject method involves contacting a sample with amplification components simultaneously with contacting the sample with a CasZ polypeptide. If all components (amplification components and detection components, e.g., CasZ polypeptid...
Claims
1. 1. A method for directing a Class 2 CRISPR / Cas endonuclease to a target sequence of a target nucleic acid, comprising: The target nucleic acid (a) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3; (b) a guide RNA comprising a guide sequence that hybridizes to a target sequence of the target nucleic acid and comprising a region that binds to the class 2 CRISPR / Cas endonuclease; contacting a non-natural complex comprising Including, the contacting is performed outside a cell in vitro, inside a cell in vitro, or ex vivo; and The method is not performed in human germ cells or human embryonic cells. The method.
2. 10. The method of claim 1, which results in modification of the target nucleic acid, modulation of transcription from the target nucleic acid, or modification of a polypeptide associated with the target nucleic acid.
3. The method of claim 2 , wherein the target nucleic acid is modified by cleavage.
4. The method according to any one of claims 1 to 3, wherein the target nucleic acid is selected from double-stranded DNA, single-stranded DNA, RNA, genomic DNA, and extrachromosomal DNA.
5. 5. The method of any one of claims 1 to 4, wherein the guide sequence and the region that binds to the class 2 CRISPR / Cas endonuclease are heterologous to each other.
6. The method of any one of claims 1 to 5, wherein the contacting results in genome editing.
7. The contacting (a) the exterior of bacterial cells, and the exterior of archaeal cells; or (b) outside a cell in vitro; or (c) Inside the target cell The method according to any one of claims 1 to 5, wherein
8. The contacting (a) the class 2 CRISPR / Cas endonuclease, or a nucleic acid encoding the class 2 CRISPR / Cas endonuclease; and (b) the guide RNA or a nucleic acid encoding the guide RNA. introducing at least one of the following into said target cell: The method of claim 7, comprising:
9. 9. The method of claim 8, wherein the nucleic acid encoding the class 2 CRISPR / Cas endonuclease is a non-native sequence that is codon-optimized for expression in the target cell.
10. The method according to any one of claims 7 to 9, wherein the target cell is a eukaryotic cell.
11. The target cell is (a) cultured in vitro; or (b) ex vivo; The method according to any one of claims 7 to 10.
12. 11. The method of claim 10, wherein the eukaryotic cell is selected from the group consisting of a plant cell, a fungal cell, a unicellular eukaryote, a mammalian cell, a reptile cell, an insect cell, an avian cell, a fish cell, a parasitic cell, an arthropod cell, an invertebrate cell, a vertebrate cell, a rodent cell, a mouse cell, a rat cell, a primate cell, a non-human primate cell, and a human cell.
13. The method of any one of claims 7 to 12, wherein said contacting further comprises introducing a DNA donor template into said target cell.
14. The method of any one of claims 1 to 13, comprising contacting the target nucleic acid with a trans-activating non-coding RNA (trancRNA).
15. (a) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3, or a nucleic acid encoding said class 2 CRISPR / Cas endonuclease; (b) a guide RNA, or a nucleic acid encoding the guide RNA, comprising a guide sequence complementary to a target sequence of a target nucleic acid and comprising a region capable of binding to the class 2 CRISPR / Cas endonuclease; A composition comprising a non-natural complex comprising: The class 2 CRISPR / Cas endonuclease interacts with the guide RNA to form a ribonucleoprotein complex that is targeted to the target sequence via base pairing between the guide RNA and the target sequence. The composition.
16. The composition of claim 15, further comprising a trans-activating non-coding RNA (trancRNA) or a nucleic acid encoding said transRNA.
17. (a) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3 or a nucleic acid encoding said class 2 CRISPR / Cas endonuclease; (b) a guide RNA, or a nucleic acid encoding the guide RNA, comprising a guide sequence complementary to a target sequence of a target nucleic acid and comprising a region capable of binding to the class 2 CRISPR / Cas endonuclease; A kit comprising a non-natural complex comprising: The class 2 CRISPR / Cas endonuclease interacts with the guide RNA to form a ribonucleoprotein complex that is targeted to the target sequence via base pairing between the guide RNA and the target sequence. The kit.
18. 18. The kit of claim 17, further comprising a trans-activating non-coding RNA (trancRNA) or a nucleic acid encoding said transRNA.
19. (a) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3, or a nucleic acid encoding said class 2 CRISPR / Cas endonuclease; and (b) a guide RNA, or a nucleic acid encoding said guide RNA, comprising a guide sequence complementary to a target sequence of a target nucleic acid and comprising a region capable of binding to said class 2 CRISPR / Cas endonuclease; 1. A genetically modified eukaryotic cell comprising: the eukaryotic cell is other than a human germ cell or a human embryonic cell, The class 2 CRISPR / Cas endonuclease interacts with the guide RNA to form a ribonucleoprotein complex that is targeted to the target sequence via base pairing between the guide RNA and the target sequence. , the eukaryotic cell.
20. 20. The eukaryotic cell of claim 19, further comprising a trans-activating non-coding RNA (trancRNA) or a nucleic acid encoding said transRNA.
21. 17. A composition according to claim 15 or 16 for use in a method of therapeutic treatment of a patient.
22. 17. The composition of claim 15 or 16, wherein the guide RNA and / or the trancRNA comprises one or more of modified nucleobases, modified backbones or non-natural internucleoside linkages, modified sugar moieties, locked nucleic acids, peptide nucleic acids, and deoxyribonucleotides.
23. 17. The composition of claim 15 or 16, wherein at least one of the class 2 CRISPR / Cas endonuclease, the nucleic acid encoding the class 2 CRISPR / Cas endonuclease, the guide RNA, the nucleic acid encoding the guide RNA, the truncRNA, and the nucleic acid encoding the truncRNA is conjugated to a heterologous moiety.
24. 1. A method for detecting target DNA in a sample, comprising: (a) subjecting the sample to (i) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3; (ii) a guide RNA comprising a region that binds to the class 2 CRISPR / Cas endonuclease and a guide sequence that hybridizes to the target DNA; and (iii) a detection DNA that is single-stranded and does not hybridize to the guide sequence of the guide RNA; and contacting the (b) measuring a detectable signal generated by cleavage of the detection DNA by the class 2 CRISPR / Cas endonuclease, thereby detecting the target DNA; The method comprising:
25. 25. The method of claim 24, wherein the target DNA is single-stranded or double-stranded.
26. 26. The method of any one of claims 24 or 25, wherein measuring the detectable signal comprises one or more of gold nanoparticle-based detection, fluorescence polarization, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based sensing.
27. The method according to any one of claims 24 to 26, wherein the detection DNA comprises a fluorescent dye pair.
28. 28. The method of claim 27, wherein the fluorescent dye pair is a quencher / fluorophore pair.
29. The method of any one of claims 24 to 28, comprising amplifying nucleic acid in the sample.
30. 1. A kit for detecting a target DNA in a sample, comprising: (a) a guide RNA, or a nucleic acid encoding said guide RNA, comprising a region that binds to a class 2 CRISPR / Cas endonuclease and a guide sequence that is complementary to a target DNA; (b) a labeled detection DNA that is single-stranded and does not hybridize to the guide sequence of the guide RNA; and (c) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3, wherein the class 2 CRISPR / Cas endonuclease interacts with the guide RNA and forms a ribonucleoprotein complex that is targeted to the target sequence via base pairing between the guide RNA and the target sequence. The kit comprising:
31. 31. The kit of claim 30, wherein the detection DNA comprises a fluorescent dye pair.
32. 32. The kit of claim 31 , wherein the fluorescent dye pair is a quencher / fluorophore pair.
33. The kit of any one of claims 30 to 32, further comprising a nucleic acid amplification component.
34. 1. A method for cleaving single-stranded DNA (ssDNA), comprising: A population of nucleic acids comprising a target DNA and a plurality of non-target ssDNAs is subjected to (i) a class 2 CRISPR / Cas endonuclease comprising the amino acid sequence of SEQ ID NO: 3; and (ii) a guide RNA comprising a region that binds to the class 2 CRISPR / Cas endonuclease and a guide sequence that hybridizes to the target DNA; contacting the the class 2 CRISPR / Cas endonuclease cleaves the plurality of non-target ssDNA; the contacting is performed outside a cell in vitro, inside a cell in vitro, or ex vivo; and The method is not performed in human germ cells or human embryonic cells. The method.
35. 35. The method of claim 34, wherein the contacting is performed inside a cell in vitro or ex vivo.
36. 36. The method of claim 34 or 35, wherein the target DNA is single-stranded or double-stranded.
37. 19. A kit according to claim 17 or 18 for use in a method of therapeutic treatment of a patient.
38. 19. The kit of claim 17 or 18, wherein the guide RNA and / or the trancRNA comprises one or more of modified nucleobases, modified backbones or non-natural internucleoside linkages, modified sugar moieties, locked nucleic acids, peptide nucleic acids, and deoxyribonucleotides.
39. 19. The kit of claim 17 or 18, wherein at least one of the class 2 CRISPR / Cas endonuclease, the nucleic acid encoding the class 2 CRISPR / Cas endonuclease, the guide RNA, the nucleic acid encoding the guide RNA, the transRNA, and the nucleic acid encoding the transRNA is conjugated to a heterologous moiety.
40. 21. A eukaryotic cell according to claim 19 or 20 for use in a method of therapeutic treatment of a patient.
41. 21. The eukaryotic cell of claim 19 or 20, wherein the guide RNA and / or the trancRNA comprises one or more of modified nucleobases, modified backbones or non-natural internucleoside linkages, modified sugar moieties, locked nucleic acids, peptide nucleic acids, and deoxyribonucleotides.
42. 21. The eukaryotic cell of claim 19 or 20, wherein at least one of the class 2 CRISPR / Cas endonuclease, the nucleic acid encoding the class 2 CRISPR / Cas endonuclease, the guide RNA, the nucleic acid encoding the guide RNA, the transRNA, and the nucleic acid encoding the transRNA is conjugated to a heterologous moiety.
43. 15. The method of any one of claims 1 to 14, wherein the guide RNA and / or the trancRNA comprises one or more of the following: modified nucleobases, modified backbones or non-natural internucleoside linkages, modified sugar moieties, locked nucleic acids, peptide nucleic acids, and deoxyribonucleotides.
44. 15. The method of any one of claims 1 to 14, wherein at least one of the class 2 CRISPR / Cas endonuclease, the nucleic acid encoding the class 2 CRISPR / Cas endonuclease, the guide RNA, the nucleic acid encoding the guide RNA, the truncRNA, and the nucleic acid encoding the truncRNA is conjugated to a heterologous moiety.
Citation Information
Cited By
Casz compositions and methods of use
EP3704239A1