Novel small V-type RNA programmable endonuclease system

The novel CRISPR Cas nuclease, B-GEn.16, addresses size and off-target challenges by offering precise gene editing in eukaryotic cells, particularly mammalian cells, through a small size and specific PAM sequence, enhancing viral vector integration and editing efficiency.

JP2025523400APending Publication Date: 2025-07-23BAYER AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024572058
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-10
Filing Date
2023-06-07
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems face challenges such as large size, limited activity in eukaryotic cells, off-target effects, immune responses, complex PAM requirements, and size constraints for viral delivery, making them unsuitable for precise gene editing in mammalian cells.

Method used

Development of a novel type V CRISPR Cas nuclease, B-GEn.16, with a small size and specific PAM sequence, enabling efficient targeting and editing in eukaryotic cells, particularly mammalian cells, using a guide RNA and tracr RNA system.

Benefits of technology

B-GEn.16 provides high activity and precision in eukaryotic cells, overcoming size and off-target issues, facilitating effective gene editing and manipulation within viral vectors like AAV.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523400000010
    Figure 2025523400000010
  • Figure 2025523400000011
    Figure 2025523400000011
  • Figure 2025523400000012
    Figure 2025523400000012
Patent Text Reader

Abstract

Novel systems for targeting, editing or manipulating DNA in a cellular or cell-free environment using novel V-type B-Gen.16 (SEQ ID NO: 1) and variants thereof, as well as methods and kits for manipulating DNA, are described herein.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to the field of molecular biology and, in particular, to a novel CRISPR Cas RNA programmable DNA endonuclease, designated B-GEn.16, for gene editing and other uses.

Background Art

[0002] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and CRISPR associated (Cas) genes, collectively known as the CRISPR-Cas or CRISPR / Cas system, are currently understood to provide bacteria and archaea with immunity against phage infection. The CRISPR-Cas systems of prokaryotic adaptive immunity are a highly diverse group of protein effector non-coding elements, like the locus structure, and some examples have been engineered and adapted to produce important biotechnology.

[0003] Components of the systems involved in host defense include one or more effector proteins capable of modifying DNA or RNA, and an RNA guide element that serves to target these protein activities to specific sequences on phage DNA or RNA. The RNA guide consists of CRISPR RNA (crRNA) and may require additional trans-acting RNA (tracrRNA) to enable targeted nucleic acid manipulation by the effector protein. The crRNA consists of a segment called the "direct repeat" that is responsible for binding the crRNA to the effector protein and a segment called the "spacer sequence" that is complementary to the desired nucleic acid target sequence. The CRISPR system can be reprogrammed to target another DNA or RNA target by modifying the spacer sequence of the crRNA.

[0004] The CRISPR-Cas system can be broadly classified into two classes: Class 1 systems consist of multiple effector proteins that come together to form a complex around the crRNA, while Class 2 systems consist of a single effector protein that forms a complex with the crRNA guide to target DNA or RNA substrates. The monomeric subunit effector compositions of Class 2 systems provide a simpler set of components for engineering and application and have been an important source of programmable effectors to date. Thus, the discovery, engineering, and optimization of novel Class 2 systems could lead to extensive and powerful programmable technologies for genome engineering and beyond. Genome editing using the RNA-guided DNA targeting principle of CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas (CRISPR-associated protein) has been widely utilized over the past few years. Five types of CRISPR-Cas systems (Type I, II and IIb, III, V, and VI) have been described. Most uses of CRISPR-Cas for genome editing have been with Type II systems. The main advantage provided by bacterial Type II CRISPR-Cas systems is the minimal requirement for programmable DNA interference: the endonuclease, Cas9, induced by a customizable dual RNA structure. As first demonstrated in the original Type II system of Streptococcus pyogenes, the trans-activating CRISPR RNA (tracrRNA) binds to the invariant repeats of the precursor CRISPR RNA (pre-crRNA) and forms a dual RNA that is essential for both RNase III-mediated co-maturation of crRNAs in the presence of Cas9 and cleavage of invading DNA by Cas9. As shown in Streptococcus pyogenes, Cas9 is induced by the double-strand formed between the mature activating tracrRNA and the targeting crRNA and introduces site-specific double-stranded DNA (dsDNA) breaks in invading homologous DNA.Cas9 is a multi-domain enzyme that uses the HNH nuclease domain to cleave the target strand (defined as complementary to the spacer sequence of the crRNA) and the RuvC-like domain to cleave the non-target strand.

[0005] In addition to type II CRISPR Cas9 nuclease, numerous different type V CRISPR Cas nucleases such as Cas12a, Cas12b, Cas12e, Cas12f, Cas13a, Cas13b have been described (Koonin et al., Curr Opin Microbiol. June 2017; 37:67-78, and Makarova et al., Nat Rev Microbiol. February 2020; 18(2):67-83). Some of these systems do not require tracrRNA (Cas12a, Cas13a, Cas13b), while Cas12b typically requires tracrRNA (Koonin et al., Curr Opin Microbiol. June 2017; 37:67-78).

[0006] Genome editing in mammalian cells is, in part, limited by the size of the various Cas9 proteins. Cas9 derived from Staphylococcus pyogenes (SpyCas9), the most widely used enzyme to date, contains approximately 4.2 kb of DNA (WO2013 / 176722), and its direct combination with a cognate single-guide RNA (sgRNA) further increases the size. Adeno-associated virus is among the vectors used for the delivery of Cas9 enzymes in gene therapy applications. However, the AAV cargo size is limited to approximately 4.5 kb. Due to size constraints, the delivery of Cas9 and its sgRNA and potential DNA repair template may prevent the use of this method. Smaller Cas9 molecules have been characterized, but most of them are plagued by protospacer adjacent motif (PAM) sequences that are not defined in the same way as those used by SpyCas9. For example, Staphylococcus aureus (SauCas9) uses the sequence "NNGRR(T)" where R = A or G, and Campylobacter jejuni (Cja) Cas9 uses the "NNNACAC" / "NNNRYAC" PAMs where Y = T or G, respectively. The ambiguity of the PAM increases the likelihood of unwanted activity of the enzyme at off-target sequences that have high or perfect sequence identity to the PAM. Since the mis-targeting of similar sites ("off-target") increases the likelihood of harmful events, the specificity of these systems remains a concern.

[0007] Existing CRISPR-Cas systems generally have one or more of the following drawbacks: a) Their size is too large to be carried within the genome of established therapeutically suitable viral delivery systems such as adeno-associated virus (AAV). b) Many of them are not substantially active in non-host environments, such as eukaryotic cells, particularly mammalian cells. c) Those nucleases can catalyze DNA strand cleavage when there are mismatches between the spacer sequence and the protospacer sequence, resulting in undesirable off-target effects, for example, making them unsuitable for gene therapy applications or other applications that require high precision. d) They can induce an immune response that may limit their use for in vivo applications in mammals. e) They require complex and / or long PAMs that limit the target selection of DNA-targeted segments. f) They have poor expression from plasmids or viral vectors. g) They require additional RNA sequences, either as part of an additional RNA sequence or as part of the guide RNA, to be active.

[0008] The following references: Nucleic Acids Research, Vol. 48, Issue 9, May 21, 2020, pp5016 - 5023, International Publication No. 2020123887, International Publication No. 2017117395 disclose very distant relatives to SEQ ID NO: 1, having an amino acid identity of about 52% or less compared to SEQ ID NO: 1.

Prior Art Documents

Patent Documents

[0009]

Patent Document 1

Patent Document 2

Patent Document 3

Non-Patent Documents

[0010]

Non-Patent Document 1

Non-Patent Document 2

[0011] The present invention relates to a novel type V CRISPR Cas nuclease called B-GEn.16 (SEQ ID NO: 1). Furthermore, the present invention relates to a polypeptide having at least 60%, preferably at least 70%, more preferably at least 80%, even more preferably at least 90%, particularly preferably at least 95%, most preferably at least 99% amino acid identity to the sequence according to SEQ ID NO: 1 over its entire length, or any nucleic acid encoding the same. Unless otherwise specified, the name B-GEn.16 includes the polypeptide of SEQ ID NO: 1 and any polypeptide defined in this paragraph.

[0012] The present invention further relates to a CRISPR Cas system comprising a suitable guide RNA or tracr RNA, and a target DNA, which includes the sequences contained in B-GEn.16, SEQ ID NOs: 8, 9, and 10.

[0013] One of the main features of B-GE.16 is its particularly small size, making it an ideal candidate for integration into viral vectors such as AAV.

[0014] Another advantageous feature of B-Gen.16 is its preferred PAM sequence ("CCN").

[0015] Furthermore, as supported by the examples herein, B-Gen.16 exhibits high activity in eukaryotic cells, particularly mammalian cells.

[0016] In one aspect, provided herein is a method for targeting, editing, modifying, or manipulating target DNA at one or more positions intracellularly or in vitro, the method comprising: (I) introducing into the intracellular or in vitro environment a heterologous B-GEn.16 polypeptide disclosed herein or a nucleic acid encoding B-GEn.16 disclosed herein; and (II) introducing into the cell or in vitro environment one or more heterologous single guide RNA(s) (sgRNA) or DNA encoding such one or more sgRNAs, wherein each sgRNA or DNA encoding an sgRNA comprises: (a) an engineered DNA targeting segment that includes RNA and is capable of hybridizing to a target sequence in a polynucleotide locus, (b) a tracr mate sequence composed of RNA, and (c) a tracr RNA sequence composed of RNA, wherein the tracr mate sequence hybridizes to the tracr sequence and (a), (b), and (c) are arranged in a 5' to 3' direction; and (III) creating one or more nicks or cuts or base edits in the target DNA, wherein the B-GEn.16 polypeptide is directed to the target DNA by the processed or unprocessed form of the sgRNA.

[0017] In one aspect, provided herein is the use of a composition for targeting, editing, modifying, or manipulating target DNA at one or more positions intracellularly or in vitro, the composition comprising: (I) the B-GEn.16 polypeptide disclosed herein or a nucleic acid encoding the same; and / or (II) one or more single guide RNAs (sgRNAs) or DNA(s) suitable for in situ generation of such one or more sgRNAs, each comprising: (a) an engineered DNA targeting segment composed of RNA and capable of hybridizing to such a target sequence in a polynucleotide locus; (b) a tracr mate sequence composed of RNA; and (c) a tracr RNA sequence composed of RNA, wherein the tracr mate sequence is capable of hybridizing to the tracr sequence, and (c), (b), and (a) are arranged in the 5' to 3' direction, respectively.

[0018] In another aspect, provided herein is a cell comprising: (I) a B-GEn.16 polypeptide as disclosed herein or a nucleic acid encoding the B-GEn.16 polypeptide as disclosed herein; and (II) one or more single guide RNAs (sgRNAs) or DNA(s) suitable for in situ generation of such one or more sgRNAs, each comprising: (a) an engineered DNA targeting segment capable of hybridizing to a target sequence in a polynucleotide locus; (b) a tracr mate sequence; and (c) a tracr RNA sequence, wherein the tracr mate sequence is capable of hybridizing to the tracr sequence, and (c), (b), and (a) are arranged in the 5' to 3' direction, respectively.

[0019] In yet another aspect, provided herein is a kit comprising: (I) a nucleic acid sequence encoding a B-GEn.16 polypeptide as disclosed herein, wherein the nucleic acid sequence encoding B-GEn.16 is operably linked to a promoter; and (II) one or more single guide RNAs (sgRNAs) or DNA(s) suitable for in situ generation of such one or more sgRNAs, each sgRNA comprising: (a) an engineered DNA targeting segment capable of hybridizing to a target sequence in a polynucleotide locus, (b) a tracr mate sequence, and (c) a tracr RNA sequence, wherein the tracr mate sequence is capable of hybridizing to the tracr sequence, and (a), (b), and (c) are arranged in a 5' to 3' direction.

[0020] The entire disclosures of each patent document and scientific paper referred to herein, as well as those patent documents and scientific papers cited thereby, are hereby expressly incorporated by reference herein for all purposes.

[0021] Additional features and advantages of the invention will be described in more detail below.

Brief Description of the Drawings

[0022]

Fig. 1a

Fig. 1b

Fig. 2

Table 1

Mode for Carrying Out the Invention

[0023] Reference to Sequence Listing The 27 sequences of SEQ ID NOs: 1 to 27 disclosed in this specification are included in the sequence listing named BHC221018-WO_sequence_LIsting.xml (ST.26. SEQ ID NO: 1 (B-GEn.16 protein)) and printed in this specification.

[0024] Summary of the Sequences Described in the Sequence Listing

Table 2

[0025] Amino Acid Sequence of SEQ ID NO: 1 (B-GEn.16) TIFF2025523400000005.tif55166

[0026] Detailed Description of the Invention This application provides a novel CRISPR-Cas nuclease and a gene editing system based on such a nuclease. The novel nuclease is referred to herein as the B-GEn nuclease or the B-GEn.16 nuclease.

[0027] Very preferably, the group of B-GEn nucleases includes the following members, which are listed in Table 1: Table 1:

Table 3

[0028] One embodiment according to the present invention is a polypeptide that is at least 60% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0029] One preferred embodiment according to the present invention is a polypeptide that is at least 70% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0030] One more preferred embodiment according to the present invention is a polypeptide that is at least 80% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0031] One even more preferred embodiment according to the present invention is a polypeptide that is at least 90% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0032] One particularly more preferred embodiment according to the present invention is a polypeptide that is at least 95% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0033] One particularly preferred embodiment according to the present invention is a polypeptide that is at least 99% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0034] One more particularly preferred embodiment according to the present invention is a polypeptide that is at least 99.5% identical at the amino acid level compared to SEQ ID NO: 1 or a nucleic acid encoding the same.

[0035] Another embodiment according to the present invention is a variant of the following B-GEn.16: (I) Variants having at least 60%, for example, at least 70%, at least 80%, at least 90%, at least 95%, at least 99% amino acid identity increasing in preferred order over the entire length to the sequences according to any of SEQ ID NO: 1. (II) a variant according to (I) that contains an additional component, e.g., a nuclear localization signal, for obtaining proper activity of the B-GEn.16 CRISPR system not only in cell-free reactions or prokaryotic cells, but also in eukaryotic cellular environments, including living organisms such as plants or animals; (III) Codon-optimized variants of B-GEn.16 and corresponding polynucleotide sequences encoding variants according to (I) and (II).

[0036] Unless otherwise stated, the term B-GEn.16 includes all variants specified under (I), (II), and (III).

[0037] B-GEn.16-based CRISPR-Cas system One embodiment according to the present invention is a composition comprising: (a) a B-GEn.16 polypeptide, or a polynucleotide encoding such a B-GEn.16; (b) a single heterologous guide RNA (sgRNA) or DNA enabling the generation of such a sgRNA in situ, comprising: i. an engineered DNA targeting segment that is composed of RNA and capable of hybridizing to a target sequence in a polynucleotide locus; ii. a tracr mate sequence composed of RNA; and iii. comprises a tracr RNA sequence composed of RNA; wherein the tracr mate sequence hybridizes to the tracr sequence, and (i), (ii) and (iii) are arranged in a 5' to 3' orientation. The composition comprises: Within the sgRNA, the tracr mate sequence and the tracr sequence are generally linked by a suitable loop sequence to form a stem-loop structure.

[0038] Suitable PAM sequences for use in CRISPR-Cas systems, including B-GEn.16 The function of B-GEn.16 typically requires a suitable protospacer adjacent motif (「PAM」) sequence on the 5' side of the target sequence. Suitable PAM sequences are listed in Table 2, where the engineered DNA targeting segment is directly adjacent at its 3' end to the PAM sequence on the target DNA segment or such a PAM sequence is part of the target DNA sequence in its 5' portion.

[0039] Table 2: PAM sequences suitable for the corresponding B-GEn.16 endonuclease

Table 4

[0040] tracr sequences suitable for the use of B-GEn.16 in the CRISPR Cas system The tracr sequences suitable for the use of B-GEn.16 in the CRISPR Cas system are provided in SEQ ID NO: 10. Alternatively, variants of this sequence can be used. Variants can include either a portion or a truncated form of such a sequence and / or a sequence having base modifications at one or more locations of this sequence.

[0041] In some embodiments, the polynucleotide encoding B-GEn.16 and the sgRNA contain a suitable promoter and / or a suitable nuclear localization signal for expression in an intracellular or in vitro environment.

[0042] Another embodiment according to the present invention is a method of targeting, editing, modifying, or manipulating target DNA at one or more positions intracellularly or in vitro, comprising the following: (a) introducing into the intracellular or in vitro environment a heterologous B-GEn.16 polypeptide or a nucleic acid encoding the same protein; and (b) introducing a single heterologous guide RNA (sgRNA) or DNA suitable for the in situ generation of such an sgRNA: An engineered DNA targeting segment composed of i.RNA that can hybridize with a target sequence in a polynucleotide locus ii. A tracr mate sequence composed of RNA, and iii. A tracr RNA sequence composed of RNA, wherein the tracr mate sequence can hybridize with the tracr sequence, and (i), (ii) and (iii) are arranged in the 5' to 3' direction; (c) Creating one or more cuts, nicks or edits in the target DNA, wherein the B-GEn.16 polypeptide is directed to the target DNA by the processed or unprocessed form of the gRNA A method comprising the steps of.

[0043] Another embodiment according to the present invention is the use of a composition for targeting, editing, modifying or manipulating target DNA at one or more positions in a cell or in vitro (a) A B-GEn.16 polypeptide, or a polynucleotide encoding such a B-GEn.16 (b) A single guide RNA (sgRNA) or DNA suitable for the generation of such an sgRNA in situ, comprising: i. An engineered DNA targeting segment composed of RNA that can hybridize with such a target sequence in a polynucleotide locus ii. A tracr mate sequence composed of RNA, and iii. A tracr RNA sequence composed of RNA, wherein the tracr mate sequence hybridizes with the tracr sequence, and (i), (ii) and (iii) are arranged in the 5' to 3' direction The use of a composition comprising.

[0044] Another embodiment according to the present invention is a cell ex vivo or in vitro (a) Heterologous B-GEn.16 polypeptide or nucleic acid encoding the same, (b) A single heterologous guide RNA (sgRNA) or DNA suitable for the in situ generation of such sgRNA, comprising: i. An engineered DNA targeting segment composed of RNA and capable of hybridizing with such a target sequence in a polynucleotide locus, ii. A tracr mate sequence composed of RNA, and iii. A tracr RNA sequence composed of RNA, wherein the tracr mate sequence hybridizes with the tracr sequence, and (i), (ii), and (iii) are arranged in the 5' to 3' direction A cell comprising the same, or such a cell whose genome has been targeted, edited, modified, or engineered using (a) and (b) above.

[0045] A further embodiment according to the present invention is a kit comprising: (a) A nucleic acid sequence encoding B-GEn.16, wherein the nucleic acid sequence encoding B-GEn.16 is operably linked to a promoter or ribosome binding site; (b) A single heterologous guide RNA (sgRNA) or DNA suitable for the in situ generation of such sgRNA, comprising: i. An engineered DNA targeting segment composed of RNA and capable of hybridizing with such a target sequence in a polynucleotide locus, ii. A tracr mate sequence composed of RNA, and iii. A tracr RNA sequence composed of RNA, wherein the tracr mate sequence hybridizes with the tracr sequence, and (i), (ii), and (iii) are arranged in the 5' to 3' direction Or a kit comprising the same, Or (a) B-GEn.16 protein; (b) One or more single heterologous guide RNAs (sgRNAs), each of which: iv. an engineered DNA targeting segment composed of RNA and capable of hybridizing to such a target sequence in a polynucleotide locus v. a tracr mate sequence composed of RNA, and vi. a tracr RNA sequence composed of RNA, wherein the tracr mate sequence hybridizes to the tracr sequence and (i), (ii) and (iii) are arranged in the 5' to 3' direction A kit comprising the same.

[0046] Yet another embodiment according to the present invention is a composition and method for targeting, editing, modifying or manipulating one or more target DNAs at one or more positions in a cell or in vitro, comprising: (a) B-GEn.16 (b) a guide RNA (gRNA) or DNA suitable for the generation of such gRNA in situ, comprising: i. an engineered DNA targeting segment composed of RNA and capable of hybridizing to such a target sequence in a polynucleotide locus ii. a tracr RNA sequence composed of RNA, wherein (i) and (ii) are in one single RNA molecule and (iii) is on another RNA molecule Comprising compositions and methods comprising the same.

[0047] Multiplexing In another aspect, provided herein is a method of editing or modifying DNA at multiple locations within a cell, comprising: i) introducing into the cell a B-GEn.16 polypeptide or a nucleic acid encoding the B-GEn.16 polypeptide; and ii) introducing into the cell a single heterologous nucleic acid comprising two or more pre-CRISPR RNAs (pre-crRNAs) as RNA or encoded as DNA and under the control of one promoter, wherein each pre-crRNA comprises a repeat-spacer array or a repeat-spacer, wherein the spacer comprises a nucleic acid sequence complementary to a target sequence in the DNA, the repeat comprises a stem-loop structure, the B-GEn.16 polypeptide cleaves two or more pre-crRNAs upstream of the stem-loop structure to generate two or more intermediate crRNAs, the two or more intermediate crRNAs are processed into two or more mature crRNAs, and the two or more mature crRNAs guide the B-GEn.16 polypeptide to introduce two or more double-strand breaks (DSBs) into the DNA. For example, one advantage of B-GEn.16 is that it is possible to introduce into the cell only one pre-crRNA containing several repeat-spacer units that are processed by B-GEn.16 upon introduction, into active repeat-spacer units targeting several different sequences on the DNA.

[0048] In another aspect, provided herein is a method for editing or modifying DNA at multiple locations within a cell, essentially comprising: i) introducing into the cell a form of B-GEn.16 with reduced endoribonuclease activity, as a polypeptide or a nucleic acid encoding the B-GEn.16 polypeptide; and ii) introducing a single heterologous nucleic acid comprising, as RNA or encoded as DNA, one or more pre-CRISPR RNAs (pre-crRNAs), intermediate crRNAs or mature crRNAs under the control of one or more promoters, wherein each crRNA comprises a repeat-spacer array, the spacer comprises a nucleic acid sequence complementary to a target sequence in the DNA, and the repeat comprises a stem-loop structure, wherein the B-GEn.16 polypeptide has reduced or abolished endoribonuclease activity and binds to one or more single heterologous RNAs having a complete endonuclease as directed by one or more spacer sequences in the single heterologous nucleic acid.

[0049] In some embodiments, the pre-crRNA sequences in the single heterologous nucleic acid are linked together at specific positions, orientations, sequences, or by specific chemical linkages to direct or differentially regulate the endonuclease activity of B-GEn.16 at each of the sites specified by the different crRNA sequences.

[0050] In another aspect, provided herein is an example of a general method for editing or modifying the structure or function of DNA at multiple locations within a cell, essentially comprising: i) introducing into the cell an RNA-guided endonuclease, such as B-GEn.16, as a polypeptide or a nucleic acid encoding the RNA-guided endonuclease; and ii) introducing a single heterologous nucleic acid comprising or encoding two or more guide RNAs under the control of one or more promoters, wherein the activity or function of the RNA-guided endonuclease is directed by the guide RNA sequences in the single heterologous nucleic acid.

[0051] Definition As used interchangeably herein, the terms "polynucleotide", "nucleic acid", and "nucleic acid" refer to polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the term includes single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids / triple helices, or polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases, but is not limited thereto.

[0052] "Oligonucleotide" generally refers to a polynucleotide between about 5 and about 100 nucleotides in single-stranded or double-stranded DNA. However, for the purposes of the present disclosure, there is no upper limit to the length of the oligonucleotide. Oligonucleotides are also known as "oligomers" or "oligos", and can be isolated from genes or chemically synthesized by methods known in the art. The terms "polynucleotide" and "nucleic acid" should be understood to include single-stranded (such as sense or antisense) and double-stranded polynucleotides where applicable to the described embodiments.

[0053] "Genomic DNA" refers to the DNA of the genome of an organism, including but not limited to bacteria, fungi, archaea, protists, viruses, plants, or animals.

[0054] The term "manipulating" DNA includes ligating, single-strand nicking, or cleaving, for example, cleaving both strands of DNA, or modifying or editing the DNA or a polypeptide associated with the DNA. Manipulating DNA can silence, activate, or regulate (increase or decrease) the expression of RNA or polypeptide encoded by the DNA, or prevent or enhance the binding of a polypeptide to the DNA.

[0055] A "stem-loop structure" refers to a nucleic acid having a secondary structure that includes a region of nucleotides, which is known or predicted to form a double strand (the stem portion) connected at one end mainly by a region of single-stranded nucleotides (the loop portion). The terms "hairpin" and "foldback" structures are also used herein to refer to the stem-loop structure. Such structures are well known in the art, and these terms are used in accordance with their well-known meanings in the art. As is known in the art, the stem-loop structure does not require exact base pairing. Thus, the stem may contain one or more base mismatches. Alternatively, the base pairing may be exact, e.g., it may contain no mismatches at all.

[0056] "Hybridizable" or "complementary" or "substantially complementary" means that a nucleic acid (e.g., RNA or DNA) is sequence-specific, antiparallel, in a manner (e.g., the nucleic acid specifically binds to complementary diffusion), and non-covalently binds to another nucleic acid under in vitro and / or in vivo conditions of appropriate temperature and solution ionic strength, e.g., forms Watson-Crick base pairs and / or G / U base pairs, "anneals", or "hybridizes". As is known in the art, standard Watson-Crick base pairing includes: pairing of adenine (A) with thymidine (T), pairing of adenine (A) with uracil (U), and pairing of guanine (G) with cytosine (C) [DNA, RNA]. In addition, hybridization between two RNA molecules (e.g., dsRNA), between a guanine (G) base pair and uracil (U) is also known in the art. For example, G / U base pairing is partially involved in the degeneracy (e.g., redundancy) of the genetic code associated with tRNA anticodon base pairing with codons in mRNA. In the context of the present disclosure, the guanine (G) of the protein-binding segment (dsRNA duplex) of the guide RNA molecule is considered to be complementary to uracil (U), and vice versa. Thus, if a G / U base pair can be made at a given nucleotide position of the protein-binding segment (dsRNA duplex) of the guide RNA molecule, that position is not considered non-complementary, but rather complementary.

[0057] Hybridization and washing conditions are well known and are exemplified in Sambrook, J., Fritsch, E. F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly in Chapter 11 and Table 11.1 thereof; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). The temperature and ionic strength conditions determine the "stringency" of the hybridization.

[0058] Hybridization allows for base mismatches but requires that the two nucleic acids contain complementary sequences. The conditions appropriate for hybridization between two nucleic acids depend on variables well known in the art, such as the length and degree of complementarity of the nucleic acids. The greater the degree of complementarity between two nucleotide sequences, the higher the melting temperature (Tm) value of the hybrid of the nucleic acids having those sequences. For hybridization between nucleic acids having a short stretch of complementarity (e.g., complementarity over 35 or fewer, 30 or fewer, 25 or fewer, 22 or fewer, 20 or fewer, or 18 or fewer nucleotides), the position of the mismatch becomes important (see Sambrook et al., supra, 11.7-11.8). Generally, the length of nucleic acids that can hybridize is at least 10 nucleotides. Exemplary minimum lengths of nucleic acids that can hybridize are at least 15 nucleotides; at least 20 nucleotides; at least 22 nucleotides; at least 25 nucleotides; and at least 30 nucleotides. Further, one of ordinary skill in the art will recognize that the temperature and salt concentration of the wash solution can be adjusted as needed according to factors such as the length of the complementary region and the degree of complementarity.

[0059] It is understood in the art that the sequence of a polynucleotide need not be 100% complementary to the sequence of its target nucleic acid in order to hybridize specifically. Further, a polynucleotide can hybridize over one or more segments such that intervening or adjacent segments do not participate in the hybridization event (e.g., loop structures or hairpin structures). A polynucleotide can comprise at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to a target region within the target nucleic acid sequence that it targets. For example, an antisense nucleic acid in which 18 of 20 nucleotides are complementary to the target region and thus hybridize specifically can represent 90 percent complementarity. In this example, the remaining non-complementary nucleotides may be clustered or interspersed with the complementary nucleotides and need not be contiguous with each other or with the complementary nucleotides. The percent complementarity between specific nucleic acid sequences within a nucleic acid can be routinely determined using the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al., J. Mol. Biol. 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), known in the art, or the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), or using the default settings of the algorithm of Smith and Waterman (Adv. Appl. Math. 1981(2)482-489).

[0060] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein and refer to a polymeric form of amino acids of any length that can include encoded and non-encoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.

[0061] As used herein, "binding" (e.g., with respect to an RNA binding domain of a polypeptide) refers to non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). In the state of non-covalent interaction, macromolecules are said to "associate" or "interact" or "bind" (e.g., when molecule X is said to interact with molecule Y, it means that molecule X binds non-covalently to molecule Y). Not all components of the binding interaction need to be sequence-specific (e.g., contacts with phosphate residues in the DNA backbone), but some of the binding interaction can be sequence-specific. The binding interaction is generally characterized by a dissociation constant (Kd) of less than 10 -6 M, less than 10 -7 M, less than 10 -8 M, less than 10 -9 M, less than 10 -10 M, less than 10 -11 M, less than 10 -12 M, less than 10 -13 M, less than 10 -14 M, or less than 10 -15 M. "Affinity" refers to the strength of binding, and an increase in binding affinity correlates with a lower Kd.

[0062] "Binding domain" means a protein domain that can bind non-covalently to another molecule. A binding domain can bind, for example, to a DNA molecule (DNA-binding protein), an RNA molecule (RNA-binding protein), and / or a protein molecule (protein-binding protein). In the case of a protein domain-binding protein, it can bind to itself (forming a homodimer, homotrimer, etc.), and / or can bind to one or more molecules of another protein(s).

[0063] The term "conservative amino acid substitution" refers to the interchangeability in proteins of amino acid residues having similar side chains. For example, the group of amino acids having aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids having aliphatic-hydroxyl side chains consists of serine and threonine; the group of amino acids having amide-containing side chains consists of asparagine and glutamine; the group of amino acids having aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids having basic side chains consists of lysine, arginine, and histidine; the group of amino acids having acidic side chains consists of glutamate and aspartate; and the group of amino acids having sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substituents are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.

[0064] A polynucleotide or polypeptide has a certain percentage of "sequence identity" to another polynucleotide or polypeptide, which means that when aligned, that percentage of bases or amino acids is the same and they are at the same relative positions when comparing the two sequences. Sequence identity can be determined in several different ways. To determine sequence identity, various methods and computer programs (e.g., BLAST, T-COFFEE, MUSCLE, MAFFT, etc.) available on the world wide web at sites including ncbi.nlm.nili.gov / BLAST, ebi.ac.uk / Tools / msa / tcoffee, ebi.Ac.Uk / Tools / msa / muscle, mafft.cbrc / alignment / software can be used to align the sequences. See, for example, Altschul et al., (1990), J. Mol. Biol. 215:403-10. In some embodiments of the present disclosure, standard sequence alignment in the art is used in accordance with the present disclosure to determine the amino acid residues in the B-GEn.16 polypeptide or its variant that "correspond" to the amino acid residues in another Cas9 endonuclease. The amino acid residues in the B-GEn.16 polypeptide or its variant that correspond to the amino acid residues of another Cas9 endonuclease appear at the same positions in the sequence alignment.

[0065] A DNA sequence that "encodes" a particular RNA is a DNA nucleic acid sequence that is transcribed into RNA. The polydeoxyribonucleotides may encode RNA (mRNA) that is translated into protein, or the polydeoxyribonucleotides may encode RNA that is not translated into protein (e.g., tRNA, rRNA, or guide RNA; also referred to as "non-coding" RNA or "ncRNA"). A "protein-coding sequence" or a sequence that encodes a particular protein or polypeptide is a nucleic acid sequence that, when placed under the control of appropriate regulatory sequences, is transcribed in vitro or in vivo (in the case of DNA) into mRNA and (in the case of mRNA) translated into a polypeptide. The boundaries of the coding sequence are determined by the start codon (N-terminus) at the 5' end and the translation termination nonsense codon (C-terminus) at the 3' end. The coding sequence can include, but is not limited to, cDNA derived from prokaryotic or eukaryotic mRNA, genomic DNA sequences derived from prokaryotic or eukaryotic DNA, and synthetic nucleic acids. The transcription termination sequence is generally located 3' of the coding sequence.

[0066] As used herein, a "promoter sequence" or "promoter" is a DNA regulatory region that can bind to RNA polymerase and initiate transcription of a downstream (3' direction) coding or non-coding sequence. As used herein, a promoter sequence is bound at its 3' end by a transcription start site and extends upstream (5' direction) to include the minimum number of bases or elements necessary to initiate transcription at a detectable level above background. Within a promoter sequence, a transcription start site, as well as protein binding domains involved in the binding of RNA polymerase, are found. Eukaryotic promoters often, but not always, contain a "TATA" box and a "CAAT" box. A variety of promoters, including inducible promoters, can be used to drive the various vectors of the present disclosure. A promoter can be a constitutively active promoter (e.g., a promoter that is in a constitutively active "ON" state), which can be an inducible promoter (e.g., a promoter whose state, active / "ON" or inactive / "OFF", is controlled by an external stimulus, such as the presence of a specific temperature, compound, or protein), which can be a spatially restricted promoter (e.g., a transcriptional control element, enhancer, etc.) (e.g., a tissue-specific promoter, a cell-type specific promoter, etc.), and which can be a temporally restricted promoter (e.g., a promoter that is in an "ON" state or an "OFF" state during a specific stage of embryonic development or during a specific stage of a biological process, such as the hair follicle cycle in a mouse). Suitable promoters can be derived from viruses and thus can be referred to as viral promoters, or can be derived from any organism, including prokaryotes or eukaryotes. Suitable promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol Ill).Exemplary promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter; the adenovirus major late promoter (Ad MLP); the herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter region (CMVIE), the Rous sarcoma virus (RSV) promoter, the human U6 small nuclear promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497 - 500 (2002)), the enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1;31(17)), the human H1 promoter (H1), and the like. Examples of inducible promoters include, but are not limited to, the T7 RNA polymerase promoter, the T3 RNA polymerase promoter, the isopropyl - β - D - thiogalactopyranoside (IPTG) regulatable promoter, the lactose inducible promoter, the heat shock promoter, the tetracycline regulatable promoter, the steroid regulatable promoter, the metal regulatable promoter, the estrogen receptor regulatable promoter, and the like. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline; RNA polymerase, such as T7 RNA polymerase; estrogen receptor; estrogen receptor fusion proteins, and the like.

[0067] In some embodiments, the promoter is a spatially restricted promoter (e.g., a cell-type specific promoter, a tissue-specific promoter, etc.) such that in a multicellular organism, the promoter is active (e.g., "ON") in a subset of specific cells. Spatially restricted promoters may also be referred to as enhancers, transcriptional control elements, regulatory sequences, etc. Any suitable spatially restricted promoter can be used, and the choice of suitable promoters (e.g., a brain-specific promoter, a promoter driving expression in a subset of neurons, a promoter driving expression in the germline, a promoter driving expression in the lung, a promoter driving expression in muscle, a promoter driving expression in pancreatic islet cells, etc.) depends on the organism. For example, various spatially restricted promoters are known for plants, insects, nematodes, mammals, mice, etc. Thus, spatially restricted promoters can be used to regulate the expression of nucleic acids encoding site-specific modifying enzymes in a wide variety of different tissues and cell types depending on the organism. Some spatially restricted promoters are also temporally restricted such that the promoter is in the "ON" or "OFF" state during specific stages of embryonic development or during specific stages of a biological process (e.g., the hair follicle cycle in mice). For illustrative purposes, examples of spatially restricted promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, etc.Neuron-specific spatially restricted promoters include, but are not limited to, the neuron-specific enolase (NSE) promoter (see, e.g., EMBL HSEN02, X51956); the aromatic amino acid decarboxylase (AADC) promoter; the neurofilament promoter (see, e.g., GenBank HUMNFL, L04147); the synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); the thy-1 promoter (see, e.g., Chen et al., (1987) Cell 51:7-19; and Llewellyn et al., (2010) Nat. Med. 16(10):1161-1166); the serotonin receptor promoter (see, e.g., GenBank S62283); the tyrosine hydroxylase promoter (TH) (see, e.g., Oh et al., (2009) Gene Ther. 16:437; Sasaoka et al., (1992) Mol. Brain Res. 16:274; Boundy et al., (1998) J. Neurosci. 18:9989; and Kaneda et al., (1991) Neuron 6:583-594); the GnRH promoter (see, e.g., Radovick et al., (1991) Proc. Natl. Acad. Sci. USA 88:3402-3406); the L7 promoter (see, e.g., Oberdick et al., (1990) Science 248:223-226); the DNMT promoter (see, e.g., Bartge et al., (1988) Proc. Natl. Acad. Sci. USA 85:3648-3652); the enkephalin promoter (see, e.g., Comb et al., (1988) EMBO J. 17:3793-3805); the myelin basic protein (MBP) promoter; the Ca2+-calmodulin-dependent protein kinase II-alpha (CamKIIα) promoter (see, e.g., Mayford et al., (1996) Proc. Natl. Acad. Sci. USA 93:13250; and Casanova et al., (2001) Genesis 31:37); the CMV enhancer / platelet-derived growth factor-β promoter (see, e.g., Liu et al., (2004) Gene Therapy 11:52-60).

[0068] As used interchangeably herein, the terms “DNA regulatory sequence,” “control element,” and “regulatory element” refer to transcriptional and translational control sequences such as promoters, enhancers, polyadenylation signals, terminators, proteolytic signals, etc., that provide and / or regulate the transcription of non-coding sequences (e.g., guide RNAs) or coding sequences (e.g., B-GEn.16 polypeptide or variants thereof), and / or regulate the translation of the encoded polypeptide.

[0069] As used herein, the terms “naturally occurring” or “unmodified” when applied to a nucleic acid, polypeptide, cell, or organism refer to a nucleic acid, polypeptide, cell, or organism found in nature. For example, a polypeptide or polynucleotide sequence present in an organism (including viruses) that can be isolated from a natural source and has not been intentionally modified by humans in the laboratory is naturally occurring.

[0070] As used herein, the term “chimeric” when applied to a nucleic acid or polypeptide refers to one entity composed of structures derived from different sources. For example, when “chimeric” is used in relation to a chimeric polypeptide (e.g., chimeric B-GEn.16 protein), the chimeric polypeptide contains amino acid sequences derived from different polypeptides. The chimeric polypeptide can include either a modified or naturally occurring polypeptide sequence (e.g., a first amino acid sequence derived from a modified or unmodified B-GEn.16 protein; and a second amino acid sequence other than the B-GEn.16 protein). Similarly, in relation to a polynucleotide encoding a chimeric polypeptide, “chimeric” includes nucleotide sequences derived from different coding regions (e.g., a first nucleotide sequence encoding a modified or unmodified B-GEn.16 protein; and a second nucleotide sequence encoding a polypeptide other than the B-GEn.16 protein).

[0071] The term "chimeric polypeptide" refers to a polypeptide that does not exist naturally and is produced, for example, by an artificial combination (e.g., "fusion") of segments of amino acid sequences isolated by two or more different methods through human intervention. A polypeptide containing a chimeric amino acid sequence is a chimeric polypeptide. Some chimeric polypeptides can be referred to as "fusion variants".

[0072] "Heterologous", as used herein, means a nucleotide or peptide that is not found in a natural nucleic acid or protein, respectively. The B-GEn.16 fusion protein described herein may contain the RNA-binding domain of the B-GEn.16 polypeptide (or a variant thereof) fused to a heterologous polypeptide sequence (e.g., a polypeptide sequence derived from a protein other than B-GEn.16). The heterologous polypeptide may exhibit an activity (e.g., enzymatic activity) also exhibited by the B-GEn.16 fusion protein (e.g., methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). A heterologous nucleic acid can be ligated (e.g., by genetic manipulation) to a naturally occurring nucleic acid (or a variant thereof) to generate a fusion polynucleotide encoding the fusion polypeptide. As another example, in a fusion variant B-GEn.16 polypeptide, the variant B-GEn.16 polypeptide can be fused to a heterologous polypeptide (e.g., a polypeptide other than B-GEn.16), which exhibits an activity also exhibited by the fusion variant B-GEn.16 polypeptide. The heterologous nucleic acid is ligated to the variant B-GEn.16 polypeptide (e.g., by genetic engineering) to generate a polynucleotide encoding the fusion variant B-GEn.16 polypeptide. "Heterologous", as used herein, further means a nucleotide or polypeptide in a cell that is not its natural cell.

[0073] The term "cognate" usually refers to two biomolecules that interact or coexist naturally.

[0074] As used herein, "recombinant" means the product of various combinations of cloning, restriction, polymerase chain reaction (PCR) and / or ligation steps that result in a construct in which a particular nucleic acid (DNA or RNA) or vector has a structural coding or non-coding sequence distinguishable from the endogenous nucleic acid found in the natural system. DNA sequences encoding polypeptides can be assembled from cDNA fragments or from a series of synthetic oligonucleotides to provide synthetic nucleic acids that can be expressed from recombinant transcription units contained in cell or cell-free transcription and translation systems. Genomic DNA containing related sequences can also be used in the formation of recombinant genes or transcription units. Sequences of non-translated DNA may be present 5' or 3' to the open reading frame, and such sequences do not interfere with the manipulation or expression of the coding region and, in fact, may act to regulate the production of the desired product by various mechanisms (see "DNA regulatory sequences" below). Additionally, or alternatively, DNA sequences encoding non-translated RNAs (e.g., guide RNAs) may also be considered recombinant. Thus, for example, the term "recombinant" nucleic acid refers to a nucleic acid that is not naturally occurring, e.g., made by the artificial combination of two other isolated segments of sequence through human intervention. This artificial combination is often accomplished by either chemical synthesis means or the artificial manipulation of isolated segments of nucleic acid, e.g., genetic engineering techniques. Such manipulations are generally done to replace codons with codons encoding the same amino acid, a conserved amino acid, or a non-conserved amino acid. Additionally, or alternatively, nucleic acid segments of desired functions are joined together to generate a combination of desired functions. This artificial combination is often accomplished by either chemical synthesis means or the artificial manipulation of isolated segments of nucleic acid, e.g., genetic engineering techniques. When a recombinant polynucleotide encodes a polypeptide, the sequence of the encoded polypeptide can be that of a naturally occurring ("wild-type") sequence or a variant of a naturally occurring sequence (e.g., a mutant).Thus, the term "recombinant" polypeptide does not necessarily mean a polypeptide whose sequence does not occur naturally. Instead, a "recombinant" polypeptide is encoded by a recombinant DNA sequence, but the sequence of the polypeptide may or may not occur naturally (e.g., "wild-type") or may not occur naturally (e.g., variant, mutant, etc.). Thus, a "recombinant" polypeptide may be the result of human intervention but may have an amino acid sequence that occurs naturally. The term "not occurring naturally" includes molecules that are significantly different from their naturally occurring counterparts, including chemically modified or mutated molecules.

[0075] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an "insert", can be attached to effect replication of the attached segment in a cell.

[0076] An "expression cassette" contains a DNA coding sequence operably linked to a promoter. "Operably linked" refers to a juxtaposition wherein the components so described are in a relationship that enables them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if the promoter affects its transcription or expression. The terms "recombinant expression vector" or "DNA construct" are used interchangeably herein and refer to a DNA molecule containing a vector and at least one insert. Recombinant expression vectors are generally produced for the purpose of expressing and / or propagating an insert or for the construction of other recombinant nucleotide sequences. The nucleic acid(s) may or may not be operably linked to a promoter sequence and may or may not be operably linked to DNA regulatory sequences.

[0077] As used herein, the term "operably linked" means a physical or functional linkage between two or more elements, such as polypeptide sequences or polynucleotide sequences, that enables them to operate in the intended manner. For example, an operable linkage between a target polynucleotide and a regulatory sequence (e.g., a promoter) is a functional linkage that enables expression of the target polynucleotide. In this sense, the term "operably linked" refers to the positioning of a regulatory region and a coding sequence such that the regulatory region is effective to regulate transcription or translation of the coding sequence of interest. In some embodiments disclosed herein, the term "operably linked" refers to a configuration in which a control sequence is positioned in an appropriate position relative to a sequence encoding a polypeptide or functional RNA such that the control sequence directs or regulates the expression or cellular localization of the mRNA encoding the polypeptide, the polypeptide, and / or the functional RNA. Thus, a promoter is operably linked to a nucleic acid sequence if it can mediate transcription of the nucleic acid sequence. The operably linked elements can be continuous or non-continuous.

[0078] A cell is "genetically modified" or "transformed" or "transfected" by exogenous DNA, such as a recombinant expression vector, when such DNA has been introduced into the cell. The presence of exogenous DNA results in a permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently linked) into the genome of the cell.

[0079] For example, in prokaryotes, yeast, and mammalian cells, the transforming DNA may be maintained on an episomal element such as a plasmid. For eukaryotic cells, stably transformed cells are those in which the transforming DNA is integrated into the chromosome so that it is inherited by daughter cells via chromosomal replication. This stability is demonstrated by the ability of eukaryotic cells to establish cell lines or clones containing a population of daughter cells that contain the transforming DNA. A "clone" is a population of cells derived from a single cell or a common ancestor by mitosis. A "cell line" is a clone of primary cells capable of stable growth in vitro over many generations.

[0080] Suitable methods of genetic modification (also referred to as "transformation") include, for example, viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun method, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, for example, Panyam et al., Adv Drug Deliv Rev. 2012 Sep 13. pp: S169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023), but are not limited thereto.

[0081] As used herein, "host cell" means a cell derived from a eukaryotic cell, prokaryotic cell (e.g., a bacterial or archaeal cell), or a multicellular organism (e.g., a cell line) cultured as a single cell entity, either in vivo or in vitro, where the eukaryotic or prokaryotic cell can be or has been used as a recipient of nucleic acid and includes progeny of the original cell that has been transformed by the nucleic acid. It is understood that progeny of a single cell may not necessarily be identical in form, or in genomic or total DNA complement, to the original parent due to natural, accidental, or intentional mutations. A "recombinant host cell" (also referred to as a "genetically modified host cell") is a host cell into which a heterologous nucleic acid, such as an expression vector, has been introduced. For example, a bacterial host cell is a bacterial host cell that has been genetically modified by introduction of an exogenous nucleic acid (e.g., a plasmid or recombinant expression vector) into a suitable bacterial host cell, and a eukaryotic host cell is a eukaryotic host cell (e.g., a mammalian germ cell) that has been genetically modified by introduction of an exogenous nucleic acid into a suitable eukaryotic host cell.

[0082] As used herein, "target DNA" is a polydeoxyribonucleotide containing a "target site" or "target sequence". The terms "target site", "target sequence", "target protospacer DNA", or "protospacer-like sequence" are used interchangeably herein to refer to a nucleic acid sequence present in the target DNA to which a DNA target segment (also called a "spacer") of a guide RNA binds when sufficient conditions for binding exist. For example, the target site (or target sequence) 5'-GAGCATATC-3' within the target DNA is targeted (or bound to, or hybridized with, or complementary to) by the RNA sequence 5'-GAUAUGCUC-3'. Suitable DNA / RNA binding conditions include physiological conditions normally present in cells. Other suitable DNA / RNA binding conditions (e.g., conditions in a cell-free system) are known in the art; see, e.g., Sambrook, supra. The strand of the target DNA that is complementary to and hybridizes with the guide RNA is called the "complementary strand", and the strand of the target DNA that is complementary to the "complementary strand" (and thus not complementary to the guide RNA) is referred to as the "non-complementary strand" or "non-complementary strand".

[0083] The term "site-specific modifying enzyme" or "RNA-binding site-specific modifying enzyme" refers to a polypeptide that binds to RNA and targets a specific DNA sequence such as the B-GEn.16 polypeptide. The site-specific modifying enzymes described herein target a specific DNA sequence by the RNA molecule to which it binds. The RNA molecule contains a sequence that binds to, hybridizes with, or is complementary to the target sequence within the target DNA, and thus targets the bound polypeptide to a specific position within the target DNA (target sequence). "Cleavage" means the cleavage of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand cleavage and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two different single-strand cleavage events. DNA cleavage can result in the generation of either blunt ends or sticky ends. In certain embodiments, a complex comprising a guide RNA and a site-specific modifying enzyme is used for targeted double-strand DNA cleavage.

[0084] The terms "nuclease" and "endonuclease" are used interchangeably herein and refer to an enzyme having endonuclease catalytic activity for polynucleotide cleavage.

[0085] The "cleavage domain" or "active domain" or "nuclease domain" of a nuclease refers to a polypeptide sequence or domain within a nuclease that has catalytic activity for DNA cleavage. The cleavage domain can be contained within a single polypeptide chain, or the cleavage activity can result from the association of two (or more) polypeptides. A single nuclease domain can consist of two or more stretches of isolated amino acids within a given polypeptide.

[0086] A "guide sequence" or "DNA targeting segment" or "DNA targeting sequence" or "spacer" includes a nucleotide sequence that is complementary to a specific sequence within a target DNA (the complementary strand of the target DNA), which is referred to herein as a "protospacer-like" sequence. A protein-binding segment (or "protein-binding sequence") interacts with a site-specific modifying enzyme. When the site-specific modifying enzyme is B-GEn.16 or a B-GEn.16-related polypeptide (described in more detail below), site-specific cleavage of the target DNA occurs at a position determined by both (i) the complementarity of base pairing between the guide RNA and the target DNA, and (ii) a short motif in the target DNA, referred to as the protospacer adjacent motif (PAM). The protein-binding segment of the guide RNA includes two complementary stretches of nucleotides that partially hybridize to each other to form a double-stranded RNA duplex (dsRNA duplex). In some embodiments, a nucleic acid (e.g., a guide RNA, a nucleic acid comprising a nucleotide sequence encoding the guide RNA; a nucleic acid encoding a site-specific modifying enzyme, etc.) includes modifications or sequences that provide additional desirable features (e.g., modified or regulated stability; intracellular targeting; tracking, e.g., fluorescent labeling; binding sites for proteins or protein complexes, etc.). Non-limiting examples include a 5' cap (e.g., a 7-methylguanylate cap (m7g)); a 3' polyadenylation tail (e.g., a 3' poly(A) tail); a riboswitch sequence (e.g., to enable regulated stability and / or regulated accessibility by a protein and / or protein complex); a stability control sequence; a sequence that forms a dsRNA duplex (e.g., a hairpin); a modification or sequence that targets the RNA to an intracellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that enables fluorescent detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including transcription activators, transcription repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.), and combinations thereof.

[0087] In some embodiments, the guide RNA includes an additional segment that provides any of the above features at either the 5' or 3' end. For example, suitable third segments include a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylation tail (e.g., a 3' poly(A) tail); a riboswitch sequence (e.g., for enabling regulated stability and / or regulated accessibility by proteins and protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (e.g., a hairpin); a sequence that targets the RNA to an intracellular location (e.g., the nucleus, mitochondria, chloroplast, etc.); a modification or sequence that provides a trace (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescence detection, a sequence that enables fluorescence detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.), and combinations thereof.

[0088] A guide RNA and a site-specific modifying enzyme, such as a B-GEn.16 polypeptide or a variant thereof, can form a ribonucleoprotein complex (e.g., bind via non-covalent interactions). The guide RNA provides target specificity to the complex by including a nucleotide sequence complementary to the sequence of the target DNA. The site-specific modifying enzyme of the complex provides endonuclease activity. In other words, the site-specific modifying enzyme is induced to a target DNA sequence (e.g., a target sequence in chromosomal nucleic acid; a target sequence in extrachromosomal nucleic acid, such as episomal nucleic acid, minicircle, etc.; a target sequence in mitochondrial nucleic acid; a target sequence in chloroplast nucleic acid; a target sequence in a plasmid, etc.) by association with the protein-binding segment of the guide RNA. RNA aptamers are known in the art and are generally synthetic versions of riboswitches. The terms "RNA aptamer" and "riboswitch" are used interchangeably herein and encompass both synthetic and natural nucleic acid sequences that provide inducible regulation of the structure (and thus the availability of a particular sequence) of the RNA molecule of which they are a part. RNA aptamers generally include sequences that fold into a particular structure (e.g., a hairpin) that specifically binds a particular drug (e.g., a small molecule). Binding of the drug causes a structural change in the folding of the RNA, which alters the characteristics of the nucleic acid of which the aptamer is a part. As non-limiting examples, (i) an activator-RNA having an aptamer may not be able to bind to cognate target-RNA unless the aptamer is bound by an appropriate drug; (ii) a target-RNA having an aptamer may not be able to bind to cognate activator-RNA unless the aptamer is bound by an appropriate drug; and (iii) a target-RNA and an activator-RNA each containing different aptamers that bind different drugs may not be able to bind to each other unless both drugs are present. As exemplified by these examples, two-molecule guide RNAs can be designed to be inducible.

[0089] Examples of aptamers and riboswitches can be found, for example, in Nakamura et al., Genes Cells. 2012 May;17(5):344-64; Vavalle et al., Future Cardiol. 2012 May;8(3):371-82; Citartan et al., Biosens Bioelectron. 2012 Apr 15;34(1):1-11; and Liberman et al., Wiley lnterdiscip Rev RNA. 2012 May-Jun;3(3):369-84, all of which are hereby incorporated by reference in their entirety.

[0090] The choice of method for gene modification generally depends on the type of cell being transformed and the circumstances in which the transformation is taking place (e.g., in vitro, ex vivo, or in vivo). General considerations of these methods can be found in Ausubel et al., Short Protocols in Molecular Biology, 3rd Edition, Wiley & Sons, 1995.

[0091] Examples of aptamers and riboswitches can be found, for example, in Nakamura et al., Genes Cells. 2012 May;17(5):344-64; Vavalle et al., Future Cardiol. 2012 May;8(3):371-82; Citartan et al., Biosens Bioelectron. 2012 Apr 15;34(1):1-11; and Liberman et al., Wiley lnterdiscip Rev RNA. 2012 May-Jun;3(3):369-84, all of which are hereby incorporated by reference in their entirety.

[0092] The term "stem cell" is used herein to refer to cells that have the ability to self-renew and the ability to generate differentiated cell types (e.g., plant stem cells, vertebrate stem cells) (see Morrison et al., (1997) Cell 88:287-298). In the context of cell ontogeny, the adjectives "differentiated" or "differentiating" are relative terms. A "differentiated cell" is a cell that has progressed further along a developmental pathway than the cell to which it is being compared. Thus, pluripotent stem cells (described below) can differentiate into lineage-restricted progenitor cells (e.g., mesodermal stem cells), which can then differentiate into more restricted cells (e.g., neuronal progenitor cells), which can differentiate into terminally differentiated cells (e.g., terminally differentiated cells such as neurons, cardiomyocytes, etc.), which play a characteristic role in a particular tissue type and may or may not retain the ability to proliferate further. Stem cells can be characterized by both the presence and absence of specific markers (e.g., proteins, RNAs, etc.). Stem cells can also be identified by both in vitro and in vivo functional assays, particularly assays related to the ability of stem cells to give rise to multiple differentiated progeny.

[0093] The stem cells of interest include pluripotent stem cells (PSCs). The term "pluripotent stem cell" or "PSC" is used herein to mean a stem cell that can give rise to all cell types of an organism. Thus, PSCs can give rise to cells of all germ layers of an organism (e.g., the endoderm, mesoderm, and ectoderm of a vertebrate). Pluripotent cells can form teratomas and can contribute to ectodermal, mesodermal, or endodermal tissues in vivo. Pluripotent stem cells of plants can give rise to all cell types of the plant (e.g., cells of roots, stems, leaves, etc.).

[0094] Animal PSCs can be induced in many different ways. For example, embryonic stem cells (ESCs) are derived from the inner cell mass of the embryo (Thomson et al., Science. 1998 Nov 6;282(5391):1145-7), while induced pluripotent stem cells (iPSCs) are derived from somatic cells (Takahashi et al., Cell. 2007 Nov 30;131(5):861-72; Takahashi et al., Nat Protoc. 2007;2(12):3081-9; Yu et al., Science. 2007 Dec 21;318(5858):1917-20. Epub 2007 Nov 20).

[0095] Since the term PSC refers to pluripotent stem cells regardless of their derivation, the term PSC encompasses the terms ESC and iPSC, as well as the term embryonic germ stem cells (EGSCs), which are another example of PSCs. PSCs may be in the form of established cell lines, obtained directly from primary embryonic tissues, or derived from somatic cells. PSCs can be the target cells of the methods described herein.

[0096] "Embryonic stem cells" (ESCs) refer to PSCs isolated from embryos, generally from the inner cell mass of blastocysts. ESC lines are listed in the NIH Human Embryonic Stem Cell Registry, for example, hESBGN-01, hESBGN-02, hESBGN-03, hESBGN-04 (BresaGen, Inc.); HES-1, HES-2, HES-3, HES-4, HES-5, HES-6 (ES Cell International); Miz-hES1 (MizMedi Hospital-Seoul National University); HSF-1, HSF-6 (University of California at San Francisco); and H1, H7, H9, H13, H14 (Wisconsin Alumni Research Foundation (WiCell Research Institute)). The stem cells of interest also include embryonic stem cells from other primates such as rhesus monkey stem cells and marmoset stem cells. Stem cells can be obtained from any mammalian species, such as humans, horses, cows, pigs, dogs, cats, rodents, such as mice, rats, hamsters, primates, etc. (Thomson et al., (1998) Science 282:1145; Thomson et al., (1995) Proc. Natl. Acad. Sci. USA 92:7844; Thomson et al., (1996) Biol. Reprod. 55:254; Shamblott et al., Proc. Natl. Acad. Sci. USA 95:13726, 1998). In culture, ESCs generally grow as flat colonies with a large nucleus-cytoplasm ratio, distinct boundaries, and prominent nucleoli. Furthermore, ESCs express SSEA-3, SSEA-4, TRA-1-60, TRA-1-81, and alkaline phosphatase, but do not express SSEA-1. Examples of methods for generating and characterizing ESCs can be found, for example, in U.S. Patent No. 7,029,913, U.S. Patent No. 5,843,780, and U.S. Patent No. 6,200,806, the disclosures of which are incorporated herein by reference.Methods for growing undifferentiated hESCs are described in WO99 / 20741, WO01 / 51616, and WO03 / 020920. “Embryonic germ stem cells” (EGSCs) or “embryonic germ cells” or “EG cells” mean PSCs derived from germ cells and / or germ cell progenitor cells, e.g., primordial germ cells, e.g., those that will become sperm and eggs. Embryonic germ cells (EG cells) are thought to have similar properties as the above-described embryonic stem cells. Examples of methods for generating and characterizing EG cells can be found, for example, in U.S. Patent No. 7,153,684; Matsui, Y. et al., (1992) Cell 70:841; Shamblott, M. et al., (2001) Proc. Natl. Acad. Sci. USA 98:113; Shamblott, M. et al., (1998) Proc. Natl. Acad. Sci. USA, 95:13726; and Koshimizu, U. et al., (1996) Development, 122:1235, the disclosures of which are incorporated herein by reference.

[0097] “Induced pluripotent stem cells” or “iPSCs” mean PSCs derived from cells that are not PSCs (e.g., cells that are differentiated compared to PSCs). iPSCs can be derived from a plurality of different cell types including terminally differentiated cells. iPSCs have an ES cell-like morphology that grows as flat colonies with a large nucleus-cytoplasm ratio, defined borders, and prominent nuclei. In addition, iPSCs express one or more important pluripotency markers known to those of skill in the art, including, but not limited to, alkaline phosphatase, SSEA3, SSEA4, Sox2, Oct3 / 4, Nanog, TRA160, TRA181, TDGF 1, Dnmt3b, Fox03, GDF3, Cyp26al, TERT, and zfp42.

[0098] Examples of methods for generating and characterizing iPSCs can be found, for example, in U.S. Patent Application Publication Nos. US20090047263, US20090068742, US20090191159, US20090227032, US20090246875, and US20090304646, the disclosures of which are incorporated herein by reference. Generally, to generate iPSCs, somatic cells are provided with reprogramming factors (e.g., Oct4, SOX2, KLF4, MYC, Nanog, Lin28, etc.) known in the art to reprogram the somatic cells to become pluripotent stem cells.

[0099] "Somatic cell" means any cell in an organism that, in the absence of experimental manipulation, does not normally give rise to all types of cells in the organism. In other words, somatic cells are cells that are sufficiently differentiated so as not to naturally give rise to cells of all three germ layers of the body, e.g., cells of the ectoderm, mesoderm, and endoderm. For example, somatic cells include both neurons and neural progenitor cells, the latter of which can naturally give rise to all or some cell types of the central nervous system, but cannot give rise to cells of the mesoderm or endoderm lineages.

[0100] "Mitotic cell" means a cell that is undergoing mitosis.

[0101] "Post-mitotic cell" means a cell that has exited mitosis, e.g., it is in a "quiescent state" and, for example, is no longer undergoing division. This quiescent state may be temporary, e.g., reversible, or permanent.

[0102] "Meiotic cell" means a cell that is undergoing meiosis.

[0103] "Recombination" means the process of exchange of genetic information between two polynucleotides. As used herein, "homologous recombination repair (HDR)" refers to a special form of DNA repair that occurs, for example, during the repair of double-strand breaks in cells. This process requires nucleotide sequence homology and uses a "donor" molecule for the template repair of a "target" molecule (e.g., one that has experienced a double-strand break), resulting in the transfer of genetic information from the donor to the target. Homologous recombination repair can result in changes (e.g., insertions, deletions, mutations) in the sequence of the target molecule when the donor polynucleotide is different from the target molecule and part or all of the sequence of the donor polynucleotide is incorporated into the target DNA. In some embodiments, the donor polynucleotide, a part of the donor polynucleotide, a copy of the donor polynucleotide, or a part of the copy of the donor polynucleotide is incorporated into the target DNA.

[0104] "Non-homologous end joining (NHEJ)" means the repair of double-strand breaks in DNA by directly ligating the cut ends to each other without the need for a homologous template (in contrast to homologous recombination repair, which requires a homologous sequence to induce repair). NHEJ often results in the loss (deletion) of nucleotide sequences near the double-strand break site.

[0105] The terms "treatment", "treating", etc. are used herein and generally mean obtaining a desired pharmacological and / or physiological effect. This effect can be prophylactic in terms of completely or partially preventing a disease or its symptoms, and / or therapeutic in terms of partially or completely curing a disease and / or the adverse effects caused by the disease. "Treatment", as used herein, encompasses any treatment of a disease or condition in a mammal, including (a) preventing the onset of a disease or condition in a subject who has a predisposition to the disease or condition but has not yet been diagnosed as having it; (b) suppressing a disease or condition, e.g., preventing its onset; or (c) alleviating a disease, e.g., causing regression of the disease. A therapeutic agent can be administered before, during, or after the onset of a disease or injury. Treatment of an ongoing disease, where the treatment stabilizes or reduces the undesirable clinical symptoms of the subject, is of particular interest. Such treatment is preferably carried out before the complete loss of function in the affected tissue. Treatment is preferably carried out during the symptomatic stage of the disease and, in some cases, after the symptomatic stage of the disease.

[0106] The terms "individual", "subject", "host", and "patient" are used interchangeably herein and refer to any mammalian subject, particularly a human, for whom diagnosis, treatment, or therapy is desired.

[0107] General methods in molecular and cellular biochemistry can be found in standard textbooks such as Molecular Cloning: A Laboratory Manual, 3rd Edition (Sambrook et al., Harbor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Edition (Ausubel et al., eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al., eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998). The disclosures of which are incorporated herein by reference.

[0108] Where a range of values is provided, unless the context clearly dictates otherwise, each intervening value between the upper and lower limits of that range to one tenth of the unit of the lower limit, and any other stated value or intervening value in the stated range, is included within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also included within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0109] The phrase "consisting essentially of" means excluding any component that is not a specified active ingredient or component of the system, or any part that is not a specified active moiety or molecular part.

[0110] In this specification, a specific range is presented together with a numerical value preceded by the term "about". In this specification, the term "about" is used to provide literal support for the exact number that it precedes, as well as for numbers that are close to, or approximately, the number that the term precedes. When determining whether a number is close to, or approximate to, a specifically recited number, a number not recited as being close to, or approximate to, may be a number that provides a substantial equivalent of the specifically recited number in the context in which it is presented.

[0111] It is understood that certain features of the disclosure that are described in connection with separate embodiments may be provided in combination in a single embodiment. Conversely, various features of the disclosure that are described in connection with a single embodiment may also be provided separately, or in any suitable sub-combination. All combinations of embodiments associated with the disclosure are specifically embraced by the disclosure and are disclosed herein as if each and every combination was individually and explicitly disclosed. Further, all sub-combinations of various embodiments and their elements are also specifically embraced by the disclosure and are disclosed herein as if each and every such sub-combination was individually and explicitly disclosed herein.

[0112] B-GEn.16 fusion polypeptide Using B-GEn.16, a fusion protein having additional domains and activities compared to the B-GEn.16 nuclease can be formed. As a non-limiting example, the Fokl domain can be fused to a B-GEn.16 polypeptide or a variant thereof, which can include a catalytically active endonuclease domain, or the Fokl domain can be fused to a B-GEn.16 polypeptide or a variant thereof, which has been modified to inactivate the B-GEn.16 endonuclease domain. Other domains that can be fused to create a fusion protein with B-GEn.16 include transcriptional regulators, epigenetic modifiers, tags and other labels or contrast agents, histones, and / or other modalities known in the art that regulate or modify the structure or activity of gene sequences.

[0113] In some embodiments, the B-GEn.16 polypeptide or a variant thereof described herein is fused to a transcriptional activator or repressor, or an epigenetic modifier, such as a methylase, demethylase, acetylase, or deacetylase.

[0114] In some embodiments, the B-GEn.16 polypeptide or a variant thereof described herein is fused to a functional protein component for detection, intermolecular interaction, translational activation, modification, or any other manipulation known in the art.

[0115] Exemplary B-GEn.16 variant polypeptides In some embodiments, the B-GEn.16 polypeptide or a variant thereof described herein retains a) the ability to bind to a target site and optionally b) its activity. In some embodiments, the activity that is retained is endonuclease activity. In certain embodiments, the endonuclease activity does not require tracrRNA.

[0116] In some embodiments, the active portion of the B-GEn.16 polypeptide or a variant thereof is modified. In some embodiments, the modification comprises an amino acid change (e.g., deletion, insertion, or substitution) that reduces or increases the nuclease activity of the B-GEn.16 polypeptide or a variant thereof. For example, in some embodiments, the modified B-GEn.16 polypeptide or a variant thereof has a nuclease activity that is less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the corresponding unmodified B-GEn.16 polypeptide or a variant thereof. In some embodiments, the modified B-GEn.16 polypeptide or a variant thereof has substantially no nuclease activity. In some embodiments, it may have a nuclease activity that is 50%, 2-fold, 4-fold, or more than 10-fold higher.

[0117] In some embodiments, the active portion of the B-GEn.16 polypeptide or a variant thereof comprises a heterologous polypeptide having DNA modification activity and / or transcription factor activity and / or DNA-binding polypeptide modification activity. In some embodiments, the heterologous polypeptide replaces a portion of the B-GEn.16 polypeptide or a variant thereof that provides nuclease activity. In some embodiments, the B-GEn.16 polypeptide or a variant thereof comprises both a portion of the B-GEn.16 polypeptide or a variant thereof that normally provides nuclease activity (and that portion can be fully active or alternatively modified to have less than 100% of the corresponding unmodified activity) and the heterologous polypeptide. In other words, in some embodiments, the B-GEn.16 polypeptide or a variant thereof can be a fusion polypeptide comprising both a portion of the B-GEn.16 polypeptide or a variant thereof that normally provides nuclease activity and the heterologous polypeptide.

[0118] For example, in the B-GEn.16 fusion protein, the B-GEn.16 polypeptide or its variant can be fused to a heterologous polypeptide sequence (e.g., a polypeptide sequence derived from a protein other than B-GEn.16). The heterologous polypeptide sequence can exhibit an activity (e.g., enzymatic activity) also exhibited by the B-GEn.16 fusion protein (e.g., methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). The heterologous nucleic acid sequence can be ligated (e.g., by genetic engineering) to another nucleic acid sequence to generate a fusion nucleotide sequence encoding the fusion polypeptide. In some embodiments, the B-GEn.16 fusion polypeptide is generated by fusing the B-GEn.16 polypeptide or its variant to a heterologous sequence that provides intracellular localization (e.g., a nuclear localization signal (NLS) for targeting to the nucleus; a mitochondrial localization signal for targeting to the mitochondria; a chloroplast localization signal for targeting to the chloroplast: an ER retention signal, etc.). In some embodiments, the heterologous sequence can provide a tag (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, etc.; a HIS tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag, etc.) to facilitate tracking or purification. In some embodiments, the heterologous sequence can provide an increase or decrease in stability. In some embodiments, the heterologous sequence can provide a binding domain (e.g., providing the ability of the B-GEn.16 fusion polypeptide to bind to another protein of interest, e.g., DNA or a histone-modifying protein, a transcription factor or a transcriptional repressor, a mobilization protein, etc.) or a nucleotide of interest (e.g., an aptamer or a target site of a nucleotide-binding protein).

[0119] In some embodiments, according to any of the B-GEn.16 polypeptide variants described herein, the B-GEn.16 polypeptide variant has reduced endodeoxyribonuclease activity. For example, a B-GEn.16 polypeptide variant suitable for use in the transcriptional regulation methods of the present disclosure exhibits endodeoxyribonuclease activity that is less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the unmodified B-GEn.16 polypeptide.

[0120] In some embodiments, the variant B-GEn.16 polypeptide substantially lacks detectable endodeoxyribonuclease activity (dB-GEn.16). In some embodiments, when the B-GEn.16 polypeptide variant has reduced catalytic activity, the polypeptide can still bind to the target DNA in a site-specific manner (since it is still directed to the target DNA sequence by the guide RNA) as long as it retains the ability to interact with the guide RNA. In some embodiments, the variant B-GEn.16 polypeptide is a nickase that can cleave the complementary strand of the target DNA but has a reduced ability to cleave the non-complementary strand of the target DNA.

[0121] In some embodiments, the variant B-GEn.16 polypeptide is a nickase that can cleave the non-complementary strand of the target DNA but has a reduced ability to cleave the complementary strand of the target DNA.

[0122] In some embodiments, the variant B-GEn.16 polypeptide has a reduced ability to cleave both the complementary and non-complementary strands of the target DNA. For example, alanine substitutions are contemplated.

[0123] In some embodiments, the variant B-GEn.16 polypeptide is a fusion polypeptide (the "variant B-GEn.16 fusion polypeptide"), for example, a fusion polypeptide comprising i) the variant B-GEn.16 polypeptide; and ii) a covalently attached heterologous polypeptide (also referred to as the "fusion partner").

[0124] The heterologous polypeptide may exhibit activities (e.g., enzymatic activities) also shown by the variant B-GEn.16 fusion polypeptide (e.g., methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.). The heterologous nucleic acid sequence can be linked to another nucleic acid sequence (e.g., by genetic engineering) to generate a fusion nucleotide sequence encoding the fusion polypeptide. In some embodiments, the variant B-GEn.16 fusion polypeptide is generated by fusing the variant B-GEn.16 polypeptide with a heterologous sequence that provides intracellular localization (e.g., the heterologous sequence is an intracellular localization sequence, e.g., a nuclear localization signal (NLS) for targeting to the nucleus; a mitochondrial localization signal for targeting to mitochondria; a chloroplast localization signal for targeting to chloroplasts; an ER retention signal, etc.). In some embodiments, the heterologous sequence can provide a tag (e.g., the heterologous sequence is a detectable label) to facilitate tracking and / or purification (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, etc.; a histidine tag, e.g., 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag, etc.). In some embodiments, the heterologous sequence can provide an increase or decrease in stability (e.g., the heterologous sequence is a stability control peptide, e.g., a degron that is controllable in some cases (e.g., a temperature-sensitive or drug-controllable degron sequence, see below)). In some embodiments, the heterologous sequence can provide an increase or decrease in transcription from a target DNA (e.g., the heterologous sequence is a transcriptional regulatory sequence, e.g., a transcription factor / activator or a fragment thereof, a protein that recruits a transcription factor / activator or a fragment thereof, a transcription repressor or a fragment thereof, a protein that recruits a transcription repressor or a fragment thereof, a small molecule / drug-responsive transcriptional regulator, etc.).In some embodiments, the heterologous sequence can provide a binding domain (e.g., the heterologous sequence is a protein-binding sequence that provides, for example, the ability of a fusion dB-GEn.16 polypeptide to bind to another protein of interest, such as a DNA or histone-modifying protein, a transcription factor or a transcription repressor, a mobilization protein, etc.).

[0125] Suitable fusion partners that provide increased or decreased stability include, but are not limited to, degron sequences. It will be readily understood by those skilled in the art that degrons are amino acid sequences that control the stability of the proteins of which they are a part. For example, the stability of a protein containing a degron sequence is at least partially controlled by the degron sequence. In some embodiments, a suitable degron is constitutive such that the degron affects protein stability independent of experimental controls (e.g., the degron is not drug-inducible, temperature-inducible, etc.). In some embodiments, the degron provides a variant B-GEn.16 polypeptide with controllable stability such that the variant B-GEn.16 polypeptide can be "on" (e.g., stable) or "off" (e.g., unstable, reduced) depending on the desired conditions. For example, if the degron is a temperature-sensitive degron, the variant B-GEn.16 polypeptide can be functional (e.g., "on", stable) below a threshold temperature (e.g., 42°C, 41°C, 40°C, 39°C, 38°C, 37°C, 36°C, 35°C, 34°C, 33°C, 32°C, 31°C, 30°C, etc.) but non-functional (e.g., "off", reduced) above the threshold temperature. As another example, if the degron is a drug-inducible degron, the presence or absence of the drug can switch the protein from an "off" (e.g., unstable) state to an "on" (e.g., stable) state or vice versa. An exemplary drug-inducible degron is derived from the FKBP12 protein. The stability of the degron is controlled by the presence or absence of a small molecule that binds to the degron.

[0126] Examples of suitable degrons include, but are not limited to, Shield-1, DHFR, auxin, and / or temperature-controlled degrons. Non-limiting examples of suitable degrons are known in the art (e.g., Dohmen et al., Science, 1994.263(5151): p. 1273-1276: Construction of temperature-sensitive mutants: a heat-inducible degron; Schoeber et al., Am J Physiol Renal Physiol. 2009 Jan;296(1):F204-11: Conditional rapid expression and function of multimeric TRPV5 channels using Shield-1; Chu et al., Bioorg Med Chem Lett. 2008 Nov 15;18(22):5941-4: Recent progress regarding FKBP-derived destabilizing domains; Kanemaki, Pflugers Arch. 2012 Dec 28: The forefront of protein expression control by conditional degrons; Yang et al., Mol Cell. 2012 Nov 30;48(4):487-8: Motivation to destruction: methyl degron; Barbour et al., Biosci Rep. 2013 Jan 18;33(1).: Characterization of a bipartite degron that controls the ubiquitin-independent degradation of thymidylate synthase; and Greussingra, J Vis Exp. 2012 Nov 10;(69): Monitoring of ubiquitin-proteasome activity in living cells using a degron (dgn) destabilized green fluorescent protein (GFP)-based reporter protein; all of which are hereby incorporated by reference in their entirety).

[0127] Exemplary degron arrays have been well-characterized and tested in both cells and animals. Thus, fusing B-GEn.16 to a degron array generates a "tunable" and "inducible" B-GEn.16 polypeptide. Any of the fusion partners described herein can be used in any desired combination. As one non-limiting example to illustrate this point, a B-GEn.16 fusion protein can include a YFP sequence for detection, a degron sequence for stability, and a transcriptional activator sequence for increasing transcription from a target DNA. Furthermore, the number of fusion partners that can be used in a B-GEn.16 fusion protein is not limited. In some embodiments, a B-GEn.16 fusion protein includes one or more (e.g., two or more, three or more, four or more, or five or more) heterologous sequences.

[0128] Suitable fusion partners include, but are not limited to, polypeptides that provide methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, crotonylation activity, decrotonylation activity, propionylation activity, depropionylation activity, myristoylation activity, or demyristoylation activity, any of which can be for the purpose of directly modifying DNA (e.g., methylation of DNA) or modifying a DNA-associating polypeptide (e.g., histone or DNA-binding protein). Further suitable fusion partners include, but are not limited to, boundary elements (e.g., CTCF), proteins and fragments thereof that provide peripheral mobilization (e.g., lamin A, lamin B, etc.), and protein docking elements (e.g., FKBP / FRB, Pil 1 / Aby 1, etc.).

[0129] The B-GEn.16 polypeptide or its variant can also be isolated and purified according to conventional recombinant synthesis methods. The lysate can be prepared from the expression host, and the lysate can be purified using HPLC, size exclusion chromatography, gel electrophoresis, affinity chromatography, or other purification techniques. In most cases, the composition used contains at least 20% by weight of the desired product, at least about 75% by weight, at least about 95% by weight, and for therapeutic purposes, typically at least 99.5% by weight, with respect to the method of preparing the product and the contaminants associated with its purification. Generally, the percentage is based on the total protein. The guide RNA and / or the B-GEn.16 polypeptide or its variant and / or the donor polynucleotide, whether introduced as nucleic acid or polypeptide, are provided to the cell for about 30 minutes to about 24 hours, for example, 1 hour, 1.5 hours, 2 hours, 2.5 hours, 3 hours, 3.5 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 12 hours, 16 hours, 18 hours, 20 hours, or any other period from about 30 minutes to about 24 hours, which can be repeated at a frequency of about daily to about every 4 days, for example, every 1.5 days, every 2 days, every 3 days, or any other frequency from about daily to about every 4 days. The agent(s) can be provided to the cell one or more times, for example, 1 time, 2 times, 3 times, or more than 3 times. After each contact event, the cell is incubated with the agent for a certain time, for example, 16 - 24 hours, and then the medium is replaced with fresh medium and the cell is further cultured. When two or more different targeting complexes are provided to the cell (for example, two different guide RNAs that are complementary to different sequences within the same or different target DNAs), the complexes can be provided simultaneously (for example, as two polypeptides and / or nucleic acids) or delivered simultaneously. Alternatively, they can be provided sequentially, for example, the targeting complex can be provided first, followed by the second targeting complex, etc., or vice versa.

[0130] Nucleic acid Guide RNA / sgRNA In some embodiments, the systems, compositions, and methods described herein use a genome-targeting nucleic acid that can direct the activity of an associating polypeptide (e.g., the B-GEn.16 polypeptide or a variant thereof) to a specific target sequence within a target nucleic acid. In some embodiments, the genome-targeting nucleic acid is RNA. The genome-targeting RNA is referred to herein as a "guide RNA" or "gRNA". The guide RNA has at least a spacer sequence that can hybridize to the target nucleic acid sequence of interest and a CRISPR repeat sequence (such CRISPR repeat sequences are also referred to as "tracr mate sequences"). In type II systems, the gRNA also has a second RNA called the tracrRNA sequence. In type II guide RNA (gRNA), the CRISPR repeat sequence and the tracrRNA sequence hybridize to each other to form a duplex. In type V guide RNA (gRNA), the crRNA forms a duplex. In both systems, the duplex binds to a site-specific polypeptide such that the guide RNA and the site-specific polypeptide form a complex. The genome-targeting nucleic acid provides target specificity to the complex by its association with the site-specific polypeptide. Thus, the genome-targeting nucleic acid directs the activity of the site-specific polypeptide.

[0131] In some embodiments, the genome-targeting nucleic acid is a bimolecular guide RNA. In some embodiments, the genome-targeting nucleic acid is a single-molecule guide RNA or single guide RNA (sgRNA). The bimolecular guide RNA has two RNA strands. The first strand has, in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence, and a minimal CRISPR repeat sequence. The second strand has a minimal tracrRNA sequence (complementary to the minimal CRISPR repeat sequence), a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. The single-molecule guide RNA (sgRNA) in type II systems has, in the 5' to 3' direction, an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single-molecule guide linker, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. Any tracrRNA extension may have elements that confer additional functionality (e.g., stability) to the guide RNA. The single-molecule guide linker links the minimal CRISPR repeat and the minimal tracrRNA sequence to form a hairpin structure. Any tracrRNA extension has one or more hairpins.

[0132] The single-molecule guide RNA (sgRNA) in type V systems has, in the 5' to 3' direction, an optional tracr extension sequence, a tracrRNA sequence, a single-molecule guide linker, a minimal CRISPR repeat sequence, a spacer sequence, and an optional spacer extension sequence.

[0133] Alternatively, the single-molecule guide RNA (sgRNA) in type V systems has, in the 5' to 3' direction, an optional extension sequence, a minimal CRISPR repeat sequence, a spacer sequence, and an optional spacer extension sequence.

[0134] In yet another alternative, the single-molecule guide RNA (sgRNA) in type V systems has, in the 5' to 3' direction, an optional extension sequence, an engineered nuclease-binding RNA sequence and a spacer sequence, and an optional spacer extension sequence.

[0135] Exemplary genome-targeting nucleic acids are described, for example, in WO2018002719.

[0136] Generally, a CRISPR repeat sequence promotes one or more of: (1) excision of a DNA targeting segment adjacent to the CRISPR repeat sequence in a cell containing the corresponding tracr sequence; and (2) formation of a CRISPR complex at a target sequence, where the CRISPR complex comprises a CRISPR repeat sequence hybridized to the tracr sequence, and thus includes any sequence having sufficient complementarity to the tracr sequence. Generally, the degree of complementarity is based on an optimal alignment along the length of the shorter of the two sequences, the CRISPR repeat sequence and the tracr sequence. The optimal alignment can be determined by any suitable alignment algorithm, and can further take into account secondary structures such as self-complementarity within the tracr sequence or the CRISPR repeat sequence. In some embodiments, the degree of complementarity along the length of the shorter of the two, between the tracr sequence and the CRISPR repeat sequence when optimally aligned, is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more. In some embodiments, the tracr sequence is about 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more, 25 or more, 30 or more, 40 or more, 50 or more, or more nucleotides in length. In some embodiments, the tracr sequence and the CRISPR repeat sequence are included within a single transcript such that hybridization between the two generates a transcript having a secondary structure such as a hairpin. In some embodiments, the transcript or transcribed polynucleotide sequence has at least two or more hairpins.

[0137] The spacer of the guide RNA contains a nucleotide sequence complementary to a sequence in the target DNA. In other words, the spacer of the guide RNA interacts with the target DNA in a sequence-specific manner via hybridization (e.g., base pairing). As such, the nucleotide sequence of the spacer can vary and determines the position within the target DNA where the guide RNA and the target DNA interact. The DNA targeting segment of the guide RNA can be modified (e.g., by genetic engineering) to hybridize to any desired sequence within the target DNA.

[0138] In some embodiments, the spacer has a length of 10 nucleotides to 30 nucleotides. In some embodiments, the spacer has a length of 13 nucleotides to 25 nucleotides. In some embodiments, the spacer has a length of 15 nucleotides to 23 nucleotides. In some embodiments, the spacer has a length of 18 nucleotides to 22 nucleotides, e.g., 20 to 22 nucleotides.

[0139] In some embodiments, the percent complementarity between the DNA targeting sequence of the spacer and the protospacer of the target DNA is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) for 20 to 22 nucleotides.

[0140] In some aspects, the protospacer is directly adjacent to a suitable PAM sequence on its 3' end, or such a PAM sequence is part of the DNA targeting sequence in its 3' portion.

[0141] Modifications of the guide RNA can be used to enhance the formation or stability of a CRISPR-Cas genome editing complex comprising the guide RNA and a Cas endonuclease, such as B-GEn.16. Modifications of the guide RNA can also be used, or alternatively, to enhance the initiation, stability or dynamics of the interaction between the genome editing complex and a target sequence in the genome, which can be used, for example, to enhance on-target activity. Modifications of the guide RNA can also be used, or alternatively, to enhance specificity, for example, the relative rate of genome editing at the on-target site compared to the effect at other (off-target) sites.

[0142] Modifications can also be used, or alternatively, to increase the stability of the guide RNA, for example, by increasing its resistance to degradation by ribonucleases (RNases) present in the cell, thereby increasing its half-life in the cell. Modifications that enhance guide RNA half-life can be particularly useful in embodiments where a Cas endonuclease, such as B-GEn.16, is introduced into a cell that is edited via an RNA that needs to be translated to produce the B-GEn.16 endonuclease, as increasing the half-life of the guide RNA introduced simultaneously with the RNA encoding the endonuclease can be used to increase the time that the guide RNA and the encoded Cas endonuclease coexist in the cell.

[0143] Donor DNA or donor template Site-specific polypeptides, such as DNA endonucleases, can introduce double-strand breaks or single-strand breaks into nucleic acids, such as genomic DNA. Double-strand breaks can stimulate the cell's endogenous DNA repair pathways (e.g., homologous recombination repair (HDR) or non-homologous end joining or alternative non-homologous end joining (A-NHEJ) or microhomology-mediated end joining (MMEJ)). NHEJ can repair the cleaved target nucleic acid without the need for a homologous template. This can sometimes result in small deletions or insertions (indels) in the target nucleic acid at the cleavage site, which can lead to disruption or changes in gene expression. HDR, also known as homologous recombination (HR), can occur when a homologous repair template or donor is available.

[0144] The homologous donor template has sequences homologous to the sequences adjacent to the target nucleic acid cleavage site. Sister chromatids are generally used by the cell as a repair template. However, for the purposes of genome editing, the repair template is often supplied as an exogenous nucleic acid, such as a plasmid, double-stranded oligonucleotide, single-stranded oligonucleotide, double-stranded oligonucleotide, or viral nucleic acid. In the case of an exogenous donor template, it is common to introduce additional nucleic acid sequences (e.g., transgenes) or modifications (e.g., changes or deletions of single or multiple bases) between the homologous flanking regions so that the additional or modified nucleic acid sequences are also incorporated into the target locus. MMEJ results in genetic outcomes similar to NHEJ in that small deletions and insertions can occur at the cleavage site. MMEJ utilizes short stretches of homology adjacent to the cleavage site to drive the preferred end-joining DNA repair outcome. In some embodiments, it may be possible to predict the possible repair outcomes based on the analysis of potential microhomology in the nuclease target region.

[0145] Thus, in some cases, homologous recombination is used to insert an exogenous polynucleotide sequence into a target nucleic acid cleavage site. The exogenous polynucleotide sequence is herein referred to as the donor polynucleotide (or donor or donor sequence or polynucleotide donor template). In some embodiments, the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is inserted into the target nucleic acid cleavage site. In some embodiments, the donor polynucleotide is an exogenous polynucleotide sequence, i.e., a sequence that does not naturally occur at the target nucleic acid cleavage site.

[0146] When an exogenous DNA molecule is supplied at a sufficient concentration into the nucleus of a cell where a double-strand break occurs, the exogenous DNA can be inserted at the double-strand break during the NHEJ repair process and thus can result in a permanent addition to the genome. These exogenous DNA molecules are in some embodiments referred to as donor templates. When the donor template contains the coding sequence of one or more system components described herein, optionally together with associated regulatory sequences such as a promoter, enhancer, polyA sequence, and / or splice acceptor sequence, the one or more system components can be expressed from the integrated nucleic acid in the genome, resulting in permanent expression during the lifetime of the cell. Furthermore, the integrated nucleic acid of the donor DNA template can be transmitted to daughter cells when the cell divides.

[0147] In the presence of a sufficient concentration of a donor DNA template (referred to as a homology arm) containing an adjacent DNA sequence having homology to the DNA sequence on either side of the double-strand break, the donor DNA template can be integrated via the HDR pathway. The homology arms act as substrates for homologous recombination between the donor template and the sequences on either side of the double-strand break. This can result in error-free insertion of the donor template where the sequences on either side of the double-strand break are not altered from the sequences in the non-modified genome.

[0148] The donors supplied for editing by HDR vary significantly but generally contain the intended sequences with small or large adjacent homologous arms to enable annealing to genomic DNA. The homology regions flanking the introduced genetic changes can be as small as 30 bp or as large as a multi-kilobase cassette that may contain a promoter, cDNA, etc. Both single-stranded and double-stranded oligonucleotide donors can be used. These oligonucleotides range in size from less than 100 nt to over many kb, but longer ssDNA can also be generated and used. Double-stranded donors, including PCR amplicons, plasmids, and minicircles, are often used. Generally, AAV vectors have been found to be a very effective means of delivering donor templates, but the packaging limit for individual donors is <5 kb. Active transcription of the donor has been shown to increase HDR three-fold, and the inclusion of a promoter can increase conversion. Conversely, CpG methylation of the donor can decrease gene expression and HDR.

[0149] In some embodiments, donor DNA can be supplied with or independent of nucleases by a variety of different methods, such as transfection, nanoparticles, microinjection, or viral transduction. In some embodiments, a range of tethering options can be used to increase the availability of the donor for HDR. Examples include binding the donor to a nuclease, binding to a DNA-binding protein that binds nearby, or binding to a protein involved in DNA end-binding or repair.

[0150] In addition to genomic editing by NHEJ or HDR, site-specific gene insertion can be performed using both the NHEJ pathway and HR. The combinatorial approach may be applicable in certain situations, perhaps including intron / exon boundaries. NHEJ can be shown to be effective for ligation in introns, while error-free HDR can be more suitable for coding regions.

[0151] vector In another aspect, provided herein are nucleic acids comprising a B-GEn.16 polypeptide or variant thereof, a gRNA, and / or any nucleic acid or proteinaceous molecule necessary to practice embodiments of the present disclosure. In some embodiments, such nucleic acids are vectors (e.g., recombinant expression vectors).

[0152] Exemplary expression vectors include, but are not limited to, viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviruses (e.g., vectors derived from retroviruses such as murine leukemia virus, spleen necrosis virus, and Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), and other recombinant vectors. Other vectors contemplated for eukaryotic target cells include, but are not limited to, vectors pXT1, pSG5, pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). Further vectors contemplated for eukaryotic target cells include, but are not limited to, vectors pCTx-1, pCTx-2, and pCTx-3. Other vectors can be used as long as they are compatible with the host cell.

[0153] In some embodiments, the vector has one or more transcriptional and / or translational control elements. Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc., can be used in the expression vector. In some embodiments, the vector is a self-inactivating vector that inactivates viral sequences or components of the CRISPR machinery or other elements.

[0154] Non-limiting examples of suitable eukaryotic promoters (i.e., promoters that function in eukaryotic cells) include the cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, retrovirus-derived long terminal repeats (LTRs), the human elongation factor-1 promoter (EF1), hybrid constructs having a cytomegalovirus (CMV) enhancer fused to the chicken-β-actin promoter (CAG), the mouse stem cell virus promoter (MSCV), the phosphoglycerate kinase-1 locus promoter (PGK), and those from the mouse metallothionein-I.

[0155] For expressing small RNAs containing guide RNAs used in connection with Cas endonucleases, various promoters such as RNA polymerase III promoters including, for example, U6 and H1, may be advantageous. Descriptions of such promoters and parameters for enhancing their use are known in the art and are regularly described; see, for example, Ma, Het. et al., Molecular Therapy-Nucleic Acids 3, e161 (2014) doi:10.1038 / mtna.2014.12.

[0156] Expression vectors can also include a ribosome binding site for translation initiation and a transcription terminator. Expression vectors can also include appropriate sequences for amplifying expression. Expression vectors can also include nucleotide sequences encoding non-native tags (e.g., histidine tags, hemagglutinin tags, green fluorescent protein, etc.) that are fused to a site-specific polypeptide and thus result in a fusion protein.

[0157] In some embodiments, the promoter is an inducible promoter (e.g., heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.). In some embodiments, the promoter is a constitutive promoter (e.g., CMV promoter, UBC promoter). In some embodiments, the promoter is a spatially and / or temporally restricted promoter (e.g., tissue-specific promoter, cell type-specific promoter, etc.). In some embodiments, the vector, after being inserted into the genome, does not have a promoter for at least one gene to be expressed in the host cell when the gene is expressed under an endogenous promoter present in the genome.

[0158] Modification of Nucleic Acids and Polypeptides In some embodiments, the polynucleotides described herein can be used, for example, as further described herein and as known in the art, to enhance activity, stability or specificity, to alter delivery, to reduce the innate immune response in host cells, to further reduce protein size, or for other enhancements, and include one or more modifications. In some embodiments, such modifications result in a B-GEn.16 polypeptide comprising an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity to the sequence of SEQ ID NO:2.

[0159] Codon Optimization In certain embodiments, the modified polynucleotide is used in the CRISPR-B-GEn.16 system described herein, where DNA or RNA comprising a polynucleotide sequence encoding a guide RNA and / or a B-GEn.16 polypeptide or variant thereof can be modified as described and exemplified below. Such modified polynucleotides can be used in the CRISPR-B-GEn.16 system to edit any one or more genomic loci. In some embodiments, such modifications in the polynucleotides of the present disclosure are achieved via codon optimization, e.g., codon optimization based on the particular host cell in which the encoded polypeptide is expressed. It is understood by those skilled in the art that any nucleotide sequence and / or recombinant nucleic acid of the present disclosure can be codon optimized for expression in any species of interest. Codon optimization is well known in the art and involves modification of a nucleotide sequence for codon usage bias using a species-specific codon usage table. The codon usage table is created based on sequence analysis of the most highly expressed genes for the species of interest. In a non-limiting example, when a nucleotide sequence is expressed in the nucleus, the codon usage table is created based on sequence analysis of highly expressed nuclear genes for the species of interest. Modification of the nucleotide sequence is determined by comparing the species-specific codon usage table with the codons present in the native polynucleotide sequence.

[0160] In some embodiments, the B-GEn.16 polypeptide or variant thereof described herein is expressed from a codon-optimized polynucleotide sequence. For example, if the intended target cell is a human cell, a human codon-optimized polynucleotide sequence encoding B-GEn.16 (or a B-GEn.16 variant, e.g., an enzymatically inactive variant) would be suitable. As another non-limiting example, if the intended host cell is a mouse cell, a mouse codon-optimized polynucleotide sequence encoding B-GEn.16 (or a B-GEn.16 variant, e.g., an enzymatically inactive variant) would be suitable.

[0161] Strategies and methodologies for codon optimization are known in the art and have been described for a variety of systems including, but not limited to, yeast (Outchkourov et al., Protein Expr Purif, 24(1):18-24 (2002)) and Escherichia coli (Feng et al., Biochemistry, 39(50):15399-15409 (2000)). In some embodiments, codon optimization was performed using GeneGPS® Expression Optimization Technology (ATUM) with the expression optimization algorithm recommended by the manufacturer. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in human cells. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in E. coli cells. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in insect cells. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in Sf9 insect cells. In some embodiments, the expression optimization algorithm used in the codon optimization procedure is defined to avoid putative polyA signals (e.g., AATAAA and ATTAAA), as well as long (more than 4) stretches of A that can cause polymerase slippage.

[0162] As is well understood in the art, codon optimization of a nucleotide sequence results in a nucleotide sequence that has less than 100% identity (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or less than 99%) to the native nucleotide sequence but still encodes a polypeptide having the same function as that encoded by the original native nucleotide sequence. Thus, in representative embodiments of the present disclosure, the nucleotide sequences and / or recombinant nucleic acids of the present disclosure can be codon optimized for expression in a particular species of interest.

[0163] In some embodiments, the codon-optimized polynucleotide sequence has at least 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.2%, 99.5%, 99.8%, 99.9%, or 100% sequence identity to SEQ ID NO:1. In some aspects, the polynucleotides of the present disclosure are codon optimized for increased expression of the encoded B-GEn.16 polypeptide in target cells. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in human cells. Generally, the polynucleotides of the present disclosure are codon optimized for increased expression in any human cell. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in E. coli cells. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in insect cells. Generally, the polynucleotides of the present disclosure are codon optimized for increased expression in any insect cell. In some embodiments, the polynucleotides of the present disclosure are codon optimized for increased expression in an Sf9 insect cell expression system.

[0164] The polyadenylation signal can also be selected to optimize expression in the intended host.

[0165] Other modifications Modifications can also be, or alternatively, used to reduce the likelihood or extent to which RNA introduced into a cell induces an innate immune response. Such responses, which are well-characterized in relation to RNA interference (RNAi), including small interfering RNA (siRNA), tend to be associated with a decrease in the half-life of the RNA and / or the induction of cytokines or other factors related to the immune response, as described hereinafter and in the art.

[0166] One or more types of modifications can also be made to RNA encoding an endonuclease such as B-GEn.16 that is introduced into a cell, which include, but are not limited to, modifications that enhance the stability of the RNA (e.g., by reducing its degradation by RNases present in the cell), modifications that enhance the translation of the resulting product (e.g., endonucleases), and / or modifications that reduce the likelihood or extent to which the RNA introduced into the cell induces an innate immune response. Combinations of modifications such as those described above and others can likewise be used. In the case of CRISPR-B-GEn.16, for example, one or more types of modifications can be made to the guide RNA (including those exemplified above), and / or one or more types of modifications can be made to the RNA encoding the B-GEn.16 endonuclease (including those exemplified above).

[0167] For example, guide RNAs used in the CRISPR-B-GEn.16 system or other smaller RNAs can be readily synthesized by chemical means, as exemplified below and described in the art, and allow for the easy incorporation of numerous modifications. While chemical synthesis procedures are constantly expanding, the purification of such RNAs by procedures such as high-performance liquid chromatography (HPLC, which avoids the use of gels such as PAGE) tends to become more difficult as the polynucleotide length increases significantly beyond about 100 nucleotides. One approach used to generate longer lengths of chemically modified RNAs is to generate two or more molecules that are ligated together. Much longer RNAs, such as those encoding the B-GEn.16 endonuclease, are more readily generated enzymatically. Fewer types of modifications are generally available for use in enzymatically produced RNAs, but as described further below and in the art, there are still modifications that can be used, for example, to enhance stability, reduce the likelihood or extent of the innate immune response, and / or enhance other attributes, and new types of modifications are being developed regularly. As examples of various types of modifications, particularly those frequently used with smaller chemically synthesized RNAs, the modifications can include one or more nucleotides modified at the 2'-position of the sugar, in some embodiments 2'-O-alkyl, 2'-O-alkyl-O-alkyl or 2'-fluoro modified nucleotides. In some embodiments, the RNA modifications include 2'-fluoro, 2'-amino, and 2'O-methyl modifications on the ribose of pyrimidines, basic residues, or reverse bases at the 3'-end of the RNA. Such modifications are routinely incorporated into oligonucleotides, and these oligonucleotides have been shown to have a higher Tm (e.g., higher target binding affinity) than 2'-deoxyoligonucleotides for a given target.

[0168] Many nucleotide and nucleoside modifications have been shown to make the oligonucleotides into which they are incorporated more resistant to nuclease digestion than natural oligonucleotides; these modified oligonucleotides survive intact for longer periods than unmodified oligonucleotides. Specific examples of modified oligonucleotides include those with modified backbones, such as phosphorothioates, phosphotriesters, methylphosphonates, short-chain alkyl or cycloalkyl sugar linkages, or short-chain heteroatom or heterocyclic sugar linkages. Some oligonucleotides have oligonucleotides with phosphorothioate backbones, heteroatom backbones, in particular CH2-NH-O-CH2, CH, -N(CH3)-O-CH2 (known as methylene(methylimino) or MMI backbone), CH2-O-N(CH3)-CH2, CH2-N(CH3)-N(CH3)-CH2 and O-N(CH3)-CH2-CH2 backbones; amide backbones [see De Mesmaeker et al., Ace. Chem. Res., 28:366-374 (1995)]; morpholino backbone structures [see Summerton and Weller, U.S. Patent No. 5,034,506]; peptide nucleic acid (PNA) backbones (wherein the phosphodiester backbone of the oligonucleotide is replaced by a polyamide backbone and the nucleic acids are attached directly or indirectly to the aza nitrogen atoms of the polyamide backbone, see Nielsen et al., Science 1991, 254, 1497).Examples of phosphorus-containing linkages include, but are not limited to, phosphorothioate, chiral phosphorothioate, phosphorodithioate, phosphotriester, aminoalkyl phosphotriester, methyl and other alkyl phosphonates (including 3'-alkylene phosphonates and chiral phosphonates), phosphinate, phosphoramidate (including 3'-aminophosphoramidate and aminoalkyl phosphoramidate), thionophosphoramidate, thionoalkyl phosphonate, thionoalkyl phosphotriester, and boranophosphate having a normal 3'-5' linkage, 2'-5' linkage analogs thereof, and those with inverted polarity where adjacent pairs of nucleoside units are linked 3'-5' and 5'-3' or 2'-5' and 5'-2'; see U.S. Patent Nos. 3,687,808; 4,469,863; 4,476,301; 5,023,243; 5,177,196; 5,188,897; 5,264,423; 5,276,019; 5,278,302; 5,286,717; 5,321,131; 5,399,676; 5,405,939; 5,453,496; 5,455,233; 5,466,677; 5,476,925; 5,519,126; 5,536,821; 5,541,306; 5,550,111; 5,563,253; 5,571,799; 5,587,361; and 5,625,050.

[0169] Morpholino-based oligomeric compounds are described in Braasch and Corey, Biochemistry, 41(14):4503-4510(2002); Genesis, Volume 30, Issue 3, (2001); Heasman, Dev. Biol., 243:209-214(2002); Nasevicius et al., Nat. Genet., 26:216-220(2000); Lacenra etc., Proc. Nat / . Acad. Sci., 97:9591-9596(2000); and U.S. Patent No. 5,034,506 issued on July 23, 1991. Cyclohexenyl nucleic acid oligonucleotide mimics are described in Wang et al., J. Am, Chem. Soc., 122:8595-8602(2000).

[0170] The modified oligonucleotide backbone that does not contain phosphorus atoms therein has a backbone formed by short-chain alkyl or cycloalkyl nucleoside linkages, mixed heteroatom and alkyl or cycloalkyl nucleoside linkages, or one or more short-chain heteroatom or heterocyclic nucleoside linkages. These include those having a morpholino linkage (formed in part from the sugar moiety of the nucleoside); a siloxane backbone; sulfide, sulfoxide and sulfone backbones; formacetyl and thioformacetyl backbones; methyleneformacetyl and thioformacetyl backbones; methyleneacetyl and thioformacetyl backbones; alkene-containing backbones; sulfamate backbones; methyleneimino and methylenehydrazino backbones; sulfonate and sulfonamide backbones; amide backbones; and others having mixed N, O, S and CH2 component moieties; see U.S. Pat. Nos. 5,034,506; 5,166,315; 5,185,444; 5,214,134; 5,216,141; 5,235,033; 5,264,562; 5,264,564; 5,405,938; 5,434,257; 5,466,677; 5,470,967; 5,489,677; 5,541,307; 5,561,225; 5,596,086; 5,602,240; 5,610,289; 5,602,240; 5,608,046; 5,610,289; 5,618,704; 5,623,070; 5,663,312; 5,633,360; 5,677,437; and 5,677,439; each of which is incorporated herein by reference.

[0171] One or more substituted sugar moieties can also be included, for example, at the 2'-position the following: OH, SH, SCH3, F, OCN, OCH3, OCH3O(CH2) n CH3, O(CH2) n NH2 or O(CH2) n CH3 (where n is 1 to 10); C1 to C 10Lower alkyl, alkoxyalkoxy, substituted lower alkyl, aralkyl or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; SOCH3; SO2CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; RNA cleavage group; reporter group; intercalator; a group for improving the pharmacokinetic properties of an oligonucleotide; a group for improving the pharmacodynamic properties of an oligonucleotide or other substituents having similar properties. In some embodiments, the modification includes 2'-methoxyethoxy (also known as 2'-O-CH2CH2OCH3, 2'-O-(2-methoxyethyl)) (Martinet a / , Helv. Chim. Acta, 1995, 78, 486). Other modifications include 2'-methoxy (2'-O-CH3), 2'-propoxy (2'-OCH2CH2CH3) and 2'-fluoro (2'-F). Similar modifications can also be made at other positions on the oligonucleotide, particularly at the 3'-position of the sugar on the 3'-terminal nucleotide and at the 5'-position of the 5'-terminal nucleotide. The oligonucleotide may also have a sugar mimic such as cyclobutyl instead of a pentofuranosyl group. In some embodiments, both the sugar of the nucleotide unit and the internucleoside linkage, e.g., the backbone, are replaced with novel groups. The base units are maintained for hybridization with a suitable nucleic acid target compound. One such oligomeric compound, an oligonucleotide mimic that has been shown to have excellent hybridization properties, is called a peptide nucleic acid (PNA). In a PNA compound, the sugar backbone of the oligonucleotide is replaced with an amide-containing backbone, e.g., an aminoethylglycine backbone. The nucleic acid bases are retained and are directly or indirectly attached to the aza nitrogen atoms of the amide portion of the backbone. Representative U.S. patents that teach the preparation of PNA compounds include, but are not limited to, U.S. Pat. Nos. 5,539,082; 5,714,331; and 5,719,262. Further teachings of PNA compounds can be found in Nielsen et al., Science, 254:1497-1500 (1991).

[0172] Guide RNAs can also additionally or alternatively include nucleobase (often simply referred to as "base" in the art) modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleobases include nucleobases that are rare or transiently found in natural nucleic acids, such as hypoxanthine, 6-methyladenine, 5-Me pyrimidines, particularly 5-methylcytosine (also called 5-methyl-2'-deoxycytosine and often referred to as 5-Me-C in the art), 5-hydroxymethylcytosine (HMC), glycosyl HMC and gentobiosyl HMC, and synthetic nucleobases such as 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolylalkyl)adenine, 2-(aminoalkylamino)adenine or other hetero-substituted alkyladenines, 2-thiouracil, 2-thiothymine, 5-bromouracil, 5-hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N6(6-aminohexyl)adenine, and 2,6-diaminopurine. Kornberg, A, DNA Replication, W.H. Freeman & Co., San Francisco, pp75-77 (1980); Gebeyehu et al., Nucl. Acids Res. 15:4513 (1997). "Universal" bases known in the art, such as inosine, can also be included. 5-Me-C substitution has been shown to increase the stability of nucleic acid duplexes by 0.6 - 1.2 °C (Sanghvi, Y.S., in Crooke, S.T. and Lebleu, B., eds., Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276-278) and is an embodiment of base substitution.

[0173] Modified nucleobases include other synthetic and natural nucleobases, 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azouracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thio, 8-thioalkyl, 8-hydroxy and other substituted adenines and guanines, 5-halo especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylquanine, 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaadenine and 3-deazaguanine and 3-deazaadenine, and the like.

[0174] Other useful nucleobases include those disclosed in U.S. Patent No. 3,687,808; those disclosed in "The Concise Encyclopedia of Polymer Science And Engineering", pages 858 - 859, edited by Kroschwitz, J.I., John Wiley & Sons, 1990; those disclosed in Englisch et al., Angewandte Chemie, International Edition, 1991, 30, page 613; and those disclosed in Sanghvi, Y.S., Chapter 15, Antisense Research and Applications, pages 289 - 302, Crooke, T. and Lebleu, B., eds., CRC Press, 1993. Some of these nucleobases are particularly useful for increasing the binding affinity of the oligomeric compounds of the present disclosure. These include 5 - substituted pyrimidines, 6 - azapyrimidines, and N - 2, N - 6, and O - 6 substituted purines (including 2 - aminopropyladenine, 5 - propynyluracil, and 5 - propynylcytosine). 5 - methylcytosine substitution has been shown to increase nucleic acid duplex stability by 0.6 - 1.2 °C (Sanghvi, Y.S., Crooke, S.T. and Lebleu, B., eds, Antisense Research and Applications, CRC Press, Boca Raton, 1993, pp. 276 - 278) and, more particularly, is an embodiment of base substitution when combined with 2’ - O - methoxyethyl sugar modification. Modified nucleobases are described in U.S. Patent Nos. 3,687,808; 5,130,302; 5,134,066; 5,175,273; 5,176; 5,175,266; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; 5,525,711; 5,552,540.

[0175] Not all positions in a given oligonucleotide need to be uniformly modified, and in fact, two or more of the above-described modifications may be incorporated into a single oligonucleotide or within a single nucleoside within the oligonucleotide.

[0176] In some embodiments, the guide RNA and / or the mRNA encoding an endonuclease such as B-GEn.16 of the present disclosure is capped using any one of current capping methods such as mCAP, ARCA, or enzymatic capping methods to produce a viable mRNA construct that maintains biological activity and avoids self / non-self intracellular responses. In some embodiments, the guide RNA and / or the mRNA encoding an endonuclease such as B-GEn.16 of the present disclosure is capped by using the CleanCap™ (Trilink) co-transcriptional capping method.

[0177] In some embodiments, the guide RNA and / or the mRNA encoding the endonuclease of the present invention comprises one or more modifications selected from the group consisting of pseudouridine, N 1 -methylpseudouridine, and 5-methoxyuridine. In some embodiments, one or more N 1 -methylpseudouridines are incorporated into the guide RNA and / or the mRNA encoding the endonuclease of the present disclosure to provide enhanced RNA stability and / or protein expression and reduced immunogenicity in animal cells such as mammalian cells (e.g., human and mouse). In some embodiments, the N 1 -methylpseudouridine modification is incorporated in combination with one or more 5-methylcytidines.

[0178] In some embodiments, the guide RNA and / or the mRNA (or DNA) encoding an endonuclease such as B-GEn.16 is chemically linked to one or more moieties or conjugates that enhance the activity, cellular distribution, or cellular uptake of the oligonucleotide.Such moieties include, but are not limited to, cholesterol moieties [Letsinger et al., Proc. Natl. Acad. Sci. USA, 86:6553-6556 (1989)]; cholic acid [Manoharan et al., Bioorg. Med. Chem. Lett., 4:1053-1060 (1994)]; thioethers such as hexyl-S-tritylthiol [Manoharan et al., Ann. N.Y Acad. Sci., 660:306-309 (1992) and Manoharan et al., Bioorg. Med. Chem. Lett., 3:2765-2770 (1993)]; thiocolesterol [Oberhauser et al., Nucl. Acids Res., 20:533-538 (1992)]; aliphatic chains such as dodecanediol or undecyl residues [Kabanov et al., FEBS Lett., 259:327-330 (1990) and Svinarchuk et al., Biochimie, 75:49-54 (1993)]; phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate [Manoharan et al., Tetrahedron Lett., 36:3651-3654 (1995) and Shea et al., Nucl. Acids Res., 18:3777-3783 (1990)]; polyamines or polyethylene glycol chains [Mancharan et al., Nucleosides & Nucleotides, 14:969-973 (1995)]; adamantaneacetic acid [Manoharan et al., Tetrahedron Lett., 36:3651-3654 (1995)]; palmitoyl moieties [Mishra et al., Biochim. Biophys. Acta, 1264:229-237 (1995)]; or octadecylamine or hexylamino-carbonyl-t oxy cholesterol moieties [Crooke et al., J. Pharmacol. Exp. Ther., 277:923-937 (1996)].See U.S. Pat. Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241; 5,391,723; 5,416,203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0179] Sugars and other moieties can be used to target proteins and complexes containing nucleotides, such as cationic polysomes and liposomes, to specific sites. For example, hepatocyte-directed uptake can be mediated via the asialoglycoprotein receptor (ASGPR); see, e.g., Hu et al., Protein Pept Lett. 21(10):1025-30 (2014). Other systems known in the art and regularly developed can be used to target the biomolecules and / or their complexes used in this case to specific target cells of interest.

[0180] These targeted moieties or conjugates can include conjugate groups covalently attached to functional groups such as primary or secondary hydroxyl groups. Suitable conjugate groups include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacodynamic properties of the oligomer, and groups that enhance the pharmacokinetic properties of the oligomer. Exemplary conjugate groups include cholesterol, lipids, phospholipids, biotin, phenazine, folate, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that can enhance pharmacodynamic properties include groups that improve uptake, enhance resistance to degradation, and / or enhance sequence-specific hybridization to the target nucleic acid. Groups that can enhance pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of the compounds of the present disclosure. Representative conjugate groups are disclosed in International Patent Application No. PCT / US92 / 09196, filed Oct. 23, 1992, and U.S. Patent No. 6,287,860, which are incorporated herein by reference. Conjugate moieties include, but are not limited to, lipid moieties such as cholesterol moieties, cholic acid, thioethers such as hexyl-5-tritylthiol, thiocolesterol, aliphatic chains such as dodecanediol or undecyl residues, phospholipids such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H phosphonate, polyamine or polyethylene glycol chains, or adamantane acetic acid, palmitoyl moieties, or octadecylamine or hexylamino-carbonyl-oxy cholesterol moieties.For example, see U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414,077; 5,486,603; 5,512,439; 5,578,718; 5,608,046; 4,587,044; 4,605,735; 4,667,025; 4,762,779; 4,789,737; 4,824,941; 4,835,263; 4,876,335; 4,904,582; 4,958,013; 5,082,830; 5,112,963; 5,214,136; 5,082,830; 5,112,963; 5,214,136; 5,245,022; 5,254,469; 5,258,506; 5,262,536; 5,272,250; 5,292,873; 5,317,098; 5,371,241; 5,391,723; 5,416,203; 5,451,463; 5,510,475; 5,512,667; 5,514,785; 5,565,552; 5,567,810; 5,574,142; 5,585,481; 5,587,371; 5,595,726; 5,597,696; 5,599,923; 5,599,928 and 5,688,941.

[0181] Even polynucleotides that are less susceptible to chemical synthesis and are generally produced by enzymatic synthesis can be modified by various means. Such modifications can include, for example, the introduction of specific nucleotide analogs, the incorporation of specific sequences or other moieties at the 5' or 3' ends of the molecule, and other modifications. As an example, the mRNA encoding B-GEn.16 is about 4 kb in length and can be synthesized by in vitro transcription. Modifications to the mRNA can be applied, for example, to increase its translation or stability (e.g., by increasing its resistance to degradation by cells), or to reduce the tendency of RNA to induce an innate immune response that is often observed in cells after the introduction of exogenous RNA, particularly longer RNAs such as those encoding B-GEn.16.

[0182] A number of such modifications have been described in the art and include, for example, the use of a polyA tail, a 5' cap analog (e.g., an anti-reverse cap analog (ARCA) or m7G(5')ppp(5')G (mCAP)), a modified 5' or 3' untranslated region (UTR), a modified base (e.g., pseudo-UTP, 2-thio-UTP, 5-methylcytidine-5'-triphosphate (5-methyl-CTP) or N6-methyl-ATP), or treatment with a phosphatase to remove the 5' end phosphate. These and other modifications are known in the art and new modifications of RNA are being developed on a regular basis.

[0183] There are a number of commercial suppliers of modified RNAs, including, for example, TriLink Biotech, Axolabs, Bio-Synthesis Inc., Dharmacon, and many others. As described by TriLink, for example, 5-methyl-CTP can be used to impart desirable features such as increased nuclease stability, increased translation, or decreased interaction of innate immune receptors with in vitro transcribed RNA. 5'-methylcytidine-5'-triphosphate (5-methyl-CTP), N6-methyl-ATP, and pseudo-UTP and 2-thio-UTP have also been shown to enhance translation while reducing innate immune stimulation in culture and in vivo, as shown in the publications by Konmann et al. and Warren et al. referred to below.

[0184] Chemically modified mRNAs delivered in vivo have been shown to be able to be used to achieve improved therapeutic effects; see, for example, Kormann et al., Nature Biotechnology 29, 154-157 (2011). Such modifications can be used, for example, to increase the stability of the RNA molecule and / or to reduce its immunogenicity. Using chemical modifications such as pseudo-U, N6-methyl-A, 2-thio-U, and 5-methyl-C, replacement of only one quarter of the uridine residues and cytidine residues with 2-thio-U and 5-methyl-C, respectively, has been found to significantly reduce the toll-like receptor (TLR)-mediated recognition of mRNA in mice. Thus, by reducing the activation of the innate immune system, these modifications can be used to effectively increase the stability and lifespan of mRNA in vivo; see, for example, Konmann et al., supra.

[0185] Furthermore, repeated administration of synthetic messenger RNAs incorporating modifications designed to bypass the innate antiviral response has also been shown to be able to reprogram differentiated human cells to pluripotency. See, for example, Warren et al., Cell Stem Cell, 7(5):618-30 (2010). Such modified mRNAs that act as primary reprogramming proteins could be an efficient means of reprogramming multiple human cell types. Such cells are referred to as induced pluripotent stem cells (iPSCs), and it has been found that enzymatically synthesized RNAs incorporating 5-methyl-CTP, pseudo-UTP, and anti-reverse cap analog (ARCA) can be used to effectively evade the cellular antiviral response; see, for example, Warren et al., supra. Other modifications of polynucleotides described in the art include, for example, the use of polyA tails, the addition of 5' cap analogs (e.g., m7G(5')ppp(5')G (mCAP)), modification of 5' or 3' untranslated regions (UTRs), or treatment with phosphatase to remove the 5' terminal phosphate, and new approaches are being developed regularly.

[0186] Many compositions and techniques applicable to the generation of modified RNAs for use herein have been developed in connection with the modification of RNA interference (RNAi), including small interfering RNAs (siRNAs). siRNAs present certain challenges in vivo because their effects on gene silencing via mRNA interference are generally transient and may require repeated dosing. In addition, siRNAs are double-stranded RNAs (dsRNAs), and mammalian cells have an immune response that has evolved to detect and neutralize dsRNAs, which are often by-products of viral infection. Thus, mammalian enzymes such as PKR (dsRNA-responsive kinase) and potentially retinoic acid-inducible gene I (RIG-I) can mediate cellular responses to dsRNA, similar to Toll-like receptors (such as TLR3, TLR7, and TLR8) being able to induce cytokine induction in response to such molecules; see, for example, reviews by Angart et al., Pharmaceuticals (Basel) 6(4):440-468 (2013); Kanasty et al., Molecular Therapy 20(3):513-524 (2012); Burnett et al., Biotechnol J. 6(9):1130-46 (2011); Judge and Maclachlan, Hum Gene Ther 19(2):111-24 (2008); and the references cited therein.

[0187] A variety of modifications have been developed to enhance RNA stability, reduce the innate immune response, and / or achieve other benefits that may be useful in the context of introducing polynucleotides into human cells as described herein; see, for example, reviews by Whitehead KA et al., Annual reviews of Chemical and Biomolecular Engineering, 2:77-96 (2011); Gaglione and Messere, Mini Rev Med Chem, 10(7):578-95 (2010); Chernolovskaya et al., Curr Opin Mol Ther., 12(2):158-67 (2010); Deleavey et al., Curr Protoc Nucleic Acid Chem Chapter 16:Unit 16.3 (2009); Behlke, Oligonucleotides 18(4):305-19 (2008): Fucini et al., Nucleic Acid Ther 22(3):205-210 (2012); Bremsen et al., Front Genet 3:154 (2012).

[0188] As described above, there are many commercial suppliers of modified RNAs, many of which specialize in modifications designed to improve the effectiveness of siRNAs. Based on the various findings reported in the literature, various approaches have been provided. For example, Dharmacon states that, as reported by Kale, Nature Reviews Drug Discovery 11:125-140 (2012), substituting non-bridging oxygen with sulfur (phosphorothioate, PS) is widely used to improve the nuclease resistance of siRNAs. Modifications at the 2'-position of ribose have been reported to improve the nuclease resistance of internucleotide phosphate linkages while increasing duplex stability (Tm), which has also been shown to provide protection from immune activation. As reported by Soutschek et al., Nature 432:173-178 (2004), the combination of moderate PS backbone modifications and small, well-tolerated 2'-substitutions (2'-O-, 2'-fluoro, 2'-hydro) is associated with highly stable siRNAs for in vivo applications; and 2'-O-methyl modifications have been reported to be effective in improving stability, as reported by Volkov, Oligonucleotides 19:191-202 (2009). With regard to reducing the induction of the innate immune response, modifying specific sequences with 2'-O-methyl, 2'-fluoro, 2'-hydro has been reported to reduce TLR7 / TLR8 interactions while generally preserving silencing activity; see, for example, Judge et al., Mol. Ther. 13:494-505 (2006); and Cekaite et al., J. Mol. Biol. 365:90-108 (2007). Further modifications such as 2-thiouracil, pseudouracil, 5-methylcytosine, 5-methyluracil, and N6-methyladenosine have also been shown to minimize the immune effects mediated by TLR3, TLR7, and TLR8; see, for example, Kariko et al., Immunity 23:165-175 (2005).

[0189] Also, as known in the art and commercially available, a number of conjugates can be applied to polynucleotides such as RNA for use herein that can enhance their delivery and / or uptake by cells, including, for example, cholesterol, tocopherol and folic acid, lipids, peptides, polymers, linkers and aptamers; see, for example, the review by Winkler, Ther. Deliv. 4:791-809 (2013) and the references cited therein.

[0190] Additional sequence In some embodiments, the guide RNA includes at least one additional segment at either the 5' or 3' end. For example, suitable additional segments include a 5' cap (e.g., 7-methylguanylate cap (m7g)); a 3' polyadenylation tail (e.g., 3' poly(A) tail); riboswitch sequences (e.g., enabling regulated stability and / or regulated accessibility by proteins and protein complexes); sequences that form dsRNA duplexes (e.g., hairpins); sequences that target the RNA to an intracellular location (e.g., nucleus, mitochondria, chloroplast, etc.); modifications or sequences that provide for tracking (e.g., direct conjugation to a fluorescent molecule, direct conjugation to a moiety that facilitates fluorescence detection, sequences that enable fluorescence detection, etc.); modifications or sequences that provide a binding site for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); modifications or sequences that provide increased, decreased, and / or controllable stability; and combinations thereof.

[0191] Stability control sequence The stability control sequence affects the stability of RNA (e.g., guide RNA). Non-limiting examples of suitable stability control sequences are transcription terminator segments (e.g., transcription termination sequences). The transcription terminator segment of the guide RNA can have a full length of 10 nucleotides to 100 nucleotides, e.g., 10 nucleotides (nt) to 20 nt, 20 nt to 30 nt, 30 nt to 40 nt, 40 nt to 50 nt, 50 nt to 60 nt, 60 nt to 70 nt, 70 nt to 80 nt, 80 nt to 90 nt, or 90 nt to 100 nt. For example, the transcription terminator segment can have a length of 15 nucleotides (nt) to 80 nt, 15 nt to 50 nt, 15 nt to 40 nt, 15 nt to 30 nt, or 15 nt to 25 nt.

[0192] In some embodiments, the transcription termination sequence is a sequence that is functional in eukaryotic cells. In some embodiments, the transcription termination sequence is a sequence that is functional in prokaryotic cells.

[0193] Examples of nucleotide sequences that can be included in the stability control sequence (e.g., transcription termination segment, or any segment of the guide RNA that provides an increase in stability) include, for example, the Rho-independent trp termination site.

[0194] Mimetic In some embodiments, the nucleic acid can be a nucleic acid mimetic. The term "mimetic" as applied to polynucleotides is intended to include polynucleotides in which only the furanose ring, or both the furanose ring and the internucleotide linkage, are replaced with non-furanose groups, and replacement of only the furanose ring is also referred to in the art as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety is maintained for hybridization with a suitable target nucleic acid. A polynucleotide mimetic that has been shown to have excellent hybridization properties in one such nucleic acid is called a peptide nucleic acid (PNA). In PNA, the sugar backbone of the polynucleotide is replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotides are retained and are attached directly or indirectly to the aza nitrogen atom of the amide portion of the backbone.

[0195] One polynucleotide mimetic that has been reported to have excellent hybridization properties is peptide nucleic acid (PNA). The backbone in a PNA compound is two or more linked aminoethylglycine units, which gives PNA an amide-containing backbone. The heterocyclic base moiety is attached directly or indirectly to the aza nitrogen atom of the amide portion of the backbone. Representative U.S. patents that describe the preparation of PNA compounds include, but are not limited to, U.S. Patent Nos. 5,539,082; 5,714,331; and 5,719,262.

[0196] Another class of polynucleotide mimics that has been studied is based on linked morpholino units (morpholino nucleic acids) having a heterocyclic base attached to a morpholino ring. A number of linking groups that connect the morpholino monomer units in morpholino nucleic acids have been reported. One class of linking groups has been selected to give nonionic oligomeric compounds. Nonionic morpholino-based oligomeric compounds are less likely to have undesirable interactions with cellular proteins. Morpholino-based polynucleotides are nonionic mimics of oligonucleotides and are less likely to form undesirable interactions with cellular proteins (Dwaine A. Braasch and David R. Corey, Biochemistry, 2002, 41(14), 4503-4510). Morpholino-based polynucleotides are disclosed in U.S. Patent No. 5,034,506. Various compounds within the morpholino class of polynucleotides having different linking groups that connect the monomer subunits have been prepared.

[0197] A further class of polynucleotide mimics is referred to as cyclohexenyl nucleic acid (GeNA). The furanose ring that is normally present in DNA / RNA molecules is replaced by a cyclohexenyl ring. GeNA DMT-protected phosphoramidite monomers were prepared and used in the synthesis of oligomeric compounds according to classical phosphoramidite chemistry. Completely modified GeNA oligomeric compounds and oligonucleotides having specific positions modified with GeNA have been prepared and studied (see Wang et al., J. Am. Chem. Soc., 2000, 122, 8595-8602). In general, the incorporation of GeNA monomers into a DNA strand increases the stability of the DNA / RNA hybrid. GeNA oligoadenylates formed complexes with RNA and DNA complexes having a stability similar to that of native complexes. Studies incorporating the GeNA structure into native nucleic acid structures have shown, by NMR and circular dichroism, facile conformational adaptation.

[0198] Further modifications include locked nucleic acids (LNAs) in which the 2'-hydroxyl group is linked to the 4'-carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage and thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group that bridges the 2'-oxygen atom and the 4'-carbon atom, where n is 1 or 2 (Singh et al., Chem. Commun., 1998, 4, 455-456). LNAs and LNA analogs exhibit very high duplex thermal stability (Tm = +3 to +10 °C), stability against 3'-exonuclease degradation, and good solubility properties with complementary DNA and RNA. Potent and non-toxic antisense oligonucleotides containing LNAs have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 2000, 97, 5633-5638).

[0199] The synthesis and preparation of the LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine, and uracil, along with their oligomerization and nucleic acid recognition properties, have been described (Koshkin et al., Tetrahedron, 1998, 54, 3607-3630). LNA and its preparation are also described in WO98 / 39352 and WO99 / 14226.

[0200] Modified sugar moiety Nucleic acids can also contain one or more substituted sugar moieties. Suitable polynucleotides include sugar substituents selected from OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S- or N-alkynyl; or O-alkyl-O-alkyl, where alkyl, alkenyl and alkynyl can be substituted or unsubstituted C1-C10 alkyl or C2-C10 alkenyl and alkynyl. Particularly suitable are O((CH2)nO)mCH3, O(CH2)nOCH3, O(CHz)nNH2, O(CH2)CH3, O(CH2)nONH2, and O(CH2)nON((CH2)nCH3)2, where n and m are from 1 to about 10. Other suitable polynucleic acids include C1 to C10 lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage groups, reporter groups, intercalators, groups for improving the pharmacokinetic properties of oligonucleotides, or groups for improving the pharmacodynamic properties of oligonucleotides, and other substituents having similar properties. Suitable modifications include 2'-methoxyethoxy 2'-O-CH2-CH2OCH3, -2'-O-(2-methoxyethyl) or also known as 2'-MOE (Martin et al., Hely.Chim.Acta, 1995, 78, 486-504), for example, alkoxyalkoxy groups are included. Further suitable modifications are 2'-dimethylaminooxyethoxy, for example, the O(CH2)2ON(CH3)2 group (2'-DMAOE) as described in the following examples herein, and 2'-dimethylaminoethoxyethoxy (also known in the art as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), for example, 2'-O-CH2-O-CH2-N(CH3)2.

[0201] Other suitable sugar substituents include methoxy (-O-CH3), aminopropoxy (-O-CH2CH2CH2NH2), allyl (-CH2-CH=CH2), -O-allyl (-O-CH2-CH=CH2), and fluoro (F). The 2'-sugar substituent can be in the arabino (up) or ribo (down) position. A preferred 2'-arabino modification is 2'-F. Similar modifications can also be made at other positions on the oligomeric compound, particularly at the 3'-end nucleoside or at the 3'-position of the sugar in a 2'-5'-linked oligonucleotide and at the 5'-position of the 5'-end nucleotide. The oligomeric compound can also have a sugar mimic such as a cyclobutyl moiety instead of a pentofuranosyl sugar.

[0202] Modification and substitution of bases Nucleic acids can also include nucleic acid base (often simply referred to as "base" in the art) modifications or substitutions. As used herein, "unmodified" or "natural" nucleic acid bases include the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleic acid bases include 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C≡CH) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azauracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine and other synthetic and natural nucleic acid bases. Further modified nucleic acid bases include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3’,2’:4,5)pyrrolo(2,3-d)pyrimidin-2-one).

[0203] The heterocyclic base moiety may also include those in which the purine or pyrimidine base is replaced by another heterocycle, such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone. Further nucleobases include those disclosed in U.S. Patent No. 3,687,808, those disclosed in The Concise Encyclopedia Of Polymer Science And Engineering, pages 858-859, Kroschwitz, J.I., editor, John Wiley & Sons, 1990, those disclosed by Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613, and those disclosed by Sanghvi, Y.S., Chapter 15, Antisense Research and Applications, pages 289-302, Crooke, S.T. and Lebleu, B., editors, CRC Press, 1993. Some of these nucleobases are useful for increasing the binding affinity of oligomeric compounds. These include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines (including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine). 5-Methylcytosine substitution has been shown to increase the stability of nucleic acid duplexes by 0.6-1.2 °C (Sanghvi et al., editors, Antisense Research and Applications, CRC Press, Boca Raton, 1993, pages 276-278), and is a suitable base substitution, for example, when combined with 2'-O-methoxyethyl sugar modification.

[0204] "Complementary" refers to the ability to pair between two sequences containing natural or non-natural (e.g., modified as described above) bases (nucleosides) or analogs thereof through base stacking and specific hydrogen bonding. For example, if the base at one position of a nucleic acid can hydrogen bond with the base at the corresponding position of a target, the bases are considered to be complementary to each other at that position. Nucleic acids can include universal bases or inert abasic spacers that do not contribute positively or negatively to hydrogen bonding. Base pairing can include both standard Watson-Crick base pairing and non-Watson-Crick base pairing (e.g., wobble base pairing and Hoogsteen base pairing).

[0205] With respect to complementary base pairing, adenosine-type bases (A) are complementary to thymidine-type bases (T) or uracil-type bases (U), cytosine-type bases (C) are complementary to guanosine-type bases (G), and universal bases such as 3-nitropyrrole or 5-nitroindole can hybridize with any A, C, U, or T and are considered to be complementary thereto. See Nichols et al., Nature, 1994; 369:492-493 and Loakes et al., Nucleic Acids Res, 1994; 22:4039-4043. Inosine (I) is also considered in the art to be a universal base and is considered to be complementary to any A, C, U, or T. See Watkins and Santalucia, Nucl. Acids Research, 2005; 33(19):6258-6267.

[0206] Conjugate Another possible modification of the nucleic acid involves chemically linking to the polynucleotide one or more moieties or conjugates that enhance the activity, cellular distribution, or cellular uptake of the oligonucleotide. These moieties or conjugates can include conjugate groups covalently bound to a functional group such as a primary or secondary hydroxyl group. Conjugate groups include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycol, polyethers, groups that enhance the pharmacodynamic properties of the oligomer, and groups that enhance the pharmacokinetic properties of the oligomer. Suitable conjugate groups include, but are not limited to, cholesterol, lipids, phospholipids, biotin, phenazine, folate, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that enhance the pharmacodynamic properties include groups that improve uptake, enhance resistance to degradation, and / or enhance sequence-specific hybridization to the target nucleic acid. Groups that enhance the pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of the nucleic acid.

[0207] The conjugate moiety includes, but is not limited to, a cholesterol moiety (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, 6553-6556), a cholic acid moiety (Manoharan et al., Bioorg. Med. Let., 1994, 4, 1053-1060), a thioether, such as hexyl-S-tritylthiol (Manoharan et al., Ann. N.Y. Acad. Sci., 1992, 660, 306-309; Manoharan et al., Bioorg. Med. Chem. Let., 1993, 3, 2765-2770), thiocolesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, 533-538), an aliphatic chain, such as dodecanediol or undecyl residue (Saison-Behmoaras et al., EMBO J., 1991, 10, 1111-1118; Kabanov et al., FEBS Lett., 1990, 259, 327-330; Svinarchuk et al., Biochimie, 1993, 75, 49-54), a phospholipid, such as di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651-3654; Shea et al., Nucl. Acids Res., 1990, 18, 3777-3783), a polyamine or a polyethylene glycol chain (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, 969-973), or an adamantaneacetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651 3654), a palmitoyl moiety (Mishra et al., Biochim. Biophys. Acta, 1995, 1264, 229-237), or an octadecylamine or hexylamino-carbonyl-oxy cholesterol moiety (Crooke et al., J. Pharmacal. Exp. Ther., 1996, 277, 923-937).

[0208] The conjugate may include a "protein transduction domain" or PTD (also known as a CPP cell membrane permeable peptide), which may refer to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates passage through a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD conjugated to another molecule can span from small polar molecules to large macromolecules and / or nanoparticles, for example facilitating the passage of molecules through membranes from the extracellular space into the intracellular space or from the cytosol into an organelle. In some embodiments, the PTD is covalently attached to the amino terminus of an exogenous polypeptide (e.g., the B-GEn.16 polypeptide or a variant thereof). In some embodiments, the PTD is covalently attached to the C-terminus or N-terminus of an exogenous polypeptide (e.g., the B-GEn.16 polypeptide or a variant thereof). In some embodiments, the PTD is covalently attached to a nucleic acid (e.g., a guide RNA, a polynucleotide encoding a guide RNA, a polynucleotide encoding the B-GEn.16 polypeptide or a variant thereof, etc.). Exemplary PTDs include, but are not limited to, the minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT containing YGRKKRRQRRR); polyarginine sequences containing a sufficient number of arginines to directly penetrate cells (e.g., 3, 4, 5, 6, 7, 8, 9, 10 or 10-50 arginines); the VP22 domain (Zender et al., (2002) Cancer Gene Ther. 9(6):489-96); the Drosophila antennapedia protein transduction domain (Noguchi et al., (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al., (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al., (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); in some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al., (2009) lntegr Biol (Camb) June;1(5-6):371-381).ACPP contains a polycationic CPP (e.g., Arg9 or "R9") connected via a cleavable linker to a matching polyanion (e.g., Glu9 or "E9"), which reduces the net charge to near zero, thereby inhibiting adhesion and uptake into cells. When the linker is cleaved, the polyanion is released, locally unmasking polyarginine and its inherent adhesiveness, and thus "activating" the ACPP to cross the membrane. In some embodiments, the PTD is chemically modified to increase the bioavailability of the PTD. See, for example, Expert Opin Drug Deliv. 2009 Nov;6(11):1195-205.

[0209] Polypeptide modification The B-GEn.16 polypeptide or variant thereof expressed from a codon-optimized polynucleotide sequence can be produced in vitro, or by eukaryotic cells, by prokaryotic cells, or by in vitro transcription and translation (IVTT), and it can be further processed by unfolding, such as by heat denaturation, OTT reduction, etc., and can be refolded further using methods known in the art.

[0210] Modifications for the purpose of not changing the primary sequence include chemical derivatization of the polypeptide, such as acylation, acetylation, carboxylation, amidation, etc. Also included are modifications of glycosylation, such as by modifying the glycosylation pattern of the polypeptide during its synthesis and processing, or in further processing steps, such as by exposing the polypeptide to enzymes that affect glycosylation, such as mammalian glycosylation or deglycosylation enzymes. Also included are sequences having phosphorylated amino acid residues, such as phosphotyrosine, phosphoserine, or phosphothreonine.

[0211] In some embodiments, the B-GEn.16 polypeptide or a variant thereof is modified using conventional molecular biological techniques and synthetic chemistry to improve their resistance to proteolysis, alter target sequence specificity, optimize solubility characteristics, change protein activity (e.g., transcriptional regulatory activity, enzymatic activity, etc.), or make them more suitable as therapeutic agents. Analogs of such polypeptides include those containing residues other than the naturally occurring L-amino acids, such as O-amino acids or unnatural synthetic amino acids. D-amino acids may be substituted for some or all of the amino acid residues. The B-GEn.16 polypeptide or a variant thereof can be prepared by in vitro synthesis using conventional methods known in the art. Various commercially available synthetic devices, such as automated synthesizers from Applied Biosystems, Inc., Beckman, etc., are available. By using a synthesizer, natural amino acids can be substituted with unnatural amino acids. The specific sequence and preparation method can be determined by convenience, economy, required purity, etc.

[0212] Optionally, various groups that enable linkage to other molecules or surfaces can be introduced into the peptide during synthesis or expression. Thus, cysteine can be used to create, for example, a thiol ether, histidine for linking to a metal ion complex, a carboxyl group for forming an amide or ester, an amino group for forming an amide, etc.

[0213] Recombinant cells In some embodiments, the codon-optimized B-GEn.16 system described herein can be used in eukaryotes, such as mammalian cells, such as human cells. Any human cell is suitable for use with the codon-optimized B-GEn.16 system disclosed herein.

[0214] In some embodiments, the cells are ex vivo or in vitro and comprise (a) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding the B-GEn.16 polypeptide or variant described herein, or a B-GEn.16 polypeptide or variant expressed from the nucleic acid; and (b) a gRNA or nucleic acid encoding a gRNA, wherein the gRNA comprises a gRNA or nucleic acid capable of directing the B-GEn.16 polypeptide or variant to a target polynucleotide sequence. In some embodiments, the cells comprise a nucleic acid comprising a codon-optimized polynucleotide sequence. In some embodiments, the cells comprise a gRNA. In some embodiments, the cells comprise a nucleic acid encoding a gRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the cells comprise one or more additional gRNAs or nucleic acids encoding one or more additional gRNAs. In some embodiments, the cells further comprise a donor template.

[0215] In one aspect, some embodiments disclosed herein relate to a method of transforming a cell comprising introducing a nucleic acid provided herein into a host cell such as an animal cell, and selecting or screening for a transformed cell. The terms "host cell" and "recombinant host cell" are used interchangeably herein. It is understood that such terms refer not only to a particular target cell but also to the progeny or potential progeny of such a cell. Since certain modifications may occur in succeeding generations either as a result of mutations or environmental influences, such progeny may not in fact be identical to the parent cell, but are still included within the scope of the terms used herein. A variety of techniques for transforming such host cells and species are known in the art and are described in the technical and scientific literature. Accordingly, a cell culture comprising at least one recombinant cell disclosed herein is also within the scope of this application. Methods and systems suitable for generating and maintaining cell cultures are known in the art.

[0216] In related aspects, some embodiments relate to recombinant host cells, e.g., recombinant animal cells comprising the nucleic acids described herein. The nucleic acid can be stably integrated into the host genome, or can be replicated episomally, or can be present in the recombinant host cell as a mini-circle expression vector for stable or transient expression. Thus, in some embodiments disclosed herein, the nucleic acid is maintained and replicated in the recombinant host cell as an episomal unit. In some embodiments, the nucleic acid is stably integrated into the genome of the recombinant cell. In some embodiments, the nucleic acid is present in the recombinant host cell as a mini-circle expression vector for stable or transient expression.

[0217] In some embodiments, the host cell can be genetically engineered (e.g., transduced or transformed or transfected) using the vector constructs of the present application, which can be, for example, a vector for homologous recombination comprising a nucleic acid sequence homologous to a portion of the genome of the host cell, or an expression vector for the expression of any or a combination of genes of interest. The vector can be in the form of, for example, a plasmid, virus particle, phage, etc. In some embodiments, the vector for the expression of a polypeptide of interest can also be designed for integration into the host, e.g., by homologous recombination.

[0218] In some embodiments, the present disclosure provides a genetically modified host cell, e.g., an isolated genetically modified host cell, wherein the genetically modified host cell comprises: 1) an exogenous guide RNA; 2) an exogenous nucleic acid comprising a nucleotide sequence encoding the guide RNA; 3) an exogenous nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof; 4) an exogenous B-GEn.16 polypeptide or a variant thereof expressed from a nucleic acid comprising the codon-optimized polynucleotide sequence; or 5) a combination of any of the above. In some embodiments, the genetically modified cell is generated, for example, by genetically modifying a host cell with: 1) an exogenous guide RNA; 2) an exogenous nucleic acid comprising a nucleotide sequence encoding the guide RNA; 3) an exogenous nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof; 4) an exogenous B-GEn.16 polypeptide or a variant thereof expressed from a nucleic acid comprising the codon-optimized polynucleotide sequence; or 5) a combination of any of the above.

[0219] All cells suitable as the target cells as described above are also suitable as genetically modified host cells. For example, the genetically modified host cell of interest can be a cell derived from any organism, such as a bacterial cell, an archaeal cell, a cell of a unicellular eukaryote, a plant cell, an algal cell (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens (C. Agardh), etc.), a fungal cell (e.g., a yeast cell), an animal cell, a cell derived from an invertebrate (e.g., Drosophila, cnidarian, echinoderm, nematode, etc.), a cell derived from a vertebrate (e.g., fish, amphibian, reptile, bird, mammal), a cell derived from a mammal (e.g., pig, cow, goat, sheep, rodent, rat, non-human primate, human, etc.). In some embodiments, the genetically modified host cell can be any cell derived from a human.

[0220] In some embodiments, the genetically modified host cell of the present disclosure is genetically modified with an exogenous nucleic acid comprising a nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. In some embodiments, the genetically modified host cell is genetically modified with an exogenous nucleic acid comprising a nucleotide sequence encoding a B-GEn.16 polypeptide or variant as described herein. The DNA of the genetically modified host cell can be targeted for modification by introducing into the cell a guide RNA (or DNA encoding a guide RNA that determines the genomic location / sequence to be modified) and any donor nucleic acid. In some embodiments, the nucleotide sequence encoding the B-GEn.16 polypeptide or a variant thereof is operably linked to an inducible promoter (e.g., heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.). In some embodiments, the codon-optimized nucleotide sequence encoding the B-GEn.16 polypeptide or a variant thereof is operably linked to a spatially limited and / or temporally limited promoter (e.g., tissue-specific promoter, cell-type specific promoter, cell cycle-specific promoter). In some aspects, the codon-optimized nucleotide sequence encoding the B-GEn.16 polypeptide or a variant thereof is operably linked to a constitutive promoter.

[0221] In some embodiments, the genetically modified host cell is in vitro. In some embodiments, the genetically modified host cell is in vivo. In some embodiments, the genetically modified host cell is a prokaryotic cell or is derived from a prokaryotic cell. In some embodiments, the genetically modified host cell is a bacterial cell or is derived from a bacterial cell. In some embodiments, the genetically modified host cell is an archaeal cell or is derived from an archaeal cell. In some embodiments, the genetically modified host cell is a eukaryotic cell or is derived from a eukaryotic cell. In some embodiments, the genetically modified host cell is a plant cell or is derived from a plant cell. In some embodiments, the genetically modified host cell is an animal cell or is derived from an animal cell. In some embodiments, the genetically modified host cell is an invertebrate cell or is derived from an invertebrate cell. In some embodiments, the genetically modified host cell is a vertebrate cell or is derived from a vertebrate cell. In some embodiments, the genetically modified host cell is a mammalian cell or is derived from a mammalian cell. In some embodiments, the genetically modified host cell is a rodent cell or is derived from a rodent cell. In some embodiments, the genetically modified host cell is a human cell or is derived from a human cell. In some embodiments, the genetically modified host cell is a human cell or is derived from a human cell.

[0222] The present disclosure further provides progeny of the genetically modified cell, which may contain the same exogenous nucleic acid or polypeptide as the genetically modified cell from which it is derived. The present disclosure further provides, in some embodiments, a composition comprising the genetically modified host cell.

[0223] In some embodiments, the genetically modified host cell is a genetically modified stem cell or progenitor cell. Suitable host cells include, for example, stem cells (such as adult stem cells, embryonic stem cells, iPS cells, etc.) and progenitor cells (such as cardiac progenitor cells, neural progenitor cells, etc.). Other suitable host cells include mammalian stem cells and progenitor cells, such as rodent stem cells, rodent progenitor cells, human stem cells, human progenitor cells, etc. Other suitable host cells include in vitro host cells, such as isolated host cells. In some embodiments, the genetically modified host cell contains an exogenous guide RNA nucleic acid. In some embodiments, the genetically modified host cell contains an exogenous nucleic acid comprising a nucleotide sequence encoding a guide RNA. In some embodiments, the genetically modified host cell contains an exogenous B-GEn.16 polypeptide or a variant thereof expressed from a codon-optimized nucleotide sequence. In some embodiments, the genetically modified host cell contains an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. In some embodiments, the genetically modified host cell contains an exogenous nucleic acid comprising 1) a nucleotide sequence encoding a guide RNA, and 2) a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof.

[0224] Non-human genetically modified organism In some embodiments, the genetically modified host cell is genetically modified with an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. When such a cell is a eukaryotic unicellular organism, the modified cell can be considered a genetically modified organism. In some embodiments, the non-human genetically modified organism is a B-GEn.16 transgenic multicellular organism.

[0225] In some embodiments, a genetically modified non-human host cell (e.g., a cell genetically modified with an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof) can generate a genetically modified non-human organism (e.g., a mouse, fish, frog, fly, nematode, etc.). For example, when the genetically modified host cell is a pluripotent stem cell (e.g., PSC) or a germ cell (e.g., sperm, oocyte, etc.), the entire genetically modified organism can be derived from the genetically modified host cell. In some embodiments, the genetically modified host cell is any pluripotent stem cell (e.g., ESC, iPSC, pluripotent plant stem cells, etc.) or germ cell (e.g., spermatocyte, oocyte, etc.) that can give rise to a genetically modified organism, either in vivo or in vitro. In some embodiments, the genetically modified host cell is a vertebrate PSC (e.g., ESC, iPSC, etc.) and is used to generate a genetically modified organism (e.g., injecting PSC into a blastocyst to generate a chimeric / mosaic animal, which can then be mated to generate a non-chimeric / non-mosaic genetically modified organism; in the case of plants, by transplantation, etc.). Any suitable method / protocol for generating a genetically modified organism, including the methods described herein, is suitable for generating a genetically modified host cell comprising an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. Methods for producing genetically modified organisms are known in the art. See, for example, Cho et al., Curr Protoc Cell Biol. 2009 Mar; Chapter 19: Unit 19.11: Generation of Transgenic Mice; Gama et al., Brain Struct Funct. 2010 Mar; 214(2-3):91-109. Epub 2009 Nov 25: Gene Transfer in Animals: An Overview; Husaini et al., GM Crops. 2011 Jun-Dec; 2(3):150-62. Epub 2011 Jun 1: Approaches for Gene Targeting and Targeted Gene Expression in Plants.

[0226] In some embodiments, the genetically modified organism comprises target cells for the methods of the present disclosure and can thus be considered a source of target cells. For example, when producing a genetically modified organism using genetically modified cells comprising an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, the cells of the genetically modified organism comprise an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. In some such embodiments, the DNA of the cell or the cells of the genetically modified organism can be targeted for modification by introducing guide RNA (or DNA encoding guide RNA) and optionally donor nucleic acid into the cell or cells. For example, introduction of guide RNA (or DNA encoding guide RNA) into a subset of cells of a genetically modified organism (e.g., brain cells, intestinal cells, kidney cells, lung cells, blood cells, etc.) can target the DNA of such cells for modification, and its genomic location depends on the DNA targeting sequence of the introduced guide RNA.

[0227] In some embodiments, the genetically modified organism is a source of target cells for the methods of the present disclosure. For example, a genetically modified organism comprising cells genetically modified with an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof can provide a source of genetically modified cells, such as PSCs (e.g., ESCs, iPSCs, sperm, oocytes, etc.), neurons, progenitor cells, cardiomyocytes, etc.

[0228] In some embodiments, the genetically modified cell is a PSC comprising an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. Thus, the PSC can be a target cell such that the DNA of the PSC can be targeted for modification by introduction of a guide RNA (or DNA encoding a guide RNA) and any donor nucleic acid into the PSC, and the genomic location of the modification depends on the DNA targeting sequence of the introduced guide RNA. Thus, in some embodiments, the methods described herein can be used to modify the DNA of PSCs derived from genetically modified organisms (e.g., deleting and / or replacing any desired genomic location). Such modified PSCs can then be used to generate an organism having both (i) an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof and (ii) a DNA modification introduced into the PSC.

[0229] In some embodiments, the exogenous nucleic acid can be under the control of (e.g., operably linked to) an unknown promoter (e.g., if the nucleic acid is randomly integrated into the host cell genome), or can be under the control of (e.g., operably linked to) a known promoter. Suitable known promoters can be any known promoter, including constitutively active promoters (e.g., CMV promoter), inducible promoters (e.g., heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.), spatially restricted and / or temporally restricted promoters (e.g., tissue-specific promoter, cell type-specific promoter, etc.).

[0230] A genetically modified organism (e.g., an organism whose cells contain a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof) can be any organism, including, for example, plants; algae; invertebrates (e.g., cnidarians, echinoderms, worms, flies, etc.); vertebrates (e.g., fish (e.g., zebrafish, pufferfish, goldfish, etc.), amphibians (e.g., salamanders, frogs, etc.), reptiles, birds, mammals, etc.); ungulates (e.g., goats, pigs, sheep, cows, etc.); rodents (e.g., mice, rats, hamsters, guinea pigs); lagomorphs (e.g., rabbits, etc.).

[0231] In some embodiments, the active moiety is an RNase domain. In some embodiments, the active moiety is a DNase domain.

[0232] Transgenic non-human animals As described above, in some embodiments, a nucleic acid (e.g., a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof) or a recombinant expression vector is used as a transgene to generate a transgenic animal that produces a B-GEn.16 polypeptide or a variant thereof. Accordingly, the present disclosure further provides a transgenic non-human animal comprising a transgene comprising a nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof as described above. In some embodiments, the genome of the transgenic non-human animal comprises a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. In some embodiments, the transgenic non-human animal is homozygous for the genetic modification. In some embodiments, the transgenic non-human animal is heterozygous for the genetic modification. In some embodiments, the transgenic non-human animal is a vertebrate, such as a fish (e.g., zebrafish, goldfish, pufferfish, loach, etc.), an amphibian (e.g., frog, salamander, etc.), a bird (e.g., chicken, pigeon, etc.), a reptile (e.g., snake, lizard, etc.), a mammal (e.g., ungulates, such as pigs, cows, goats, sheep, etc., lagomorphs (e.g., rabbits), rodents (e.g., rats, mice), non-human primates, etc.).

[0233] In some embodiments, the nucleic acid is an exogenous nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof. In some embodiments, the exogenous nucleic acid can be under the control of (e.g., operably linked to) an unknown promoter (e.g., when the nucleic acid is randomly integrated into the host cell genome), or under the control of (e.g., operably linked to) a known promoter. Suitable known promoters can be any known promoter, including constitutively active promoters (e.g., CMV promoter), inducible promoters (e.g., heat shock promoter, tetracycline-regulated promoter, steroid-regulated promoter, metal-regulated promoter, estrogen receptor-regulated promoter, etc.), spatially limited and / or temporally limited promoters (e.g., tissue-specific promoters, cell type-specific promoters, etc.).

[0234] Introduction of Nucleic Acid into Host Cell In some embodiments, the methods of the disclosure include introducing one or more nucleic acids comprising a nucleotide sequence encoding a guide RNA and / or a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof into a host cell (or a population of host cells). In some embodiments, the cells containing the target DNA are in vitro. In some embodiments, the cells containing the target DNA are in vivo. In some embodiments, the nucleotide sequence encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof is operably linked to an inducible promoter. In some embodiments, the nucleotide sequence encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof is operably linked to a constitutive promoter.

[0235] A nucleic acid comprising a guide RNA, or a nucleotide sequence encoding the same, can be introduced into a host cell by any of a variety of well-known methods. Similarly, if the method involves introducing into a host cell a nucleic acid comprising a codon-optimized nucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, such a nucleic acid can be introduced into the host cell by any of a variety of well-known methods. The guide polynucleotide (RNA or DNA) and / or the B-GEn.16 polynucleotide (RNA or DNA) can be delivered by viral or non-viral delivery vehicles known in the art.

[0236] Methods for introducing nucleic acids into host cells are known in the art, and any known method can be used to introduce a nucleic acid (e.g., an expression construct) into a stem cell or progenitor cell. Suitable methods include, for example, viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun method, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery (see, e.g., Panyam et al., Adv Drug Deliv Rev. 2012 Sep 13. pii: 50169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023), including but not limited to exosome delivery.

[0237] The polynucleotide can be delivered by non-viral delivery vehicles including, but not limited to, nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small molecule RNA conjugates, aptamer-RNA chimeras, and RNA-fusion protein complexes. Some exemplary non-viral delivery vehicles are described in Peer and Lieberman, Gene Therapy, 18:1127-1133 (2011) (which focuses on non-viral delivery vehicles for siRNA which are also useful for delivery of other polynucleotides).

[0238] Suitable systems and techniques for delivering the nucleic acids (e.g., mRNA and sgRNA) of the present disclosure for gene editing include lipid nanoparticles (LNPs). As used herein, the term "lipid nanoparticles" includes liposomes as described for the introduction of nucleic acids and / or polypeptides into cells, regardless of their lamellarity, shape or structure and lipoplexes. These lipid nanoparticles can complex with biologically active compounds (e.g., nucleic acids and / or polypeptides) and are useful as in vivo delivery vehicles. In general, any method known in the art can be applied to prepare lipid nanoparticles containing one or more nucleic acids of the present disclosure and to prepare a complex of the biologically active compound and the lipid nanoparticles. Examples of such methods are disclosed, for example, in Biochim Biophys Acta 1979, 557:9; Biochim et Biophys Acta 1980, 601:559; Liposomes: A practical approach (Oxford University Press, 1990); Pharmaceutica Acta Helvetiae 1995, 70:95; Current Science 1995, 68:715; Pakistan Journal of Pharmaceutical Sciences 1996, 19:65; Methods in Enzymology 2009, 464:343).Particularly suitable systems and techniques for preparing LNP formulations comprising one or more nucleic acids and / or polypeptides of the present disclosure include, but are not limited to, those developed by Intellia (see, e.g., WO2017173054A1), Alnylam (see, e.g., WO2014008334A1), Modernatx (see, e.g., WO2017070622A1 and WO2017099823A1), TranslateBio, Acuitas (see, e.g., WO2018081480A1), Genevant Sciences, Arbutus Biopharma, Tekmira, Arcturus, Merck (see, e.g., WO2015130584A2), Novartis (see, e.g., WO2015095340A1), and Dicerna; all of which are hereby incorporated by reference in their entirety.

[0239] Suitable nucleic acids comprising a nucleotide sequence encoding a guide RNA and / or a B-GEn.16 polypeptide or variant thereof include an expression vector, where the expression vector comprises a nucleotide sequence encoding the guide. In some embodiments, the expression vector is a viral construct, e.g., a recombinant adeno-associated virus construct (see, e.g., U.S. Patent No. 7,078,387), a recombinant adenovirus construct, a recombinant lentivirus construct, a recombinant retrovirus construct, etc.Suitable expression vectors include, but are not limited to, viral vectors (e.g., vaccinia virus; poliovirus; adenovirus (see, e.g., Li et al., Invest Opthalmol Vis Sci 35:2543 - 2549, 1994; Borras et al., Gene Ther 6:515 - 524, 1999; Li and Davidson, PNAS 92:7700 - 7704, 1995; Sakamoto et al., Hum Gene Ther 5:1088 - 1097, 1999; WO94 / 12649, WO93 / 03769; WO93 / 19191; WO 94 / 28938; WO 95 / 11984 and WO95 / 00655); adeno - associated virus (e.g., Ali et al., Hum Gene Ther 9:81 - 86, 1998, Flannery et al., PNAS 94:6916 - 6921, 1997; Bennett et al., Invest Opthalmol Vis Sci 38:2857 - 2863, 1997; Jomary et al., Gene Ther 4:683 - 690, 1997, Rolling et al., Hum Gene Ther 10:641 - 648, 1999; Ali et al., Hum Mol Genet 5:591 - 594, 1996; Srivastava in WO 93 / 09239, Samulski et al., J. Vir. (1989) 63:3822 - 3828; Mendelson et al., Viral. (1988) 166:154 - 165; and Flotte et al., PNAS (1993) 90:10613 - 10617); SV40; herpes simplex virus; human immunodeficiency virus (see, e.g., Miyoshi et al., PNAS 94:10319 - 23, 1997; Takahashi et al., J Virol 73:7812 - 7816, 1999); viral vectors based on retroviral vectors (e.g., vectors derived from retroviruses such as murine leukemia virus, spleen necrosis virus, and Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus)).

[0240] Adeno - associated virus (AAV) Recombinant adeno-associated virus (AAV) vectors can be used for delivery. Known techniques for producing rAAV particles in the art involve providing a polynucleotide delivered between two AAV inverted terminal repeats (ITRs), the AAV rep and cap genes, and helper virus functions to cells. Production of rAAV requires the following components to be present within a single cell (referred to herein as a packaging cell): a polynucleotide of interest between two ITRs, AAV rep and cap genes separated from the AAV genome (i.e., not present therein), and helper virus functions. The AAV rep and cap genes can be derived from any AAV serotype from which a recombinant virus can be derived, and can be derived from an AAV serotype different from the ITRs on the packaged polynucleotide, including but not limited to AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11, AAV-12, AAV-13, and AAV rh.74. Production of pseudotyped rAAV is disclosed, for example, in WO01 / 83692.

Table 5

[0241] The method of producing packaging cells is to create a cell line that stably expresses all the components necessary for AAV particle production. For example, a plasmid (or plasmids) containing the polynucleotide of interest between AAV ITRs, the AAV rep and cap genes separated from the AAV genome, and a selectable marker such as the neomycin resistance gene is integrated into the cell's genome. The AAV genome is introduced into a bacterial plasmid by procedures such as GC tailing (Samulski et al., 1982, Proc. Natl. Acad. Sci. USA, 79:2077-2081), addition of synthetic linkers containing restriction endonuclease cleavage sites (Laughlinra, 1983, Gene, 23:65-73), or direct blunt-end ligation (Senapathy & Carter, 1984, J. Bioi. Chem., 259:4661-4666). The packaging cell line is then infected with a helper virus such as adenovirus. The advantage of this method is that the cells are selectable and suitable for large-scale production of rAAV. Another example of a suitable method is to use an adenovirus or baculovirus instead of a plasmid to introduce the rAAV genome and / or the rep and cap genes into the packaging cells.

[0242] The general principles of rAAV production are outlined, for example, in Carter, 1992, Current Opinions in Biotechnology, 1533-539; and Muzyczka, 1992, Curr. Topics in Microbial. and Immunol., 158:97-129). Various approaches are described in Ratschin et al., Mol. Cell. Biol. 4:2072 (1984); Hermonat et al., Proc. Natl. Acad. Sci. USA, 81:6466 (1984); Tratschin et al., Mol. Cell. Biol. 5:3251 (1985); Mclaughlin et al., J. Virol., 62:1963 (1988); and Lebkowski et al., 1988 Mol. Cell. Biol., 7:349 (1988), Samulski et al. (1989, J. Virol., 63:3822-3828); U.S. Patent No. 5,173,414; WO95 / 13365 and corresponding U.S. Patent No. 5,658,776; WO95 / 13392; WO96 / 17947; PCT / US98 / 18600; WO97 / 09441 (PCT / US96 / 14423); WO97 / 08298 (PCT / US96 / 13872); WO97 / 21825 (PCT / US96 / 20777); WO97 / 06243 (PCT / FR96 / 01064); WO 99 / 11764; Perrin et al. (1995) Vaccine 13:1244-1250; Paul et al. (1993) Human Gene Therapy 4:609-615; Clark et al. (1996) Gene Therapy 3:1124-1132; U.S. Patent No. 5,786,211; U.S. Patent No. 5,871,982; and U.S. Patent No. 6,258,595.

[0243] The AAV vector serotype used for transduction depends on the target cell type. For example, the following exemplary cell types are known to be transduced, among other things, by the indicated AAV serotypes.

Table 6

[0244] A number of suitable expression vectors are known to those skilled in the art, and many are commercially available. The following vectors are provided as examples for eukaryotic host cells: pXT1, pSG5 (Stratagene), pSVK3, pBPV, pMSG, and pSVLSV40 (Pharmacia). However, any other vector can be used as long as it is compatible with the host cell.

[0245] Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc., can be used in the expression vector (see, for example, Bitter et al. (1987) Methods in Enzymology, 153:516-544).

[0246] In some embodiments, the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof can be provided as RNA. In such cases, the RNA encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof can be produced by direct chemical synthesis or transcribed in vitro from DNA encoding the guide RNA. Methods for synthesizing RNA from a DNA template are well known in the art. In some embodiments, the RNA encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof is synthesized in vitro using an RNA polymerase enzyme (e.g., T7 polymerase, T3 polymerase, SP6 polymerase, etc.). Once synthesized, the RNA can come into direct contact with the target DNA or can be introduced into cells by any of the well-known techniques for introducing nucleic acids into cells (e.g., microinjection, electroporation, transfection, etc.).

[0247] The guide RNA (introduced as DNA or RNA) and / or the B-GEn.16 polypeptide or variant thereof (introduced as DNA or RNA) and / or the nucleotide encoding the donor polynucleotide can be provided to cells using well-developed transfection techniques; see, for example, Angel and Yanik (2010) PLoS ONE 5(7):e11756, as well as the commercially available TransMessenger® reagent from Qiagen, the Stemfect™ RNA transfection kit from Stemgent, and the TransiT®-mRNA transfection kit from Mims Bio. See also Beumer et al. (2008) Efficient gene targeting in Drosophila by direct embryo injection with zinc-finger nucleases. PNAS 105(50):19821-19826. In addition, or alternatively, the guide RNA and / or the B-GEn.16 polypeptide or variant thereof, and / or the B-GEn.16 fusion polypeptide or variant thereof, and / or the nucleic acid encoding the donor polynucleotide can be provided on a DNA vector. Many vectors useful for transferring nucleic acids into target cells are available, such as plasmids, cosmids, minicircles, phages, viruses, etc. Vectors containing the nucleic acid may be maintained episomally, for example, as viruses such as plasmids, minicircle DNA, cytomegalovirus, adenovirus, etc., or they may be integrated into the target cell genome via homologous recombination or random integration, for example, via retrovirus-derived vectors such as MMLV, HIV-1, ALV, etc.

[0248] The vector can be provided directly to the cell. In other words, the cell is contacted with a vector containing a nucleic acid encoding a guide RNA and / or a B-GEn.16 polypeptide or variant thereof and / or a B-GEn.16 fusion polypeptide or variant thereof and / or a donor polynucleotide such that the vector is taken up by the cell. Methods for contacting cells with nucleic acid vectors that are plasmids, including electroporation, calcium chloride transfection, microinjection, and lipofection, are well known in the art. For viral vector delivery, the cell is contacted with viral particles containing a nucleic acid encoding a guide RNA and / or a B-GEn.16 polypeptide or variant thereof and / or a B-GEn.16 fusion polypeptide or variant thereof and / or a donor polynucleotide. Retroviruses, such as lentiviruses, are particularly suitable for the methods of the present disclosure. Commonly used retroviral vectors are "defective," i.e., they are unable to produce viral proteins required for productive infection. Rather, replication of the vector requires growth in a packaging cell line. To generate viral particles containing the nucleic acid of interest, the retroviral nucleic acid containing the nucleic acid is packaged into a viral capsid by a packaging cell line. Different packaging cell lines provide different envelope proteins (homotropic, amphotropic, or xenotropic) that are incorporated into the capsid, and this envelope protein determines the specificity of the viral particle for the cell (homotropic for mouse and rat; amphotropic for most mammalian cell types including human, dog, and mouse; xenotropic for most mammalian cell types except mouse cells). Appropriate packaging cell lines can be used to ensure that the cell is targeted by the packaged viral particles. Methods for introducing a retroviral vector containing a nucleic acid encoding a reprogramming factor into a packaging cell line and methods for collecting the viral particles produced by the packaging cell line are well known in the art. The nucleic acid can also be introduced directly by microinjection (e.g., injection of RNA into a zebrafish embryo).

[0249] Vectors used to provide cells with a nucleic acid encoding a guide RNA and / or a B-GEn.16 polypeptide or variant thereof and / or a B-GEn.16 fusion polypeptide or variant thereof and / or a donor polynucleotide generally contain a suitable promoter to drive expression of the nucleic acid of interest, i.e., transcriptional activation. In other words, the nucleic acid of interest is operably linked to the promoter. This can include a ubiquitously active promoter, such as the CMV-13-actin promoter, or an inducible promoter, such as a promoter that is active in a particular cell population or that responds to the presence of a drug such as tetracycline. By transcriptional activation, transcription is intended to increase at least 10-fold, at least 100-fold, and more typically at least 1000-fold above basal levels in the target cells. Further, vectors used to provide cells with a guide RNA and / or a B-GEn.16 polypeptide or variant thereof and / or a B-GEn.16 fusion polypeptide or variant thereof and / or a donor polynucleotide can include a nucleic acid sequence encoding a selectable marker within the target cell to identify cells that have taken up the guide RNA and / or a B-GEn.16 polypeptide or variant thereof and / or a B-GEn.16 fusion polypeptide or variant thereof and / or a donor polynucleotide.

[0250] A guide RNA and / or a B-GEn.16 polypeptide or a variant thereof and / or a B-GEn.16 fusion polypeptide or a variant thereof can be used instead to contact DNA or be introduced into cells as RNA. Methods for introducing RNA into cells are known in the art and can include, for example, direct injection, transfection, or any other method used for introducing DNA. The B-GEn.16 polypeptide or a variant thereof can instead be provided to cells as a polypeptide. Such a polypeptide can optionally be fused to a polypeptide domain that increases the solubility of the product. The domain can be linked to the polypeptide via a defined protease cleavage site, for example, a TEV sequence that is cleaved by TEV protease. The linker can also include one or more flexible sequences, for example, 1 to 10 glycine residues. In some embodiments, cleavage of the fusion protein is carried out in a buffer that maintains the solubility of the product, for example, in the presence of 0.5 to 2 M urea, in the presence of a polypeptide and / or polynucleotide that increases solubility, etc. Domains of interest include endosomal degradation domains, such as the influenza HA domain; and other polypeptides that assist in production, such as the IF2 domain, GST domain, GRPE domain, etc. The polypeptide can be formulated for improved stability. For example, the peptide can be PEGylated, in which case the polyethyleneoxy groups provide an improved lifespan in the bloodstream.

[0251] Additionally, or alternatively, a B-GEn.16 polypeptide or variant thereof can be fused to a polypeptide permeable domain to facilitate cellular uptake. A number of permeable domains are known in the art and can be used in the non-integrating polypeptides of the present disclosure, including peptides, peptidomimetics, and non-peptide carriers. For example, a permeable peptide can be derived from the third α-helix of the Drosophila melanogaster transcription factor Antennapedia, called penetratin, which contains the amino acid sequence RQIKIWFQNRRMKWKK (this sequence is not disclosed under this patent application). As another example, a permeable peptide can contain the HIV-1 tat basic region amino acid sequence, which can include, for example, amino acids 49-57 of the naturally occurring tat protein.

[0252] Other permeable domains include polyarginine motifs such as the region of amino acids 34 to 56 of the HIV-1 rev protein, nonaarginine, actaaarginine, etc. (See, for example, Futaki et al. (2003) Curr Protein Pept Sci. 2003 Apr;4(2):87-9 and 446; and Wender et al. (2000) Proc. Natl. Acad. Sci. U.S.A. 2000 Nov. 21;97(24):13003-8; US Patent Application Publication No. 20030220334; No. 20030083256; No. 20030032593; and No. 20030022831, where the teachings of translocation peptides and peptoids are specifically incorporated herein by reference). The nonaarginine (R9) sequence is one of the more efficient PTDs that have been characterized (Wender et al., 2000; Uemura et al., 2002). The site at which the fusion is made can be selected to optimize the biological activity, secretion or binding properties of the polypeptide. The optimal site can be determined by routine experimentation. In some embodiments, the polypeptide permeable domain is chemically modified to increase the bioavailability of the PTD. See, for example, Expert Opin Drug Deliv. 2009 Nov;6(11):1195-205.

[0253] Generally, an effective amount of a guide RNA and / or a B-GEn.16 polypeptide or a variant thereof and / or a donor polynucleotide is provided to a target DNA or cell to induce a target modification. The effective amount of the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof and / or the donor polynucleotide is an amount that induces an increase of at least 2-fold in the amount of targeted modification observed with the gRNA as compared to cells contacted with a negative control, such as an empty vector or an irrelevant polypeptide. That is, the effective amount or dosage of the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof and / or the donor polynucleotide induces an increase of 2-fold, 3-fold, 4-fold or more in the amount of target modification observed in the target DNA region, and in some embodiments an increase of 5-fold, 6-fold or more, sometimes 7-fold or 8-fold or more in the amount of recombination observed, for example, an increase of 10-fold, 50-fold or 100-fold or more in the amount of recombination observed in some embodiments, an increase of 200-fold, 500-fold, 700-fold or 1000-fold or more in the amount of recombination observed in some embodiments, for example an increase of 5000-fold or 10,000-fold. The amount of target modification can be measured by any suitable method. For example, a split reporter construct containing a complementary sequence to the spacer of the guide RNA adjacent to the homologous sequence, which can reconstitute a nucleic acid encoding an active reporter in the cell when recombined, can be co-transfected into the cell, and after contact with the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof and / or the donor polynucleotide, for example, 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 36 hours, 48 hours, 72 hours or more after contact with the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof and / or the donor polynucleotide, the amount of reporter protein is evaluated.As another example, the degree of recombination in a target genomic DNA region containing a target DNA sequence, in a more sensitive assay, e.g., after contact with a guide RNA and / or a B-GEn.16 polypeptide or a variant thereof and / or a donor polynucleotide, e.g., 2 hours, 4 hours, 8 hours, 12 hours, 24 hours, 36 hours, 48 hours, 72 hours or more after contact with the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof and / or the donor polynucleotide, can be evaluated by PCR or Southern hybridization of the region.

[0254] Contacting the cell with the guide RNA and / or the B-GEn.16 polypeptide or variant thereof and / or the donor polynucleotide can occur in any culture medium under any culture conditions that promote cell survival. For example, the cells can be suspended in a suitable nutrient medium such as lscove's modified DMEM or RPMI 1640 supplemented with fetal bovine serum or heat-inactivated fetal bovine serum (about 5-10%), L-glutamine, thiols, particularly 2-mercaptoethanol, and antibiotics such as penicillin and streptomycin. The culture may contain growth factors to which the cells respond. A growth factor, as defined herein, is a molecule that can promote cell survival, growth, and / or differentiation either in culture or in intact tissue through its specific effect on transmembrane receptors. Growth factors include polypeptide and non-polypeptide factors. Conditions that promote cell survival generally permit non-homologous end joining and homologous recombination repair. In applications where it is desirable to insert a polynucleotide sequence into a target DNA sequence, a polynucleotide containing the donor sequence to be inserted is also provided to the cell. A "donor sequence" or "donor polynucleotide" means a nucleic acid sequence that is inserted into a cleavage site induced by the B-GEn.16 polypeptide or variant thereof. The donor polynucleotide has sufficient sequence homology to the adjacent genomic region of the cleavage site, e.g., 70%, 80%, 85%, 90%, 95%, or 100% sequence identity with the nucleotide sequence adjacent to the cleavage site, e.g., within about 50 bases of the cleavage site, e.g., within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or immediately adjacent to the cleavage site, and supports homologous recombination repair between it and the genomic sequence with which it is homologous. About 25, 50, 100, or 200 nucleotides, or more than 200 nucleotides (or any integer value between 10 and 200 nucleotides or more) of homologous sequence between the donor and the genomic sequence support homologous recombination repair. The donor sequence can be of any length, e.g., 10 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more, etc.

[0255] The donor sequence is generally not identical to the genomic sequence it replaces. Rather, the donor sequence can include at least one or more single nucleotide substitutions, insertions, deletions, inversions, or rearrangements with respect to the genomic sequence, as long as there is sufficient sequence identity to support homologous recombination repair. In some embodiments, the donor sequence includes a non-homologous sequence flanked by two regions (also called homologous arms) that are homologous to the target DNA region, such that homologous recombination repair between the target DNA region and the two flanking homologous arms results in the insertion of the non-homologous sequence into the target region. The donor sequence may also include a vector backbone containing a sequence that is not homologous to the DNA region of interest and is not intended for insertion into the DNA region of interest. Generally, the homologous regions of the donor sequence have at least 50% sequence identity to the genomic sequence where recombination is desired. In certain embodiments, there is 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.9% sequence identity. Depending on the length of the donor polynucleotide, any value of sequence identity from 1% to 100% can exist. The donor sequence may include specific sequence differences compared to the genomic sequence, such as restriction sites, nucleotide polymorphisms, selection markers (e.g., drug resistance genes, fluorescent proteins, enzymes, etc.), which may be used to assess successful insertion of the donor sequence at the cleavage site, or in some cases for other purposes (e.g., to indicate expression at the target genomic locus). In some embodiments, when located in the coding region, such nucleotide sequence differences do not change the amino acid sequence or result in amino acid changes that do not substantially affect the structure or function of the protein. Alternatively, these sequence differences may include adjacent recombination sequences such as FLP, loxP sequences, etc., that can be later activated for removal of the marker sequence.

[0256] The donor sequence can be provided to the cell as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. It can be introduced into the cell in linear or circular form. When introduced in linear form, the ends of the donor sequence can be protected by methods known to those skilled in the art (e.g., from exonuclease degradation). For example, one or more dideoxynucleotide residues are added to the 3’ end of the linear molecule, and / or self-complementary oligonucleotides are ligated to one or both ends. See, for example, Chang et al., (1987) Proc. Natl. Acad. Sci. USA 84:4959-4963; Nehls et al., (1996) Science 272:886-889. Additional methods for protecting exogenous polynucleotides from degradation include the addition of terminal amino groups, and the use of modified internucleotide linkages, such as phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues, but are not limited thereto. Instead of protecting the ends of the linear donor sequence, additional lengths of sequence that can be degraded without affecting recombination can be included outside of the homology arms. The donor sequence can be introduced into the cell as part of a vector molecule having additional sequences, such as an origin of replication, a promoter, and a gene encoding antibiotic resistance. Further, as described above for the guide RNA and / or the B-GEn.16 polypeptide or variants thereof and / or the nucleic acid encoding the donor polynucleotide, the donor sequence can be introduced as a naked (e.g., unmodified) nucleic acid, as a nucleic acid complexed with an agent such as a liposome or a poloxamer, or can be delivered by a virus (e.g., an adenovirus, AAV).

[0257] Following the above method, the DNA region of interest can be cleaved and modified, e.g., "genetically modified", ex vivo. In some embodiments, when a selectable marker is inserted into the DNA region of interest, the cell population can be enriched for those containing the genetic modification by separating the genetically modified cells from the remaining population. Prior to enrichment, the "genetically modified" cells may constitute only about 1% or more (e.g., 2% or more, 3% or more, 4% or more, 5% or more, 6% or more, 7% or more, 8% or more, 9% or more, 10% or more, 15% or more, or 20% or more) of the cell population. Separation of the "genetically modified" cells can be achieved by any suitable separation technique appropriate for the selectable marker used. For example, if a fluorescent marker is inserted, the cells can be separated by fluorescence-activated cell sorting, while if a cell surface marker is inserted, the cells can be separated from the heterogeneous population by affinity separation techniques such as magnetic separation, affinity chromatography, "panning" using an affinity reagent bound to a solid matrix, or other suitable techniques. Techniques that provide accurate separation include fluorescence-activated cell sorters of varying degrees of sophistication. For example, multi-color channels, low-angle and obtuse-angle light scatter detection channels, impedance channels, etc. Cells can be selected against dead cells by using a dye that associates with dead cells (e.g., propidium iodide). Any technique that does not unduly adversely affect the viability of the genetically modified cells can be used. A highly enriched cell composition for cells containing the modified DNA is thus achieved. "Highly enriched" means that the genetically modified cells are 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, e.g., about 95% or more, or 98% or more of the cell composition. In other words, the composition may be a substantially pure composition of genetically modified cells.

[0258] The genetically modified cells produced by the method described herein can be used immediately. Additionally, or alternatively, the cells can be frozen at liquid nitrogen temperature, stored for a long period of time, thawed, and reused. In such cases, the cells are generally frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered medium, or other solutions commonly used in the art for storing cells at such freezing temperatures, and thawed by methods generally known in the art for thawing frozen cultured cells.

[0259] The genetically modified cells can be cultured in vitro under various culture conditions. The cells can be grown in culture, for example, under conditions that promote their growth, such as by growing in a culture. The culture medium can be liquid or semi-solid, containing, for example, agar, methylcellulose, etc. The cell population can be suspended in a suitable nutrient medium such as lscove's modified DMEM or RPMI 1640, usually supplemented with fetal bovine serum (about 5 - 10%), L-glutamine, thiols, particularly 2-mercaptoethanol, and antibiotics such as penicillin and streptomycin. The culture may contain growth factors to which each cell responds. Growth factors as defined herein are molecules that can promote cell survival, growth, and / or differentiation either in culture or in intact tissue, through a specific effect on transmembrane receptors. Growth factors include polypeptide and non-polypeptide factors. Thus genetically modified cells can be transplanted into a subject for purposes such as gene therapy, for example, to treat a disease, or as an anti-viral, anti-pathogenic, or anti-cancer therapeutic agent, for the production of genetically modified organisms in agriculture, or for biological research. The subject can be a neonate, juvenile, or adult. Particularly interesting are mammalian subjects. Mammalian species that can be treated by the method of the present invention include canines and felines; horses; cattle; sheep, etc.; and primates, particularly humans. Animal models, particularly small mammals (e.g., mice, rats, guinea pigs, hamsters, lagomorphs (e.g., rabbits), etc.) can be used for experimental studies.

[0260] Cells can be provided to a subject alone or, for example, together with a suitable substrate or matrix to support their growth and / or organization in the tissue into which they are transplanted. Generally, at least 1×10 3 cells, for example, 5×10 3 cells, 1×10 4 cells, 5×10 4 cells, 1×10 5 cells, 1×10 6 cells or more are administered. Cells can be introduced into the subject via any of the following routes: parenteral, subcutaneous, intravenous, intracranial, intraspinal, intraocular, or intrathecal. Cells can be introduced by injection, catheter, etc. Examples of methods of local delivery, i.e., delivery to the site of injury, include, for example, via an Ommaya reservoir, for example, for intrathecal delivery (see, for example, U.S. Pat. Nos. 5,222,982 and 5,385,582, incorporated herein by reference), by bolus injection, for example, with a syringe, for example, injection into a joint; for example, by continuous infusion, for example, by cannula insertion, for example, with convection (see, for example, U.S. Patent Application Publication No. 20070254842, incorporated herein by reference), or by implanting a device in which the cells are reversibly immobilized (see, for example, U.S. Patent Application Publication Nos. 20080081064 and 20090196903, incorporated herein by reference), etc. Also, for the purpose of producing transgenic animals (e.g., transgenic mice), cells may be introduced into an embryo (e.g., a blastocyst).

[0261] In some embodiments, the nucleotide sequence encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof is operably linked to a control element, such as a transcriptional control element, e.g., a promoter. The transcriptional control element is generally functional in either eukaryotic cells, such as mammalian cells (e.g., human cells); or prokaryotic cells (e.g., bacterial cells or archaeal cells). In some embodiments, the nucleotide sequence encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof is operably linked to a plurality of control elements that enable the expression of the nucleotide sequence encoding the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof in both prokaryotic and eukaryotic cells.

[0262] The promoter can be a constitutively active promoter (e.g., a promoter that is constitutively present in an active “ON” state), which can be an inducible promoter (e.g., a promoter whose state, active / “ON” or inactive / “OFF”, is controlled by an external stimulus, such as the presence of a specific temperature, compound, or protein), which can be a spatially restricted promoter (e.g., a transcriptional control element, enhancer, etc.) (e.g., a tissue-specific promoter, a cell-type specific promoter, etc.), and which can be a temporally restricted promoter (e.g., the promoter is in an active / “ON” or inactive / “OFF” state during a specific stage of embryonic development or during a specific stage of a biological process, e.g., the hair follicle cycle in a mouse).

[0263] Suitable promoters can be derived from viruses and can thus be referred to as viral promoters, or can be derived from any organism, including prokaryotes or eukaryotes. Suitable promoters can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Exemplary promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus long terminal repeat (LTR) promoter; the adenovirus major late promoter (Ad MLP); the herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter region (CMVIE), the Rous sarcoma virus (RSV) promoter, the human U6 small nuclear promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497 - 500 (2002)), the enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1; 31(17)), the human H1 promoter (H1), and the like.

[0264] Examples of inducible promoters include, but are not limited to, the T7 RNA polymerase promoter, the T3 RNA polymerase promoter, the isopropyl - β - D - thiogalactopyranoside (IPTG) - regulatable promoter, the lactose - inducible promoter, the heat - shock promoter, the tetracycline - regulatable promoter (e.g., TetON, Tet - OFF, etc.), the steroid - regulatable promoter, the metal - regulatable promoter, the estrogen receptor - regulatable promoter, and the like. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline; RNA polymerase, e.g., T7 RNA polymerase; estrogen receptor; estrogen receptor fusions, and the like.

[0265] In some embodiments, the promoter is a spatially restricted promoter (e.g., a cell-type specific promoter, a tissue-specific promoter, etc.), such that in a multicellular organism, the promoter is active (e.g., "ON") in a specific subset of cells. Spatially restricted promoters may also be referred to as enhancers, transcriptional control elements, regulatory sequences, etc. Any suitable spatially restricted promoter can be used, and the choice of suitable promoter (e.g., a brain-specific promoter, a promoter driving expression in a subset of neurons, a promoter driving expression in the germline, a promoter driving expression in the lung, a promoter driving expression in muscle, a promoter driving expression in pancreatic islet cells, etc.) depends on the organism. For example, various spatially restricted promoters are known for plants, flies, worms, mammals, mice, etc. Thus, a spatially restricted promoter can be used to regulate the expression of a nucleic acid encoding a B-GEn.16 polypeptide or a variant thereof in a wide variety of different tissues and cell types, depending on the organism. Some spatially restricted promoters are also temporally restricted such that the promoter is in the "ON" or "OFF" state during specific stages of embryonic development or during specific stages of a biological process (e.g., the hair follicle cycle in mice).

[0266] For illustrative purposes, examples of spatially limited promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, and the like.Neuron-specific spatially restricted promoters include the neuron-specific enolase (NSE) promoter (see, for example, EMBL HSEN02, X51956); the aromatic amino acid decarboxylase (AADC) promoter; the neurofilament promoter (see, for example, GenBank HUMNFL, L04147); the synapsin promoter (see, for example, GenBank HUMSYNIB, M55301); the thy-1 promoter (see, for example, Chen et al. (1987) Ce / 151:7-19; and Llewellyn et al. (2010) Nat. Med. 16(10):1161-1166); the serotonin receptor promoter (see, for example, GenBank S62283); the tyrosine hydroxylase promoter (TH) (see, for example, Oh et al. (2009) Gene Ther 16:437; Sasaoka et al. (1992) Mol. Brain Res. 16:274; Boundy et al. (1998) J. Neurosci. 18:9989; and Kaneda et al. (1991) Neuron 6:583-594); the GnRH promoter (see, for example, Radovick et al. (1991) Proc. Natl. Acad. Sci. USA 88:3402-3406); the L7 promoter (see, for example, Oberdick et al. (1990) Science 248:223-226); the DNMT promoter (see, for example, Bartge et al. (1988) Proc. Nat / . Acad. Sci. USA 85:3648-3652); the enkephalin promoter (see, for example, Comb et al. (1988) EMBO J. 17:3793-3805); the myelin basic protein (MBP) promoter; the Ca2+-calmodulin-dependent protein kinase II-α (CamKIIa) promoter (see, for example, Mayford et al. (1996) Proc. Nat / . Acad. Sci. USA 93:13250; and Casanova et al. (2001) Genesis 31:37); the CMV enhancer / platelet-derived growth factor-β promoter (see, for example, Liu et al. (2004) Gene Therapy 11:52-60), etc., but are not limited thereto.

[0267] Examples of adipocyte-specific spatially restricted promoters include the aP2 gene promoter / enhancer, such as the region from -5.4 kb to +21 bp of the human aP2 gene (see, for example, Tozzo et al. (1997) Endocrinol. 138:1604; Ross et al. (1990) Proc. Natl. Acad. Sci. USA 87:9590; and Pavjani et al. (2005) Nat. Med. 11:797); the glucose transporter-4 (GLUT4) promoter (see, for example, Knight et al. (2003) Proc. Natl. Acad. Sci. USA 100:14725); the fatty acid translocase (FAT / CD36) promoter (see, for example, Kuriki et al. (2002) Biol. Pharm. Bull. 25:1476; and Sato et al. (2002) J. Biol. Chem. 277:15703); the stearoyl-CoA desaturase-1 (SCD1) promoter (see, for example, Tabor et al. (1999) J. Biol. Chem. 274:20603); the leptin promoter (see, for example, Mason et al. (1998) Endocrinol. 139:1013; and Chen et al. (1999) Biochem. Biophys. Res. Comm. 262:187); the adiponectin promoter (see, for example, Kita et al. (2005) Biochem. Biophys. Res. Comm. 331:484; and Chakrabarti (2010) Endocrinol. 151:2408); the adipsin promoter (see, for example, Platt et al. (1989) Proc. Natl. Acad. Sci. USA 86:7490); the resistin promoter (see, for example, Seo et al. (2003) Malec. Endocrinol. 17:1522); and the like, but are not limited thereto.

[0268] Cardiac cell-specific spatially restricted promoters include, but are not limited to, control sequences derived from the following genes: myosin light chain-2, α-myosin heavy chain, AE3, cardiac troponin C, cardiac actin, etc. Franz et al. (1997) Cardiovasc. Res. 35:560-566; Robbins et al. (1995) Ann. N.Y. Acad. Sci. 752:492-505; Linn et al. (1995) Circ. Res. 76:584-591; Parmacek et al. (1994) Mol. Cell. Biol. 14:1870-1885; Hunter et al. (1993) Hypertension 22:608-617; and Sartorelli et al. (1992) Proc. Natl. Acad. Sci. USA 89:4047-4051.

[0269] Smooth muscle cell-specific spatially restricted promoters include, but are not limited to, the SM22α promoter (see, for example, Akyilrek et al. (2000) Mol. Med. 6:983; and U.S. Patent No. 7,169,874); the smooth muscle myosin promoter (see, for example, WO 2001 / 018048 pamphlet); the α-smooth muscle actin promoter, etc. For example, a 0.4 kb region of the SM22α promoter, which contains two CArG elements, has been shown to mediate vascular smooth muscle cell-specific expression (see, for example, Kim et al., (1997) Mol. Cell. Biol. 17, 2266-2278; Li et al., (1996) J. Cell Biol. 132, 849-859; and Moessler et al., (1996) Development 122, 2415-2425).

[0270] Examples of photoreceptor-specific spatially restricted promoters include, but are not limited to, the rhodopsin promoter; the rhodopsin kinase promoter (Young et al. (2003) Ophthalmol. Vis. Sci. 44:4076); the β-phosphodiesterase gene promoter (Nicoud et al. (2007) J. Gene. Med. 9:1015); the retinitis pigmentosa gene promoter (Nicoud et al. (2007) as described above); the interphotoreceptor retinoid-binding protein (IRBP) gene enhancer (Nicoud et al. (2007) as described above); the IRBP gene promoter (Yokoyama et al. (1992) Exp Eye Res. 55:225); and the like.

[0271] Compositions comprising guide RNA In some embodiments, compositions comprising guide RNA are provided herein. In addition to the guide RNA, the composition can comprise one or more of: salts such as NaCl, MgCl2, KCl, MgSO4, etc.; buffers such as Tris buffer, N-(2-hydroxyethyl)piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), MES sodium salt, 3-(N-morpholino)propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS), etc.; solubilizing agents; detergents such as nonionic detergents such as Tween-20, etc.; nuclease inhibitors; and the like. For example, in some embodiments, the composition comprises a buffer for stabilizing the guide RNA and the nucleic acid.

[0272] In some embodiments, the guide RNA present in the composition is pure, e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or greater than 99% pure, where “% purity” means the stated percentage that the guide RNA is free of other macromolecules or contaminants that may be present during the production of the guide RNA.

[0273] Composition comprising the B-GEn.16 polypeptide In some embodiments, provided herein are compositions comprising a B-GEn.16 polypeptide or a variant thereof expressed from a codon-optimized polynucleotide sequence. In addition to the B-GEn.16 polypeptide or a variant thereof, one or more of: salts such as NaCl, MgCl2, KCl, MgSO4, etc.; buffers such as Tris buffer, HEPES, MES, sodium salt of MES, MOPS, TAPS, etc.; solubilizing agents; detergents such as non-ionic surfactants such as Tween-20, etc.; protease inhibitors; reducing agents (e.g., dithiothreitol); etc. can be included.

[0274] In some embodiments, the B-GEn.16 polypeptide or a variant thereof present in the composition is pure, e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or greater than 99% pure, where “% purity” means the percent as described that the B-GEn.16 polypeptide or a variant thereof does not contain other proteins, other macromolecules, or contaminants that may be present during the production of the B-GEn.16 polypeptide or a variant thereof.

[0275] Composition comprising a guide RNA and a site-specific modifying polypeptide In some embodiments, provided herein are compositions comprising (i) a guide RNA or a polynucleotide encoding a guide RNA; and (ii) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid. In some embodiments, the B-GEn.16 polypeptide or a variant thereof exhibits enzymatic activity to modify a target DNA. In some embodiments, the B-GEn.16 polypeptide or a variant thereof exhibits enzymatic activity to modify a polypeptide encoded by a target DNA. In some embodiments, the B-GEn.16 polypeptide or a variant thereof regulates transcription from a target DNA.

[0276] In some embodiments, the components of the composition are individually pure, e.g., each of the components is at least 75%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 99% pure. In some embodiments, the individual components of the composition are pure prior to being added to the composition.

[0277] Kit In some embodiments, a kit for implementing the methods described herein is provided. The kit can include one or more: for example, a B-GEn.16 polypeptide or a variant thereof expressed from a codon-optimized polynucleotide sequence; a guide RNA; a nucleic acid comprising a nucleotide sequence encoding the guide RNA. The kit can include a complex comprising two or more: a B-GEn.16 polypeptide or a variant thereof; a nucleic acid comprising a nucleotide encoding the B-GEn.16 polypeptide or a variant thereof; a guide RNA; a nucleic acid comprising a nucleotide sequence encoding the guide RNA. In some embodiments, the kit includes a B-GEn.16 polypeptide or a variant thereof, or a polynucleotide encoding the same. In some embodiments, the active portion of the B-GEn.16 polypeptide or a variant thereof exhibits reduced or inactivated nuclease activity. In some embodiments, the B-GEn.16 polypeptide or a variant thereof is a B-GEn.16 fusion protein.

[0278] In some embodiments, the kit comprises: (a) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (b) a gRNA or a nucleic acid encoding a gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence. In some embodiments, the kit comprises a nucleic acid comprising a codon-optimized polynucleotide sequence. A kit comprising a B-GEn.16 polypeptide or a variant thereof expressed from the codon-optimized polynucleotide sequence, or a nucleic acid comprising the codon-optimized polynucleotide sequence, can further comprise one or more additional reagents, wherein such additional reagents are selected from, for example, buffers for introducing the B-GEn.16 polypeptide or a variant thereof into cells; wash buffers; control reagents; control expression vectors or polynucleotides; reagents for in vitro production of the B-GEn.16 polypeptide or a variant thereof from DNA, and the like.

[0279] In some embodiments of any of the kits described herein, the kit comprises an sgRNA. In some embodiments, the kit comprises two or more sgRNAs.

[0280] In some embodiments of any of the kits described herein, the gRNA (e.g., comprising two or more guide RNAs) can be provided as an array (e.g., an array of RNA molecules, an array of DNA molecules encoding guide RNAs, etc.). Such kits can be useful, for example, in combination with the above-described genetically modified host cells comprising a B-GEn.16 polypeptide or a variant thereof.

[0281] In some embodiments of any of the kits described herein, the kit further comprises a donor polynucleotide for effecting a desired genetic modification.

[0282] The components of the kit may be in separate containers or may be combined in a single container.

[0283] Any of the kits described herein may further comprise one or more additional reagents, where such additional reagents may be selected from, for example, dilution buffers; reconstitution solutions; wash buffers; control reagents; control expression vectors or polynucleotides; reagents for the in vitro production of a B-GEn.16 polypeptide or variant thereof from DNA, and the like.

[0284] In addition to the above components, the kit may further comprise instructions for using the components of the kit to carry out the method. The instructions for carrying out the method are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate such as paper or plastic. As such, the instructions may be present in the kit as an accompanying document, in the labeling of the kit's container or its components (e.g., associated with a package or subpackage). In some embodiments, the instructions are present as an electronic storage data file on a suitable computer-readable storage medium, such as a CD-ROM, diskette, flash drive, etc. In yet other embodiments, no actual instructions are present in the kit, but means are provided for obtaining the instructions from a remote source, for example via the Internet. An example of this embodiment is a kit that includes a web address from which the instructions can be viewed and / or downloaded. Similar to the instructions, this means for obtaining the instructions is recorded on a suitable substrate.

[0285] The disclosed method A method of modifying a target DNA and / or a polypeptide encoded by the target DNA In some embodiments, methods are provided herein for modifying a target DNA and / or a polypeptide encoded by the target DNA. In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 1 or a variant thereof having at least 90% sequence identity with SEQ ID NO: 1 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) providing a gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (a "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0286] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 2 or a variant thereof having at least 90% sequence identity with SEQ ID NO: 2 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) providing a gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (a "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0287] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 3 or a variant thereof having at least 90% sequence identity to SEQ ID NO: 3 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) a gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (a "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0288] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 4 or a variant thereof having at least 90% sequence identity to SEQ ID NO: 4 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) a gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (a "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0289] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 5 or a variant thereof having at least 90% sequence identity with SEQ ID NO: 5 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) a gRNA or a nucleic acid encoding a gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA (the "targeting complex") is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0290] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 6 or a variant thereof having at least 90% sequence identity with SEQ ID NO: 6 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) a gRNA or a nucleic acid encoding a gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA (the "targeting complex") is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0291] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 7 or a variant thereof having at least 90% sequence identity to SEQ ID NO: 7 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from such nucleic acid; and (ii) a gRNA or a nucleic acid encoding a gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (the "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0292] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 8 or a variant thereof having at least 90% sequence identity to SEQ ID NO: 8 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from such nucleic acid; and (ii) a gRNA or a nucleic acid encoding a gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (the "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0293] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 9 or a variant thereof having at least 90% sequence identity to SEQ ID NO: 9 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) a gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (a "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0294] In some embodiments, the method comprises: (i) a nucleic acid encoding SEQ ID NO: 133 or a variant thereof having at least 90% sequence identity to SEQ ID NO: 133 encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (ii) a gRNA or a nucleic acid encoding the gRNA, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence, such that a complex (a "targeting complex") comprising the B-GEn.16 polypeptide or a variant thereof and the gRNA is formed and comes into contact with a target DNA comprising the target polynucleotide sequence.

[0295] In some embodiments, provided herein is a method of targeting, editing, modifying, or manipulating target DNA at one or more positions in a cell or in an in vitro environment, comprising: (a) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding, for example, a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and (b) introducing a gRNA or a nucleic acid encoding the gRNA into the cell or in vitro environment, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence in the target DNA. In some embodiments, the method comprises introducing a nucleic acid comprising the codon-optimized polynucleotide sequence into the cell or in vitro environment. In some embodiments, the method comprises introducing a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid into the cell or in vitro environment. In some embodiments, the B-GEn.16 polypeptide comprises (or consists of) the amino acid sequence of SEQ ID NO:1. In some embodiments, the method comprises introducing the gRNA into the cell or in vitro environment. In some embodiments, the method comprises introducing a nucleic acid encoding the gRNA into the cell or in vitro environment. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the method comprises introducing one or more additional gRNAs or a nucleic acid encoding one or more additional gRNAs that target the target DNA into the cell or in vitro environment. In some embodiments, the method further comprises introducing a donor template into the cell or in vitro environment.

[0296] In some embodiments, provided herein is a method of targeting, editing, modifying, or manipulating target DNA at one or more positions in a cell or in an in vitro environment, comprising: (a) a nucleic acid encoding a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from such a nucleic acid; and (b) introducing a gRNA or a nucleic acid encoding a gRNA into the cell or in the in vitro environment, wherein the gRNA is capable of directing the B-GEn.16 polypeptide or a variant thereof to a target polynucleotide sequence in the target DNA. In some aspects, the method comprises introducing a B-GEn.16 polypeptide or a variant thereof expressed from a nucleic acid into the cell or in the in vitro environment. In some embodiments, the B-GEn.16 polypeptide or a variant thereof comprises the amino acid sequence of SEQ ID NO:1, or a variant thereof having at least 95% sequence identity to those amino acid sequences. In some embodiments, the method comprises introducing a gRNA into the cell or in the in vitro environment. In some embodiments, the method comprises introducing a nucleic acid encoding a gRNA into the cell or in the in vitro environment. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the method comprises introducing one or more additional gRNAs or nucleic acids encoding one or more additional gRNAs that target the target DNA into the cell or in the in vitro environment. In some embodiments, the method further comprises introducing a donor template into the cell or in the in vitro environment.

[0297] As described above, the gRNA or sgRNA and the B-GEn.16 polypeptide or variant thereof can form a ribonucleoprotein complex. The guide RNA provides target specificity to the complex by including a nucleotide sequence complementary to the sequence of the target DNA. The B-GEn.16 polypeptide or variant thereof of the complex provides endonuclease activity. In some embodiments, the complex modifies the target DNA, resulting in, for example, DNA cleavage, DNA methylation, DNA damage, DNA repair, etc. In some embodiments, the complex modifies a target polypeptide associated with the target DNA (e.g., histone, DNA-binding protein, etc.), resulting in, for example, histone methylation, histone acetylation, histone ubiquitination, etc. The target DNA can be, for example, naked DNA in vitro (e.g., not bound by a DNA-binding protein), chromosomal DNA in cells in vitro, chromosomal DNA in cells in vivo, etc.

[0298] The nuclease activity of the B-GEn.16 polypeptide or its variant described in this specification can cleave target DNA to generate double-strand breaks. These breaks are then repaired by the cell in one of two ways: non-homologous end joining and homologous recombination repair. In non-homologous end joining (NHEJ), double-strand breaks are repaired by directly ligating the broken ends to each other. In this process, several base pairs can be inserted or deleted at the cleavage site. In homologous recombination repair, a donor polynucleotide having homology to the cleaved target DNA sequence is used as a template for the repair of the cleaved target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. As such, new nucleic acid material can be inserted / copied at that site. In some embodiments, the target DNA is contacted with the donor polynucleotide. In some embodiments, the donor polynucleotide is introduced into the cell. Modification of the target DNA by NHEJ and / or homologous recombination repair can result in, for example, gene correction, gene substitution, gene tagging, transgene insertion, nucleotide deletion, nucleotide insertion, gene disruption, gene mutation, sequence substitution, etc. Thus, cleavage of DNA by the B-GEn.16 polypeptide or its variant can be used to delete nucleic acid material from the target DNA sequence by allowing the cell to repair the sequence in the absence of an exogenously provided donor polynucleotide (e.g., to disrupt a gene that makes a cell susceptible to infection (e.g., the CCRS or CXCR4 gene that makes it susceptible to HIV infection), to remove trinucleotide repeat sequences that cause disease in neurons, to create gene knockouts and mutations as disease models in research, etc.). Thus, the method can be used to knock out a gene (resulting in a complete absence of transcription / translation or a change in transcription / translation) or to knock in genetic material at a selected locus in the target DNA.

[0299] In addition to, or alternatively to, when the guide RNA and the B-GEn.16 polypeptide or a variant thereof are co-administered to a cell with a donor polynucleotide sequence comprising a segment having at least homology to the target DNA sequence, the subject method can be used for adding nucleic acid material to the target DNA sequence, for example, for inserting or replacing (e.g., for "knocking in" a nucleic acid encoding a protein, siRNA, miRNA, etc.), for adding a tag (e.g., 6xHis, a fluorescent protein (e.g., green fluorescent protein, yellow fluorescent protein, etc.), hemagglutinin (HA), FLAG, etc.), for adding regulatory sequences to a gene (e.g., a promoter, a polyadenylation signal, an internal ribosome entry sequence (IRES), a 2A peptide, a start codon, a stop codon, a splice signal, a localization signal, etc.), for modifying a nucleic acid sequence (e.g., introducing a mutation), etc. Thus, the complex comprising the guide RNA and the B-GEn.16 polypeptide or a variant thereof is useful for any in vitro or in vivo application where it is desirable to site-specifically modify DNA in a "targeted" manner, for example, for gene knockout, gene knockin, gene editing, gene tagging, sequence replacement, etc., such as for treating a disease, for example, as used in gene therapy, etc., or as an antiviral, anti-pathogenic, or anti-cancer therapeutic agent or for the production of genetically modified organisms in agriculture, for the large-scale production of proteins by cells for therapeutic, diagnostic, or research purposes, for the induction of iPS cells, for biological research, for targeting the genes of pathogens for deletion or replacement, etc.

[0300] In some embodiments, the methods described herein use a B-GEn.16 polypeptide or a variant thereof that includes a heterologous sequence (e.g., a B-GEn.16 fusion polypeptide). In some embodiments, the heterologous sequence can provide for the intracellular localization of the B-GEn.16 polypeptide or a variant thereof (e.g., a nuclear localization signal (NLS) for targeting to the nucleus; a mitochondrial localization signal for targeting to the mitochondria; a chloroplast localization signal for targeting to the chloroplast; an ER retention signal, etc.). In some embodiments, the heterologous sequence can provide a tag (e.g., a fluorescent protein, e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, etc.; a histidine tag, e.g., a 6XHis tag; a hemagglutinin (HA) tag; a FLAG tag; a Myc tag, etc.) to facilitate tracking or purification. In some embodiments, the heterologous sequence can provide for increased or decreased stability.

[0301] In some embodiments, the methods described herein use a guide RNA and a B-GEn.16 polypeptide or variant thereof as an inducible system for blocking gene expression in a target cell. In some embodiments, a nucleic acid encoding an appropriate guide RNA and / or an appropriate B-GEn.16 polypeptide or variant thereof is integrated into the chromosome of the target cell and is under the control of an inducible promoter. When the guide RNA and / or the B-GEn.16 polypeptide or variant thereof is induced, the target DNA is cleaved (or otherwise modified) at the desired location (e.g., a target gene on another plasmid), and a complex is formed when both the guide RNA and the B-GEn.16 polypeptide or variant thereof are present. Thus, in some embodiments, a target cell is engineered (e.g., under the control of an inducible promoter) to contain a nucleic acid sequence encoding an appropriate B-GEn.16 polypeptide or variant thereof in the genome and / or an appropriate guide RNA on a plasmid, enabling experiments in which the expression of any target gene (expressed from another plasmid introduced into the strain) can be controlled by inducing the expression of the guide RNA and the B-GEn.16 polypeptide or variant thereof. In some embodiments, the B-GEn.16 polypeptide or variant thereof has enzymatic activity to modify the target DNA in a way other than introducing a double-strand break. The desired enzymatic activity that can be used to modify the target DNA (e.g., by fusing a heterologous polypeptide having enzymatic activity to the B-GEn.16 polypeptide or variant thereof, thereby generating a B-GEn.16 fusion polypeptide or variant thereof) includes, but is not limited to, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deaminase activity, dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity.Methylation and demethylation are recognized in the art as important modes of epigenetic gene regulation, while DNA damage repair is essential for cell survival and proper genome maintenance in response to environmental stress. Thus, the methods herein find use in the epigenetic modification of target DNA and can be used to control the epigenetic modification of target DNA at any position of the target DNA by introducing a desired sequence into the spacer region of the guide RNA. The methods herein also find use in the intentional and controlled damage of DNA at any desired position within the target DNA. The methods herein also find use in the sequence-specific and controlled repair of DNA at any desired position within the target DNA. Methods for targeting DNA-modifying enzyme activity to specific positions of target DNA find use in both research and clinical applications.

[0302] In some embodiments, multiple guide RNAs are used to simultaneously modify different positions on the same target DNA or on different target DNAs. In some embodiments, two or more guide RNAs target the same gene or transcript or locus. In some embodiments, two or more guide RNAs target different non-related loci. In some embodiments, two or more guide RNAs target different but related loci.

[0303] In some embodiments, the B-GEn.16 polypeptide or variant thereof is provided directly as a protein. As one non-limiting example, a fungus (e.g., yeast) can be transformed with an exogenous protein and / or nucleic acid using spheroplast transformation (see Kawai et al., Bioeng Bugs. 2010 Nov-Dec;1(6):395-403: “transformation of Saccharomyces cerevisiae and other fungi: methods and possible underlying mechanism”; and Tanka et al., Nature. 2004 Mar 18;428(6980):323-8: “Conformational variations in an infectious protein determine prion strain differences”; both of which are incorporated herein by reference in their entirety). Thus, the B-GEn.16 polypeptide or variant thereof can be incorporated into a spheroplast (with or without a nucleic acid encoding a guide RNA, with or without a donor polynucleotide), and the spheroplast can be used to introduce the contents into a yeast cell. The B-GEn.16 polypeptide or variant thereof can be introduced into (provided to) the cell by any suitable method; such methods are known to those of skill in the art. As another non-limiting example, the B-GEn.16 polypeptide or variant thereof can be directly injected into a cell (with or without a nucleic acid encoding a guide RNA, with or without a donor polynucleotide), such as a cell of a zebrafish embryo, the pronucleus of a fertilized mouse oocyte, etc.

[0304] Method of regulating transcription In some embodiments, provided herein is a method of regulating the transcription of a target nucleic acid in a host cell. This method generally involves contacting the target nucleic acid with an enzymatically inactive B-GEn.16 polypeptide and a guide RNA. These methods are also useful in the various applications provided.

[0305] The transcription regulation method of the present disclosure overcomes some of the drawbacks of methods involving RNAi. The transcription regulation method of the present disclosure finds use in a variety of applications, including research applications, drug discovery (e.g., high-throughput screening), target validation, industrial applications (e.g., crop engineering; microbial engineering, etc.), diagnostic applications, therapeutic applications, and imaging technologies.

[0306] In some embodiments, provided herein is a method for selectively regulating the transcription of a target DNA in a host cell, such as a human cell. The method generally comprises a) introducing into the host cell i) a guide RNA, or a nucleic acid comprising a nucleotide sequence encoding the guide RNA, and ii) a B-GEn.16 polypeptide or a variant thereof, or a nucleic acid comprising a nucleotide sequence encoding the B-GEn.16 polypeptide or a variant thereof, wherein the B-GEn.16 polypeptide or a variant thereof exhibits reduced endodeoxyribonuclease activity. The guide RNA and the B-GEn.16 polypeptide or a variant thereof form a complex in the host cell; the complex selectively regulates the transcription of the target DNA in the host cell.

[0307] In some embodiments, the methods described herein use a modified form of the B-GEn.16 protein. In some aspects, the modified form of the B-GEn.16 protein comprises an amino acid change (e.g., a deletion, insertion, or substitution) that reduces the nuclease activity of the B-GEn.16 protein. For example, in some embodiments, the modified form of the B-GEn.16 protein has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding unmodified B-GEn.16 polypeptide. In some embodiments, the modified form of the B-GEn.16 polypeptide has substantially no nuclease activity. When the B-GEn.16 polypeptide or a variant thereof is a modified form of the B-GEn.16 polypeptide having substantially no nuclease activity, it can be referred to as "dB-GEn.16".

[0308] In some embodiments, the transcription regulation methods described herein enable the selective regulation (e.g., decrease or increase) of a target nucleic acid in a host cell. For example, a “selective” decrease in the transcription of a target nucleic acid reduces the transcription of the target nucleic acid by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or more than 90% compared to the level of transcription of the target nucleic acid in the absence of a complex of a guide RNA / B-GEn.16 polypeptide or a variant thereof. A selective decrease in the transcription of a target nucleic acid decreases the transcription of the target nucleic acid but does not substantially decrease the transcription of non-target nucleic acids; for example, the transcription of non-target nucleic acids decreases by less than 10%, if at all, compared to the level of transcription of non-target nucleic acids in the absence of a complex of a guide RNA / B-GEn.16 polypeptide or a variant thereof.

[0309] In some embodiments, the B-GEn.16 polypeptide or a variant thereof has activity to regulate the transcription of target DNA (e.g., in the case of a B-GEn.16 fusion polypeptide or a variant thereof). In some embodiments, a B-GEn.16 fusion polypeptide or a variant thereof comprising a heterologous polypeptide (e.g., a transcriptional activator or transcriptional repressor polypeptide) that exhibits the ability to increase or decrease transcription is used to increase or decrease the transcription of target DNA at a specific position in the target DNA induced by the spacer of the guide RNA. Examples of source polypeptides for providing a B-GEn.16 fusion polypeptide or a variant thereof having transcriptional regulatory activity include, but are not limited to, a light-inducible transcriptional regulator, a small molecule / drug-responsive transcriptional regulator, a transcription factor, a transcriptional repressor, and the like. In some embodiments, the method is used to control the transcription of a target gene-encoded RNA (protein-encoding mRNA) and / or a target non-coding RNA (e.g., tRNA, rRNA, snoRNA, siRNA, miRNA, long ncRNA, etc.). In some embodiments, the B-GEn.16 polypeptide or a variant thereof has enzymatic activity to modify a polypeptide that associates with DNA (e.g., a histone). In some embodiments, the enzymatic activity is methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity (e.g., ubiquitination activity), deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, glycosylation activity (e.g., from GlcNAc transferase) or deglycosylation activity. The enzymatic activities listed herein catalyze covalent modification to proteins. Such modifications are known in the art to change the stability or activity of the target protein (e.g., phosphorylation by kinase activity can stimulate or silence protein activity depending on the target protein). Of particular interest as protein targets are histones.It is known in the art that histone proteins bind to DNA and form a complex known as a nucleosome. Histones can be modified (e.g., by methylation, acetylation, ubiquitination, phosphorylation) to induce structural changes in the surrounding DNA, and thus control the accessibility of potentially large portions of the DNA to interacting factors such as transcription factors, polymerases. A single histone can be modified in many different ways and in many different combinations (e.g., trimethylation of lysine 27 of histone 3, H3K27, associates with DNA regions of repressed transcription, while trimethylation of lysine 4 of histone 3, H3K4, associates with DNA regions of active transcription). Thus, a B-GEn.16 fusion polypeptide or variant thereof having histone modification activity can find use in the site-specific control of chromosome structure and can be used to alter the histone modification pattern in selected regions of target DNA. Such methods find use in both research and clinical applications.

[0310] Increased transcription "Selective" increased transcription of the target DNA can increase transcription from the target DNA by at least 1.1-fold (e.g., at least 1.2-fold, at least 1.3-fold, at least 1.4-fold, at least 1.5-fold, at least 1.6-fold, at least 1.7-fold, at least 1.8-fold, at least 1.9-fold, at least 2-fold, at least 2.5-fold, at least 3-fold, at least 3.5-fold, at least 4-fold, at least 4.5-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 12-fold, at least 15-fold, or at least 20-fold) compared to the level of transcription from the target DNA in the absence of a complex of guide RNA / B-GEn.16 polypeptide or a variant thereof. The selective increase in transcription of the target DNA increases transcription from the target DNA but does not substantially increase transcription of non-target DNA. For example, transcription of non-target DNA is less than about 5-fold (e.g., less than about 4-fold, less than about 3-fold, less than about 2-fold, less than about 1.8-fold, less than about 1.6-fold, less than about 1.4-fold, less than about 1.2-fold, or less than about 1.1-fold), if any, compared to the level of transcription of non-target DNA in the absence of a complex of guide RNA / B-GEn.16 polypeptide or a variant thereof.

[0311] As a non-limiting example, increased transcription can be achieved by fusing dB-GEn.16 to a heterologous sequence. Suitable fusion partners include, but are not limited to, polypeptides that provide an activity that indirectly increases transcription by acting directly on the target DNA or a polypeptide that associates with the target DNA (e.g., a histone or other DNA-binding protein). Suitable fusion partners include, but are not limited to, polypeptides that provide methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, desumoylation activity, ribosylation activity, deribosylation activity, crotonylation, decrotonylation, propionylation, depropionylation, myristoylation activity, or demyristoylation activity.

[0312] Additional suitable fusion partners include, but are not limited to, polypeptides that directly provide for an increase in the transcription of a target nucleic acid (e.g., a transcriptional activator or a fragment thereof, a protein that recruits a transcriptional activator or a fragment thereof, a small molecule / drug-responsive transcriptional regulator, etc.).

[0313] Non-limiting examples of methods of using the dB-GEn.16 fusion protein to increase transcription in prokaryotes include modifications of the bacterial one-hybrid (B1H) or two-hybrid (B2H) systems. In the B1H system, the DNA-binding domain (BD) is fused to a bacterial transcriptional activation domain (AD, e.g., the α subunit of Escherichia coli RNA polymerase (RNAPa)). Thus, dB-GEn.16 can be fused to a heterologous sequence containing the AD. When the dB-GEn.16 fusion protein reaches the upstream region of a promoter (targeted by a guide RNA), the AD (e.g., RNAPa) of the dB-GEn.16 fusion protein recruits the RNAP holoenzyme, resulting in transcriptional activation. In the B2H system, the BD is not directly fused to the AD; instead, their interaction is mediated by a protein-protein interaction (e.g., the GAL11P-GAL4 interaction). To modify such a system for use in the present method, dB-GEn.16 can be fused to a first protein sequence that provides a protein-protein interaction (e.g., the yeast GAL11P and / or GAL4 protein), and RNAa can be fused to a second protein sequence that completes the protein-protein interaction (e.g., GAL4 when GAL11P is fused to dB-GEn.16, GAL11P when GAL4 is fused to dB-GEn.16, etc.). The binding affinity between GAL11P and GAL4 increases the binding efficiency and the transcriptional firing rate.

[0314] Non-limiting examples of methods of using the dB-GEn.16 fusion protein to increase transcription in eukaryotes include fusions to the activation domain (AD) of dB-GEn.16 (e.g., GAL4, herpes virus activation protein VP16 or VP64, human nuclear factor NF-κB p65 subunit, etc.). To make the system inducible, the expression of the dB-GEn.16 fusion protein can be controlled by an inducible promoter (e.g., Tet-ON, Tet-OFF, etc.). Guide RNAs can be designed to target known transcriptional response elements (e.g., promoters, enhancers, etc.), known upstream activation sequences (UASs), sequences of unknown or known function suspected of being able to control the expression of the target DNA, and the like.

[0315] Additional fusion partners Non-limiting examples of fusion partners to achieve increased or decreased transcription include, but are not limited to, transcription activators and transcription repressor domains (e.g., Krueppel-associated box (KRAB or SKD); Mad mSIN3 interaction domain (SID); ERF repressor domain (ERD), etc.). In some such cases, the dB-GEn.16 fusion protein is targeted to a specific location (e.g., sequence) in the target DNA by a guide RNA, blocking RNA polymerase binding to the promoter (selectively inhibiting the transcription activator function) and / or modifying the local chromatin state (e.g., when a fusion sequence is used to modify the target DNA or a polypeptide associated with the target DNA), exerting locus-specific control such as this. In some embodiments, the change is transient (e.g., transcription repression or activation). In some embodiments, the change is heritable (e.g., when epigenetic modifications are made to the target DNA or a protein that binds to the target DNA, such as a nucleosome histone). In some embodiments, the heterologous sequence can be fused to the C-terminus of the dB-GEn.16 polypeptide. In some embodiments, the heterologous sequence can be fused to the N-terminus of the dB-GEn.16 polypeptide. In some embodiments, the heterologous sequence can be fused to an internal portion (e.g., a portion other than the N-terminus or C-terminus) of the dB-GEn.16 polypeptide. The biological effects of the methods using the dB-GEn.16 fusion protein can be detected by any suitable method (e.g., gene expression assays; chromatin-based assays, such as chromatin immunoprecipitation (ChIP), chromatin in vivo assay (CiA), etc.).

[0316] In some embodiments, the method involves the use of two or more different guide RNAs. For example, two different guide RNAs can be used in a single host cell, where the two different guide RNAs target two different target sequences in the same target nucleic acid. In some embodiments, the use of two different guide RNAs that target two different target sequences in the same target nucleic acid provides increased regulation (e.g., decrease or increase) of the transcription of the target nucleic acid.

[0317] As another example, two different guide RNAs can be used in a single host cell, where the two different guide RNAs target two different target nucleic acids. Thus, for example, the method of transcriptional regulation can further include introducing into the host cell a second guide RNA, or a nucleic acid comprising a nucleotide sequence encoding the second guide RNA.

[0318] In some embodiments, the nucleic acid (e.g., guide RNA, e.g., single molecule guide RNA; donor polynucleotide; nucleic acid encoding a B-GEn.16 polypeptide or variant thereof, etc.) comprises a modification or sequence that provides additional desirable features (e.g., modified or regulated stability; intracellular targeting; tracking, e.g., fluorescent labeling; binding sites for proteins or protein complexes, etc.). Non-limiting examples include a 5' cap (e.g., 7-methylguanylate cap (m7G)); a 3' polyadenylation tail (e.g., 3' poly(A) tail); riboswitch sequences or aptamer sequences (e.g., for enabling regulated stability and / or regulated accessibility by proteins and / or protein complexes); terminator sequences; sequences that form dsRNA duplexes (e.g., hairpins); modifications or sequences that target the RNA to subcellular locations (e.g., nucleus, mitochondria, chloroplasts, etc.); modifications or sequences that provide tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, sequences that enable fluorescent detection, etc.); modifications or sequences that provide binding sites for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); modifications of such RNA that change its structure and result in a B-GEn.16 ribonucleoprotein; and combinations thereof.

[0319] Multiple simultaneous guide RNAs In some embodiments, multiple guide RNAs are used simultaneously in the same cell to regulate transcription simultaneously at different positions on the same target DNA or on different target DNAs. In some embodiments, two or more guide RNAs target the same gene or transcript or locus. In some embodiments, two or more guide RNAs target different non-related loci. In some embodiments, two or more guide RNAs target different but related loci.

[0320] Because guide RNAs are small and robust, they can be present simultaneously on the same expression vector and, if desired, can be under the same transcriptional control. In some embodiments, two or more (e.g., three or more, four or more, five or more, ten or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, thirty-five or more, forty or more, forty-five or more, or fifty or more) guide RNAs are expressed simultaneously in a target cell (from the same or different vectors / from the same or different promoters). In some embodiments, multiple guide RNAs can be encoded in an array that mimics the naturally occurring CRISPR array of a targeter RNA. The targeting segments are encoded as sequences about 30 nucleotides in length (which can be from about 16 to about 100 nt) and are separated by CRISPR repeat sequences. The array can be introduced into the cell as DNA encoding the RNA or as RNA.

[0321] To express multiple guide RNAs, an artificial RNA processing system mediated by the Csy4 endoribonuclease can be used. For example, multiple guide RNAs can be linked in a tandem array on a precursor transcript (e.g., expressed from the U6 promoter) and separated by Csy4-specific RNA sequences. The co-expressed Csy4 protein cleaves the precursor transcript into multiple guide RNAs. Advantages of using the RNA processing system include, first, that there is no need to use multiple promoters and, second, that because all guide RNAs are processed from the precursor transcript, their concentrations are normalized for similar dB-GEn.16 binding.

[0322] Csy4 is a small endoribonuclease (RNase) protein derived from the bacterium Pseudomonas aeruginosa. Csy4 specifically recognizes RNA hairpins of at least 17 bp and exhibits rapid (<1 minute) and highly efficient (>99.9%) RNA cleavage. Unlike most RNases, the cleaved RNA fragments are stable and functionally active. Csy4-based RNA cleavage can be reused in an artificial RNA processing system. In this system, a 17-bp RNA hairpin is inserted between multiple RNA fragments transcribed as precursor transcripts from a single promoter. Co-expression of Csy4 is effective for generating individual RNA fragments.

[0323] host cell In some embodiments, the methods of the disclosure can be used to induce transcriptional regulation in mitotic or post-mitotic cells in vivo and / or ex vivo and / or in vitro. In some embodiments, the methods of the disclosure can be used to induce DNA cleavage, DNA modification, and / or transcriptional regulation in mitotic or post-mitotic cells in vivo and / or ex vivo and / or in vitro (e.g., to produce genetically modified cells that can be reintroduced into an individual).

[0324] Because the guide RNA provides specificity by hybridizing to the target DNA, the mitotic and / or post-mitotic cells can be any of a variety of host cells, where the host cells include bacterial cells; archaeal cells; unicellular eukaryotes; plant cells; algal cells, such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorela pyrenoidosa, Sargassum patens, C. Agardh, etc.; fungal cells; animal cells; cells derived from invertebrates (e.g., insects, cnidarians, echinoderms, nematodes, etc.); eukaryotic parasites (e.g., Plasmodium, such as Plasmodium falciparum); helminths, etc.; cells derived from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals, etc.); mammalian cells, such as rodent cells, human cells, non-human primate cells, etc., but are not limited thereto. Suitable host cells include naturally occurring cells; genetically modified cells (e.g., cells genetically modified in the laboratory, e.g., by "human hands"); and cells manipulated in vitro by any method. In some embodiments, the host cell is isolated.

[0325] Any type of cell (e.g., stem cells, such as embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells; somatic cells, such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells; any stage, such as in vitro or in vivo embryonic cells at the 1-cell stage, 2-cell stage, 4-cell stage, 8-cell stage, etc., zebrafish embryos; etc.) can be the subject. The cells may be from an established cell line or they may be primary cells, where "primary cells", "primary cell line", and "primary culture" are used interchangeably herein and refer to cells and cell cultures derived from the subject and grown in vitro during a limited number of passages, e.g., during division. For example, a primary culture may have been passaged 0, 1, 2, 4, 5, 10, or 15 times, but includes cultures that have not passed through a sufficient number of passages to reach the crisis stage. A primary cell line can be maintained in vitro for less than 10 passages. The target cells are, in some embodiments, single-celled organisms or are grown in culture.

[0326] When the cells are primary cells, such cells can be harvested from an individual by any suitable method. For example, white blood cells can be appropriately harvested by apheresis, leukapheresis, density gradient separation, etc., while cells from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, stomach, etc. are most preferably harvested by biopsy. For the dispersion or suspension of the recovered cells, a suitable solution can be used. Such solutions are generally balanced salt solutions, such as physiological saline, phosphate buffered saline (PBS), Hank's balanced salt solution, etc., which are supplemented with fetal bovine serum or other naturally occurring factors in combination with an acceptable buffer at a low concentration, for example, 5 - 25 mM. Suitable buffers include HEPES, phosphate buffer, lactate buffer, etc. The cells may be used immediately or may be stored and frozen for a long time and can be thawed and reused. In such cases, the cells are generally frozen in 10% dimethyl sulfoxide (DMSO), 50% serum, 40% buffered medium, or other solutions commonly used in the art for storing cells at such freezing temperatures, and are thawed by methods generally known in the art for thawing frozen cultured cells.

[0327] Use Methods for modulating transcription according to the present disclosure find use in a variety of applications, which are also provided. The applications include research applications, diagnostic applications, industrial applications, and therapeutic applications.

[0328] Research applications include, for example, determining the effect of decreasing or increasing the transcription of a target nucleic acid on, for example, development, metabolism, expression of downstream genes, etc. High-throughput genomic analysis can be carried out using a transcription regulation method that only requires changing the spacer of the guide RNA, while the protein-binding segment and the transcription termination segment can be kept constant (in some cases). A library containing a plurality of nucleic acids used for genomic analysis includes a promoter operably linked to a nucleotide sequence encoding the guide RNA, and each nucleic acid includes a common protein-binding segment, different spacers, and a common transcription termination segment. The chip is 5×104 It can contain unique guide RNAs exceeding. Applications include large-scale phenotyping, gene-function mapping, and metagenomic analysis.

[0329] The methods disclosed herein find use in the field of metabolic engineering. Since the transcription level can be efficiently and predictably controlled by designing appropriate guide RNAs as disclosed herein, the activity of a metabolic pathway (e.g., a biosynthetic pathway) can be precisely controlled and regulated by controlling the level of a specific enzyme within the metabolic pathway of interest (e.g., via increased or decreased transcription). The metabolic pathway of interest includes those used for the production of chemicals (fine chemicals, fuels, antibiotics, toxins, agonists, antagonists, etc.) and / or drugs.

[0330] The biosynthetic pathways for the purpose include: (1) the mevalonate pathway (e.g., the HMG-CoA reductase pathway) (which converts acetyl-CoA into dimethylallyl pyrophosphate (DMAPP) and isopentenyl pyrophosphate (IPP), and these are used in the biosynthesis of a wide variety of biomolecules including terpenoids / isoprenoids), (2) the non-mevalonate pathway (e.g., the "2-C-methyl-D-erythritol 4-phosphate / 1-deoxy-D-xylulose 5-phosphate pathway" or "MEP / DOXP pathway" or "DXP pathway") (which, instead, converts pyruvate and glyceraldehyde 3-phosphate into DMAPP and IPP via an alternative pathway to the mevalonate pathway), (3) the polyketide synthesis pathway (which generates various polyketides through various polyketide synthases. Polyketides include natural small molecules used in chemotherapy (e.g., tetracycline and macrolides), and industrially important polyketides include rapamycin (an immunosuppressant), erythromycin (an antibiotic), lovastatin (an anti-cholesterol drug), and epothilone B (an anti-cancer drug)), (4) the fatty acid synthesis pathway, (5) the DAHP (3-deoxy-D-arabinose heptulosonate 7-phosphate) synthesis pathway, (6) pathways for generating potential biofuels (such as short-chain alcohols and alkanes, fatty acid methyl esters and fatty alcohols, isoprenoids, etc.), but are not limited thereto.

[0331] Networks and Cascades The methods disclosed herein can be used to design an integrated network of controls (e.g., a cascade or multiple cascades). For example, guide RNAs and B-GEn.16 polypeptides or variants thereof can be used to control (e.g., regulate, e.g., increase, decrease) the expression of another DNA-targeting RNA or another B-GEn.16 polypeptide or variant thereof. For example, a first guide RNA can be designed to target the regulation of transcription of a second fusion dB-GEn.16 polypeptide having a function different from that of the first B-GEn.16 polypeptide or variant thereof (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, etc.). In some embodiments, the second fusion dB-GEn.16 polypeptide can be selected so as not to interact with the first guide RNA. In some embodiments, the second fusion dB-GEn.16 polypeptide can be selected so as to interact with the first guide RNA. In some such cases, the activities of two (or more) dB-GEn.16 proteins can compete (e.g., if the polypeptides have opposing activities) or synergize (e.g., if the polypeptides have similar or synergistic activities). Similarly, as described above, any of the complexes (e.g., guide RNA / dB-GEn.16 polypeptide) in the network can be designed to control other guide RNAs or dB-GEn.16 polypeptides. Since guide RNAs and B-GEn.16 polypeptides or variants thereof can target any desired DNA sequence, the methods described herein can be used to control and regulate the expression of any desired target. The integrated networks (e.g., cascades of interactions) that can be designed range from very simple to very complex and are not limited.

[0332] In a network where two or more components (e.g., guide RNA and dB-GEn.16 polypeptide) are each under the regulatory control of a separate guide RNA / dB-GEn.16 polypeptide complex, the expression level of one component of the network can affect (e.g., increase or decrease) the expression level of another component of the network. Through this mechanism, the expression of one component can affect the expression of different components in the same network, and the network can include a mixture of components that increase the expression of other components, as well as components that decrease the expression of other components. As will be readily understood by those skilled in the art, the above examples where the expression level of one component can affect the expression level of one or more different components are for illustrative purposes only and are not limiting. When one or more components are modified such that they are operable (e.g., under experimental control, e.g., temperature control; drug control, e.g., drug-inducible control; light control, etc.), additional complex layers can optionally be introduced into the network.

[0333] As one non-limiting example, a first guide RNA can bind to the promoter of a second guide RNA that controls the expression of a target therapeutic / metabolic gene. In such a case, the conditional expression of the first guide RNA indirectly activates the therapeutic / metabolic gene. This type of RNA cascade is useful, for example, for easily converting a repressor to an activator and can be used to control the logic or dynamics of the expression of a target gene.

[0334] Transcriptional regulatory methods can also be used for drug discovery and target validation.

[0335] Method of treating a disease or condition In some aspects of the present disclosure, the guide RNA and / or the B-GEn.16 polypeptide or a variant thereof and / or the donor polynucleotide are used to modify cellular DNA in vivo for purposes such as gene therapy, e.g., to treat a disease, or as an antiviral, anti-pathogenic, or anti-cancer therapeutic agent, for the production of genetically modified organisms in agriculture, or for biological research. In these in vivo embodiments, (i) a nucleic acid encoding a guide RNA or gRNA; (ii) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding the B-GEn.16 polypeptide or a variant thereof, or the B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and / or (iii) a component of the CRISPR / B-GEn.16 system comprising a donor polynucleotide are administered to an individual. The administration can be by any method well known in the art for the administration of peptides, small molecules, and nucleic acids to a subject. The components of the CRISPR / B-GEn.16 system can be incorporated into various formulations. More specifically, the components of the CRISPR / B-GEn.16 system of the present disclosure can be formulated into a pharmaceutical composition by combination with a suitable pharmaceutically acceptable carrier or diluent.

[0336] In some embodiments, provided herein is a pharmaceutical preparation or composition comprising components of a CRISPR / B-GEn.16 system that includes (i) a nucleic acid encoding a guide RNA or gRNA; (ii) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and / or (iii) a donor polynucleotide present in a pharmaceutically acceptable vehicle. A “pharmaceutically acceptable vehicle” can be a vehicle approved by a federal or state government regulatory agency or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeias for use in mammals such as humans. The term “vehicle” refers to a diluent, adjuvant, excipient, or carrier with which the compounds of the disclosure are formulated for administration to a mammal. Such pharmaceutical vehicles can be lipids, such as liposomes, such as liposome dendrimers; liquids, such as water and oils, such as peanut oil, soybean oil, mineral oil, sesame oil and the like of petroleum, animal, vegetable or synthetic origin, saline; acacia gum, gelatin, starch paste, talc, keratin, colloidal silica, urea and the like. Additionally, adjuvants, stabilizers, thickeners, lubricants and coloring agents can be used. The pharmaceutical composition can be formulated in solid, semi-solid, liquid or gaseous forms, such as, for example, tablets, capsules, powders, granules, ointments, solutions, suppositories, injections, inhalants, gels, microspheres, and aerosol formulations. Accordingly, administration of the components of the CRISPR / B-GEn.16 system can be accomplished by a variety of methods including oral, buccal, rectal, parenteral, intraperitoneal, intradermal, transdermal, intratracheal, intraocular, and the like. The active agent can be systemic after administration or can be localized by use of an implant that acts to maintain an active dose at the site of local, intramural, or transplantation administration. The active agent can be formulated for immediate activity or for sustained release.

[0337] For some conditions, particularly those of the central nervous system, it may be necessary to formulate agents that pass through the blood-brain barrier (BBB). One strategy for drug delivery across the BBB involves disruption of the BBB either by osmotic means such as mannitol or leukotrienes, or by biochemical means such as the use of vasoactive substances such as bradykinin. The use of BBB opening to target specific agents against brain tumors is also an option. The BBB disrupting agent can be co-administered with the therapeutic composition of the present disclosure when the composition is administered by intravascular injection. Other strategies for passing through the BBB can involve the use of endogenous transport systems, including caveolin-1-mediated transcytosis, carrier-mediated transporters such as glucose and amino acid transporters, receptor-mediated transcytosis of insulin or transferrin, and active efflux transporters such as asp-glycoproteins. The active transport moiety can also be conjugated to the therapeutic compound for use in the present disclosure to facilitate transport across the endothelial wall of the blood vessel. Additionally, or alternatively, drug delivery of the therapeutic agent behind the BBB can be by local delivery, such as intrathecal delivery, for example, through an Ommaya reservoir (see, e.g., U.S. Pat. Nos. 5,222,982 and 5,385,582, which are incorporated herein by reference); bolus injection, such as by syringe, for example, intravitreally or intracranially; continuous infusion, such as by cannula insertion, for example, with convection (see, e.g., U.S. Patent Application Publication No. 2007 / 0254842, which is incorporated herein by reference); or by implanting a device to which the agent is reversibly immobilized (see, e.g., U.S. Patent Application Publication Nos. 2008 / 0081064 and 2009 / 0196903, which are incorporated herein by reference).

[0338] Generally, an effective amount of the components of the CRISPR / B-GEn.16 system is provided, including (i) a nucleic acid encoding a guide RNA or gRNA; (ii) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and / or (iii) a donor polynucleotide. As described above for ex vivo methods, an effective amount or dosage of the components of the CRISPR / B-GEn.16 system in vivo is an amount that induces an increase of at least two-fold in the amount of recombination observed between two homologous sequences, compared to cells contacted with a negative control, such as an empty vector or an irrelevant polypeptide. The amount of recombination can be measured by any suitable method, for example, as described above and known in the art. Calculation of the effective amount or dosage of the components of the CRISPR / B-GEn.16 system to be administered is within the skill of the art and routine for one of ordinary skill in the art. The final amount to be administered depends on the route of administration and the nature of the disorder or condition being treated.

[0339] The effective amount given to a particular subject depends on various factors, some of which vary from subject to subject. A skilled clinician can determine the effective amount of a therapeutic agent to administer to a subject to halt or reverse the progression of a disease state, as needed. Using LD50 animal data and other information available for the agent, the clinician can determine the maximum safe dosage for an individual, depending on the route of administration. For example, considering the larger volume of body fluid into which a therapeutic composition is administered, the dosage administered intravenously may be more than the dosage administered intrathecally. Similarly, a composition that is rapidly removed from the body can be administered at a higher dosage or in repeated dosages to maintain a therapeutic concentration. Using ordinary techniques, a skilled clinician can optimize the dosage of a particular therapeutic agent during the course of routine clinical trials.

[0340] For inclusion in a medicament, the components of the CRISPR / B-GEn.16 system can be obtained from suitable commercial sources. As a general proposition, the pharmaceutically effective total amount of the components of the CRISPR / B-GEn.16 system administered parenterally per dose is within the range that can be measured by a dose-response curve.

[0341] Preparations for therapeutic administration based on the components of the CRISPR / B-GEn.16 system, for example, (i) a nucleic acid encoding a guide RNA or gRNA; (ii) a nucleic acid comprising a codon-optimized polynucleotide sequence encoding a B-GEn.16 polypeptide or a variant thereof, or a B-GEn.16 polypeptide or a variant thereof expressed from the nucleic acid; and / or (iii) a donor polynucleotide must be sterile. Sterility is readily achieved by filtration through a sterile filtration membrane (e.g., a 0.2 micrometer membrane). The therapeutic composition is placed in a container having a sterile outlet, for example, an intravenous fluid bag or vial having a stopper penetrable by a hypodermic needle. Treatments based on the components of the CRISPR / B-GEn.16 system can be stored as an aqueous solution or as a lyophilized formulation for reconstitution in a unit or multi-dose container, for example, a sealed ampoule or vial. As an example of a lyophilized formulation, a 10 ml vial is filled with a 1% (w / v) aqueous solution of the sterile-filtered compound in an amount of 5 ml, and the resulting mixture is lyophilized. An infusion is prepared by reconstituting the lyophilized compound with bacteriostat...

Claims

**Claim 1** The polypeptide according to SEQ ID NO: 1, or any polypeptide sequence that is at least 60% identical to any of the above, or any nucleic acid encoding the same. **Claim 2** The polypeptide according to SEQ ID NO: 1, or any polypeptide sequence that is at least 80% identical to any of the above, or any nucleic acid encoding the same. **Claim 3** The polypeptide according to SEQ ID NO: 1, or any polypeptide sequence that is at least 95% identical to any of the above, or any nucleic acid encoding the same. **Claim 4** (i) A polypeptide according to any one of claims 1 to 3, and (ii) One or more single guide RNA(s) (sgRNA) or DNA(s) enabling the in situ generation of such one or more sgRNA(s), wherein each sgRNA or DNA encoding an sgRNA comprises: a. An engineered DNA targeting segment capable of hybridizing to a target sequence in a polynucleotide locus, b. A tracr mate sequence, and c. A tracr RNA sequence where the tracr mate sequence is capable of hybridizing to the tracr sequence, and (a), (b), and (c) are arranged in the 5' to 3' direction A composition comprising the same. **Claim 5** The engineered DNA targeting segment is directly adjacent at its 3' end to a PAM sequence on the target DNA segment, or such a PAM sequence is part of the 5' portion of the target DNA sequence, where the PAM sequence comprises a sequence motif selected from "CYN" or "CNN", "N" represents any base, and "Y" represents "C" or "T". The composition according to claim 4. **Claim 6** A method of targeting, editing, modifying, or manipulating target DNA at one or more positions in a cell or in vitro, the method comprising: (i) Introducing the polypeptide according to any one of claims 1 to 3 or a nucleic acid encoding the same into a cell or an in vitro environment; and (ii) Introducing one or more single guide RNA(s) (sgRNA) or DNA(s) encoding such one or more sgRNA(s) into the cell or in vitro environment, wherein each sgRNA or DNA encoding an sgRNA comprises: a. An engineered DNA targeting segment that contains RNA and is capable of hybridizing to a target sequence in a polynucleotide locus, b. A tracr mate sequence composed of RNA, and c. A tracr RNA sequence composed of RNA comprising, wherein the tracr mate sequence hybridizes to the tracr sequence, and steps (a), (b), and (c) are arranged in the 5' to 3' direction; and (iii) creating one or more nicks or cuts or base edits in the target DNA, wherein the polypeptide is directed to the target DNA by the processed or unprocessed form of the sgRNA, steps comprising a method. **Claim 7** For targeting, editing, modifying, or manipulating target DNA at one or more positions intracellularly or in vitro, i. The polypeptide according to any one of claims 1 to 3 or the nucleic acid encoding the same; and / or ii. One or more single guide RNAs (sgRNAs) or DNA(s) suitable for in situ generation of such one or more sgRNAs, each of which: a. An engineered DNA targeting segment composed of RNA and capable of hybridizing to such a target sequence in a polynucleotide locus, b. A tracr mate sequence composed of RNA, and c. A tracr RNA sequence composed of RNA comprising, wherein the tracr mate sequence hybridizes to the tracr sequence, and (a), (b), and (c) are arranged in the 5' to 3' direction, comprising the use of a composition. **Claim 8** (i) The polypeptide according to any one of claims 1 to 3, or the nucleic acid encoding the same; and (ii) One or more single guide RNAs (sgRNAs) or DNA(s) suitable for in situ generation of such one or more sgRNAs, each of which: a. An engineered DNA targeting segment capable of hybridizing to a target sequence in a polynucleotide locus, b. A tracr mate sequence, and c. A tracr RNA sequence comprising, wherein the tracr mate sequence is capable of hybridizing to the tracr sequence, and (a), (b), and (c) are arranged in the 5' to 3' orientation, comprising a cell. **Claim 9** i. A nucleic acid encoding the polypeptide according to any one of claims 1 to 3, wherein the nucleic acid encoding such a peptide is operably linked to a promoter; and ii. One or more single guide RNAs (sgRNAs) or DNA(s) suitable for generating such one or more sgRNAs in situ, each sgRNA comprising: a. An engineered DNA targeting segment capable of hybridizing to a target sequence in a polynucleotide locus, b. A tracr mate sequence, and c. A tracr RNA sequence wherein the tracr mate sequence is capable of hybridizing to the tracr sequence, and (a), (b), and (c) are arranged in a 5' to 3' orientation A kit comprising the same.

Citation Information

Patent Citations

  • Fastener systems that provide EME protection

    WO2013176722A2

  • Novel crispr-associated transposases and uses thereof

    WO2017117395A1

  • Novel crispr-CAS systems for genome editing

    WO2020123887A2