Methods and compositions relating to sequences that guide CAS9 targeting

The CRISPR-Cas system is enhanced with specific nucleotide sequences to improve efficiency and specificity in genome editing, allowing precise DNA cleavage through a chimeric nucleic acid construct with Cas9 nuclease.

JP7788744B2Active Publication Date: 2025-12-19NORTH CAROLINA STATE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024004564
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-01-24
Filing Date
2024-01-16
Publication Date
2025-12-19
Estimated Expiration
2035-01-23

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems for genome editing lack efficiency and specificity, limiting their effectiveness in targeted DNA cleavage and application.

Method used

The invention enhances CRISPR-Cas systems by incorporating specific nucleotide sequences and structures, such as anti-zipper, bulge, anti-stitch, and hairpin sequences, to improve hybridization and targeting precision, using a chimeric nucleic acid construct with Cas9 nuclease for site-specific DNA cleavage.

Benefits of technology

The enhanced CRISPR-Cas system achieves improved efficiency and specificity in genome editing, enabling precise targeting and cleavage of double-stranded DNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007788744000002
    Figure 0007788744000002
  • Figure 0007788744000003
    Figure 0007788744000003
  • Figure 0007788744000004
    Figure 0007788744000004
Patent Text Reader

Abstract

To provide methods and compositions for genome editing and DNA targeting of proteins.SOLUTION: A synthetic trans-encoded CRISPR(tracr) nucleic acid (e.g., tracrRNA / DNA) construct comprises: in the direction from 5' to 3', an optional anti-zipper sequence comprising at least three nucleotides; a bulge sequence comprising at least three nucleotides; an anti-stitch sequence comprising a nucleotide sequence of NNANN; a nexus sequence comprising a nucleotide sequence, such as TNANNC, T(A / C)A(A / G)(G / A)C; and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, the hairpin sequence comprising at least three matched base pairs, the anti-zipper sequence being located immediately upstream of the bulge sequence, the bulge sequence being located immediately upstream of the anti-stitch sequence, the anti-stitch sequence being located immediately upstream of the nexus sequence, the nexus sequence being located immediately upstream of the hairpin sequence.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [Priority Statement] This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Application No. 61 / 931,515, filed January 24, 2014, the entire contents of which are incorporated herein by reference. FIELD OF THE INVENTION The present invention relates to a synthetic CRISPR-cas system and methods for its use for genome editing. [Background technology]

[0002] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR), in combination with associated sequences (cas), comprise the CRISPR-Cas system, which confers adaptive immunity in many bacteria. CRISPR-mediated immunity occurs through the incorporation of DNA from invasive genetic elements such as plasmids and phages as a novel "spacer."

[0003] The CRISPR-Cas system consists of an array of short DNA repeats separated by hypervariable sequences and flanked by cas genes, providing adaptive immunity against invasive genetic agents such as phages and plasmids through sequence-specific targeting and interference (Barrangou et al. 2007. Science. 315:1709-1712; Brouns et al. 2008. Science 321:960-4; Horvath and Barrangou 2010. Science. 327:167-70; Marraffini and Sontheimer 2008. Science. 322:1843-1845; Bhaya et al. 2011. Annu. Rev. Genet. 45:273-297; Terns and Terns 2011. Curr. Opin. Microbiol. 14:321-327; Westra et al. 2011. Annu. Rev. Genet. 45:273-297; Terns and Terns 2011. Curr. Opin. Microbiol. 14:321-327; Westra et al. 2011. Annu. Rev. Genet. 45:273-297). (Barrangou et al. 2012. Annu. Rev. Genet. 46:311-339; Barrangou R. 2013. RNA. 4:267-278). Typically, invasive DNA sequences are acquired as novel "spacers" (Barrangou et al. 2007. Science. 315:1709-1712), each paired with a CRISPR repeat and inserted as a novel repeat-spacer unit at the CRISPR locus. The repeat-spacer array is then transcribed as a long pre-CRISPR RNA (pre-crRNA) (Brouns et al. 2008. Science 321:960-4), which is then processed into a short interfering CRISPR RNA (crRNA) that drives sequence-specific recognition.Specifically, crRNA guides nucleases toward complementary targets for sequence-specific nucleic acid cleavage mediated by Cas endonucleases (Garneau et al. 2010. Nature. 468:67-71; Haurwitz et al. 2010. Science. 329:1355-1358; Sapranauskas et al. 2011. Nucleic Acid Res. 39:9275-9282; Jinek et al. 2012. Science. 337:816-821; Gasiunas et al. 2012. Proc. Natl. Acad. Sci. 109:E2579-E2586; Magadan et al. 2012. PLoS One. 7:e40913; Karvelis et al. 2013. RNA Biol. 10:841-851). These widespread systems occur in nearly half of bacteria (approximately 46%) and the majority of archaea (approximately 90%). They are classified into three major CRISPR-Cas system types (Makarova et al. 2011. Nature Rev. Microbiol. 9:467-477; Makarova et al. 2013. Nucleic Acid Res. 41:4360-4377) based on the cas gene content, the organization and diversity of the biochemical processes driving crRNA biogenesis, and the Cas protein complexes that mediate target recognition and cleavage. In types I and II, specialized Cas endonucleases process the pre-crRNA, which then assembles into large multi-Cas protein complexes capable of recognizing and cleaving nucleic acids complementary to the crRNA. A distinct process is involved in type II CRISPR-Cas systems, in which the pre-crRNA is processed by a mechanism in which a trans-activating crRNA (tracrRNA) hybridizes to the repeat region of the crRNA. The hybridized crRNA-tracrRNA is cleaved by RNase III, followed by a second event that removes the 5' end of each spacer, producing mature crRNA that remains bound to both tracrRNA and Cas9.The mature complex then seeks out a target dsDNA sequence (the "protospacer" sequence) that is complementary to the spacer sequence within the complex and cleaves both strands. Target recognition and cleavage by the complex in type II systems not only requires complementary sequences between the spacer sequence on the crRNA-tracrRNA complex and the target "protospacer" sequence, but also a protospacer adjacent motif (PAM) sequence located at the 3' end of the protospacer sequence. The exact PAM sequence required can differ between different type II systems. Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure provides methods and compositions that enhance the efficiency and specificity of synthetic Type II CRISPR-Cas systems, improving the efficiency and specificity of genome editing and other applications. [Means for solving the problem]

[0005] One embodiment of the present invention comprises, in the 5' to 3' direction, an optional anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; T(A / C)A(A / G)(G / A)C (or U(A / C)A(A / G)(G / A)C)), TCAAAC, (or UCAAAC), TAAGGC (or UAAGGC), GATAAGG (or GAUAAGG), GATAAGGCTT (or GAUAAGGCUU), TCAAG (or UCAAG), TCAAGCAA (or UCAAGCAA), and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, wherein the hairpin comprises at least three matching base pairs; Here, the anti-zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence.

[0006] A second aspect of the invention is a tracrRNA comprising, from 3' to 5', an optional zipper sequence comprising at least about 3 nucleotides (which, if present, hybridizes to the anti-zipper of the tracrRNA), a bulge sequence comprising at least two nucleotides (e.g., a nucleotide sequence of (-NN-)), a stitch sequence comprising a nucleotide sequence of NNTNN (or NNUNN) that hybridizes to the anti-stitch of the tracrRNA, a G sequence comprising the nucleotides G or GTT, and a G sequence comprising the nucleotides G. R1and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity to the target DNA, wherein the zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is located immediately upstream of the G R1 Located just upstream of, and G R1 is located immediately upstream of the spacer sequence.

[0007] A third aspect of the present invention provides a synthetic CRISPR nucleic acid array comprising a nucleotide sequence encoding two or more CRISPR nucleic acid constructs of the invention, wherein the two or more CRISPR nucleic acid constructs are located immediately adjacent to each other on the nucleotide sequence, the zipper sequences of the two or more CRISPR nucleic acid constructs, if present, are identical, the stitching sequences of the two or more CRISPR nucleic acid constructs are identical, and the spacer sequences of the two or more CRISPR nucleic acid constructs are identical or non-identical.

[0008] A fourth aspect of the present invention provides a chimeric nucleic acid construct comprising a synthetic tracr nucleic acid construct of the invention and a synthetic CRISPR nucleic acid construct of the invention, wherein, if present, the zipper sequence of the synthetic CRISPR nucleic acid construct is at least about 70% complementary to and hybridizes with the anti-zipper sequence of said synthetic tracr nucleic acid construct, the stitch sequence of the synthetic CRISPR nucleic acid construct is 100% complementary to and hybridizes with the anti-stitch sequence of said synthetic tracr nucleic acid construct, and the bulge sequence of the synthetic CRISPR nucleic acid construct and the bulge sequence of the synthetic CRISPR nucleic acid construct are non-complementary.

[0009] A fifth aspect of the present invention provides a method for site-specific cleavage of double-stranded target DNA, comprising contacting a chimeric nucleic acid construct of the present disclosure or an expression cassette comprising said chimeric nucleic acid construct with the target DNA in the presence of Cas9 nuclease, thereby causing site-specific cleavage of the target DNA within a region defined by hybridization of the spacer sequence to the target DNA.

[0010] A sixth aspect of the present invention provides a method for site-specific cleavage of double-stranded target DNA, comprising contacting a transcoding CRISPR (tracr) nucleic acid molecule and a CRISPR nucleic acid molecule with the target DNA in the presence of a Cas9 nuclease; wherein (a) the tracr nucleic acid molecule comprises, in a 5' to 3' direction, an optional anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, T a nexus sequence comprising the nucleotide sequence CAAGCAAAGC, or TCAAACAAAGCTTCAGC; and a hairpin sequence comprising a nucleotide sequence having at least two hairpins, each hairpin comprising at least three matching base pairs, wherein the anti-zipper sequence is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; and (b) The CRISPR nucleic acid molecule may include, in the 3' to 5' direction, an optional zipper sequence comprising at least about 3 nucleotides, a bulge sequence comprising a nucleotide sequence having at least two nucleotides (e.g., a nucleotide sequence of (-NN-)), a stitch sequence comprising a nucleotide sequence of NNTNN (or NNUNN), a G sequence comprising the nucleotides G or GTT, and a G sequence comprising the nucleotides GTT.R1 and a spacer sequence having a 5' end and a 3' end, the spacer sequence having at least 7 nucleotides at its 3' end having 100% identity with the target DNA, wherein the zipper sequence is located immediately upstream of the bulge sequence, and the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is encoded by a nucleotide sequence comprising: R1 Located just upstream of, and G R1 is located immediately upstream of the spacer sequence, and Further herein, when present, the anti-zipper sequence and zipper sequence hybridize to each other, the anti-stitch sequence and stitch sequence hybridize to each other, and the spacer sequence of the CRISPR nucleic acid molecule hybridizes to at least a portion of the target DNA (e.g., at least about 7 contiguous nucleotides (e.g., about 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, etc., and any range or variation therein) of the target DNA (e.g., a protospacer adjacent motif on the target DNA). The CRISPR nucleic acid molecule hybridizes to and is at least about 80% complementary to a spacer sequence adjacent to the PAM (PAM), thereby causing site-specific cleavage of the target DNA within a region defined by complementary binding of the spacer sequence of the CRISPR nucleic acid molecule to the target DNA. Thus, in exemplary embodiments, the spacer sequence of the CRISPR nucleic acid molecule hybridizes to a portion of the target DNA sequence adjacent to the PAM, and the target sequence can comprise, consist essentially of, or consist of about 7 to about 20 contiguous nucleotides of the target DNA sequence.

[0011] A seventh aspect of the invention provides a method for site-specific cleavage of double-stranded target DNA comprising contacting the double-stranded target DNA with a chimeric nucleic acid comprising: (a) an optional anti-zipper sequence comprising, in the 5' to 3' direction, at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC or TCAAACAAAGCTTCAGC; and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, the hairpin comprising at least three matching base pairs, and, if present, the anti-zipper sequence is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; (b) from 3' to 5', an optional zipper sequence (if present, hybridizes to the anti-zipper sequence of the first nucleotide sequence) comprising at least about 3 nucleotides, a bulge sequence comprising a nucleotide sequence having at least 2 nucleotides (e.g., a nucleotide sequence of (-NN-)), a stitch sequence comprising a nucleotide sequence of NNTNN (or NNUNN), a G sequence comprising the nucleotides G or GTT, R1 and a second nucleotide sequence comprising a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% complementarity with the target DNA (wherein the zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is located immediately upstream of the G R1 Located just upstream of, and G R1 is located immediately upstream of the spacer sequence); and (c) a third nucleotide sequence encoding an amino acid sequence having at least 80% identity to the amino acid sequence encoding the Cas9 nuclease; wherein the anti-zipper sequence and zipper sequence, if present, hybridize to each other, the anti-stitch sequence hybridizes to the stitch sequence, and the spacer sequence of the second nucleotide sequence hybridizes to at least a portion of the target DNA (e.g., at least about 7 consecutive nucleotides, preferably up to about 20 consecutive nucleotides, of the target DNA) (adjacent to a protospacer adjacent motif (PAM) on the target DNA), thereby causing site-specific cleavage of the target DNA within the region defined by the complementary binding of the spacer sequence of the second nucleotide sequence to the target DNA.

[0012] An eighth aspect of the present invention comprises a method for site-specific targeting of a polypeptide of interest to double-stranded (ds) target DNA, comprising contacting the target DNA with a chimeric nucleic acid construct of the present disclosure or an expression cassette comprising said chimeric nucleic acid construct, thereby targeting a polypeptide of interest fused to Cas9 to a specific site on the target DNA, said site being defined by hybridization of a spacer sequence to the target DNA.

[0013] A ninth aspect of the present invention comprises a method for site-specific targeting of a polypeptide of interest to double-stranded (ds) target DNA, comprising: contacting a transcoding CRISPR (tracr) nucleic acid molecule and a CRISPR nucleic acid molecule with target DNA in the presence of a Cas9 nuclease, wherein: (a) The tracr nucleic acid molecule comprises, in the 5' to 3' direction, an optional anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC or TCAAACAAAGCTTCAGC; and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, said hairpin comprising at least three matching base pairs, wherein the anti-zipper sequence, if present, is located immediately upstream of the bulge sequence, which is located immediately upstream of the anti-stitch sequence, which is located immediately upstream of the nexus sequence, and which is located immediately upstream of the hairpin sequence; and (b) the CRISPR nucleic acid molecule comprises, in the 3' to 5' direction, an optional zipper sequence comprising at least about 3 nucleotides that hybridize to an anti-zipper sequence; a bulge sequence comprising a nucleotide sequence having at least two nucleotides (e.g., a nucleotide sequence of (-NN-)); a stitch sequence comprising a nucleotide sequence of NNTNN; a G sequence comprising the nucleotides G or GTT; R1 and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity with the target DNA, wherein the zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is encoded by a nucleotide sequence comprising a G R1 Located just upstream of, and G R1 is located immediately upstream of the spacer sequence, and Further herein, the Cas9 nuclease comprises a mutation in the HNH active site motif and a mutation in the RuvC active site motif and is fused to a polypeptide of interest, wherein the anti-zipper sequence and the zipper sequence, if present, hybridize to each other, the anti-stitch sequence hybridizes to the stitch sequence, and the spacer sequence hybridizes to at least a portion of the target DNA adjacent to a protospacer adjacent motif (PAM) on the target DNA, thereby resulting in site-specific targeting of the polypeptide of interest to the target DNA within a region defined by hybridization of the spacer sequence of the CRISPR nucleic acid molecule to the target DNA.

[0014] Further provided herein are expression cassettes, cells and kits comprising the nucleic acid constructs, nucleic acid arrays, nucleic acid molecules and / or nucleotide sequences of the invention.

[0015] These and other aspects of the present invention are set forth in detail in the following description of the invention. [Brief explanation of the drawings]

[0016] [Figure 1] A multiple sequence alignment for the nexus module is shown. [Figure 2] 1 shows the maximum likelihood tree for the nexus module. [Figure 3] Figures 3A to 3D show consensus sequences for the nexus module: Figure 3A shows the consensus sequence for the Sth Cr1 group, Figure 3B shows the consensus sequence for the Sth Cr3 group, Figure 3C shows the consensus sequence for the Lrh group, and Figure 3D shows the consensus sequence for the Lbu group. [Figure 4] A maximum likelihood tree for the Cas9 nuclease is shown. [Figure 5] A multiple sequence alignment for the anti-stitch module is shown. [Figure 6]Figures 6A-6D show the consensus sequences for the anti-stitch module: Figure 6A shows the consensus sequence for the Sth Cr1 group, Figure 6B shows the consensus sequence for the Sth Cr3 group, Figure 6C shows the consensus sequence for the Lrh group, and Figure 6D shows the consensus sequence for the Lbu group. [Figure 7] A multiple sequence alignment for the bulge module is shown. [Figure 8] Figures 8A-8D show consensus sequences for the bulge module: Figure 8A shows the consensus sequence for the Sth Cr1 group, Figure 8B shows the consensus sequence for the Sth Cr3 group, Figure 8C shows the consensus sequence for the Lrh group, and Figure 8D shows the consensus sequence for the Lbu group. [Figure 9] A multiple sequence alignment for the zipper module is shown. [Figure 10] 1 shows a maximum likelihood tree for the zipper module. [Figure 11] A multiple sequence alignment for the bulge, anti-stitch and nexus modules is shown. [Figure 12] Sequence and structural details of the CRISPR-Cas system elements of Streptococcus thermophilus CR3 are shown, representing the Sth CR1 group. [Figure 13] 1 shows the sequence and structure details of the CRISPR-Cas system elements of Lactobacillus buchneri, representing the Lbu group. [Figure 14] The sequence and structure details of the CRISPR-Cas system elements of Streptococcus thermophilus CR1 are shown, representing the Sth CR1 group. [Figure 15] Sequence and structural details of the CRISPR-Cas system elements of Streptococcus pyrogenes M1 GAS are shown, representing the Sth CR3 group. [Figure 16]Sequence and structural details of the CRISPR-Cas system elements of Lactobacillus rhamnosus are shown, representing the Lrh group. [Figure 17] Sequence and structural details of the CRISPR-Cas system elements of Lactobacillus animalis are shown, representing the Lan group. [Figure 18] Sequence and structural details of the CRISPR-Cas system elements of Lactobacillus casei are shown, representing the Lca group. [Figure 19] 1 shows the sequence and structure details of the CRISPR-Cas system elements of Lactobacillus gasseri, representing the Lga group. [Figure 20] Sequence and structural details of the CRISPR-Cas system elements of Lactobacillus jensenii are shown, representing the Lje group. [Figure 21] Sequence and structural details of the CRISPR-Cas system elements of Lactobacillus pentosus are shown, representing the Lpe group. [Figure 22] Details of the sequences and structures of the CRISPR-Cas system elements of Streptococcus pyrogenes M1 GAS are shown. [Figure 23] Congruence between tracrRNA (left), CRISPR repeat (center), and Cas9 (right) sequence clustering is shown, with consistent grouping into three families across three sequence-based trees. [Figure 24]Figures 24A-B show the Cas9:sgRNA family. Figure 24A shows a phylogenetic tree based on Cas9 protein sequences from various Streptococcus and Lactobacillus species. Sequences are clustered into three families in blue, orange, and green. Figure 24B shows the predicted guide RNA consensus sequence and secondary structure for each family. Each consensus RNA consists of a crRNA (left) base-paired with a tracrRNA. Fully conserved bases are colored, variable bases are represented by black (two possible bases) or black dots (at least three possible bases), and base positions that are not always present are circled. Circles between positions indicate base pairings that are present in only some family members. [Figure 25] CRISPR repeat sequence alignments are shown. For each cluster, CRISPR repeat sequence alignments are shown, with conserved consensus nucleotides designated under each family for the Sth3 (top), Sth1 (middle), and Lb (bottom) families. [Figure 26] tracrRNA sequence alignments are shown. For each cluster, experimentally determined or computationally predicted tracrRNA sequence alignments are shown, with conserved consensus nucleotides specified under each family for the Sth3 (top), Sth1 (middle), and Lb (bottom) families. [Figure 27] The sgRNA nexus sequence alignment is shown. Universally conserved residues are colored red. The complementary nucleotides that make up the nexus stem, summarized in Figure 24B, are underlined. The nucleotides that make up the nexus loop are in the center of the gap. [Figure 28] Figures 28A-B show the self-targeting assay scheme. Orthogonal Cas9 proteins were provided via pCas9 plasmids (Figure 28A) and used as described in Figure 3. Various sgRNA chimeras were provided via psgRNA plasmids (Figure 28B) and used in combination with each of the desired Cas9s described in Figure 3. [Figure 29]Figures 29A-C show sgRNA orthogonality. Figure 29A shows the sgRNA sequences of Streptococcus thermophilus CRISPR3-Cas9 (top, blue) and S. thermophilus CRISPR1-Cas9 (bottom, orange). Figure 29B shows the protospacer-targeting scheme. The predicted PAM for each sgRNA is shown. Triangles indicate the predicted cleavage site for each Cas9. Figure 29C shows Cas9:chimeric-sgRNA orthogonality in E. coli. Chimeric sgRNAs. Each sgRNA (left) was subjected to a transformation assay (right) in E. coli expressing SthCRISPR3 Cas9 (blue) and / or SthCRISPR1 Cas9 (orange). Low transformation efficiency indicates a functional Cas9:sgRNA pair via lethal self-targeting of the E. coli genome. Values ​​reflect the SEM and geometric mean of three independent experiments. [Figure 30] CRISPR interference with complementary DNA is shown as the inability to transform Lactobacillus gasseri with a plasmid containing a protospacer sequence that matches the initial wild-type CRISPR spacer sequence. The bolded sequence flanked by the PAM (light gray italicized nucleotides) and its variants (single nucleotide polymorphisms (SNPs); black underlined nucleotides) is the protospacer. The low number of transformants indicates an active Lga CRISPR system that prevents transformation of complementary target DNA. [Figure 31] CRISPR interference with complementary DNA is shown as the inability to transform Lactobacillus casei with a plasmid containing a protospacer sequence matching the initial wild-type CRISPR spacer sequence. The bolded sequence flanked by the PAM (light gray italicized nucleotides) and its variants (SNPs, black underlined nucleotides) is the protospacer. The low number of transformants indicates an active Lca CRISPR system that prevents transformation of complementary target DNA. [Figure 32]CRISPR interference with complementary DNA is shown as the inability to transform Lactobacillus rhamnosus with a plasmid containing a protospacer sequence matching the initial wild-type CRISPR spacer sequence. The bolded sequence flanked by the PAM (light gray italicized nucleotides) and its variants (SNPs, black underlined nucleotides) is the protospacer. The low number of transformants indicates an active Lra CRISPR system, which prevents transformation of complementary target DNA. DETAILED DESCRIPTION OF THE INVENTION

[0017] The present invention will now be described below with reference to the accompanying drawings and examples, in which embodiments of the invention are shown. The description is not intended to be a detailed listing of all the different ways in which the invention may be practiced or all the features that may be added to the invention. For example, features described with respect to one embodiment may be incorporated into other embodiments, and features described with respect to a particular embodiment may be omitted from that embodiment. Accordingly, the present invention anticipates that, in some embodiments of the invention, any feature or combination of features described herein may be excluded or omitted. Additionally, numerous modifications and additions to the various embodiments suggested herein will be apparent to those skilled in the art in light of this disclosure and do not depart from the invention. Therefore, the following description is intended to describe some particular embodiments of the invention, but is not intended to exhaustively identify all permutations, combinations, and variations thereof.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The terminology used in the description of the present invention herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present invention.

[0019] All publications, patent applications, patent documents and other references cited herein are incorporated by reference in their entirety for the teachings relevant to the sentence and / or paragraph in which the reference appears.

[0020] Unless the context dictates otherwise, it is expressly intended that the various features of the invention described herein can be used in any combination. Moreover, the present invention also anticipates that in some embodiments of the invention, any feature or combination of features set forth herein can be excluded or omitted. For illustrative purposes, if the specification states that a composition comprises components A, B, and C, it is expressly intended that any of A, B, or C, or any combination thereof, alone or in any combination, can be omitted and waived.

[0021] As used in the description of this invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise.

[0022] Also, as used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, and, when interpreted in the alternative ("or"), the lack of a combination.

[0023] As used herein, the term "about," when referring to a measurable value such as a dose or duration, refers to a variation of ±20%, ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the stated amount.

[0024] As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "between about X and Y" mean "between about X and about Y", and phrases such as "from about X to Y" mean "from about X to about Y".

[0025] As used herein, the term "bulge sequence" refers to a non-complementary (non-hybridizing) nucleotide sequence contained in a synthetic tracr nucleic acid construct and a synthetic CRISPR nucleic acid construct / CRISPR nucleic acid array. In a synthetic tracr nucleic acid construct, the bulge sequence is located between the anti-zipper and anti-stitch sequences and is non-complementary (100% non-identical) to the corresponding bulge sequence in the synthetic CRISPR nucleic acid construct / CRISPR nucleic acid array, and consists of about 3 to about 6 nucleotides (e.g., about 3, 4, 5, 6 nucleotides; e.g., about 3 to about 6 nucleotides, about 3 to about 5 nucleotides, about 3 to about 4 nucleotides, etc.). The bulge sequence of a synthetic CRISPR nucleic acid construct / CRISPR nucleic acid array is located between the zipper and stitch sequences and comprises, consists essentially of, or consists of at least two nucleotides (e.g., a nucleotide sequence of (-NN-)) (e.g., about 2, 3, 4, 5, 6 nucleotides; e.g., about 2 to about 6 nucleotides, about 2 to about 5 nucleotides, about 2 to about 4 nucleotides, about 3 to about 6 nucleotides, about 3 to about 5 nucleotides, etc.). The nucleotide composition of the bulge sequence can be any series of at least two (synthetic CRISPR nucleic acid construct / CRISPR nucleic acid array) or three or more (synthetic tracr nucleic acid construct) nucleotides, so long as they are not complementary (e.g., 100% non-identical) and therefore do not hybridize to each other. As a result of the non-complementarity of the bulge sequences on the synthetic tracr nucleic acid construct and the synthetic CRISPR nucleic acid construct / CRISPR nucleic acid array, when the anti-zipper to zipper sequence and anti-stitch to stitch sequence align and hybridize (as in a chimeric nucleic acid construct), a protrusion or bulge is formed on the synthetic tracr nucleic acid construct side of the chimeric nucleic acid construct (see, e.g., Figures 12-16). Without wishing to be bound by any particular theory, it is believed that the bulge structure may be involved in the function of the CRISPR-Cas system.

[0026] As used herein, the terms "comprise", "comprises" and "comprising" specify the presence of stated features, integers, steps, operations, factors, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, factors, components, and / or groups thereof.

[0027] The transitional phrase "consisting essentially of" as used herein means that the scope of a claim is to be interpreted to include the specific materials or steps recited in the claim and which do not materially affect the basic and novel feature(s) of the claimed invention. Thus, the term "consisting essentially of" when used in the claims of the present invention is not intended to be interpreted as equivalent to "comprising."

[0028] "Cas9 nuclease" refers to a large group of endonucleases that catalyze double-stranded DNA cleavage in the CRISPR Cas system. These polypeptides are well known in the art, and many of their structures (sequences) have been characterized (see, e.g., WO2013 / 176772; WO / 2013 / 188638). The domains responsible for catalyzing dsDNA cleavage are the RuvC domain and the HNH domain. The RuvC domain is responsible for nicking the (-) strand, and the HNH domain is responsible for nicking the (+) strand (see, e.g., Gasiunasetal. PNAS 109(36):E2579-E2586 (September 4, 2012)).

[0029] As used herein, "chimera" refers to a nucleic acid molecule or polypeptide in which at least two components are derived from different sources (eg, different organisms, different coding regions).

[0030] As used herein, "complementary" can mean 100% complementarity or identity to a comparator nucleotide sequence, or it can mean less than 100% complementarity (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc. complementarity).

[0031] As used herein, the term "complementary" or "complementarity" refers to the natural binding of polynucleotides under permissive salt and temperature conditions by base pairing. For example, the sequence "AGT" binds to the complementary sequence "TCA." Complementarity between two single-stranded molecules can be "partial," where only a portion of the nucleotides bind, or it can be complete, where there is total complementarity between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant effect on the efficiency and strength of hybridization between nucleic acid strands.

[0032] As used herein, the terms "contact," "contacting," "contacted," and grammatical variations thereof refer to placing components of a desired reaction together under conditions suitable for carrying out the desired reaction (e.g., site-specific targeting of a polypeptide of interest, integration, transformation, site-specific cleavage (nicking, cleavage), amplification, etc.). Methods and conditions for carrying out such reactions are well known in the art (see, for example, Gasiunas et al. (2012) Proc. Natl. Acad. Sci. 109: E2579-E2586; M.R. Green and J. Sambrook (2012) Molecular Cloning: A Laboratory Manual. 4th Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).

[0033] A "fragment" or "portion" of a nucleotide sequence of the present invention is understood to mean a nucleotide sequence that is shortened in length (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides shortened) relative to a reference nucleic acid or nucleotide sequence, and that comprises, consists essentially of, and / or consists of a nucleotide sequence of contiguous nucleotides that is identical or nearly identical (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical) to the reference nucleic acid or nucleotide sequence. Such nucleic acid fragments or portions of the present invention may, where appropriate, be included in a larger polynucleotide of which they are a component. Thus, for example, hybridizing to at least a portion of a target DNA (or hybridize to, and other grammatical variations thereof) refers to hybridization to a nucleotide sequence identical or substantially identical to a stretch of contiguous nucleotides of the target DNA.

[0034] As used herein, "G R1 " is a single nucleotide (G) or a short three-nucleotide sequence (GTT) included in the repeat portion of a synthetic CRISPER nucleic acid construct or crRNA. G R1 does not hybridize to the anti-repeat of the synthetic tracr nucleic acid construct or tracrRNA of the present disclosure. However, in a non-standard Watson-Crick base pairing scheme, G R1 can form a wobble base pair with U at the end of the anti-CRISPR repeat portion of tracrRNA.

[0035] As used herein, the term "gene" refers to a nucleic acid molecule that can be used to produce mRNA, antisense RNA, miRNA, anti-microRNA antisense oligodeoxyribonucleotide (AMO), etc. A gene may or may not be capable of being used to produce a functional protein or gene product. A gene may include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions). A gene may be "isolated," by which is meant a nucleic acid that is substantially or essentially free from components normally found associated with the nucleic acid in its natural state. Such components include other cellular material, culture medium in recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acids.

[0036] As used herein, a "hairpin sequence" is a nucleotide sequence that comprises a hairpin. A hairpin (e.g., stem-loop, feedback) refers to a nucleic acid molecule having a secondary structure that includes a region of nucleotides that forms a double strand flanked on either side by further single-stranded regions. Such structures are well known in the art. As known in the art, the double-stranded region may contain some mismatches in base pairing or may be fully complementary. In some embodiments of the present disclosure, the hairpin sequence of the nucleic acid construct is located immediately downstream of the "nexus sequence" at the 3' end of the synthetic tracr nucleic acid construct. Without being bound to any particular theory, it is believed that the hairpin may be involved in Cas9 binding to the crRNA-tracrRNA complex (e.g., synthetic CRISPR nucleic acid construct-synthetic).

[0037] A "heterologous" or "recombinant" nucleotide sequence is a nucleotide sequence that is not naturally associated with a host cell into which it is introduced, and includes non-naturally occurring multiple copies of a naturally occurring nucleotide sequence.

[0038] Different nucleic acids or proteins that share homology are referred to herein as "homologs." The term homolog includes homologous sequences from the same and other species and orthologous sequences from the same and other species. "Homology" refers to the level of similarity between two or more nucleic acid and / or amino acid sequences in terms of the percentage of positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Thus, the compositions and methods of the present invention further include homologs to the nucleotide and polypeptide sequences of the present invention. As used herein, "orthologous" refers to homologous nucleotide and / or amino acid sequences in different species that arise from a common ancestral gene during speciation. Homologs of the nucleotide sequences of the invention have substantial sequence identity (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and / or 100%) to said nucleotide sequences of the invention. Thus, for example, homologs of Cas9 polypeptides useful in the invention may be about 70% or more homologous to any one of the Cas9 sequences provided herein.

[0039] As used herein, hybridization, hybridize, hybridizing, and grammatical variations thereof refer to the binding of two completely complementary nucleotide sequences or substantially complementary sequences, in which some mismatched base pairs may exist. Hybridization conditions are well known in the art and vary based on the length of the nucleotide sequences and the degree of complementarity between the nucleotide sequences. In some embodiments, hybridization conditions may be high stringency, or they may be medium stringency or low stringency, depending on the length and amount of complementarity of the hybridized sequences. Conditions that constitute low, medium, and high stringency for purposes of hybridization between nucleotide sequences are well known in the art (see, e.g., Gasiunas et al. (2012) Proc. Natl. Acad. Sci. 109:E2579-E2586; M.R. Green and J. Sambrook (2012) Molecular Cloning: A Laboratory Manual. 4th Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).

[0040] As used herein, the terms "increase," "increasing," "increased," "enhance," "enhanced," "enhancing," and "enhancement" (and grammatical variations thereof) refer to an increase of at least about 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500% or more compared to a control.

[0041] The term "invasive exogenous genetic element," "invasive exogenous nucleic acid," or "invasive exogenous DNA" refers to DNA that is exogenous to a bacterium (e.g., genetic elements from pathogens, including but not limited to viruses, bacteriophages, and / or plasmids).

[0042] A "native" or "wild-type" nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide sequence, polypeptide, or amino acid sequence. Thus, for example, a "wild-type mRNA" is an mRNA that is naturally occurring or endogenous in an organism. A "homologous" nucleic acid sequence is a nucleotide sequence that is naturally associated with a host cell into which it is introduced.

[0043] As used herein, a "nexus sequence" refers to a nucleotide sequence located immediately downstream of the "anti-stitch sequence" within a synthetic tracr nucleic acid construct. The nexus is approximately 6-10 nucleotides in length and contains the highly conserved sequence TNANNC. In some embodiments, the nexus may be the nucleotide sequence T(A / C)A(A / G)(G / A)C (or U(A / C)A(A / G)(G / A)C)), GATAAGGCTT (or GAUAAGGCUU), TCAAGCAA (or UCAAGCAA), or T(C / A)AA(A / C)(C / A)(A / G)(A / T) (or U(C / A)AA(A / C)(C / A)(A / G)(A / U)). Without being bound by any particular theory, based on sequence conservation, it is believed that the nexus may be important in Cas9 orthogonality and recognition.

[0044] As used herein, the terms "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "nucleotide sequence," and "polynucleotide" refer to linear or branched, single- or double-stranded RNA or DNA, or a hybrid thereof. The terms also encompass RNA / DNA hybrids. When dsRNA is produced synthetically, unusual bases, such as inosine, 5-methylcytosine, 6-methyladenine, hypoxanthine, and others, can also be used for antisense, dsRNA, and ribozyme pairing. For example, polynucleotides containing C-5 propyne analogs of cytidine and uridine have been shown to bind to RNA with high affinity and to be potent antisense inhibitors of gene expression. Other modifications, such as modifications to the phosphodiester backbone or the 2'-hydroxyl in the ribose sugar group of RNA, can also be used. The nucleic acid constructs of the present disclosure may be DNA or RNA, but are preferably DNA. Therefore, the nucleic acid constructs of the present invention may be described and used in the form of DNA, but may also be described and used in the form of RNA depending on the intended use.

[0045] As used herein, a "synthetic" nucleic acid or nucleotide sequence refers to a nucleic acid or nucleotide sequence that is not found in nature but is constructed by the hand of man (and is therefore not a product of nature).

[0046] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or the sequence of these nucleotides from the 5' to 3' end of a nucleic acid molecule, including DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which may be single-stranded or double-stranded. The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "oligonucleotide," and "polynucleotide" are used interchangeably herein to also refer to a heteropolymer of nucleotides. Unless otherwise indicated, nucleic acid molecules and / or nucleotide sequences provided herein are presented from left to right in the 5' to 3' direction and are represented using the standard code for representing nucleotide characteristics set forth in the U.S. Sequencing Rules (37 CFR Sections 1.821-1.825) and the World Intellectual Property Organization (WIPO) Standard ST.25. As used herein, "5' region" may refer to the region of a polynucleotide closest to the 5' end. Therefore, for example, the element in the 5' region of polynucleotide can be located anywhere from the first nucleotide at the 5' end of polynucleotide to the nucleotide located in the middle of polynucleotide.As used herein, " 3' region " can refer to the region of polynucleotide that is closest to the 3' end.Therefore, for example, the element in the 3' region of polynucleotide can be located anywhere from the first nucleotide at the 3' end of polynucleotide to the nucleotide located in the middle of polynucleotide.

[0047] As used herein, the term "percent sequence identity" or "percent identity" refers to the percentage of identical nucleotides within a linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complement) compared to a test ("subject") polynucleotide molecule (or its complement) when the two sequences are optimally aligned. In some embodiments, "percent identity" may refer to the percentage of identical amino acids within an amino acid sequence.

[0048] "Protospacer sequence" refers to the target double-stranded DNA, and specifically refers to the portion of the target DNA that is perfectly or substantially complementary to (and hybridizes with) the spacer sequence of a synthetic CRISPR nucleic acid construct.

[0049] As used herein, the terms "reduce," "reduced," "reducing," "reduction," "diminish," "suppress," and "decrease" (and grammatical variations thereof) refer to a decrease of at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100%, for example, as compared to a control. In certain embodiments, the decrease may result in no, or essentially no, detectable activity or amount (i.e., an insignificant amount, e.g., less than about 10% or even less than 5%). Thus, in some embodiments, mutations in the Cas9 nuclease may reduce the nuclease activity of Cas9 by at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% compared to a control (e.g., wild-type Cas9).

[0050] As used herein, "repeat sequence" refers to, for example, repeat sequences of a wild-type CRISPR locus or a synthetic CRISPR nucleic acid construct separated by a "spacer sequence." Repeat sequences can be complementary (e.g., 100% base pair match) or substantially complementary, e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher) to the corresponding anti-repeat sequence.

[0051] The "repeat sequences" of the synthetic CRISPR nucleic acid constructs of the present disclosure include optional "zipper sequences," "bulge sequences," "stitch sequences," and "spacer sequences." In some embodiments, the synthetic CRISPR nucleic acid constructs include G R1 (which in other embodiments may be included in the stitching sequence).

[0052] As used herein, "zipper sequence" refers to an optional portion of a repeat sequence located 3' or immediately upstream (in the 3' to 5' direction) of a bulge sequence in a synthetic CRISPR nucleic acid construct, and comprises, consists of, or consists essentially of at least about 3 nucleotides (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides, or any range or value therein). In some embodiments, the zipper sequence may be referred to as the "upper stem." The "zipper sequence" shares sufficient complementarity with a corresponding, optional "anti-zipper sequence" located on the synthetic tracr nucleic acid construct so that, if present, the zipper and anti-zipper sequences can hybridize to each other when they contact, thereby linking the two nucleic acid constructs together. In some embodiments, the zipper / anti-zipper sequence can be referred to as the "upper stem." Zipper sequences can be perfectly complementary (e.g., 100% base pair match) or substantially complementary, e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more) to the corresponding anti-zipper sequences. Thus, the anti-zipper sequences of the synthetic tracr nucleic acid constructs of the invention comprise, consist of, or consist essentially of at least about 3 nucleotides (e.g., about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides, or any range or value therein) that are fully complementary or substantially complementary to the corresponding zipper sequence in the synthetic CRISPR nucleic acid construct or synthetic CRISPR nucleic acid array.The anti-zipper sequence is an RNase III binding site and therefore comprises nucleotide sequences well known in the art to be involved in RNase III binding (see, e.g., Pertzev and Nicholson, Nucleic Acids Res. 34(13):3708-3721 (2006)).

[0053] "Sequence identity," as used herein, refers to the degree to which two optimally aligned polynucleotide or peptide sequences are invariant over the window of alignment of the components (eg, nucleotides or amino acids). "Identity" can be readily calculated by known methods, including but not limited to those described in: Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, AM, and Griffin, HG, eds.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Stockton Press, New York (1991).

[0054] As used herein, a "spacer sequence" is a nucleotide sequence that is complementary to a target DNA (e.g., a "protospacer sequence"). A spacer sequence can be fully complementary or substantially complementary (e.g., at least 70% complementary (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more)) to the target DNA. In an exemplary embodiment, the spacer sequence has 100% complementarity with the target DNA. In further embodiments, the 3' region of the spacer sequence is 100% complementary to the target DNA, but the 5' region of the spacer is less than 100%, and thus the overall complementarity of the spacer sequence to the target DNA is less than 100%. Thus, for example, the first 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, etc. nucleotides (seed sequence) in the 3' region of a 20-nucleotide spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA. In some embodiments, the first 7-12 nucleotides of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA. In other embodiments, the first 7-10 nucleotides of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA. In an exemplary embodiment, the first 7 nucleotides of the spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary (e.g., at least about 70% complementary) to the target DNA.

[0055] As used herein, a "stitch sequence" refers to a nucleotide sequence that comprises, consists essentially of, or consists of about 5 nucleotides in length and has a consensus nucleotide sequence of NNTNN. The "stitch sequence" is located (in the 5' to 3' direction) on the synthetic CRISPR nucleic acid construct immediately upstream of the "bulge sequence" and the "G R1 ". The "stitch sequence" tends to have a high AT content and hybridizes to the "anti-stitch sequence" located within the synthetic tracr nucleic acid construct. In certain embodiments, the stitch sequence comprises, consists essentially of, or consists of the nucleotide sequence (from 5' to 3'): NNTNN, TTTGT, TTTTA, (T / C)(T / C)T(T / C)(T / G), TTTTA, TTTCA.

[0056] As used herein, "anti-stitch sequence" refers to a nucleotide sequence that is perfectly complementary to and hybridizes with the stitch sequence (e.g., NNANN, ACAAA, TAAAA, (T / C)(A / G)T(A / G)(A / G), TAAAA, TGAAA). The anti-stitch sequence is located (5' to 3' direction) immediately downstream of the bulge sequence and immediately upstream of the "nexus sequence" on the synthetic tracr nucleic acid construct. Without wishing to be bound by any particular theory, it is believed that hybridization of the stitch sequence of a synthetic crRNA construct with the anti-stitch sequence of the synthetic tracrRNA construct involves reconstructed base pairing after the "bulge sequence." In some embodiments, the stitch / anti-stitch can be referred to as the "lower stem."

[0057] As used herein, the phrase "substantially identical" or "substantial identity" in the context of two nucleic acid molecules, nucleotide sequences, or protein sequences refers to two or more sequences or subsequences that, when compared and aligned for maximum correspondence, have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and / or 100% nucleotide or amino acid residue identity, as determined using one of the following sequence comparison algorithms or by visual inspection. In some embodiments of the invention, the substantial identity exists over a region of the sequences that is at least about 50 residues to about 150 residues in length. Thus, in some embodiments of the invention, substantial identity exists over a region of the sequence that is at least about 3 to about 15 (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 residues in length, etc., or any value or range therein), at least about 5 to about 30, at least about 10 to about 30, at least about 16 to about 30, at least about 18 to at least about 25, at least about 18, at least about 22, at least about 25, at least about 30, at least about 40, at least about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, or more residues in length, and any range therein. In representative embodiments, the sequences may be substantially identical over at least about 22 nucleotides. In certain embodiments, the sequences are substantially identical over at least about 150 residues. In some embodiments, sequences of the invention may be about 70% to about 100% identical over at least about 16 to about 25 nucleotides. In some embodiments, sequences of the invention may be about 75% to about 100% identical over at least about 16 to about 25 nucleotides. In further embodiments, sequences of the invention may be about 80% to about 100% identical over at least about 16 to about 25 nucleotides.In further embodiments, sequences of the invention may be about 80% to about 100% identical over at least about 7 nucleotides to about 25 nucleotides. In some embodiments, sequences of the invention may be about 70% identical over at least about 18 nucleotides. In other embodiments, sequences may be about 85% identical over about 22 nucleotides. In still other embodiments, sequences may be 100% identical over about 16 nucleotides. In further embodiments, sequences are substantially identical over the entire length of the coding region. Moreover, in representative embodiments, substantially identical nucleotide or protein sequences perform substantially identical functions (e.g., Cas9 HNH and / or RuvC nickase activity).

[0058] For sequence comparison, typically, one sequence serves as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and comparison sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity of the test sequence(s) relative to the reference sequence based on the designated program parameters.

[0059] Optimal sequence alignment for aligning a comparison window is widely known to those skilled in the art and can be performed using tools such as the Smith and Waterman local homology algorithm, the Needleman and Wunsch homology alignment algorithm, the Pearson and Lipman similarity search method, and optionally using computer implementations of algorithms such as GAP, BESTFIT, FASTA, and TFASTA, available as part of the GCG® Wisconsin Package® (Accelrys Inc., San Diego, CA). The "fractional identity" of an aligned section of a test sequence and a reference sequence is the number of identical elements shared by the two aligned sequences divided by the total number of elements in the reference sequence segment (i.e., the entire reference sequence or a smaller, defined portion of the reference sequence). The percent sequence identity is expressed as the fractional identity multiplied by 100. Comparison of one or more polynucleotide sequences can be performed against a full-length polynucleotide sequence or a portion thereof, or against a longer polynucleotide sequence. For purposes of the present invention, "percent identity" may be determined using BLASTX version 2.0 for translated nucleotide sequences and BLASTN version 2.0 for polynucleotide sequences.

[0060] Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information. The algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that, when aligned with words of the same length in a database sequence, match or meet some positive threshold score, T. T is referred to as the neighborhood word score threshold (Altschul et al., 1990). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using the parameters M (benefit score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0) for nucleotide sequences. For amino acid sequences, a scoring matrix is ​​used to calculate the cumulative score. Extension of word hits in each direction is halted when the cumulative alignment score falls by an amount X from its maximum achieved value, resulting in a cumulative score of 0 or less due to the accumulation of one or more negative-scoring residue alignments, or the end of either sequence. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) of 10, a cutoff of 100, M=5, N=-4, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)).

[0061] In addition to calculating percent sequence identity, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin & Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873 5787 (1993)). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability that a match between two nucleotide or amino acid sequences would occur by chance. For example, a test nucleic acid sequence is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleotide sequence to the reference nucleotide sequence is less than about 0.1 to less than about 0.001. Thus, in some embodiments of the present invention, the smallest sum probability in a comparison of the test nucleotide sequence to the reference nucleotide sequence is less than about 0.001.

[0062] Two nucleotide sequences can also be considered to be substantially complementary if the two sequences hybridize to each other under stringent conditions. In some exemplary embodiments, two nucleotide sequences considered to be substantially complementary hybridize to each other under highly stringent conditions.

[0063] "Stringent hybridization conditions" and "stringent hybridization wash conditions" in the context of nucleic acid hybridization experiments such as Southern and Northern hybridization are sequence-dependent and will vary under different environmental parameters. An extensive guide to nucleic acid hybridization can be found in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes part I chapter 2 "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York (1993). Generally, highly stringent hybridization and wash conditions are those that meet the melting temperature (T) for a specific sequence at a defined ionic strength and pH. m ) is chosen to be approximately 5°C lower than

[0064] T m is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are those where the T mAn example of stringent hybridization conditions for hybridization of complementary nucleotide sequences with more than 100 complementary residues on filters in Southern or Northern blots is 50% formamide containing 1 mg heparin at 42°C, with hybridization occurring overnight. An example of highly stringent wash conditions is 0.15 M NaCl at 72°C for approximately 15 minutes. An example of stringent wash conditions is a 0.2×SSC wash at 65°C for 15 minutes (see Sambrook, infra, for a description of SSC buffers). A low stringency wash often precedes a high stringency wash to remove background probe signal. For example, an example of a moderate stringency wash for duplexes of more than 100 nucleotides is 1×SSC at 45°C for 15 minutes. For example, an example of low stringency washing for duplexes of more than 100 nucleotides is 4-6x SSC at 40°C for 15 minutes. For short probes (e.g., about 10-50 nucleotides), stringent conditions typically include a salt concentration of less than about 1.0 M Na ion, typically about 0.01-1.0 M Na ion (or other salt) (pH 7.0-8.3), and a temperature typically of at least about 30°C. Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Generally, a signal-to-noise ratio of 2x (or higher) than that observed for an unrelated probe in a particular hybridization assay indicates detection of specific hybridization. Nucleotide sequences that do not hybridize to each other under stringent conditions are still substantially identical if the proteins they encode are substantially identical. This can occur, for example, when a copy of a nucleotide sequence is created using the maximum codon degeneracy permitted by the genetic code.

[0065] The following are exemplary sets of hybridization / wash conditions that can be used to clone homologous nucleotide sequences that are substantially identical to the reference nucleotide sequences of the present invention. In one embodiment, the reference nucleotide sequence hybridizes to a "test" nucleotide sequence in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO, 1 mM EDTA at 50°C with a wash of 2x SSC, 0.1% SDS at 50°C. In another embodiment, the reference nucleotide sequence hybridizes to a "test" nucleotide sequence in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO, 1 mM EDTA at 50°C with a wash of 1x SSC, 0.1% SDS at 50°C, or in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO, 1 mM EDTA at 50°C with a wash of 0.5x SSC, 0.1% SDS at 50°C. In yet a further embodiment, the reference nucleotide sequence is hybridized to the "test" nucleotide sequence in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO, 1 mM EDTA at 50°C with a wash of 0.1x SSC, 0.1% SDS at 50°C, or in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO, 1 mM EDTA at 50°C with a wash of 0.1x SSC, 0.1% SDS at 65°C.

[0066] Any nucleotide sequence and / or recombinant nucleic acid molecule of the present invention can be codon-optimized for expression in any species of interest. Codon optimization is widely known in the art and involves modifying a nucleotide sequence for codon usage bias using species-specific codon usage. The codon usage is generated based on sequence analysis of the most highly expressed genes for the species of interest. If the nucleotide sequence is expressed in the nucleus, the codon usage is generated based on sequence analysis of highly expressed nuclear genes for the species of interest. The modification of the nucleotide sequence is determined by comparing the species-specific codon usage to the codons present in the native polynucleotide sequence. As is understood in the art, codon optimization of a nucleotide sequence results in a nucleotide sequence that has less than 100% identity (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc.) to a naturally occurring nucleotide sequence, but still encodes a polypeptide having the same function as that encoded by the original naturally occurring nucleotide sequence. Thus, in an exemplary embodiment of the present invention, the nucleotide sequences and / or recombinant nucleic acid molecules of the present invention can be codon optimized for expression in a particular species of interest.

[0067] In some embodiments, the recombinant nucleic acid molecules, nucleotide sequences, and polypeptides of the invention are "isolated." An "isolated" nucleic acid molecule, "isolated" nucleotide sequence, or "isolated" polypeptide is a nucleic acid molecule, nucleotide sequence, or polypeptide that exists apart from its natural environment by the hand of man and is therefore not a product of nature. An isolated nucleic acid molecule, nucleotide sequence, or polypeptide may exist in a purified form, at least partially separated from at least some of the other components of the organism or virus in which it naturally occurs, e.g., nucleic acids that are normally found in association with cellular or viral structural components or other polypeptides or polynucleotides. In representative embodiments, the isolated nucleic acid molecule, isolated nucleotide sequence, and / or isolated polypeptide is at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more pure.

[0068] In other embodiments, an isolated nucleic acid molecule, nucleotide sequence, or polypeptide can exist in a non-native environment (e.g., a recombinant host cell). Thus, for example, with respect to a nucleotide sequence, the term "isolated" means separate from the chromosome and / or cell in which it naturally occurs. A polynucleotide is also isolated when it is separated from the chromosome and / or cell in which it naturally occurs and then inserted into a genetic context, chromosome, and / or cell in which it does not naturally occur (e.g., a different host cell, different regulatory sequences, and / or a location in the genome different from that in which it is found in nature). Thus, recombinant nucleic acid molecules, nucleotide sequences, and their encoded polypeptides are "isolated" in that they exist, by the hand of man, apart from their natural environment and, therefore, are not products of nature, but can, in some embodiments, be introduced and exist within a recombinant host cell.

[0069] In any of the embodiments described herein, the nucleotide sequences and / or recombinant nucleic acid molecules of the invention can be operably associated with a variety of promoters and other regulatory elements for expression in a variety of biological cells. Thus, in representative embodiments, a recombinant nucleic acid of the invention can further comprise one or more promoters operably linked to one or more nucleotide sequences.

[0070] As used herein, "operably linked" or "operably associated" means that the indicated elements are functionally associated with each other, and generally also physically associated. Thus, as used herein, the terms "operably linked" or "operably associated" refer to nucleotide sequences on a single nucleic acid molecule that are functionally related. Thus, a first nucleotide sequence operably linked to a second nucleotide sequence refers to a situation in which the first nucleotide sequence is placed in a functional relationship with the second nucleotide sequence. For example, a promoter is operably associated with a nucleotide sequence if it effects the transcription or expression of that nucleotide sequence. Those skilled in the art will understand that a control sequence (e.g., a promoter) need not be contiguous with the nucleotide sequence with which it is operably associated, so long as the control sequence functions to direct its expression. Thus, for example, intervening untranslated, yet transcribed, sequences can be present between the promoter and the nucleotide sequence, and the promoter can still be considered "operably linked" to the nucleotide sequence.

[0071] A "promoter" is a nucleotide sequence (i.e., a coding sequence) that controls or regulates the transcription of a nucleotide sequence operably associated with the promoter. The coding sequence may encode a polypeptide and / or functional RNA. Typically, a "promoter" refers to a nucleotide sequence containing a binding site for RNA polymerase II, which directs the initiation of transcription. Promoters are generally found 5' or upstream to the start of the coding region of a corresponding coding sequence. The promoter region may contain other elements that act as regulators of gene expression. These include the TATA box consensus sequence and often contain the CAAT box consensus sequence (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box may be replaced by an AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith, and A. Hollaender (eds.), Plenum Press, pp. 211-227).

[0072] Promoters may include, for example, constitutive, inducible, temporally regulated, developmentally regulated, chemically regulated, tissue-preferred and / or tissue-specific promoters for use in preparing recombinant nucleic acid molecules (i.e., "chimeric genes" or "chimeric polynucleotides"). These various types of promoters are known in the art.

[0073] The choice of promoter varies depending on the time and space requirements for expression, and also varies depending on the host cell to be transformed.Promoters for many different organisms are well known in the art.Based on the extensive knowledge existing in the art, an appropriate promoter can be selected for a specific target host organism.Therefore, for example, much is known about the promoters upstream of highly constitutively expressed genes in model organisms, and such knowledge can be easily utilized and implemented in other systems as needed.

[0074] In some embodiments, the nucleic acid constructs of the present invention may be "expression cassettes" or may be contained within an expression cassette. As used herein, "expression cassette" refers to a recombinant nucleic acid molecule comprising a nucleotide sequence of interest (e.g., a nucleic acid construct of the present invention (e.g., a synthetic tracr nucleic acid construct, a synthetic CRISPR nucleic acid construct, a synthetic CRISPR array, a chimeric nucleic acid construct; a nucleotide sequence encoding a polypeptide of interest, a nucleotide sequence encoding a cas9 nuclease)), wherein the nucleotide sequence is operably associated with at least a control sequence (e.g., a promoter). Thus, some aspects of the present invention provide expression cassettes designed to express the nucleotide sequences of the present invention.

[0075] An expression cassette containing a nucleotide sequence of interest may be chimeric, meaning that at least one of its components is heterologous to at least one of its other components, or the expression cassette may be of natural origin but derived in recombinant form to be useful for heterologous expression.

[0076] The expression cassette may also optionally contain a transcriptional and / or translational termination region (i.e., a termination region) functional in the selected host cell. A variety of transcriptional terminators are available for use in the expression cassette and are responsible for terminating transcription beyond the heterologous nucleotide sequence of interest and correcting polyadenylation of the mRNA. The termination region may be native to the transcriptional initiation region, native to the operably linked nucleotide sequence of interest, native to the host cell, or derived from another source (i.e., exogenous or heterologous to the promoter, to the nucleotide sequence of interest, to the host, or any combination thereof).

[0077] The expression cassette can also include a nucleotide sequence for a selectable marker that can be used to select transformed host cells. As used herein, a "selectable marker" refers to a nucleotide sequence that, when expressed, confers a distinguishable phenotype on host cells expressing the marker, thus allowing such transformed cells to be distinguished from those that do not possess the marker. Such a nucleotide sequence can encode either a selectable marker or a screenable marker, depending on whether the marker confers a trait that can be selected for by chemical means, such as by using a selection agent (e.g., an antibiotic), or whether the marker is a trait that can simply be identified through observation or testing, for example, by screening (e.g., fluorescence). Of course, many examples of suitable selectable markers are known in the art and can be used in the expression cassettes described herein.

[0078] In addition to expression cassettes, the nucleic acid molecules and nucleotide sequences described herein can be used in the context of vectors. The term "vector" refers to a composition for transferring, delivering, or introducing a nucleic acid (or nucleic acids) into a cell. A vector includes a nucleic acid molecule that contains the nucleotide sequence(s) to be transferred, delivered, or introduced. Vectors for use in transforming host organisms are well known in the art. Non-limiting examples of general classes of vectors include, but are not limited to, viral vectors, plasmid vectors, phage vectors, phagemid vectors, cosmid vectors, fosmid vectors, bacteriophages, artificial chromosomes, or Agrobacterium binary vectors, in double-stranded or single-stranded, linear, or circular forms, and may or may not be capable of self-transmission or mobilization. Vectors as defined herein can transform prokaryotic or eukaryotic hosts by integration into the cellular genome, or can exist extrachromosomally (e.g., as a self-replicating plasmid with an origin of replication). Further included are shuttle vectors, which refer to DNA vehicles capable of replication, naturally or by design, in two different host organisms, which may be selected from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plants, mammals, yeast, or fungal cells). In some exemplary embodiments, the nucleic acid in the vector is operably linked and under the control of an appropriate promoter or other regulatory elements for transcription in the host cell. The vector may be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, it may contain its own promoter or other regulatory elements, and in the case of cDNA, it may be under the control of an appropriate promoter or other regulatory elements for expression in the host cell. Thus, the nucleic acid molecules and / or expression cassettes of the present invention may be contained within vectors, as described herein and known in the art.

[0079] "Introducing," "introduce," "introduced" (and grammatical variations thereof) in the context of a polynucleotide of interest means presenting a nucleotide sequence of interest to a host organism or a cell of said organism (e.g., a host cell) in such a way that the nucleotide sequence is accessible to the interior of the cell. When more than one nucleotide sequence is introduced, these nucleotide sequences can be assembled as part of a single polynucleotide or nucleic acid construct or as separate polynucleotides or nucleic acid constructs, and can be located on the same or different expression constructs or transformation vectors. Thus, these polynucleotides can be introduced into a cell in a single transformation event, in separate transformation / transfection events, or, for example, they can be incorporated into an organism by conventional propagation protocols. Thus, in some aspects of the present invention, one or more nucleic acid constructs of the present invention (e.g., synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, synthetic CRISPR arrays, chimeric nucleic acid constructs; nucleotide sequences encoding polypeptides of interest, nucleotide sequences encoding cas9 nuclease, etc.) can be introduced into a host organism or cells of said host organism.

[0080] As used herein, the terms "transformation" or "transfection" refer to the introduction of heterologous nucleic acid into a cell. Cellular transformation can be stable or transient. Thus, in some embodiments, a host cell or host organism is stably transformed with a nucleic acid molecule of the invention. In other embodiments, a host cell or host organism is transiently transformed with a recombinant nucleic acid molecule of the invention.

[0081] "Transiently transformed" in the context of a polynucleotide means that the polynucleotide is introduced into a cell and does not integrate into the genome of the cell.

[0082] By "stably introducing" or "stably introduced" in the context of a polynucleotide introduced into a cell is intended that the introduced polynucleotide is stably incorporated into the genome of the cell, and thus the cell is stably transformed with the polynucleotide.

[0083] As used herein, "stable transformation" or "stably transformed" means that a nucleic acid molecule is introduced into a cell and integrated into the genome of the cell. Thus, the integrated nucleic acid molecule can be inherited by its progeny, more specifically, by the progeny of multiple subsequent generations. As used herein, "genome" also includes the nuclear and plastid genomes, and therefore includes, for example, the integration of a nucleic acid into the chloroplast or mitochondrial genome. As used herein, stable transformation can also refer to a transgene that is maintained extrachromosomally, for example, as a minichromosome or a plasmid.

[0084] Transient transformation can be detected by enzyme-linked immunosorbent assays (ELISAs) or Western blots, which can detect the presence of peptides or polypeptides encoded by one or more transgenes introduced into an organism. Stable transformation of cells can be detected by Southern blot hybridization assays of the genomic DNA of the cells using nucleic acid sequences that specifically hybridize with the nucleotide sequence of the transgene introduced into an organism (e.g., plant, mammal, insect, archaea, bacteria, etc.). Stable transformation of cells can be detected by Northern blot hybridization assays of the RNA of the cells using nucleic acid sequences that specifically hybridize with the nucleotide sequence of the transgene introduced into a plant or other organism. Stable transformation of cells can also be detected by polymerase chain reaction (PCR) or other amplification reactions well known in the art, for example, using specific primer sequences that hybridize with the target sequence(s) of the transgene, resulting in amplification of the transgene sequences, which can be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols well known in the art.

[0085] Thus, in some embodiments, the nucleotide sequences, constructs, expression cassettes may be transiently expressed and / or they may be stably integrated into the genome of the host organism.

[0086] The recombinant nucleic acid molecules / polynucleotides of the invention can be introduced into cells by any method known to those of skill in the art. In some embodiments of the invention, transformation of a cell comprises nuclear transformation. In other embodiments, transformation of a cell comprises plastid transformation (e.g., chloroplast transformation). In still further embodiments, the recombinant nucleic acid molecules / polynucleotides of the invention can be introduced into cells via conventional breeding techniques.

[0087] Transformation procedures for both eukaryotes and prokaryotes are widely known and routine in the art and are described throughout the literature (see, e.g., Jiang et al. 2013. Nat. Biotechnol. 31:233-239; Ranetal. Nature Protocols 8:2281-2308 (2013)).

[0088] Thus, nucleotide sequences can be introduced into a host organism or its cells by any number of methods well known in the art. The methods of the present invention are not dependent on a particular method, as long as one or more nucleotide sequences are introduced into the organism and gain access to the interior of at least one cell of the organism. When more than one nucleotide sequence is introduced, they can be assembled as part of a single nucleic acid construct or as separate nucleic acid constructs, and can be located on the same or different nucleic acid constructs. Thus, the nucleotide sequences can be introduced into the cells of interest in a single transformation event or in separate transformation events, or, if relevant, the nucleotide sequences can be incorporated into plants as part of a breeding protocol.

[0089] The present invention relates to compositions and methods with increased efficiency and increased specificity for site-specific nicking, cleaving and / or modifying target DNA and for site-specific targeting of a polypeptide of interest to target DNA.

[0090] The nucleic acid constructs and nucleotide sequences of the present invention can be expressed in alternative ways that do not affect the overall structure or function of the construct or sequence. Thus, for example, in some cases, a synthetic CRISPR nucleic acid (crRNA) can be expressed in a manner that does not affect the overall structure or function of the construct or sequence. R1 Alternatively, the crRNA can include a stitching sequence that includes a (G / U) wobble base, thus forming a G R1The sequence is not further represented. In some cases, GR1 is unpaired (i.e., G / A is mismatched with Lje, Figure 20). Additional equivalences include those shown in the equivalence table provided below (Table 1) (see also Figure 22).

[0091] [Table 1]

[0092] Thus, in one aspect of the present invention, a synthetic transcoding CRISPR (tracr) nucleic acid (e.g., tracrRNA, tracrDNA) construct is provided, the construct comprising, in a 5' to 3' direction, an anti-zipper sequence comprising, consisting essentially of, or consisting of at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; TNANNC, T(A / C)A(A / G)(G / A)C (or U(A / C)A(A / G)(G / A)C)), TCAAAC, (or UCAAAC), TAAGGC (or UAAGGC), GATAAGG (or GAUAAGG), GATAAGGCTT (or GAUAAGGCUU), TCAAG (or UCAAG), TCAAGCAA (or and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, said hairpin comprising at least three matching base pairs, wherein the anti-zipper sequence is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence. In some embodiments, the anti-stitch sequence comprises, consists essentially of, or consists of the nucleotide sequence NNANN, ACAAA, TAAAA, (T / C)(A / G)T(A / G)(A / G), TAAAA, TGAAA.

[0093] In a further embodiment, from the 5' to 3' direction, an optional anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANNC, T(A / C)A(A / G)(G / A)C (or U(A / C)A(A / G)(G / A)C)), TCAAAC, (or UCAAAC), TAAGGC (or UAAGGC), GATAAGG (or GAUAAGG), GATAAGGCTT (or GAUAAGGCUU), TCAAG (or UCAAG), TCAAGCAA (or UCAAGCAA), T(C / A)AA(A / C)(C / A)(A / G)(A / T) (or U(C / A)AA(A / C)(C / A)(A / G)(A / U)). , GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC, or TCAAACAAAGCTTCAGC; and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, wherein the hairpin comprises at least three matching base pairs, and wherein the anti-zipper sequence, if present, is located immediately upstream of the bulge sequence, which is located immediately upstream of the anti-stitch sequence, which is located immediately upstream of the nexus sequence, and which is located immediately upstream of the hairpin sequence.

[0094] In some embodiments, the bulge sequence of the synthetic tracr nucleic acid construct comprises, consists essentially of, or consists of at least about 3 nucleotides. In some embodiments, the bulge sequence of the synthetic tracr nucleic acid construct comprises, consists essentially of, or consists of at least about 4 nucleotides. In other embodiments, the bulge sequence of the synthetic tracr nucleic acid construct comprises, consists essentially of, or consists of 5 nucleotides. In other embodiments, the hairpin sequence of the synthetic tracr nucleic acid construct comprises, consists essentially of, or consists of at least two hairpins, each hairpin containing at least three matching base pairs.

[0095] In a further aspect, the present invention provides an optional zipper sequence comprising, consisting essentially of, or consisting of, in the 3' to 5' direction, at least about 3 nucleotides; a bulge sequence comprising, consisting essentially of, or consisting of a nucleotide sequence having at least two nucleotides (e.g., a nucleotide sequence of (-NN-)); a stitch sequence comprising, consisting essentially of, or consisting of a nucleotide sequence of NNTNN (or NNUNN); a G sequence comprising, consisting essentially of, or consisting of the nucleotides G or GTT; R1 and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising, or consisting essentially of, a spacer sequence having at least 7 nucleotides at its 3' end that have 100% identity to the target DNA, wherein the zipper sequence, if present, is located immediately upstream of the bulge sequence, and the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is located immediately upstream of the zipper sequence, and the zipper sequence is located immediately upstream of the bulge ... zipper sequence, and the stitch sequence is located immediately upstream of the zipper sequence, and the zipper sequence is located immediately upstream of the zipper sequence, and the stitch sequence is located immediately upstream of the zipper sequence, and the zipper sequence is located immediately upstream of the zipper sequence, and the stitch sequence is located immediately upstream of the zipper sequence, and the zipper sequence is located immediately upstream of the zipper sequence, and the stitch sequence is R1 Located just upstream of, and G R1 is located immediately upstream of the spacer sequence.

[0096] In a further embodiment, a synthetic CRISPR nucleic acid (e.g., crRNA, crDNA) construct is provided that comprises, consists essentially of, or consists of, in the 3' to 5' direction: an optional zipper sequence comprising a nucleotide sequence having at least 3 nucleotides that hybridize to an anti-zipper; a bulge sequence comprising a nucleotide sequence of at least about 2 nucleotides; a stitch sequence comprising a nucleotide sequence of NNUNN; and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides having 100% identity to a target DNA at its 3' end, wherein the zipper sequence, if present, is located immediately upstream of the bulge sequence, which is located immediately upstream of the stitch sequence, and which is located immediately upstream of the spacer sequence.

[0097] In some embodiments, a synthetic CRISPR nucleic acid array is provided, the synthetic CRISPR nucleic acid array comprising a nucleotide sequence encoding two or more CRISPR nucleic acid constructs of the invention, wherein the two or more CRISPR nucleic acid constructs are located immediately adjacent to each other on the nucleotide sequence, the stitch sequences of the two or more CRISPR nucleic acid constructs are identical, the spacer sequences of the two or more CRISPR nucleic acid constructs are identical or non-identical, and, if present, the zipper sequences of the two or more CRISPR nucleic acid constructs are identical.

[0098] In other aspects, chimeric nucleic acid constructs (or guide nucleic acid constructs) are provided comprising a synthetic tracr nucleic acid construct and a synthetic CRISPR nucleic acid construct of the invention, wherein the zipper sequence of the synthetic CRISPR nucleic acid construct is at least about 70% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109%, 1110, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151, 152, 153%, 154, 155%, 156%, 157%, 158%, 159%, 160%, 161%, 162%, 163 %, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc.) complementary to and hybridize with the anti-stitch sequence of the synthetic CRISPR nucleic acid construct, the stitch sequence of the synthetic CRISPR nucleic acid construct is 100% identical to and hybridizes with the anti-stitch sequence of the synthetic tracr nucleic acid construct, and the bulge sequence of the synthetic CRISPR nucleic acid construct and the bulge sequence of the synthetic CRISPR nucleic acid construct are non-complementary.

[0099] In another embodiment, a chimeric nucleic acid construct is provided comprising, consisting essentially of, or consisting of a synthetic tracr nucleic acid construct and a synthetic CRISPR nucleic acid construct of the invention, wherein the stitch sequence NNUNN of the synthetic CRISPR nucleic acid construct is 100% complementary to and hybridizes with the anti-stitch sequence NNUNN of the synthetic tracr nucleic acid construct, the (G) of the stitch sequence forms a wobble base pair with the U of the anti-stitch sequence, the bulge sequence of the synthetic CRISPR nucleic acid construct and the bulge sequence of the synthetic CRISPR nucleic acid construct are non-complementary, and, when a zipper sequence and an anti-zipper sequence are present, the zipper sequence of the synthetic CRISPR nucleic acid construct hybridizes to the anti-zipper sequence of the synthetic tracr nucleic acid construct.

[0100] In some embodiments, the chimeric nucleic acid construct can optionally further comprise a nucleotide at the end of the hybridized sequence (distal to the bulge sequence) that links the hybridized zipper and anti-zipper sequences. In further embodiments, if the zipper and anti-zipper sequences are absent, the chimeric nucleic acid construct can optionally further comprise a nucleotide that links the bulge sequence of the synthetic trac nucleic acid sequence to the bulge sequence of the synthetic CRISPR nucleic acid. The linking nucleotide can be any nucleotide (e.g., T, A, G, C), and the number of nucleotides linking the zipper sequence and anti-zipper sequence or bulge sequence can be from about 3 to about 7.

[0101] In further embodiments, the synthetic tracr nucleic acid construct, synthetic CRISPR nucleic acid construct, CRISPR nucleic acid array, or chimeric nucleic acid construct of the invention can further comprise a nucleotide sequence encoding a Cas9 nuclease, an amino acid sequence having at least 70% identity (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc.) to an amino acid sequence encoding a Cas9 nuclease, or an amino acid sequence having at least 70% identity to an amino acid sequence encoding a Cas9 nuclease. The Cas9 nuclease useful in the present invention may be any Cas9 nuclease known to catalyze DNA cleavage in a CRISPR-Cas system. As known in the art, such Cas9 nucleases contain an HNH motif and a RuvC motif (see, e.g., WO2013 / 176772; WO2013 / 188638). In some embodiments, the HNH motif or the RuvC motif may contain a mutation that reduces or eliminates their activity compared to a wild-type Cas9 nuclease. In some embodiments, only one motif (e.g., either the HNH motif or the RuvC motif) is mutated. In other embodiments, both motifs are mutated to reduce or eliminate both activities. Any type of mutation, including missense mutations, nonsense mutations, frameshift mutations, etc., can be used to reduce or eliminate the activity of the HNH motif and / or the RuvC motif in a Cas9 nuclease.

[0102] This disclosure identifies various CRISPR-Cas systems and groupings of Cas9 nucleases. These groupings include the Streptococcus thermophilus CRISPR 1 (Sth CR1) group of Cas9 nucleases, the Streptococcus thermophilus CRISPR 3 (Sth CR3) group of Cas9 nucleases, the Lactobacillus buchneri CD034 (Lb) group of Cas9 nucleases, and the Lactobacillus rhamnosus GG (Lrh) group of Cas9 nucleases. Non-limiting examples of Sth CR1 group Cas9 nucleases include Cas9 nucleases encoded by the polypeptide sequences of SEQ ID NOS: 1-9 and 51. Non-limiting examples of Sth CR3 group Cas9 nucleases include Cas9 nucleases encoded by the polypeptide sequences of SEQ ID NOS: 10-23. Non-limiting examples of Lb group Cas9 nucleases include those encoded by the polypeptide sequences of SEQ ID NOs: 28, 30-33, 35, 43, 44, 47, 50, and 52. Non-limiting examples of Lrh group Cas9 nucleases include those encoded by the polypeptide sequences of SEQ ID NOs: 24-27, 29, 34, 36-42, 45, and 53. Additional Cas9 nucleases include, but are not limited to, those from Lactobacillus curvatus CRL 705. Additional Cas9 nucleases useful in the present invention include, but are not limited to, Cas9s from Lactobacillus animalis KCTC 3501 and Lactobacillus farciminis WP 010018949.1.

[0103] Thus, in some embodiments, the Cas9 nuclease may comprise, consist essentially of, or consist of a Cas9 from the Streptococcus thermophilus CRISPR 1 (Sth CR1) group of Cas9 nucleases, a Cas9 from the Streptococcus thermophilus CRISPR 3 (Sth CR3) group of Cas9 nucleases, a Cas9 nuclease from the Lactobacillus buchneri CD034 (Lb) group of Cas9 nucleases, and / or a Cas9 nuclease from the Lactobacillus rhamnosus GG (Lrh) group of Cas9 nucleases. In further embodiments, the amino acid sequence encoding the Cas9 nuclease may be the amino acid sequence of any one of SEQ ID NOs: 1 to 53. In yet further embodiments, the Cas9 nuclease useful in the synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, synthetic CRISPR nucleic acid arrays, and / or chimeric nucleic acid constructs of the present disclosure comprises, consists essentially of, or consists of a nucleotide sequence encoding an amino acid sequence having at least 70% identity to the amino acid sequence of any one of SEQ ID NOs:1 through 53.

[0104] Additionally, in certain embodiments, the Cas9 nuclease can be encoded by a nucleotide sequence that is codon-optimized for the organism containing the target DNA. In yet other embodiments, the Cas9 nuclease can comprise at least one nuclear localization sequence.

[0105] The inventors have surprisingly discovered functional pairings between specific groups of Cas9 nucleases and the nexus sequences of synthetic tracr nucleic acid constructs (tracrRNA, tracrDNA). Thus, in some embodiments, if the nexus sequence is GATAAGGC or GATAAGGCCATGCC, the Cas9 nuclease is from the Streptococcus thermophilus CRISPR 1 (STh CR1) group of Cas9 nucleases; if the nexus sequence is TAAGGC or TAAGGCTAGTCC, the Cas9 nuclease is from the Streptococcus thermophilus CRISPR 3 (Sth CR3) group of Cas9 nucleases; if the nexus sequence is TCAAGC or TCAAGCAAAGC, the Cas9 nuclease is from the Lactobacillus buchneri CD034 (Lb) group of Cas9 nucleases; or if the nexus sequence is TCAAAC or TCAAACAAAGCTTCAGC, the Cas9 nuclease is from the Lactobacillus rhamnosus GG (Lrh) group of Cas9 nucleases.

[0106] As described herein, Cas9 nucleases useful in the present invention may contain mutations in the HNH motif and / or RuvC motif, thereby reducing or eliminating the activity of the respective motifs. As known in the art, mutations in the HNH motif reduce / eliminate site-specific nicking of the (+) strand of double-stranded target DNA, and mutations in the RucV active site reduce / eliminate site-specific nicking of the (-) strand of double-stranded target DNA. Mutations in both active sites reduce / eliminate DNA cleavage (i.e., reduce / eliminate site-specific cleavage of target DNA). Thus, in some embodiments, the synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, CRISPR nucleic acid arrays, and / or chimeric nucleic acid constructs of the present disclosure comprise Cas9 nucleases with mutations in the RuvC active site motif. In other embodiments, the synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, CRISPR nucleic acid arrays, and / or chimeric nucleic acid constructs of the present disclosure comprise a Cas9 nuclease with a mutation in an HNH active site motif. In yet further embodiments, the synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, CRISPR nucleic acid arrays, and / or chimeric nucleic acid constructs of the present disclosure comprise a Cas9 nuclease with a mutation in an HNH active site motif and in a RuvC motif.

[0107] In yet further embodiments, the Cas9 nuclease having mutations within the HNH and RuvC motifs, thereby reducing or eliminating nuclease activity, further comprises a polypeptide of interest fused to the Cas9 nuclease, and such a Cas9-polypeptide of interest fusion protein can be used to target or direct the polypeptide of interest to a specific target DNA.

[0108] Further provided herein are methods for using the synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, CRISPR nucleic acid arrays, and / or chimeric nucleic acid constructs of the present disclosure. Accordingly, in some embodiments, a method for site-specific cleavage of double-stranded target DNA is provided, comprising contacting a chimeric nucleic acid construct of the present disclosure or an expression cassette comprising a chimeric nucleic acid construct of the present disclosure with target DNA in the presence of a Cas9 nuclease (e.g., SEQ ID NOs: 1-53), thereby causing site-specific cleavage of the target DNA within a region defined by complementary hybridization of the spacer sequence to the target DNA. In some embodiments, the site-specific cleavage can be site-specific nicking of the (+) strand of the double-stranded target DNA, wherein the Cas9 nuclease comprises a mutation in a RuvC active site motif, thereby cleaving the (+) strand of the double-stranded target to create a site-specific nick in the (+) strand of the double-stranded target DNA. In other embodiments, the site-specific cleavage is a site-specific nicking of the negative strand of the double-stranded target DNA, and the Cas9 nuclease comprises a point mutation within an HNH active site motif, thereby cleaving the negative strand of the double-stranded target DNA and generating a site-specific nick in the negative strand of the double-stranded target DNA.

[0109] In a further embodiment, a method for site-specific cleavage of double-stranded target DNA is provided, comprising contacting a transcoding CRISPR (tracr) nucleic acid molecule and a CRISPR nucleic acid molecule with target DNA in the presence of a Cas9 nuclease (e.g., SEQ ID NOs: 1-53), wherein (a) the tracr nucleic acid molecule comprises, in the 5' to 3' direction: an anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC, or and (b) a CRISPR nucleic acid molecule encoded by a nucleotide sequence comprising: a nexus sequence comprising the nucleotide sequence TCAAACAAAGCTTCAGC; and a hairpin sequence comprising a nucleotide sequence comprising, consisting essentially of, or consisting of at least one hairpin, wherein the hairpin comprises at least three matching base pairs, and the anti-zipper sequence is located immediately upstream of the bulge sequence, which is located immediately upstream of the anti-stitch sequence, which is located immediately upstream of the nexus sequence, and which is located immediately upstream of the hairpin sequence; and (b) a CRISPR nucleic acid molecule encoded by a nucleotide sequence comprising, in the 3' to 5' direction, a zipper sequence comprising at least about 3 nucleotides, a bulge sequence comprising a nucleotide sequence having at least two nucleotides, a stitch sequence comprising the nucleotide sequence NNTNN (or NNUNN), a G comprising the nucleotides G or GTT. R1 and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity with the target DNA, wherein the zipper sequence is located immediately upstream of the bulge sequence, and the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is R1 Located just upstream of, and G R1are located immediately upstream of the spacer sequence, and further wherein the anti-zipper sequence and anti-stitch sequence of the tracr nucleic acid molecule are at least about 70% complementary to and hybridize with the zipper sequence and stitch sequence of the CRISPR nucleic acid molecule, respectively, and the spacer sequence of the CRISPR nucleic acid molecule is at least about 80% complementary to and hybridizes with a portion of the target DNA (adjacent to a protospacer adjacent motif (PAM) on the target DNA), thereby resulting in site-specific cleavage of the target DNA within a region defined by the complementary binding of the spacer sequence of the CRISPR nucleic acid molecule to the target DNA.

[0110] In another embodiment, a method for site-specific cleavage of double-stranded target DNA is provided, comprising contacting the double-stranded target DNA with a chimeric nucleic acid comprising, in the 5' to 3' direction, an anti-zipper sequence comprising at least about 9 nucleotides; a bulge sequence comprising at least about 4 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANNC; a nexus sequence comprising the nucleotide sequence of TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC, or TCAAACAAAGCTTCAGC; and at least one hairpin. (b) a hairpin sequence comprising, in a 3' to 5' direction, a zipper sequence comprising at least about 3 nucleotides, a bulge sequence comprising a nucleotide sequence having at least 2 nucleotides, a stitch sequence comprising the nucleotide sequence NNTNN (or NNUNN), a G comprising the nucleotides G or GTT, and a nucleotide sequence having the nucleotides GTT, ... R1and a second nucleotide sequence comprising a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity with the target DNA (the zipper sequence is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is located immediately upstream of the G R1 Located just upstream of, and G R1 is located immediately upstream of the spacer sequence; and (c) a third nucleotide sequence encoding an amino acid sequence having at least 80% identity to an amino acid sequence encoding a Cas9 nuclease (e.g., SEQ ID NOs: 1-53), wherein the anti-zipper and anti-stitch sequences of the first nucleotide sequence hybridize to the zipper and stitch sequences of the second nucleotide sequence, and the spacer sequence of the second nucleotide sequence hybridizes to a portion of the target DNA (adjacent to a protospacer adjacent motif (PAM) on the target DNA), thereby effecting site-specific cleavage of the target DNA within a region defined by complementary binding of the spacer sequence of the second nucleotide sequence to the target DNA.

[0111] In a further embodiment, a method for site-specific targeting of a polypeptide of interest to double-stranded (ds) target DNA is provided, comprising contacting a transcoding CRISPR (tracr) nucleic acid molecule and a CRISPR nucleic acid molecule with the target DNA in the presence of a Cas9 nuclease (e.g., SEQ ID NOs: 1-53), wherein (a) the tracr nucleic acid molecule comprises, in a 5' to 3' direction: an anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTC and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, the hairpin comprising at least three matching base pairs, an anti-zipper sequence located immediately upstream of a bulge sequence, the bulge sequence located immediately upstream of an anti-stitch sequence, the anti-stitch sequence located immediately upstream of the nexus sequence, and the nexus sequence located immediately upstream of the hairpin sequence; and (b) the CRISPR nucleic acid molecule is encoded by a nucleotide sequence comprising, in the 3' to 5' direction, a zipper sequence comprising at least about 3 nucleotides, a bulge sequence comprising at least two nucleotides (e.g., a nucleotide sequence of (-NN-)), a stitch sequence comprising a nucleotide sequence of NNTNN, a G comprising the nucleotides G or GTT. R1 and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity with the target DNA, wherein the zipper sequence is located immediately upstream of the bulge sequence, and the bulge sequence is located immediately upstream of the stitch sequence, and the stitch sequence is R1 Located just upstream of, and G R1is located immediately upstream of the spacer sequence, and further wherein the Cas9 nuclease comprises a mutation in the HNH active site motif, a mutation in the RuvC active site motif, and is fused to a polypeptide of interest, wherein the anti-zipper sequence is about 70% complementary to and hybridizes with the zipper sequence, the stitch sequence is 100% complementary to and hybridizes with the stitch sequence, and the spacer sequence is about 80% complementary to and hybridizes with the target DNA adjacent to a protospacer adjacent motif (PAM) on the target DNA, thereby resulting in site-specific targeting of the polypeptide of interest to the target DNA at a region defined by complementary binding of the spacer sequence of the CRISPR nucleic acid molecule to the target DNA.

[0112] In representative embodiments, with respect to synthetic tracr nucleic acid constructs as described herein, the bulge sequence or first nucleotide sequence of the synthetic tracr nucleic acid molecule can comprise, consist essentially of, or consist of about 3, 4, or 5 nucleotides. In other embodiments, the bulge sequence can comprise, consist essentially of, or consist of 5 nucleotides, and the hairpin sequence can comprise, consist essentially of, or consist of at least two hairpins, where each hairpin contains at least three matching base pairs.

[0113] In a further embodiment, the invention provides a method for site-specific cleavage of double-stranded target DNA comprising contacting a transcoding CRISPR (tracr) nucleic acid molecule and a CRISPR nucleic acid molecule with the target DNA in the presence of a Cas9 nuclease, wherein: (a) the tracr nucleic acid molecule is encoded by a nucleotide sequence comprising, in the 5' to 3' direction, an optional anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; a nexus sequence comprising the nucleotide sequence of TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC, or TCAAACAAAGCTTCAGC; a hairpin sequence comprising a nucleotide sequence having at least one hairpin, wherein the hairpin comprises at least three matching base pairs; and the anti-zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; and (b) the CRISPR nucleic acid molecule is encoded by a nucleotide sequence comprising, in a 3' to 5' direction, an optional zipper sequence comprising a nucleotide sequence having at least three nucleotides that hybridize to the anti-zipper, a bulge sequence comprising a nucleotide sequence of (-NN-), a stitch sequence comprising a nucleotide sequence of NNUNN, and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end that have 100% identity to the target DNA; and the zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the stitch sequence, the stitch sequence is located immediately upstream of the spacer sequence; and Further herein, when an anti-zipper and zipper sequence are present, the anti-zipper sequence hybridizes to the zipper sequence, the NNANN of the anti-stitch sequence is complementary to and hybridizes to the NNUNN of the stitch sequence, and the spacer sequence of the CRISPR nucleic acid molecule hybridizes to a portion of the target DNA (adjacent to a protospacer adjacent motif (PAM) on the target DNA), thereby resulting in site-specific cleavage of the target DNA within a region defined by hybridization of the spacer sequence of the CRISPR nucleic acid molecule to the target DNA.

[0114] In still further embodiments, a method for site-specific cleavage of double-stranded target DNA is provided, the method comprising contacting the double-stranded target DNA with a chimeric nucleic acid comprising: (a) a first nucleotide sequence comprising, in the 5' to 3' direction, an optional anti-zipper sequence comprising at least about three nucleotides; a bulge sequence comprising at least about three nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; a nexus sequence comprising the nucleotide sequence of TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC, or TCAAACAAAGCTTCAGC; and a hairpin sequence comprising a nucleotide sequence having at least one hairpin, said hairpin comprising at least three matching base pairs; wherein the anti-zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; (b) a second nucleotide sequence comprising, in the 3' to 5' direction, an optional zipper sequence comprising a nucleotide sequence having at least three nucleotides that hybridize to the anti-zipper, a bulge sequence comprising a nucleotide sequence having at least two nucleotides (e.g., a nucleotide sequence of (-NN-)), a stitching sequence comprising a nucleotide sequence of NNUNN, and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity to the target DNA; and If present, the zipper sequence is located immediately upstream of the bulge sequence, which is located immediately upstream of the stitch sequence, which is located immediately upstream of the spacer sequence; and (c) a third nucleotide sequence encoding an amino acid sequence having at least 80% identity to an amino acid sequence encoding a Cas9 nuclease (e.g., SEQ ID NOs: 1-53); Here, when a zipper sequence and an anti-zipper sequence are present, the zipper sequence hybridizes to the anti-zipper sequence, the NNANN of the anti-stitch sequence is complementary to and hybridizes with the NNUNN of the stitch sequence, and the spacer sequence of the second nucleotide sequence hybridizes to a portion of the target DNA (adjacent to a protospacer adjacent motif (PAM) on the target DNA), thereby resulting in site-specific cleavage of the target DNA within a region defined by hybridization of the spacer sequence of the second nucleotide sequence to the target DNA.

[0115] In a further embodiment, a method for site-specific targeting of a polypeptide of interest to double-stranded (ds) target DNA is provided, comprising contacting a transcoding CRISPR (tracr) nucleic acid molecule and a CRISPR nucleic acid molecule with the target DNA in the presence of a Cas9 nuclease (e.g., SEQ ID NOs: 1-53); wherein (a) the tracr nucleic acid molecule is encoded by a nucleotide sequence comprising, in the 5' to 3' direction, an optional anti-zipper sequence comprising at least about 3 nucleotides; a bulge sequence comprising at least about 3 nucleotides; an anti-stitch sequence comprising the nucleotide sequence of NNANN; a nexus sequence comprising the nucleotide sequence of TNANNC, T(A / C)A(A / G)(G / A)C, TCAAAC, TAAGGC, GATAAGG, GATAAGGCTT, TCAAG, TCAAGCAA, T(C / A)AA(A / C)(C / A)(A / G)(A / T), GATAAGGCCATGCC, TAAGGCTAGTCC, TCAAGCAAAGC, or TCAAACAAAGCTTCAGC; a hairpin sequence comprising a nucleotide sequence having at least one hairpin, wherein the hairpin comprises at least three matching base pairs; and the anti-zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; and (b) the CRISPR nucleic acid molecule is encoded by a nucleotide sequence comprising, in the 3' to 5' direction, an optional zipper sequence comprising a nucleotide sequence having at least three nucleotides that hybridize to the anti-zipper, a bulge sequence comprising a nucleotide sequence having at least two nucleotides (e.g., a nucleotide sequence of (-NN-)), a stitch sequence comprising a nucleotide sequence of NNUNN, and a spacer sequence having a 5' end and a 3' end, the spacer sequence comprising at least 7 nucleotides at its 3' end having 100% identity to the target DNA; and the zipper sequence, if present, is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the stitch sequence, the stitch sequence is located immediately upstream of the spacer sequence, and Further herein, the Cas9 nuclease comprises a mutation in the HNH active site motif and a mutation in the RuvC active site motif and is fused to a polypeptide of interest, wherein, if a zipper sequence and an anti-zipper sequence are present, the zipper sequence hybridizes to the anti-zipper sequence, the anti-stitch sequence NNANN is complementary to and hybridizes to the stitch sequence NNUNN, and the spacer sequence hybridizes to the target DNA adjacent to a protospacer adjacent motif (PAM) on the target DNA, thereby resulting in site-specific targeting of the polypeptide of interest to the target DNA within a region defined by hybridization of the spacer sequence of the CRISPR nucleic acid molecule to the target DNA.

[0116] In some embodiments, when the anti-zipper and zipper sequences or the first and second nucleotide sequences of the tracr nucleic acid molecule and the aCRISPR nucleic acid molecule hybridize, the hybridized sequence can optionally further comprise an additional nucleotide at the end of the hybridized sequence distal to the bulge sequence, thereby linking the hybridized zipper and anti-zipper sequences. In further embodiments, when the zipper and anti-zipper are absent, the chimeric nucleic acid construct can optionally further comprise a nucleotide linking the bulge sequence of the synthetic trac nucleic acid sequence to the bulge sequence of the synthetic CRISPR nucleic acid. The linking nucleotide can be any nucleotide (e.g., T, A, G, C), and the number of nucleotides linking the zipper and anti-zipper sequences or bulge sequences can be from about 3 to about 7.

[0117] Any wild-type, mutant, or codon-optimized Cas9 nuclease, including but not limited to SEQ ID NOs: 1-53, or that contains at least one nuclear localization sequence described herein, can be used in the methods of the invention.

[0118] Further provided herein are expression cassettes and vectors comprising the nucleic acid constructs, nucleic acid arrays, nucleic acid molecules and / or nucleotide sequences of the invention, which can be used in the methods of the disclosure.

[0119] In further aspects, the nucleic acid constructs, nucleic acid arrays, nucleic acid molecules, and / or nucleotide sequences of the present invention can be introduced into cells of a host organism. Any cell / host organism for which the present invention is useful can be used. Exemplary host organisms include, but are not limited to, plants, bacteria, archaea, fungi, animals, mammals, insects, birds, fish, amphibians, cnidarians, humans, or non-human primates. In certain embodiments, the host organism may be, but is not limited to, Homo sapiens, Drosophila melanogaster, Mus musculus, Rattus norvegicus, Caenorhabditis elegans, Saccharomyces pombe, Saccharomyces cerevisiae, Glycine max, Zeae maydis, Gossypium hirsutum, or Arabidopsis thaliana. In further embodiments, cells useful in the present invention may be, but are not limited to, stem cells, somatic cells, germ cells, plant cells, animal cells, bacterial cells, archaeal cells, fungal cells, mammalian cells, insect cells, avian cells, fish cells, amphibian cells, cnidarian cells, human cells, or non-human primate cells. In other embodiments, cells useful in the present invention include, but are not limited to, cells from Homo sapiens, Drosophila melanogaster, Mus musculus, Rattus norvegicus, Caenorhabditis elegans, Saccharomyces pombe, Saccharomyces cerevisiae, Glycine max, Zeae maydis, Gossypium hirsutum, or Arabidopsis thaliana.

[0120] In further embodiments of the invention, the polypeptide of interest may include, but is not limited to, a helicase, a nuclease, a methyltransferase, a gyrase, a demethylase, a kinase, a dismutase, an integrase, a transposase, a telomerase, a recombinase, an acetyltransferase, a deacetylase, a polymerase, a phosphatase, a ligase, a ubixin ligase, a photolyase, or a glycosylase. In other embodiments of the invention, the polypeptide of interest has depurination activity, oxidation activity, pyrimidine dimer formation activity, alkylation activity, DNA repair activity, DNA damage activity, deubiquitination activity, adenylation activity, deadenylation activity, sumoylation activity, desumoylation activity, ribosylation activity, deribosylation activity, myristoylation activity, demyristoylation activity, or telomere repair activity, or deamination activity. In exemplary embodiments, the polypeptide of interest may be a polypeptide having kinase activity, nuclease activity, methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, phosphatase activity, ubixin ligase activity, deubiquitinating activity, or telomere repair activity.

[0121] Further provided herein are kits comprising the nucleic acid constructs, nucleic acid molecules, and / or nucleotide sequences of the invention.

[0122] Thus, in one embodiment, a kit for site-specific cleavage of double-stranded DNA is provided, the kit comprising a synthetic tracr nucleic acid construct, a synthetic CRISPR nucleic acid construct, a CRISPR nucleic acid array, or a chimeric nucleic acid construct of the invention. In another embodiment, a kit for site-specific targeting of a polypeptide of interest to double-stranded (ds) target DNA is provided, the kit comprising a synthetic tracr nucleic acid construct, a synthetic CRISPR nucleic acid construct, a CRISPR nucleic acid array, or a chimeric nucleic acid construct of the invention. In some embodiments, the kit may comprise a synthetic tracr nucleic acid construct, a synthetic CRISPR nucleic acid construct, a CRISPR nucleic acid array, and / or a chimeric nucleic acid construct of the invention contained in one or more expression cassettes. In yet further embodiments, the kit may further comprise a Cas9 nuclease (e.g., SEQ ID NOs: 1-53) for use with the nucleic acid constructs, nucleic acid arrays, nucleic acid molecules, and / or nucleotide sequences of the invention described herein.

[0123] In further aspects, the kit may include primers that include portions of the CRISPR repeat sequences in both directions, hi other embodiments, the kit may include primers designed to extend through the CRISPR repeat sequences in both directions, including the boundaries of the CRISPR array (i.e., the leader end at one end and the trailer end at the other). In a further embodiment, the kit may further comprise instructions for use.

[0124] The present invention will now be described with reference to the following examples. It should be understood that these examples are not intended to limit the scope of the claims to this invention, but are intended to be illustrative of particular embodiments. Any variations in the exemplified methods that occur to those skilled in the art are intended to be within the scope of the present invention. [Example]

[0125] Example 1. Evaluation of the functional role of modules identified within guide sequences Clustered regularly interspaced short palindromic repeats (CRISPR) and associated Cas proteins confer adaptive immunity against invasive genetic agents in bacteria and archaea 1 In the type II CRISPR-Cas system, the endonuclease Cas9, guided by a signature RNA, specifically targets sequences complementary to the CRISPR spacer and generates double-stranded DNA breaks (DSBs) using two nickase domains (Makarova, K. et al. Nat Rev Microbiol 9, 467-477 (2011); Garneau, J. et al. Nature 468, 67-71 (2010); Sapranauskas, R. et al. Nucleic Acids Res 39, 9275-9282 (2011); Gasiunas, G. et al. Proc Natl Acad Sci U S A 109, E2579-E2586 (2012); Jinek, M. et al. Science 337, 816-821 (2012)). Any DNA sequence can be targeted as long as it is flanked by a Cas9-specific protospacer-adjacent motif (PAM). Garneau, J. et al. Nature 468, 67-71 (2010); Sapranauskas, R. et al. Nucleic Acids Res 39, 9275-9282 (2011); Gasiunas, G. et al. Proc Natl Acad Sci U S A 109, E2579-E2586 (2012); Jinek, M. et al., Science 337, 816-821 (2012); Sternberg, SH et al. Nature 507, 62 (2014)). Targeting and cleavage by the Cas9 system relies on an RNA duplex consisting of CRISPR RNA (crRNA) and transactivating crRNA (tracrRNA). 8This natural complex can be replaced by a synthetic single-guide RNA (sgRNA) chimera that mimics the crRNA:tracrRNA duplex (Jinek, M. et al., Science 337, 816-821 (2012)). sgRNA combined with Cas9 creates a convenient, compact, and portable sequence-specific targeting system suitable for engineering and heterologous transfer into a variety of model systems of industrial and translational interest.

[0126] Thus, Cas9:sgRNA technology, which provides a compact and practical means for generating double-strand breaks (DSBs), has revolutionized genome organization (Mali, P. et al., Science 339, 823-826 (2013); Cong, L. et al., Science 339, 819-823 (2013); Jiang, W. et al., Nat. Biotechnol. 31, 233-239 (2013); Sander, JD & Joung, JK, Nature Biotechnol. 32, 347-355 (2014)), opening up new avenues for high-throughput genome-wide genetic screening. 13,14, expanding the transcriptional regulation toolbox (Qi, LS et al. Cell 152, 1173-1183 (2013); Gilbert, LA et al. Cell 154, 442-451 (2013)). Furthermore, the absence of interactions between evolutionarily distant Cas9:sgRNAs (Chylinski, K. et al. RNA biology 10, 726-737 (2013); Fonfara, I. et al. Nucleic Acids Res (2013); Esvelt, K. et al. Nature Methods (2013)) has enabled multiple independent targeting to be achieved in cells when coexisting functional type II CRISPR-Cas systems function simultaneously (Barrangou, R. et al. Science 315, 1709-1712 (2007); Horvath, P. et al. J Bacteriol. 190, 1401-1412 (2008)). Despite the widespread use of these molecular mechanisms, critical properties of sgRNA guides and their involvement in defining functionally orthologous Cas9 endonucleases remain to be characterized. Indeed, early attention on Cas9 targeting and cleavage focused on spacer:target complementarity and PAM sequence sensitivity, while information defining the factors that drive Cas9:sgRNA interactions and direct orthogonality between type II CRISPR-Cas systems remains scarce. Therefore, we set out to identify and characterize the properties within sgRNAs that confer Cas9 targeting and cleavage specificity, paving the way for new engineering avenues for CRISPR technology.

[0127] Therefore, to evaluate the functional roles and implications of the various modules identified within the guide sequences described herein (e.g., synthetic tracr nucleic acid constructs, synthetic CRISPR nucleic acid constructs, chimeric nucleic acid constructs (tracr nucleic acid-synthetic CRISPR nucleic acid constructs)), we designed guide mutational variants containing modifications or deletions of each of the aforementioned functional modules. We selected the SthCRISPR3 system as a representative functional model and first established a positive functional control (wild type, WT) using a native guide sequence. We then used a "stitch-deleted" " We then tested a "stitch" mutant, which lacks the bulge, and observed a loss of function. We then tested a "bulge" mutant, which lacks the nexus, and observed a loss of function. We then tested a "nexus" mutant, which lacks the first hairpin, and observed a loss of function. We then tested a "hairpin" mutant, which lacks the first hairpin, and observed a loss of function. We then tested sequence specificity and variability sensitivity, establishing that the nexus sequence is specific in the mutant constructs, while other functional modules, particularly the zipper, bulge, hairpin, and stitch, are variability tolerant. The results of these experiments are shown in Figures 1-21.

[0128] Therefore, Figure 1 shows the multiple sequence alignment for the nexus module, and Figure 2 provides a maximum likelihood tree for the nexus module developed through this study. Figures 3A-3D show the consensus sequences for the nexus modules of the Sth Cr1 group (Figure 3A), Sth Cr3 group (Figure 3B), Lrh group (Figure 3C), and Lbu group (Figure 3D).

[0129] Figure 5 shows a multiple sequence alignment for the anti-stitch module, while Figures 6A-6D show the consensus sequences for the anti-stitch modules of the Sth Cr1 group (Figure 6A), the Sth Cr3 group (Figure 6B), the Lrh group (Figure 6C), and the Lbu group (Figure 6D).

[0130] Similarly, a multiple sequence alignment for the bulge modules is provided in Figure 7, and the consensus sequence for the bulge modules of the Sth Cr1 group is shown in Figure 8A. Figure 8B shows the consensus sequence for the bulge modules of the Sth Cr3 group. Figure 8C shows the consensus sequence for the Lrh group, and Figure 8D shows the consensus sequence for the bulge modules of the Lbu group.

[0131] A multiple sequence alignment for the zipper module is shown in FIG. 9, and FIG. 10 shows a maximum likelihood tree for the zipper module.

[0132] Figure 11 shows a multiple sequence alignment for the bulge, anti-stitch, and nexus modules. Figure 4 shows a maximum likelihood tree for the Cas9 nuclease.

[0133] Figures 12-21 show guide sequences and targeting for various cRNA:tracRNA constructs, including Streptococcus thermophilus CR3 representing the Sth CR1 group (Figure 12); Lactobacillus buchneri representing the Lbu group (Figure 13); Streptococcus thermophilus CR1 representing the Sth CR1 group (Figure 14); Streptococcus pyrogenes M1 GAS representing the Sth CR3 group (Figure 15); Lactobacillus rhamnosus representing the Lrh group (Figure 16); Lactobacillus animalis representing the Lan group (Figure 17); Lactobacillus casei representing the Lca group (Figure 18); Lactobacillus gasseri representing the Lga group (Figure 19); Lactobacillus jensenii representing the Lje group (Figure 20); and Lactobacillus pentosus representing the Lpe group (Figure 21). The bottom of each of Figures 12-21 represents a target dsDNA containing the target sequence (open structure) and the adjacent (3') PAM. The top of each figure represents the CRISPR RNA (crRNA), consisting of a 5' portion complementary to the target sequence and a 3' portion derived from the CRISPR repeat; and also shows the tracrRNA, consisting of the anti-CRISPR repeat portion, and the nexus and 3' hairpin. As shown in each of Figures 12-21, the complementary portions of the crRNA:tracrRNA duplex consist of the lower stem (lower complementary portion), the bulge (herniated mismatch), and the upper stem (upper complementary portion).

[0134] Example 2. Characterization of guide RNA sequences in various additional Type II systems The findings described in Example 1 establish the critical modules in the sgRNA required to support Streptococcus pyrogenes Cas9 (SpyCas9) activity. However, despite its widespread use in genome editing, SpyCas9 is only one of many Cas9 orthologs found in nature (Chylinski, K. et al. RNA biology 10, 726-737 (2013); Fonfara, I. et al. Nucleic Acids Res (2013)). Therefore, we next investigated whether the same sgRNA sequence characteristics occur in other type II CRISPR-Cas systems. We investigated the sgRNA sequences in Streptococcus and Lactobacillus genomes (where type II systems preferentially occur). 2 We sampled 41 Cas9 sequences from the genome of Cas9 gene expression vectors (CRISPR-like repeats) and identified their corresponding CRISPR repeats to predict tracrRNA sequences. Cas9 protein sequences clustered into three major sequence groups (Figure 23). As expected, similar groupings were observed when clustering was performed using either CRISPR repeats or predicted tracrRNA sequences (Figure 24A, Figure 23, Figure 25, Figure 26), confirming the presence of anti-CRISPR repeats within tracrRNA and the close molecular relationship between Cas9 and the crRNA:tracrRNA pair (Makarova, K. et al. Nat Rev Microbiol 9, 467-477 (2011); Deltcheva, E. et al. Nature 471, 602-607 (2011); Fonfara, I. et al. Nucleic Acids Res (2013)). Within the tracrRNA sequences, we consistently observed the functional modules identified for SpyCas9 (Figure 24B), with conservation of the overall sgRNA / crRNA:tracrRNA structure across families and a high level of sequence conservation within clusters.

[0135] The presence of a bulge with a twisted orientation between the lower stem (i.e., stitch / anti-stitch) and the upper stem (i.e., zipper / anti-zipper) was consistently observed across the diversity of systems. The length of the lower stem was highly conserved within families and variable between families. Interestingly, the highest level of conservation was observed for the nexus sequence (Figure 24B, Figure 27). The general nexus shape, with a GC-rich stem and offset uracil, was shared between the two Streptococcus families. In contrast, the unique double-stem nexus (Figure 24A-B) was unique and ubiquitous in the Lactobacillus system. Surprisingly, several bases within the nexus were strictly conserved even between different families, including A52 and C55 (Figure 24A-B), further highlighting the crucial role of this module. Indeed, A52 interacts with the backbone of residues 1103-1107 near the 5' end of the target strand in the crystal structure of SpyCas9, suggesting that interactions between the protein backbone and the nexus may be required for PAM binding.

[0136] Determining the relationship of guide RNA structure to the orthogonality of the Cas9 protein.The findings described herein suggest a potential relationship between sgRNA sequence and structure and the diversity of Cas9 proteins. This observation prompted us to determine the sgRNA modules that define the group of Cas9 orthologs. Therefore, we selected endonucleases from two naturally coexisting orthologous S. thermophilus type II systems, Sth1 Cas9 and Sth3 Cas9 (Horvath, P. et al. J Bacteriol. 190, 1401-1412 (2008)), to investigate the relationship between sgRNA composition and Cas9 orthogonality. A series of experiments was designed based on the self-targeting activity in Escherichia coli (Figure 28A-B) to test whether specific mutations within the sgRNA could promote cleavage activity in previously orthologous Cas9s. We identified regions within the E. coli genome containing overlapping Cas9 target sites for the Sth1Cas9 and Sth3Cas9 systems, confirming that cleavage occurs within a single nucleotide (Figure 29B) and that the PAM sequences overlap appropriately. We created chimeric versions of the two sgRNA backbones, exchanging the spacer, lower stem (i.e., stitch / antistitch)-bulge-upper stem (i.e., zipper / anti-zipper), nexus, and hairpin (Figure 29C), and tested their ability to drive self-targeting by either Sth1Cas9 or Sth3Cas9 (Gomaa, AA et al. MBio. 5, e00928-13 (2014)). We first confirmed that these two systems are indeed orthogonal in this assay system, and that each guide alone drives targeting by its cognate Cas9 (Figure 29C). We then showed that exchanging the spacer sequence resulted in an sgRNA with a CRISPR3 spacer and a CRISPR1 backbone capable of supporting Sth1Cas9 cleavage activity, but the converse was not true for Sth3Cas9 activity.sgRNAs containing the CRISPR1 spacer and CRISPR3 backbone do not support Sth3Cas9 activity (Figure 29C). We hypothesize that this unidirectional reciprocal functionality is due to the flexibility of the spacing requirements between the protospacer and PAM in the SthCRISPR1 system (Chen et al. J. Biol. Chem. doi:10.1074 / jbc.M113.539726.(2014)) (Figure 29C, upper panel). We then demonstrated that the functionality between sgRNA and Cas9 can be altered simply by swapping the nexus-hairpin combination between the two orthogonal systems. A major consequence of reprogramming the sgRNA is that its ability to guide the original Cas9 is lost in the process (Figure 29C, lower panel). This contrasts with the standard idea that CRISPR repeat sequences play a key role in defining orthologous CRISPR-Cas systems. However, these results demonstrate that chimeric sgRNAs with altered nexus sequences can reprogram orthogonality in a predictable and unidirectional manner, which is crucial for further exploiting orthogonal Cas9 proteins associated with different PAMs (Esvelt, K. et al. Nature Methods (2013)).

[0137] Recent structural and biochemical data have begun to shed light on the mechanism of DNA recognition and cleavage by Cas9 (Jinek, M. et al., Science 337, 816-821 (2012); Jinek, M. et al., Science 343, 6176 (2014); Nishimasu, H. et al. Cell 156, 935 (2014)). Electron micrographs of apo-, RNA-bound, and protein / RNA / DNA complexes have shown that upon binding to guide RNA, Cas9 undergoes dramatic conformational changes to facilitate target DNA binding and cleavage (Jinek, M. et al., Science 343, 6176 (2014)). Consistent with the electron microscopy images, crystal structures show that the SpyCas9:sgRNA:DNA:complex and apo-SpyCas9 occupy significantly different structures, with substantial rearrangements of the RNA- and DNA-binding domains occurring between the two structures (Jinek, M. et al., Science 343, 6176 (2014); Nishimasu, H. et al. Cell 156, 935 (2014)). The nexus occupies a critical position in the SpyCas9-sgRNA:DNA complex, aligning many key components of the protein and sgRNA and properly positioning both the protein and RNA to receive the target DNA duplex for cleavage. Upon binding to the sgRNA:DNA, an arginine-rich bridge helix binds to the base and lower stem of the nexus. In addition, the nexus interacts with two small regions from the two lobes of SpyCas9, which we propose to establish as Nexus Interacting Region 1 (NIR1) 446-497 and Nexus Interacting Region 2 (NIR2) 1105-1138. Both of these regions are disordered within the apoSpyCas9 structure and are involved in, among other things, PAM recognition. 21NIR2 contains two tryptophan residues identified as important in the nexus. NIR2 also directly interacts with the lower stem, and the face opposite the nexus-binding site is located very close to the 3' end of the target strand, suggesting that interaction with the nexus may be necessary to prime the PAM recognition site. Notably, in the Actinomyces naeslundii Cas9 (AnaCas9) apo-structure (Jinek, M. et al., Science 343, 6176 (2014)), NIR2 is primed and contains an insertion of approximately 50 amino acids. There is reason to believe that AnaCas9 may recognize larger nexuses (and possibly PAM sequences).

[0138] However, these results reveal six distinct features within guide RNAs and establish the bulge and nexus as structure- and sequence-specific properties that guide Cas9 targeting and cleavage. This provides a basis for optimizing sgRNA composition and the opportunity to engineer shorter, minimal guide RNAs that contain smaller regions of double-stranded RNA that potentially trigger innate immune responses and are more suitable for packaging into viruses such as adeno-associated viruses. This understanding of Type II CRISPR-Cas systems is supported by Briner et al. (Mol. Cell 56:333-339 (2014)) and Nishimasu et al. (Cell 156:935-949 (2014)), who showed that sequence modifications affect the ability of the modified sequence to guide Cas9. Notably, these studies confirm that the nexus sequence within the guide is critical for guiding Cas9 toward complementary DNA and subsequent cleavage.

[0139] The ability to reprogram Cas9 orthogonally using chimeric sgRNAs with engineered nexus sequences opens new avenues for the development of novel Cas9 proteins with the potential to exploit the diversity of natural Cas9 orthologs, including short Cas9 variants for convenient packaging and delivery. Additionally, the ability to reprogram Cas9 with chimeric sgRNAs will enable increased use of various PAMs for flexible management of target frequency (short PAMs are frequent) and specificity through reduced off-target cleavage (longer PAMs are rare). This also opens up multiplexing opportunities by using a single Cas9 with various chimeric guides or by simultaneously using orthogonal systems with different combinations of standard or chimeric sgRNAs. Collectively, our findings open new avenues for Cas9-dependent DNA targeting and set the stage for the development of next-generation CRISPR-based technologies.

[0140] Example 3. The cas9 genes from the CRISPR1 and CRISPR3 loci were PCR-amplified from genomic DNA from S. thermophilus LMD-9 and cloned into pwtCas9-bacteria (Addgene #44250) (Qi, LS et al. Cell 152, 1173-1183 (2013)). To construct the sgRNA-expressing plasmids, the SpeI restriction site in the pdCas9-bacteria plasmid (Addgene #44249) (Id.) was removed, and a gBlock (IDT) encoding a zraP-targeting sgRNA based on the CRISPR1 or CRISPR3 locus was combined with the PCR-amplified backbone of the pgRNA-bacteria plasmid (Addgene #44251) (Id.). For transformation analysis, E. coli K-12 was used. Transformation efficiency was calculated by dividing the number of transformants for the tested sgRNA plasmid by the number of transformants for the psgRNA-C1-T4 control plasmid, as previously described (Gomaa, AA et al. MBio. 5, e00928-13 (2014)).

[0141] Plasmid construction.To construct the Cas9-expressing plasmid, the Cas9 genes from the CRISPR1 locus (Sth1-Cas9) and CRISPR3 locus (Sth3-Cas9) were PCR-amplified from genomic DNA extracted from S. thermophilus LMD-9. Each PCR product was combined with the PCR-amplified backbone of pwtCas9-bacteria (Addgene #44250) (Qi, L. et al. Cell 152, 1173-1183 (2013)) by Gibson assembly. To construct the sgRNA-expressing plasmid, the SpeI restriction site in the pdCas9-bacteria plasmid (Addgene #44249) (Id.) was removed by digesting the plasmid with SpeI, blunting the ends, and religating to produce the pdCas9ΔSpeI plasmid. Separately, gBlocks (IDT) encoding zraP-targeting sgRNAs based on the CRISPR1 (C1) or CRISPR3 (C3) locus in S. thermophilus LMD-9 were combined with the PCR-amplified backbone of the pgRNA-bacterial plasmid (Addgene #44251) (Id.) by Gibson assembly, thereby replacing the original S. pyogenes sgRNA sequence with the designed sgRNA sequence. The resulting sgRNA plasmid and pdCas9ΔSpeI backbone were then digested with AatII and XhoI, and the gel-extracted fragments of each sgRNA plasmid and pdCas9ΔSpeI plasmid were ligated together to form psgRNA-SthC1 and psgRNA-SthC3. To modify the sgRNA sequence, 5' phosphorylated oligonucleotides were annealed and ligated to the SpeI / KpnI or KpnI / HindIII sites of the psgRNA-C1 or psgRNA-C3 plasmid. All plasmid modifications were confirmed by sequencing.

[0142] Strains and growth conditions.E. coli K-12 subst. MG1655 (genotype: E. coli K-12 F λ- ilvG- rfb-5C rph-1) was used for all transformation analyses. Strains were grown aerobically at 37°C and 250 RPM in LB medium (10 g / L tryptone, 5 g / L yeast extract, 10 g / L sodium chloride) unless otherwise indicated. Medium was supplemented with antibiotics (34 μg / ml chloramphenicol, 50 μg / ml ampicillin) as needed.

[0143] Transformation assay. Frozen stocks of cells carrying the indicated Cas9-expressing plasmids were streaked for isolation, and individual colonies were inoculated into 3 ml of LB medium and grown overnight. The resulting cultures were back-diluted into 45 ml of LB medium and measured in a Nanodrop 2000c spectrophotometer (Thermo Scientific) to determine the ABS. 600The cells were grown until the ΔΨ of sgRNA was 0.6–0.8. The cultures were then washed twice with ice-cold 10% glycerol before being pipetted and resuspended in 200–400 μl of 10% glycerol. The suspended cells (50 μl) were transformed with 25 ng of the indicated sgRNA-expressing plasmid using a MicroPulser Electroporator (BioRad) and allowed to recover for 1 hour in 300 μl of SOC medium (Quality Biological). After recovery, 200 μl of the culture containing different amounts of LB medium was plated on LB agar containing 100 ng / ml anhydrotetracycline. Transformation efficiency was calculated by dividing the number of transformants for the tested dgRNA plasmid by the number of transformants for the psgRNA-C1-T4 control plasmid, as previously described (Gomaa, AA et al. MBio. 5, e00928-13 (2014)). To reduce experiment-to-experiment variability in transformation efficiency, the tested sgRNA plasmids and control plasmids were transformed into the same batch of electrocompetent cells. Similarly, with reference to Figures 30-32, transformation assays were used to test the ability of the plasmids to be electroporated into Lactobacillus strains harboring an active CRISPR-Cas system (Lbu - Lactobacillus buchneri, Figure 30; Lrh - Lactobacillus rhamnosus, Figure 32; Lca - Lactobacillus casei, Figure 31. Plasmids were designed to contain a protospacer sequence identical to the first spacer sequence in the host CRISPR locus. Various plasmids were designed to flank the protospacer with a complete PAM (NTAAC for Lga; NNGAA for Lca; NGAAA for Lrh; the PAM region is shown in Figures 30-32 with underlined nucleotides and their mutated variants (nucleotides tested for efficiency are italicized). Experiments also included a control non-targeting sequence (the penultimate entry for each experiment shown in Figures 30-32) and a control plasmid lacking the targeting sequence (the last entry for each experiment shown in Figures 30-32).The ability of the native CRISPR-Cas system to prevent plasmid uptake by DNA targeting is measured as the difference in transformation efficiency between the test sequence and the two aforementioned controls.

Claims

1. A single guide RNA (sgRNA) comprising: (a) a synthetic transcoding CRISPR (tracr) nucleic acid construct portion; and (b) a synthetic CRISPR nucleic acid construct portion, the synthetic tracr nucleic acid construct portion comprises, in the 5' to 3' direction, nucleotides 4 to 93 of SEQ ID NO:63; wherein the synthetic tracr nucleic acid construct portion is an anti-zipper sequence comprising at least four nucleotides, wherein said at least four nucleotides comprise nucleotides 22-25 of SEQ ID NO:63; a bulge sequence comprising nucleotides 30-32 of SEQ ID NO:63; an anti-stitch sequence comprising nucleotides 33-39 of SEQ ID NO:63; a nexus sequence comprising nucleotides 40 to 69 of SEQ ID NO: 63, said nexus forming a hairpin structure comprising a double-stranded duplex stem; and a hairpin sequence comprising a stem with at least three matching base pairs; Including, wherein the anti-zipper sequence is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; and The synthetic CRISPR nucleic acid construct portion comprises, in a 3' to 5' direction: a zipper sequence comprising at least four nucleotides, wherein said at least four nucleotides comprise the nucleotide sequence GAUG; A stitching sequence comprising the nucleotide sequence GACUCUG; and a spacer sequence having a 5' end and a 3' end, the 3' end of which contains at least 7 nucleotides having 100% complementarity with the target DNA; Including, the zipper sequence is located immediately upstream of the stitch sequence, and the stitch sequence is located immediately upstream of the spacer sequence; where: the stitch sequence of the synthetic CRISPR nucleic acid construct portion is 100% complementary to and hybridizes with the anti-stitch sequence of the synthetic tracr nucleic acid construct portion, and the zipper sequence of the synthetic CRISPR nucleic acid construct portion is at least 90% complementary to and hybridizes with the anti-zipper sequence of the synthetic tracr nucleic acid construct portion; wherein the sgRNA functions with a Cas9 polypeptide having at least 90% identity to the amino acid sequence of SEQ ID NO: 42; sgRNA.

2. The sgRNA of claim 1, The sgRNA can form a nucleic acid-protein complex with a Cas9 polypeptide having at least 90% identity to the amino acid sequence of SEQ ID NO: 42; sgRNA.

3. 3. The sgRNA of claim 1 or 2, The Cas9 polypeptide is a Cas9 nuclease from the Lactobacillus rhamnosus (Lrh) group; sgRNA.

4. The sgRNA of any one of claims 1 to 3, the Cas9 polypeptide comprises a mutation within a RuvC active site motif; sgRNA.

5. The sgRNA of any one of claims 1 to 4, the Cas9 polypeptide comprises a mutation within an HNH active site motif. sgRNA.

6. The sgRNA of any one of claims 1 to 5, The Cas9 polypeptide contains mutations within an HNH active site motif and within a RuvC active site motif and is fused to a polypeptide of interest. sgRNA.

7. Encoding the sgRNA of any one of claims 1 to 6. Expression cassette.

8. A cell (excluding human germ cells, human fertilized eggs, and human embryos) comprising the expression cassette of claim 7.

9. 1. A non-medical method for site-specific cleavage of double-stranded target DNA, comprising: The sgRNA of any one of claims 1 to 6 is contacted with target DNA in the presence of Cas9 nuclease, and the spacer sequence of the synthetic CRISPR nucleic acid construct portion hybridizes to a portion of the target DNA adjacent to a protospacer adjacent motif (PAM) on the target DNA, thereby causing site-specific cleavage of the target DNA within a region defined by complementary hybridization of the spacer sequence to the target DNA. Methods (excluding applications to human germ cells, human fertilized eggs, and human embryos).

10. 10. The method of claim 9, the site-specific cleavage is a site-specific nicking of the (+) strand of the double-stranded target DNA; the Cas9 nuclease comprises a mutation within a RuvC active site motif, thereby cleaving the (+) strand of the double-stranded target to generate a site-specific nick within the (+) strand of the double-stranded target DNA; or the site-specific cleavage is a site-specific nicking of the negative strand of the double-stranded target DNA; the Cas9 nuclease contains a point mutation within an HNH active site motif, thereby cleaving the negative strand of the double-stranded target DNA to generate a site-specific nick in the negative strand of the double-stranded target DNA; method.

11. 1. A non-medical method for site-specific targeting of a polypeptide of interest to double-stranded (ds) target DNA, comprising:

7. The method of claim 6, wherein the sgRNA of claim 6 is contacted with the target DNA to target the polypeptide of interest fused to the Cas9 to a specific site on the target DNA, the site being defined by complementary hybridization of the spacer sequence to the target DNA. Methods (excluding applications to human germ cells, human fertilized eggs, and human embryos).

12. 11. The method of claim 9 or 10, The PAM comprises a nucleotide sequence of GAAA, CCCC, CAAA, GAAC, GACC, CAAC, or GCCC; method.

13. 13. The method of claim 9, 10 or 12, The Cas9 nuclease is encoded by a nucleotide sequence that is codon-optimized for the organism containing the target DNA. method.

14. 14. The method of claim 9, 10, 12 or 13, The Cas9 nuclease comprises at least one nuclear localization sequence. method.

15. A system for modifying a target nucleic acid, comprising: an sgRNA comprising (a) a synthetic transcoding CRISPR (tracr) nucleic acid construct portion, and (b) a synthetic CRISPR nucleic acid construct portion; the synthetic tracr nucleic acid construct portion comprises, in the 5' to 3' direction, nucleotides 4 to 93 of SEQ ID NO:63; wherein the synthetic tracr nucleic acid construct portion is an anti-zipper sequence comprising at least four nucleotides, wherein said at least four nucleotides comprise nucleotides 22-25 of SEQ ID NO:63; a bulge sequence comprising nucleotides 30-32 of SEQ ID NO:63; an anti-stitch sequence comprising nucleotides 33-39 of SEQ ID NO:63; a nexus sequence comprising nucleotides 40 to 69 of SEQ ID NO: 63, said nexus forming a hairpin structure comprising a double-stranded duplex stem; and a hairpin sequence comprising a stem with at least three matching base pairs; Including, wherein the anti-zipper sequence is located immediately upstream of the bulge sequence, the bulge sequence is located immediately upstream of the anti-stitch sequence, the anti-stitch sequence is located immediately upstream of the nexus sequence, and the nexus sequence is located immediately upstream of the hairpin sequence; and The synthetic CRISPR nucleic acid construct portion comprises, in a 3' to 5' direction: a zipper sequence comprising at least four nucleotides, wherein said at least four nucleotides comprise the nucleotide sequence GAUG; A stitching sequence comprising the nucleotide sequence GACUCUG; and a spacer sequence having a 5' end and a 3' end, the 3' end of which contains at least 7 nucleotides having 100% complementarity with the target DNA; Including, the zipper sequence is located immediately upstream of the stitch sequence, and the stitch sequence is located immediately upstream of the spacer sequence; where: the stitch sequence of the synthetic CRISPR nucleic acid construct portion is 100% complementary to and hybridizes with the anti-stitch sequence of the synthetic tracr nucleic acid construct portion, the zipper sequence of the synthetic CRISPR nucleic acid construct portion is at least 90% complementary to and hybridizes with the anti-zipper sequence of the synthetic tracr nucleic acid construct portion, and the hybridized zipper sequence and anti-zipper sequence are linked by a nucleotide distal to the bulge sequence; and, (c) a Cas9 nuclease having at least 90% identity to the amino acid sequence of SEQ ID NO: 42; Including, wherein the Cas9 nuclease can bind to the sgRNA and form a complex, and the spacer sequence of the synthetic CRISPR nucleic acid construct portion hybridizes to a portion of the target nucleic acid adjacent to a protospacer adjacent motif (PAM) on the target nucleic acid, thereby directing the Cas9 to the target nucleic acid and modifying the target nucleic acid. system.

16. 16. The system of claim 15, The Cas9 nuclease having at least 90% identity to the amino acid sequence of SEQ ID NO: 42 is a Cas9 derived from the Lactobacillus rhamnosus (Lrh) group; system.

17. 17. The system of claim 15 or 16, The Cas9 nuclease comprises a mutation within a RuvC active site motif. system.

18. The system of any one of claims 15 to 17, the Cas9 nuclease comprises a mutation within an HNH active site motif; system.

19. The system of any one of claims 15 to 18, The modification of the nucleic acid is site-specific cleavage of double-stranded target DNA. system.

20. The system of any one of claims 15 to 19, The Cas9 nuclease contains mutations within an HNH active site motif and within a RuvC active site motif and is fused to a polypeptide of interest. system.

21. The system of any one of claims 15 to 20, The PAM comprises a nucleotide sequence of GAAA, CCCC, CAAA, GAAC, GACC, CAAC, or GCCC; system.