Crispr-cas component systems, methods and compositions for sequence manipulation

JP2024170441A5Active Publication Date: 2025-05-14THE BROAD INST INC +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024138377
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2013-06-17
Filing Date
2024-08-20
Publication Date
2025-05-14
Estimated Expiration
2033-12-12

AI Technical Summary

Technical Problem

There is a need for alternative and robust genome engineering techniques that facilitate targeting of multiple locations within the nuclear genome, as existing methods like designer zinc fingers and TALEs are not scalable or versatile.

Method used

The use of CRISPR/Cas systems, which utilize a short RNA molecule to program a Cas enzyme for specific DNA targeting, combined with vector systems that include regulatory elements and nuclear localization sequences to enhance targeting efficiency in eukaryotic cells.

Benefits of technology

This approach simplifies genome editing methodologies, enabling precise and efficient targeting of multiple genomic locations, accelerating the classification and mapping of genetic factors associated with biological functions and diseases.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide systems, methods, and compositions for manipulation of sequences and / or activities of target sequences.SOLUTION: Provided are vectors and vector systems, some of which encode one or more components of a CRISPR complex, as well as methods for the design and use of such vectors. Also provided are methods for directing CRISPR complex formation in eukaryotic cells and methods for selecting specific cells by introducing precise mutations utilizing the CRISPR / Cas system.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Applications and Incorporation by Reference This application claims priority to U.S. Provisional Patent Applications Nos. 61 / 736,527, 61 / 748,427, 61 / 768,959, 61 / 791,409, and 61 / 835,931, all of which are entitled SYSTEMS AND METHODS OF USE, all bearing Broad Reference Nos. BI-2011 / 008 / WSGR Docket No. 44063-701.101, BI-2011 / 008 / WSGR Docket No. 44063-701.102, BI-2011 / 008 / VP Docket No. 44790.01.2003, BI-2011 / 008 / VP Docket No. 44790.02.2003, and BI-2011 / 008 / VP Docket No. 44790.03.2003, respectively. and METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed on December 12, 2012, January 2, 2013, February 25, 2013, March 15, 2013, and June 17, 2013, respectively. Reference is made to U.S. Provisional Patent Applications Nos. 61 / 758,468; 61 / 769,046; 61 / 802,174; 61 / 806,375; 61 / 814,263; 61 / 819,803 and 61 / 828,130, each entitled ENGINEERING AND OPTIMIZATION OF SYSTEMS, METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed on January 30, 2013; February 25, 2013; March 15, 2013; March 28, 2013; April 20, 2013; May 6, 2013 and May 28, 2013, respectively. See also U.S. Provisional Patent Applications Nos. 61 / 835,936, 61 / 836,127, 61 / 836,101, 61 / 836,080, 61 / 836,123, and 61 / 835,973, each filed on June 17, 2013.See also U.S. Provisional Patent Application No. 61 / 842,322 and U.S. Patent Application No. 14 / 054,414, each bearing Broad Reference No. BI-2011 / 008A, entitled CRISPR-CAS SYSTEMS AND METHODS FOR ALTERING EXPRESSION OF GENE PRODUCTS, filed July 2, 2013 and October 15, 2013, respectively.

[0002] The above applications, and all documents cited in those applications or during their prosecution ("application cited documents"), and all documents cited or referenced in those application cited documents, as well as all documents cited or referenced herein ("herein cited documents"), and all documents cited or referenced in the herein cited documents, together with any manufacturer's instructions, manuals, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated by reference and may be used in the practice of this invention. More specifically, all references are incorporated by reference to the same extent as if each individual document were individually and specifically indicated to be incorporated by reference.

[0003] The present invention relates generally to systems, methods, and compositions used in the control of gene expression, including sequence targeting, e.g., genomic perturbation or gene editing, which may use vector systems related to Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and its components.

[0004] STATEMENT REGARDING FEDERALLY FUNDED RESEARCH This invention was made with government support under NIH Pioneer Award DP1MH100706 awarded by the National Institutes of Health. The U.S. Government has certain rights in this invention. [Background technology]

[0005] Recent advances in genome sequencing technology and analytical methods have significantly accelerated the ability to classify and map genetic factors related to a wide range of biological functions and diseases.Accurate genome targeting technology is needed to enable the systematic reverse engineering of causal gene mutations by enabling the selective perturbation of individual genetic elements, and to advance synthetic biology, biotechnology and pharmaceutical applications.Genome editing technology, such as designer zinc finger, transcription activator-like effector (TALE), or homing meganuclease, can be used to produce targeted genome perturbations, but there is still a need for new genome engineering technology that is inexpensive, easy to set up, scalable, and easy to target multiple locations in eukaryotic genomes. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] U.S. Patent No. 4,873,316 Summary of the Invention [Problem to be solved by the invention]

[0007] There is an urgent need for alternative, robust systems and technologies for versatile sequence targeting. The present invention addresses this need and provides related advantages. CRISPR / Cas or CRISPR-Cas systems (both terms are used interchangeably throughout this application) do not require the generation of customized proteins to target specific sequences, but rather allow a single Cas enzyme to be programmed by a short RNA molecule to recognize a specific DNA target; in other words, the Cas enzyme can be recruited to a specific DNA target using the short RNA molecule. The addition of CRISPR-Cas systems to the repertoire of genome sequencing technologies and analytical methods significantly simplifies methodology and accelerates the ability to classify and map genetic factors associated with a diverse range of biological functions and diseases. To effectively utilize CRISPR-Cas systems for genome editing without adverse effects, it is important to understand the engineering and optimization aspects of these genome engineering tools, which are an embodiment of the claimed invention. [Means for solving the problem]

[0008] In one aspect, the present invention provides a vector system, comprising one or more vectors.In some embodiments, the system comprises: (a) a first regulatory element, which is operably linked to a tracr mate sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr mate sequence (when expressed, the guide sequence directs the sequence-specific binding of CRISPR complex to target sequence in eukaryotic cells, and the CRISPR complex comprises (1) the guide sequence hybridized to target sequence, and (2) the tracr mate sequence hybridized to the tracr sequence); and (b) a second regulatory element, which is operably linked to the enzyme coding sequence encoding the CRISPR enzyme, which comprises a nuclear localization sequence; component (a) and (b) are on the same or different vectors of the system.In some embodiments, component (a) further comprises the tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, and each of the two or more guide sequences, when expressed, directs the sequence-specific binding of CRISPR complexes to different target sequences in eukaryotic cells.In some embodiments, the system comprises a third regulatory element, for example, a tracr sequence under the control of a polymerase III promoter.In some embodiments, the tracr sequence shows at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned.Determining optimal alignment is within the skill of those skilled in the art.For example, there are publicly available and commercially available alignment algorithms and programs, such as, but not limited to, ClustalW, Smith-Waterman in MATLAB, Bowtie, Geneious, BioPython, and SeqMan. In some embodiments, the CRISPR complex comprises one or more nuclear localization sequences of sufficient strength to drive accumulation of said CRISPR complex in a detectable amount in the nucleus of a eukaryotic cell.Without being bound by theory, it is believed that nuclear localization sequences are not necessary for CRISPR complex activity in eukaryotes, but including such sequences improves the activity of the system, particularly for targeting nucleic acid molecules in the nucleus. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae (S. pneumoniae), Streptococcus pyogenes (S. pyogenes), or S. thermophilus (S. thermophilus) Cas9, and may include mutant Cas9s derived from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs cleavage of one or two strands at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides, or 10-30, or 15-25, or 15-20 nucleotides in length. Generally, and throughout this specification, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, but no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, for example, by standard molecular cloning techniques.Another type of vector is a viral vector, in which virally derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. Some vectors can autonomously replicate in the host cell into which they are introduced (e.g., bacterial vectors with a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors) are integrated into the genome of the host cell upon introduction into the host cell, thereby replicating along with the host genome. Furthermore, some vectors can direct the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." Common expression vectors useful in recombinant DNA technology are often in the form of plasmids.

[0009] A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector can be selected based on the host cell to be used for expression and contains one or more regulatory elements operably linked to the nucleic acid sequence to be expressed. In the context of a recombinant expression vector, "operably linked" means that the nucleotide sequence of interest is linked to a regulatory element in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0010] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, e.g., polyadenylation signals and poly-U sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired target tissue, e.g., muscle, nerve cells, bone, skin, blood, specific organs (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). The regulatory element can also direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell type-specific. In some embodiments, the vector contains one or more PolIII promoters (e.g., 1, 2, 3, 4, 5, or more PolIII promoters), one or more PolII promoters (e.g., 1, 2, 3, 4, 5, or more PolII promoters), one or more PolI promoters (e.g., 1, 2, 3, 4, 5, or more PolI promoters), or a combination thereof. Examples of PolIII promoters include, but are not limited to, U6 and H1 promoters.Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) [see, e.g., Boshart et al., Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term "regulatory element" also encompasses enhancer elements, such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), p. 466-472, 1988); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), p. 1527-31, 1981). Those skilled in the art will recognize that the design of the expression vector may depend on factors such as the choice of host cell to be transformed, the desired expression level, and the like. The vector can be introduced into a host cell to produce a protein, including a transcript, a fusion protein, or a peptide, encoded by the nucleic acid described herein (e.g., a clustered regularly interspaced short repeat (CRISPR) transcript, protein, enzyme, mutant thereof, fusion protein thereof, etc.).

[0011] Advantageous vectors include lentiviruses and adeno-associated viruses, and such vector types can also be selected for targeting specific cell types.

[0012] In one aspect, the present invention provides a vector comprising a regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising one or more nuclear localization sequences. In some embodiments, the regulatory element drives the transcription of the CRISPR enzyme in eukaryotic cells so that the CRISPR enzyme accumulates in the nucleus of the eukaryotic cell in detectable amounts. In some embodiments, the regulatory element is a polymerase II promoter. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae (S. pneumoniae), Streptococcus pyogenes (S. pyogenes), or S. thermophilus (S. thermophilus) Cas9, and may include mutant Cas9s derived from these organisms. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs cleavage of one or two strands at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity.

[0013] In one aspect, the present invention provides a CRISPR enzyme comprising one or more nuclear localization sequences of sufficient strength to drive the accumulation of a detectable amount of the CRISPR enzyme in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae (S. pneumoniae), Streptococcus pyogenes (S. pyogenes), or S. thermophilus (S. thermophilus) Cas9, and may include mutant Cas9s derived from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme lacks the ability to cleave one or more strands of the target sequence to which it binds.

[0014] In one aspect, the present invention provides a eukaryotic host cell comprising: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr mate sequence (the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising a CRISPR enzyme complexed with (1) the guide sequence hybridized to the target sequence and (2) the tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme-coding sequence encoding the CRISPR enzyme comprising a nuclear localization sequence. In some embodiments, the host cell comprises components (a) and (b). In some embodiments, components (a), (b), or (a) and (b) are stably integrated into the genome of the host eukaryotic cell. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of which, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in the eukaryotic cell. In some embodiments, the eukaryotic host cell further comprises a third regulatory element operably linked to the tracr sequence, such as a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences strong enough to drive the accumulation of a detectable amount of the CRISPR enzyme in the nucleus of the eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, and may include mutant Cas9s from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs one or two strand cleavage at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides, or 10-30, 15-25, or 15-20 nucleotides in length. In one aspect, the invention provides a non-human eukaryotic organism, preferably a multicellular eukaryotic organism, comprising a eukaryotic host cell according to any of the described embodiments. In another aspect, the invention provides a eukaryotic organism, preferably a multicellular eukaryotic organism, comprising a eukaryotic host cell according to any of the described embodiments. In some embodiments of these aspects, the organism may be an animal, e.g., a mammal. The organism may also be an arthropod, e.g., an insect. The organism may also be a plant. Furthermore, the organism may be a fungus.

[0015] In one aspect, the present invention provides a kit comprising one or more of the components described herein. In some embodiments, the kit comprises a vector system and instructions for use of the kit. In some embodiments, the vector system comprises: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr mate sequence (the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with the tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme comprising a nuclear localization sequence. In some embodiments, the kit comprises components (a) and (b) on the same or different vectors of the system. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr mate sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of which, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in a eukaryotic cell. In some embodiments, the system further comprises a third regulatory element operably linked to the tracr sequence, such as a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences strong enough to drive the accumulation of a detectable amount of the CRISPR enzyme in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme.In some embodiments, the Cas9 enzyme is S. pneumoniae, S. pyogenes, or S. thermophilus Cas9, and may include mutant Cas9s from these organisms. The enzyme may be a Cas9 homolog or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs one or two strand cleavage at the location of the target sequence. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides, or 10-30, 15-25, or 15-20 nucleotides in length.

[0016] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method comprises binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide, thereby modifying the target polynucleotide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, the guide sequence in turn being bound to a tracr mate sequence hybridized to a tracr sequence. In some embodiments, the cleavage comprises cleaving one or both strands at the location of the target sequence by the CRISPR enzyme. In some embodiments, the cleavage results in reduced transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence. In some embodiments, the method further comprises delivering one or more vectors to the eukaryotic cell, wherein the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a tracr mate sequence, and a tracr sequence. In some embodiments, the vectors are delivered into the eukaryotic cell in a subject. In some embodiments, the modification is performed in the eukaryotic cell in cell culture. In some embodiments, the method further comprises isolating the eukaryotic cell from a subject before the modification. In some embodiments, the method further comprises returning the eukaryotic cell and / or cells derived therefrom to the subject.

[0017] In one aspect, the present invention provides a method for modifying the expression of polynucleotide in eukaryotic cells.In some embodiments, the method comprises: allowing CRISPR complex to bind to polynucleotide, whereby said binding causes the expression of said polynucleotide to increase or decrease; said CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence in said polynucleotide, and said guide sequence is then linked to a tracr mate sequence that hybridizes to a tracr sequence.In some embodiments, the method further comprises delivering one or more vectors into said eukaryotic cell, and said one or more vectors drive the expression of one or more of CRISPR enzyme, the guide sequence that is linked to a tracr mate sequence, and the tracr sequence.

[0018] In one aspect, the present invention provides a method for generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes (a) introducing one or more vectors into a eukaryotic cell, where the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a tracr mate sequence, and a tracr sequence; and (b) binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide within the disease gene, where the CRISPR complex comprises (1) a guide sequence hybridized to a target sequence within the target polynucleotide, and (2) a CRISPR enzyme complexed with a tracr mate sequence hybridized to a tracr sequence, thereby generating a model eukaryotic cell containing a mutated disease gene. In some embodiments, the cleavage includes cleavage of one or both strands at the location of the target sequence by the CRISPR enzyme. In some embodiments, the cleavage results in decreased transcription of the target gene. In some embodiments, the method further comprises repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides in the target polynucleotide. In some embodiments, the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.

[0019] In one aspect, the present invention provides a method for developing a bioactive agent that modulates a cell signaling event associated with a disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes: (a) contacting a test compound with a model cell of any one of the described embodiments; and (b) detecting a change in a readout that indicates a reduction or increase in a cell signaling event associated with the mutation in the disease gene, thereby developing the bioactive agent that modulates the cell signaling event associated with the disease gene.

[0020] In one aspect, the present invention provides a recombinant polynucleotide comprising a guide sequence upstream of a tracr mate sequence, wherein the guide sequence, when expressed, directs the sequence-specific binding of a CRISPR complex to the corresponding target sequence present in eukaryotic cells.In some embodiments, the target sequence is a viral sequence present in eukaryotic cells.In some embodiments, the target sequence is a proto-oncogene or an oncogene.

[0021] In one aspect, the present invention provides a method for selecting one or more prokaryotic cells by introducing one or more mutations into a gene in one or more prokaryotic cells, the method comprising: introducing one or more vectors into the prokaryotic cells (the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a tracr mate sequence, a tracr sequence, and an editing template; the editing template comprises one or more mutations that prevent CRISPR enzyme cleavage); allowing the editing template to homologously recombine with a target polynucleotide in the cell to be selected; and allowing a CRISPR complex to bind to the target polynucleotide to cause cleavage of the target polynucleotide within the gene (the CRISPR complex comprises (1) a guide sequence hybridized to a target sequence within the target polynucleotide, and (2) a tracr mate sequence hybridized to the tracr sequence, wherein binding of the CRISPR complex to the target polynucleotide induces cell death), thereby enabling selection of one or more prokaryotic cells into which one or more mutations have been introduced. In a preferred embodiment, the CRISPR enzyme is Cas9. In another embodiment of the invention, the cells to be selected can be eukaryotic cells. This embodiment of the invention allows for the selection of defined cells without requiring a selectable marker or a two-step process that may include a counterselection system.

[0022] Accordingly, it is not the object of the present invention to encompass any previously known products, methods of making the products, or methods of using the products, and therefore, applicants reserve the right to, and hereby disclose, a disclaimer of any previously known products, methods of making the products, or methods of using the products. It is further noted that the present invention does not encompass within its scope any products, methods, or methods of making or using the products that do not meet the description and enablement requirements of the USPTO (35 U.S.C. Section 112, first paragraph) or the EPO (European Patent Convention, Article 83), and therefore, applicants reserve the right to, and hereby disclose, a disclaimer of any previously described products, methods of making the products, or methods of using the products.

[0023] In this disclosure, and particularly in the claims and / or paragraphs, terms such as "comprises," "comprised," "comprising," and the like, may have the meaning ascribed to them in U.S. Patent Law; for example, they may mean "includes," "included," "including," and the like; and terms such as "consisting essentially of" and "consists essentially of" have the meaning ascribed to them in U.S. Patent Law, for example, it is noted that they allow for elements not expressly recited, but exclude elements found in the prior art or that affect a basic or novel characteristic of the invention. These and other embodiments are disclosed or obvious from the following detailed description and are encompassed thereby.

[0024] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings. [Brief explanation of the drawings]

[0025] [Figure 1] A schematic model of the CRISPR system is shown. The Cas9 nuclease from Streptococcus pyogenes (yellow) is targeted to genomic DNA by a synthetic guide RNA (sgRNA) consisting of a 20-nt guide sequence (blue) and a scaffold (red). The guide sequence base pairs with the DNA target (blue) immediately upstream of a required 5'-NGG protospacer adjacent motif (PAM; magenta), and Cas9 mediates a double-strand break (DSB) (red triangle) approximately 3 bp upstream of the PAM. [Figure 2A] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2B] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2C] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2D] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2E] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 2F] Exemplary CRISPR systems, possible mechanisms of action, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing nuclear localization and CRISPR activity are presented. [Figure 3] 1 shows exemplary expression cassettes for expression of CRISPR system elements in eukaryotic cells, the predicted structure of exemplary guide sequences, and CRISPR system activity measured in eukaryotic and prokaryotic cells. [Figure 4A] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 4B] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 4C] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 4D] 1 shows the results of an assessment of SpCas9 specificity for exemplary targets. [Figure 5A] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5B] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5C] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5D] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5E] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5F] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 5G] An exemplary vector system and results for its use in directing homologous recombination in eukaryotic cells are presented. [Figure 6] A table of protospacer sequences is provided, and modification efficiency results for protospacer targets and corresponding PAMs designed based on exemplary S. pyogenes and S. thermophilus CRISPR systems for loci in the human and mouse genomes are summarized. Cells were transfected with Cas9 and either pre-crRNA / tracrRNA or chimeric RNA and analyzed 72 hours post-transfection. Percent indels were calculated based on Surveyor assay results from the indicated cell lines (N=3 for all protospacer targets; errors are standard error of the mean (SEM) where ND indicates not detectable using the Surveyor assay and NT indicates not tested in this study). [Figure 7A] 1 shows a comparison of different tracrRNA transcripts for Cas9-mediated gene targeting. [Figure 7B] 1 shows a comparison of different tracrRNA transcripts for Cas9-mediated gene targeting. [Figure 7C] 1 shows a comparison of different tracrRNA transcripts for Cas9-mediated gene targeting. [Figure 8] Figure 1 shows a schematic diagram of the surveyor nuclease assay for the detection of double-strand break-induced microinsertions and deletions. [Figure 9]1 shows an exemplary bicistronic expression vector for expression of CRISPR system elements in eukaryotic cells. [Figure 10] 1 shows the bacterial plasmid transformation interference assay, the expression cassettes and plasmids used therein, and the cell transformation efficiency used therein. [Figure 11] A histogram of the distance between the adjacent S. pyogenes SF370 locus 1 PAM (NGG) (Figure 10A) and S. thermophilus LMD9 locus 2 PAM (NNAGAAW) (Figure 10B) in the human genome; as well as the distance for each PAM in chromosomes (Chr) (Figure 10C). [Figure 12] 1 shows an exemplary CRISPR system, exemplary adaptations for expression in eukaryotic cells, and results of studies assessing CRISPR activity. [Figure 13] 1 shows an exemplary operation of the CRISPR system for targeting genomic loci in mammalian cells. [Figure 14] 1 shows the results of Northern blot analysis of crRNA processing in mammalian cells. [Figure 15] Exemplary selection of protospacers in the human PVALB and mouse Th loci is shown. [Figure 16] 1 shows an exemplary protospacer and corresponding PAM sequence target of the S. thermophilus CRISPR system in the human EMX1 locus. [Figure 17] A table of sequences is provided for the primers and probes used in Surveyor, RFLP, genomic sequencing, and Northern blot assays. [Figure 18A] 1 shows an exemplary engineering of a CRISPR system with chimeric RNA and the results of a SURVEYOR assay for system activity in eukaryotic cells. [Figure 18B] 1 shows an exemplary engineering of a CRISPR system with chimeric RNA and the results of a SURVEYOR assay for system activity in eukaryotic cells. [Figure 18C]1 shows an exemplary engineering of a CRISPR system with chimeric RNA and the results of a SURVEYOR assay for system activity in eukaryotic cells. [Figure 19] 1 shows a graphical representation of the results of a SURVEYOR assay for CRISPR system activity in eukaryotic cells. [Figure 20] FIG. 1 shows an exemplary visualization of several Streptococcus pyogenes (S. pyogenes) Cas9 target sites in the human genome using the UCSC genome browser. [Figure 21] 1 shows the predicted secondary structure for an exemplary chimeric RNA containing a guide sequence, a tracr mate sequence, and a tracr sequence. [Figure 22] 1 shows an exemplary bicistronic expression vector for expression of CRISPR system elements in eukaryotic cells. [Figure 23] This study demonstrates that Cas9 nuclease activity against endogenous targets can be exploited for genome editing. (a) Concept of genome editing using the CRISPR system. CRISPR targeting constructs direct cleavage at chromosomal loci and are cotransformed with an editing template that recombines with the target to prevent cleavage. Kanamycin-resistant transformants that survived CRISPR attack contained modifications induced by the editing template: tracr, a transactivating CRISPR RNA; and aphA-3, a kanamycin resistance gene. (b) Transformation of crR6M DNA into R68232.5 cells with no editing template, the R6 wild-type srtA, or the R6370.1 editing template. Recombination of either R6srtA or R6370.1 prevented cleavage by Cas9. Transformation efficiency was calculated as colony-forming units (cfu) per μg of crR6M DNA; mean values ​​with standard deviations from at least three independent experiments are shown. PCR analysis was performed on eight clones for each transformation. "Un." indicates the unedited srtA locus of strain R68232.5; "Ed." indicates the edited template. The R68232.5 and R6370.1 targets are distinguished by restriction with EaeI. [Figure 24]Analysis of PAM and seed sequences that eliminate Cas9 cleavage is shown. (a) PCR products with randomized PAM or randomized seed sequences were transformed into crR6 cells. These cells expressed Cas9 loaded with crRNA targeting a chromosomal region of R68232.5 cells (highlighted in pink) that is absent from the R6 genome. More than 2 × 10 chloramphenicol-resistant transformants carrying inactive PAM or seed sequences were combined for amplification and deep sequencing of the target region. (b) Relative ratio of reads after transformation of random PAM constructs in crR6 cells (compared to reads in R6 transformants). Relative abundance for each 3-nucleotide PAM sequence is shown. Severely underrepresented sequences (NGG) are shown in red; partially underrepresented ones are shown in orange (NAG). (c) Relative ratio of reads after transformation of random seed sequence constructs in crR6 cells (compared to reads in R6 transformants). The relative abundance of each nucleotide for each position of the first 20 nucleotides of the protospacer sequence is shown. High abundance indicates the lack of cleavage by Cas9, i.e., CRISPR-inactive mutation. The gray line indicates the level of the WT sequence. The dotted line represents the level at which mutation significantly disrupts cleavage (see the section "Deep sequencing data analysis" in Example 5). [Figure 25]Introduction of single and multiple mutations using the CRISPR system in S. pneumoniae is shown. (a) Nucleotide and amino acid sequences of wild-type and edited (green nucleotides; underlined amino acid residues) bgaA. The protospacer, PAM, and restriction sites are indicated. (b) Transformation efficiency of cells transformed with the targeting construct in the presence of the editing template or control. (c) PCR analysis of eight transformants from each editing experiment followed by digestion with BtgZI (R → A) and TseI (NE → AA). The deletion of bgaA was revealed as a smaller PCR product. (d) Miller assay to measure β-galactosidase activity in the WT and edited strains. (e) For single-step double deletion, the targeting construct contained two spacers (in this case, matching srtA and bgaA) and was cotransformed with two different editing templates. (f) PCR analysis of eight transformants to detect deletions of the srtA and bgaA loci. Six of eight transformants deleted both genes. [Figure 26] We present a mechanism underlying editing using the CRISPR system. (a) A stop codon was introduced into the erythromycin resistance gene ermAM to generate strain JEN53. The stop codon was targeted with the CRISPR::ermAM(stop) construct, and the wild-type sequence can be restored by using the ermAM wild-type sequence as the editing template. (b) Mutant and wild-type ermAM sequences. (c) Percentage of erythromycin-resistant (ermR) cfu calculated from total or kanamycin-resistant (kanR) cfu. (d) Percentage of total cells acquiring both the CRISPR construct and the editing template. Co-transformation of the CRISPR targeting construct produced more transformants (t-test, p=0.011). In all cases, values ​​represent the mean ± standard deviation for three independent experiments. [Figure 27]Genome editing using the CRISPR system in Escherichia coli (E. coli) is illustrated. (a) A kanamycin resistance plasmid (pCRISPR) carrying a gene-targeting and editing CRISPR array was cotransformed with the mutation-specifying oligonucleotides into the HME63 recombineering strain containing a chloramphenicol resistance plasmid (pCas9) carrying cas9 and tracr. (b) A K42T mutation conferring streptomycin resistance was introduced into the rpsL gene. (c) The percentage of streptomycin-resistant (strepR) cfu calculated from total or kanamycin-resistant (kanR) cfu. (d) The percentage of total cells acquiring both the pCRISPR plasmid and the editing oligonucleotide. Cotransformation of the pCRISPR targeting plasmid produced more transformants (t-test, p=0.004). In all cases, values ​​represent the mean ± standard deviation of three independent experiments. [Figure 28]We demonstrate that transformation of crR6 genomic DNA results in editing of the targeted locus. (a) The IS1167 element of S. pneumoniae R6 was replaced with the CRISPR01 locus of S. pyogenes SF370 to generate the crR6 strain. This locus encodes the Cas9 nuclease, a CRISPR array with six spacers, the tracrRNA required for crRNA biogenesis, and the proteins Cas1, Cas2, and Csn2, which are not required for targeting. Strain crR6M contains a minimal functional CRISPR system lacking cas1, cas2, and csn2. The aphA-3 gene encodes kanamycin resistance. Protospacers from streptococcal bacteriophages φ8232.5 and φ370.1 were fused to the chloramphenicol resistance gene (cat) and integrated into the srtA gene of strain R6 to generate strains R68232.5 and R6370.1. (b) Left panel: Transformation of crR6 and crR6M genomic DNA into R68232.5 and R6370.1. As a control for cell competence, a streptomycin resistance gene was also transformed. Right panel: PCR analysis of eight R68232.5 transformants with crR6 genomic DNA. Primers amplifying the srtA locus were used for PCR. Seven of the eight genotyped colonies had replaced the R68232.5 srtA locus with the WT locus from crR6 genomic DNA. [Figure 29]Chromatograms of the DNA sequences of edited cells obtained in this study are provided. In all cases, the wild-type and mutant protospacer and PAM sequences (or their reverse complements) are shown. Where relevant, the amino acid sequence encoded by the protospacer is provided. For each editing experiment, all strains in which PCR and restriction analysis confirmed the introduction of the desired modification were sequenced. Representative chromatograms are shown. (a) Chromatogram for introduction of a PAM mutation into the R68232.5 target (Figure 23d). (b) Chromatogram for introduction of R>A and NE>AA mutations into β-galactosidase (bgaA) (Figure 25c). (c) Chromatogram for introduction of a 6664 bp deletion within the bgaA ORF (Figures 25c and 25f). The dotted lines indicate the limits of the deletion. (d) Chromatogram for introduction of a 729 bp deletion within the srtA ORF (Figure 25f). The dotted lines indicate the limits of the deletion. (e) Chromatogram for the generation of a premature stop codon in ermAM (Figure 33). (f) rpsL editing in E. coli (Figure 27). [Figure 30] CRISPR immunity against random S. pneumoniae targets containing different PAMs is illustrated. (a) Location of 10 random targets on the S. pneumoniae R6 genome. Selected targets have different PAMs and are present on both strands. (b) Spacers corresponding to the targets were cloned in a minimal CRISPR array on the plasmid pLZ12 and transformed into strain crR6Rc, which supplies the processing and targeting machinery in trans. (c) Transformation efficiency of different plasmids in strains R6 and crR6Rc. No colonies were recovered for transformation of pDB99-108 (T1-T10) in crR6Rc. The dotted line represents the detection limit of the assay. [Figure 31]A general scheme for targeted genome editing is provided. To facilitate targeted genome editing, crR6M was further engineered to contain only one repeat of tracrRNA, Cas9, and the CRISPR array, followed by a kanamycin resistance marker (aphA-3), generating strain crR6Rk. DNA from this strain was used as a template for PCR with primers designed to introduce a new spacer (green box indicated by N). The left- and right-hand PCRs were assembled using the Gibson method to create the targeting construct. Both the targeting and editing constructs were then transformed into strain crR6Rc, which is equivalent to crR6Rk, except that the kanamycin resistance marker has been replaced with a chloramphenicol resistance marker (cat). Approximately 90% of the kanamycin-resistant transformants contained the desired mutation. [Figure 32] The distribution of distances between PAMs is illustrated. NGG and CCN are considered valid PAMs. Data are shown for the S. pneumoniae R6 genome and a random sequence with identical length and identical GC content (39.7%). The dotted line represents the average distance between PAMs in the R6 genome (12). [Figure 33]We demonstrate CRISPR-mediated editing of the ermAM locus using genomic DNA as a targeting construct. Because genomic DNA is used as the targeting construct, it is necessary to circumvent CRISPR autoimmunity; therefore, a spacer for a sequence not present in the chromosome must be used (in this case, the ermAM erythromycin resistance gene). (a) Nucleotide and amino acid sequences of the wild-type and mutant (red text) ermAM genes. The protospacer and PAM sequences are shown. (b) Schematic diagram of CRISPR-mediated editing of the ermAM locus using genomic DNA. A construct carrying the ermAM targeting spacer (blue box) was generated by PCR and Gibson assembly and transformed into strain crR6Rc, generating strain JEN37. Genomic DNA from JEN37 was then used as a targeting construct and cotransformed with the editing template into JEN38, a strain in which the srtA gene has been replaced with a wild-type copy of ermAM. The kanamycin-resistant transformant contains the edited genotype (JEN43). (c) The number of kanamycin-resistant cells obtained after cotransformation of targeting and editing or control templates. 5.4 × 103 cfu / ml was obtained in the presence of the control template, and 4.3 × 105 cfu / ml was obtained with the edited template. This difference indicates an editing efficiency of approximately 99% [(4.3 × 105 - 5.4 × 103) / 4.3 × 105]. (d) To confirm the presence of edited cells, seven kanamycin-resistant clones and JEN38 were streaked on agar plates with (erm+) or without (erm-) erythromycin. Only the positive control showed resistance to erythromycin. The ermAMmut genotype of one of these transformants was also confirmed by DNA sequencing (Figure 29e). [Figure 34]Sequential introduction of mutations by CRISPR-mediated genome editing is illustrated. (a) Schematic diagram of sequential introduction of mutations by CRISPR-mediated genome editing. First, R6 is engineered to generate crR6Rk. crR6Rk is cotransformed with an editing construct for the ΔsrtA in-frame deletion, along with an srtA targeting construct fused to cat for chloramphenicol selection of edited cells. Strain crR6ΔsrtA is generated by chloramphenicol-based selection. Subsequently, the ΔsrtA strain is cotransformed with an editing construct containing the ΔbgaA in-frame deletion, and a bgaA targeting construct fused to aphA-3 for kanamycin selection of edited cells. Finally, the engineered CRISPR locus can be erased from the chromosome by first cotransformation of R6 DNA containing the wild-type IS1167 locus and a plasmid (pDB97) carrying the bgaA protospacer, followed by spectinomycin-based selection. (b) PCR analysis of eight chloramphenicol (Cam)-resistant transformants to detect deletion of the srtA locus. (c) β-galactosidase activity measured by Miller assay. In S. pneumoniae, this enzyme is anchored to the cell wall by sortase A. Deletion of the srtA gene results in the release of β-galactosidase into the supernatant. The ΔbgaA mutant shows no activity. (d) PCR analysis of eight spectinomycin (Spec)-resistant transformants to detect replacement of the CRISPR locus with wild-type IS1167. [Figure 35]The background mutation frequency of CRISPR in S. pneumoniae is illustrated. (a) Transformation of CRISPR::φ or CRISPR::erm(stop) targeting constructs into JEN53 with or without the ermAM editing template. The difference in kanRCFU between CRISPR::φ and CRISPR::erm(stop) indicates that Cas9 cleavage kills non-edited cells. Mutants that escape CRISPR interference in the absence of editing template are observed at a frequency of 3 × 10-3. (b) PCR analysis of the CRISPR locus of escapers shows that 7 out of 8 have spacer deletions. (c) Escaper #2 carries a point mutation in cas9. [Figure 36] We describe the reconstitution of essential elements of the S. pyogenes CRISPR locus 1 in Escherichia coli (E. coli) using pCas9. The plasmid contained the tracrRNA, Cas9, and a leader sequence driving the crRNA array. The pCRISPR plasmid contained only the leader and array. A spacer can be inserted into the crRNA array between the BsaI sites using annealed oligonucleotides. The oligonucleotide design is shown below. pCas9 carries chloramphenicol resistance (CmR) and is based on the low-copy pACYC184 plasmid backbone. pCRISPR is based on the high-copy pZE21 plasmid. Two plasmids were required because pCRISPR plasmids containing a spacer targeting the E. coli chromosome cannot be constructed using this organism as a cloning host if Cas9 is also present (it would kill the host). [Figure 37]CRISPR-directed editing in E. coli MG1655 is illustrated. An oligonucleotide (W542) carrying a point mutation conferring streptomycin resistance and abrogating CRISPR immunity was cotransformed with a plasmid targeting rpsL (pCRISPR::rpsL) or a control plasmid (pCRISPR::φ) into wild-type E. coli strain MG1655 containing pCas9. Transformants were selected on media containing either streptomycin or kanamycin. The dotted line indicates the detection limit of the transformation assay. [Figure 38] Background mutation frequency of CRISPR in E. coli HME63 is illustrated. (a) Transformation of pCRISPR::φ or pCRISPR::rpsL plasmids into HME63 competent cells. Mutants that escape CRISPR interference were observed at a frequency of 2.6 × 10-4. (b) Amplification of the CRISPR array of escapers showed that 8 of 8 had deleted the spacer. [Figure 39A] A circular representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 39B] A circular representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 39C] A circular representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 39D] A circular representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 40A]A linear representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 40B] A linear representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 40C] A linear representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 40D] A linear representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 40E] A linear representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 40F] A linear representation of a phylogenetic analysis revealing five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) is shown. [Figure 41A] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41B] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41C] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41D] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41E] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41F] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41G] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41H] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41I] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41J] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41K] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41L] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 41M] The sequence where the mutation site is located within the SpCas9 gene is shown. [Figure 42] A schematic construct is shown in which a transcription activation domain (VP64) is fused to Cas9 with two mutations in the catalytic domain (D10 and H840). [Figure 43A] Genome editing via homologous recombination. (a) Schematic of SpCas9 nickase with a D10A mutation in the RuvC I catalytic domain. (b) Schematic depicting homologous recombination (HR) at the human EMX1 locus using either a sense or antisense single-stranded oligonucleotide as the repair template. The upper red arrow indicates the sgRNA cleavage site; PCR primers for genotyping (Tables J and K) are shown as arrows in the right panel. (c) Sequence of the region modified by HR. d, SURVEYOR assay (n=3) for wild-type (wt) and nickase (D10A) SpCas9-mediated indels at the EMX1 target locus. Arrows indicate the positions of predicted fragment sizes. [Figure 43B]Genome editing via homologous recombination. (a) Schematic of SpCas9 nickase with a D10A mutation in the RuvC I catalytic domain. (b) Schematic depicting homologous recombination (HR) at the human EMX1 locus using either a sense or antisense single-stranded oligonucleotide as the repair template. The upper red arrow indicates the sgRNA cleavage site; PCR primers for genotyping (Tables J and K) are shown as arrows in the right panel. (c) Sequence of the region modified by HR. d, SURVEYOR assay (n=3) for wild-type (wt) and nickase (D10A) SpCas9-mediated indels at the EMX1 target locus. Arrows indicate the positions of predicted fragment sizes. [Figure 43C] Genome editing via homologous recombination. (a) Schematic of SpCas9 nickase with a D10A mutation in the RuvC I catalytic domain. (b) Schematic depicting homologous recombination (HR) at the human EMX1 locus using either a sense or antisense single-stranded oligonucleotide as the repair template. The upper red arrow indicates the sgRNA cleavage site; PCR primers for genotyping (Tables J and K) are shown as arrows in the right panel. (c) Sequence of the region modified by HR. d, SURVEYOR assay (n=3) for wild-type (wt) and nickase (D10A) SpCas9-mediated indels at the EMX1 target locus. Arrows indicate the positions of predicted fragment sizes. [Figure 43D] Genome editing via homologous recombination. (a) Schematic of SpCas9 nickase with a D10A mutation in the RuvC I catalytic domain. (b) Schematic depicting homologous recombination (HR) at the human EMX1 locus using either a sense or antisense single-stranded oligonucleotide as the repair template. The upper red arrow indicates the sgRNA cleavage site; PCR primers for genotyping (Tables J and K) are shown as arrows in the right panel. (c) Sequence of the region modified by HR. d, SURVEYOR assay (n=3) for wild-type (wt) and nickase (D10A) SpCas9-mediated indels at the EMX1 target locus. Arrows indicate the positions of predicted fragment sizes. [Figure 44A] 1 shows a single vector design for SpCas9. [Figure 44B] 1 shows a single vector design for SpCas9. [Figure 45] Quantitation of NLS-Csn1 constructs NLS-Csn1, Csn1, Csn1-NLS, NLS-Csn1-NLS, NLS-Csn1-GFP-NLS and UnTFN cleavage is shown. [Figure 46] The exponential frequencies of NLS-Cas9, Cas9, Cas9-NLS, and NLS-Cas9-NLS are shown. [Figure 47] 1 shows a gel demonstrating that SpCas9 with nickase mutations do not (individually) induce double-strand breaks. [Figure 48] 1 shows the design of the oligo DNA used as the homologous recombination (HR) template in this experiment, and a comparison of the HR efficiency induced by different combinations of Cas9 protein and HR template. [Figure 49A] Conditional Cas9 and Rosa26 targeting vector maps are shown. [Figure 49B] Constitutive Cas9, Rosa26 targeting vector map is shown. [Figure 50A] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50B] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50C] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50D] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50E] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50F] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50G]The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 50H] The sequences of the elements present in the vector maps of Figures 49A-B are shown. [Figure 51] A schematic diagram of key elements in constitutive and conditional Cas9 constructs is shown. [Figure 52] Functional validation of the expression of constitutive and conditional Cas9 constructs is shown. [Figure 53] Validation of Cas9 nuclease activity by Surveyor. [Figure 54] Quantification of Cas9 nuclease activity is shown. [Figure 55] Construct design and homologous recombination (HR) strategy are shown. [Figure 56] Genomic PCR genotyping results for constitutive (right) and conditional (left) constructs at two different gel exposure times (3 min for the top row and 1 min for the bottom row) are shown. [Figure 57] 1 shows Cas9 activation in mESCs. [Figure 58] 1 shows a schematic diagram of the strategy used to mediate gene knockout via NHEJ using the nickase version of Cas9 together with two guide RNAs. [Figure 59] This demonstrates how DNA double-strand break (DSB) repair can facilitate gene editing. In the error-prone non-homologous end joining (NHEJ) pathway, the ends of DSBs are processed and rejoined together by endogenous DNA repair mechanisms, which can result in random insertion / deletion (indel) mutations at the junction site. Indel mutations occurring within the coding region of a gene can result in frameshifts and premature stop codons, resulting in gene knockout. Alternatively, the homology-directed repair (HDR) pathway can be utilized, which provides a repair template in the form of a plasmid or single-stranded oligodeoxynucleotide (ssODN) to enable high-fidelity and precise editing. [Figure 60]The experimental timeline and overview are shown. Steps include reagent design, construction, validation, and cell line propagation. Custom sgRNAs (light blue bars) for each target and genotyping primers are designed in silico via our online design tool (available at genome-engineering.org / tools). The sgRNA expression vectors are then cloned into a Cas9-containing plasmid (PX330) and verified via DNA sequencing. The completed plasmid (pCRISPR) and an optional repair template to promote homologous recombination repair are then transfected into cells and assayed for their ability to mediate targeted cleavage. Finally, the transfected cells can be clonally propagated to obtain isogenic cell lines bearing the specified mutations. [Figure 61]Target selection and reagent preparation are shown. (a) For S. pyogenes Cas9, the 20-bp target (highlighted in blue) must be followed by a 5'-NGG, which can occur on either strand of the genomic DNA. Applicants recommend using the online tool described in this protocol to assist with target selection (www.genome-engineering.org / tools). (b) Schematic diagram for co-transfection of the Cas9 expression plasmid (PX165) and a PCR-amplified U6-driven sgRNA expression cassette. Using a U6 promoter-containing PCR template and an anchored forward primer (U6Fwd), sgRNA-encoding DNA can be appended onto the U6 reverse primer (U6Rev) and synthesized as an extended DNA oligo (Ultramer oligo from IDT). Note that the guide sequence (blue N) in U6Rev is the reverse complement of the 5'-NGG-flanking target sequence. (c) Schematic diagram of scarless cloning of guide sequence oligos into a plasmid (PX330) containing Cas9 and the sgRNA scaffold. The guide oligo (blue N) contains overhangs for ligation into a pair of BbsI sites on PX330, with the top and bottom strand orientations matching those of the genomic target (i.e., the top oligo is the 20-bp sequence preceding the 5'-NGG in the genomic DNA). Digestion of PX330 with BbsI allows for direct insertion of the annealed oligo to replace the type IIs restriction site (blue box). Note the additional G placed before the first base of the guide sequence. Applicants found that the additional G before the guide sequence does not adversely affect targeting efficiency. If the optimal 20-nt guide sequence does not start with a guanine, the additional guanine ensures that the sgRNA is efficiently transcribed by the U6 promoter, which prefers a guanine in the first base of the transcript. [Figure 62]Predicted results for multiplex NHEJ are shown. (a) Schematic of the SURVEYOR assay used to determine the percentage of indels. First, genomic DNA from a heterogeneous population of Cas9-targeted cells is amplified by PCR. The amplicons are then slowly reannealed to generate heteroduplexes. The reannealed heteroduplexes are cleaved by SURVEYOR nuclease, while the homoduplexes remain intact. Cas9-mediated cleavage efficiency (% indels) is calculated based on the percentage of cleaved DNA determined by the integrated intensity of the gel bands. (b) Two sgRNAs (orange and blue bars) are designed to target the human GRIN2B and DYRK1A loci. The SURVEYOR gel shows modifications at both loci in transfected cells. Colored arrows indicate the predicted fragment sizes for each locus. (c) A pair of sgRNAs (light and green bars) is designed to excise an exon (dark blue) in the human EMX1 locus. The target sequence and PAM (red) are indicated by their respective colors, and the cleavage site is indicated by a red triangle. The predicted junction is shown below. Individual clones isolated from cell populations transfected with sgRNAs 3, 4, or both were assayed by PCR (OUT Fwd, OUT Rev) to reflect approximately 270 bp deletions. Representative clones with no (12 / 23), monoallelic (10 / 23), and biallelic (1 / 23) modifications are shown. IN Fwd and IN Rev primers were used to screen for inversion events (Figure 6d). (d) Quantification of clonal lines deleting EMX1 exons. Two pairs of sgRNAs (3.1, 3.2, left-flanking sgRNAs; 4.1, 4.2, right-flanking sgRNAs) were used to mediate deletions of variable sizes around one EMX1 exon. Transfected cells were clonally isolated and expanded for genotyping analysis for deletion and inversion events. Of the 105 clones screened, 51 (49%) and 11 (10%) carrying heterozygous and homozygous deletions, respectively. As the junctions can be variable, estimated deletion sizes are listed. [Figure 63]We demonstrate the application of ssODN and targeting vectors to mediate HR using both wild-type and nickase mutant forms of Cas9 in HEK293FT and HUES9 cells, with efficiencies ranging from 1.0 to 27%. [Figure 64] This figure shows a schematic diagram of a PCR-based method for rapid and efficient CRISPR targeting in mammalian cells. A plasmid containing the human RNA polymerase III promoter U6 is PCR-amplified using a U6-specific forward primer and a reverse primer carrying the reverse complement of a portion of the U6 promoter, an sgRNA(+85) scaffold with a guide sequence, and seven T nucleotides for transcription termination. The resulting PCR product is purified and co-delivered with a plasmid carrying Cas9 driven by the CBh promoter. [Figure 65] SURVEYOR Mutation Detection Kit results from Transgenomics are shown for each gRNA and each control. A positive SURVEYOR result is one large band corresponding to genomic PCR and two smaller bands that are products of SURVEYOR nuclease creating a double-stranded break at the mutation site. Each gRNA was validated in the mouse cell line Neuro-N2a by liposome transient co-transfection with hSpCas9. 72 hours after transfection, genomic DNA was purified using QuickExtract DNA from Epicentre. PCR was performed to amplify the locus of interest. [Figure 66] Surveyor results are shown for 38 live pups (lanes 1-38), one dead pup (lane 39), and one wild-type control pup (lane 40). Pups 1-19 were injected with gRNA Chd8.2, and pups 20-38 were injected with gRNA Chd8.3. Of the 38 live pups, 13 tested positive for mutations. The one dead pup also had a mutation. No mutations were detected in the wild-type sample. Genomic PCR sequencing was consistent with the SURVEYOR assay findings. [Figure 67]The designs of different Cas9NLS constructs are shown. All Cas9s were human codon-optimized versions of SpCas9. The NLS sequence was attached to the cas9 gene at either the N- or C-terminus. All Cas9 variants with different NLS designs were cloned into a backbone vector containing them driven by the EF1a promoter. On the same vector, there was a chimeric RNA targeting the human EMX1 locus driven by the U6 promoter, forming a two-component system. [Figure 68] Figure 1 shows the efficiency of genome cleavage induced by Cas9 variants carrying different NLS designs. The percentage indicates the fraction of human EMX1 genomic DNA cleaved by each construct. All experiments were from three biological replicates, n = 3, and the error indicates the standard error of the mean (SEM). [Figure 69A] This figure shows the design of a CRISPR-TF (transcription factor) with transcriptional activation activity. The chimeric RNA is expressed by the U6 promoter, while the human codon-optimized double mutant version of the Cas9 protein (hSpCas9m), operably linked to three NLSs and a VP64 functional domain, is expressed by the EF1a promoter. The double mutations D10A and H840A render the Cas9 protein unable to introduce any cleavage but maintain its ability to bind to target DNA when guided by the chimeric RNA. [Figure 69B]Figure 1 shows transcriptional activation of the human SOX2 gene by the CRISPR-TF system (chimeric RNA and Cas9-NLS-VP64 fusion protein). 293FT cells were transfected with plasmids carrying two components: (1) different U6-driven chimeric RNAs targeting 20-bp sequences within or surrounding the human SOX2 genomic locus, and (2) an hSpCas9m (double mutant)-NLS-VP64 fusion protein driven by EF1a. Ninety-six hours after transfection, 293FT cells were harvested, and the level of activation was measured by induction of mRNA expression using a qRT-PCR assay. All expression levels are normalized to the control group (gray bar), which represents results from cells transfected with the CRISPR-TF backbone plasmid without the chimeric RNA. The qRT-PCR probe used to detect SOX2 mRNA was the Taqman Human Gene Expression Assay (Life Technologies). All experiments represent data from three biological replicates, n=3, and error bars indicate the standard error (sem). [Figure 70] NLS architecture optimization for SpCas9 is shown. [Figure 71] A QQ plot for the NGGNN sequence is shown. [Figure 72] A histogram of data density is shown along with the fitted normal distribution (black line) and the .99 quantile (dotted line). [Figure 73] RNA-guided repression of bgaA expression by dgRNA::cas9** is shown. a. Cas9 protein binds to tracrRNA and precursor CRISPR RNA, which is processed by RNase III to form crRNA. The crRNA directs Cas9 binding to the bgaA promoter and represses transcription. b. Targets used to direct Cas9** to the bgaA promoter are shown. The putative -35, -10, and bgaA start codons are shown in bold. c. Beta-galactosidase activity measured by Miller assay in the absence of targeting and for four different targets. [Figure 74] Figure 1 shows the characterization of Cas9**-mediated repression. a. The gfpmut2 gene and its promoter are shown, along with the locations of the different target sites used in this study, including the -35 and -10 signals. b. Relative fluorescence upon targeting of the coding strand. c. Relative fluorescence upon targeting of the non-coding strand. d. Northern blot using probes B477 and B478 on RNA extracted from T5, T10, B10 or a non-targeted control strain. e. Effect of increasing numbers of mutations at the 5' end of the crRNA of B1, T5 and B10. DETAILED DESCRIPTION OF THE INVENTION

[0026] The drawings herein are for illustrative purposes only and are not necessarily drawn to scale.

[0027] The terms "polynucleotide," "nucleotide," "nucleotide sequence," "nucleic acid," and "oligonucleotide" are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide can contain one or more modified nucleotides, such as methylated nucleotides or nucleotide analogs. Modifications to the nucleotide structure, if present, can be imparted before or after assembly of the polymer. The sequence of nucleotides can be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.

[0028] In embodiments of the present invention, the terms "chimeric RNA," "chimeric guide RNA," "guide RNA," "single guide RNA," and "synthetic guide RNA" are used interchangeably to refer to a polynucleotide sequence comprising a guide sequence, a tracr sequence, and a tracr mate sequence. The term "guide sequence" refers to an approximately 20-bp sequence within the guide RNA that defines the target site, and can be used interchangeably with the terms "guide" or "spacer." The term "tracr mate sequence" can also be used interchangeably with the term "direct repeat."

[0029] The term "wild-type" as used herein is a term of the art understood by those skilled in the art and means the typical form of an organism, strain, gene or characteristic as it occurs in nature, as distinguished from mutant or variant forms.

[0030] The term "variant" as used herein should be taken to mean a display of qualities having a pattern that deviates from that occurring in nature.

[0031] The terms "non-naturally occurring" and "engineered" are used interchangeably and refer to artificial involvement. When referring to a nucleic acid molecule or polypeptide, the term means that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated or found in the natural state.

[0032] "Complementarity" refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either through classical Watson-Crick base pairing or other non-classical types. The percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" means that all consecutive residues of a nucleic acid sequence will hydrogen bond with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions.

[0033] As used herein, "stringent conditions" for hybridization refer to conditions under which a nucleic acid having complementarity to a target sequence will predominantly hybridize to the target sequence and will not substantially hybridize to non-target sequences. Stringent conditions are generally sequence-dependent and vary depending on numerous factors. Generally, the longer the sequence, the higher the temperature at which the sequence will specifically hybridize to its target sequence. Non-limiting examples of stringent conditions are detailed in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology—Hybridization With Nucleic Acid Probes Part I, Second Chapter "Overview of principles of hybridization and the strategy of nucleic acid probe assay," Elsevier, NY.

[0034] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex stabilized through hydrogen bonding between the bases of nucleotide residues. Hydrogen bonding can occur through Watson-Crick base pairing, Hoogsteen binding, or any other sequence-specific manner. The complex can contain two strands forming a double-stranded structure, three or more strands forming a multistranded complex, a single self-hybridizing strand, or any combination thereof. A hybridization reaction can constitute a step in a more extensive process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence that can hybridize to a given sequence is referred to as the "complementary strand" of the given sequence.

[0035] As used herein, "expression" refers to the process by which a polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcripts) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide can be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell.

[0036] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymer may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. The term also encompasses amino acid polymers that have undergone modifications, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both D- and L-optical isomers, as well as amino acid analogs and peptidomimetics.

[0037] The terms "subject," "individual," and "patient" are used interchangeably herein to refer to a vertebrate, preferably a mammal, more preferably a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Also included are tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro.

[0038] The terms "therapeutic agent," "therapeutic drug," or "treatment agent" are used interchangeably and refer to a molecule or compound that confers some beneficial effect upon administration to a subject. Beneficial effects include enabling a diagnostic measurement; ameliorating a disease, symptom, disorder, or pathological condition; reducing or preventing the onset of a disease, symptom, disorder, or pathological condition; and generally neutralizing a disease, symptom, disorder, or pathological condition.

[0039] As used herein, "treatment" or "treating" or "alleviating" or "ameliorating" are used interchangeably. These terms refer to an approach to obtain a benefit or desired result, for example, but not limited to, a therapeutic benefit and / or a preventive benefit. A therapeutic benefit refers to any treatment-related improvement in or effect on one or more diseases, conditions, or symptoms being treated. For preventive benefit, the composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more physiological symptoms of a disease, even if the disease, condition, or symptom has not previously manifested.

[0040] The term "effective amount" or "therapeutically effective amount" refers to an amount of a drug sufficient to produce a benefit or desired result. The therapeutically effective amount may vary depending on one or more of the subject and condition being treated, the subject's weight and age, the severity of the condition, the mode of administration, etc., which can be easily determined by those skilled in the art. This term also applies to the dose that provides an image for detection by any one of the imaging methods described herein. The specified dose may vary depending on one or more of the specific drug selected, the dosing regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, the tissue to be imaged, and the physical delivery system in which it is carried.

[0041] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of one in the art. See Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd edition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F.M.A.usubel, et al. eds., (1987)); series METHODS IN ENZYMOLOGY (Academic Press, Inc.): PCR 2: A PRACTICAL APPROACH (M.J. MacPherson, B.D. Hames, and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988), ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (R.I. Freshney, ed. (1987)).

[0042] Some embodiments of the present invention relate to vector systems containing one or more vectors, or to vectors themselves.Vector can be designed for the expression of CRISPR transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells.For example, CRISPR transcripts can be expressed in bacterial cells, such as Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells.Suitable host cells are further discussed in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990).Alternatively, recombinant expression vectors can be transcribed and translated in vitro, for example, using T7 promoter regulatory sequences and T7 polymerase.

[0043] Vectors can be introduced into and propagated within prokaryotes. In some embodiments, prokaryotes are used to amplify copies of vectors to be introduced into eukaryotic cells or as intermediate vectors in the production of vectors to be introduced into eukaryotic cells (e.g., amplifying plasmids as part of a viral vector packaging system). In some embodiments, prokaryotes are used to amplify copies of vectors, express one or more nucleic acids, and provide a source of one or more proteins, for example, for delivery to a host cell or host organism. Protein expression in prokaryotes is most often carried out in Escherichia coli using vectors containing constitutive or inducible promoters directing the expression of either fusion or non-fusion proteins. Fusion vectors add multiple amino acids to the encoded protein, for example, to the amino terminus of the recombinant protein. Such fusion vectors can serve one or more purposes, for example, (i) increasing the expression of the recombinant protein; (ii) increasing the solubility of the recombinant protein; and (iii) aiding in the purification of the recombinant protein by acting as a ligand in affinity purification. In fusion expression vectors, a proteolytic cleavage site is often introduced at the junction of the fusion moiety and the recombinant protein to allow separation of the recombinant protein from the fusion moiety after purification of the fusion protein. Such enzymes and their cognate recognition sequences include factor Xa, thrombin, and enterokinase. Exemplary fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, Mass.), and pRIT5 (Pharmacia, Piscataway, NJ), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A to the target recombinant protein, respectively.

[0044] Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET 11d (Studier et al., GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990) 60-89).

[0045] In some embodiments, the vector is a yeast expression vector. Examples of vectors for expression in the yeast Saccharomyces cerivisae include pYepSec1 (Baldari, et al., 1987. EMBO J. 6:229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30:933-943), pJRY88 (Schultz et al., 1987. Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.).

[0046] In some embodiments, the vector drives protein expression in insect cells using a baculovirus expression vector. Baculovirus vectors available for protein expression in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39).

[0047] In some embodiments, the vector may drive expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329:840) and pMT2PC (Kaufman, et al., 1987. EMBO J. 6:187-195). When used in mammalian cells, expression vector control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus type 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0048] In some embodiments, the recombinant mammalian expression vector can direct expression of the nucleic acid preferentially in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific; Pinkert, et al., 1987, Genes Dev. 1:268-277), lymphoid-specific promoters (Calame and Eaton, 1988, Adv. Immunol. 43:235-275), promoters of T-cell receptors in particular (Winoto and Baltimore, 1989, EMBO J. 8:729-733) and immunoglobulins (Baneiji, et al., 1983, Cell 33:729-740; Queen and Baltimore, 1983, Cell 33:741-748), neuron-specific promoters (e.g., the neurofilament promoter; Byrne and Ruddle, 1989, Proc. Natl. Acad. Sci. USA 86:5473-5477), pancreatic-specific promoters (Edlund, et al., 1989, Proc. Natl. Acad. Sci. USA 86:5473-5477), and the like. et al., 1985, Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoters; U.S. Pat. No. 4,873,316 and European Patent Application Publication No. 264,166). Developmentally regulated promoters, such as murine hox promoters (Kessel and Gruss, 1990, Science 249:374-379) and the alpha-fetoprotein promoter (Campes and Tilghman, 1989, Genes Dev. 3:537-546), are also encompassed.

[0049] In some embodiments, the regulatory element is operably linked to one or more elements of the CRISPR system, so as to drive the expression of one or more elements of the CRISPR system.Generally, CRISPR (Clustered Regularly Interspaced Short Repeats), also known as SPIDR (Spacer Interspersed Direct Repeat), constitutes a family of DNA loci that are usually specific to certain bacterial species.CRISPR loci include a distinct class of interspersed short sequence repeats (SSRs) recognized in Escherichia coli (E. coli) (Ishino et al., J. Bacteriol.,169:5429-5433

[1987] ; and Nakata et al., J. Bacteriol.,171:3553-3556

[1989] ) and related genes. Similar interspersed SSRs have been identified in Haloferax mediterranei, Streptococcus pyogenes, Anabaena, and Mycobacterium tuberculosis (see Groenen et al., Mol. Microbiol., 10:1057-1065

[1993] ; Hoe et al., Emerg. Infect. Dis., 5:254-263

[1999] ; Masepohl et al., Biochim. Biophys. Acta 1307:26-30

[1996] ; and Mojica et al., Mol. Microbiol., 17:85-93

[1995] ). CRISPR loci typically differ from other SSRs in the structure of their repeats, which are termed short regularly interspaced repeats (SRSRs) (Janssen et al., OMICS J. Integ. Biol., 6:23-33

[2002] ; and Mojica et al., Mol. Microbiol., 36:244-246

[2000] ). Generally, repeats are short elements that occur in clusters regularly spaced by unique intervening sequences of substantially constant length (Mojica et al.,

[2000] , supra). Repeat sequences are highly conserved between strains, butThe number of interspersed repeats and the sequence of the spacer region typically vary from strain to strain (van Embden et al., J. Bacteriol., 182:2393-2401

[2000] ). CRISPR loci have been identified in over 40 prokaryotes (e.g., Jansen et al., Mol. Microbiol., 43:1565-1575

[2002] ; and Mojica et al. See, e.g., et al.,

[2005] , including, but not limited to, the genera Aeropyrum, Pyrobaculum, Sulfolobus, Archaeoglobus, Halocarcula, Methanobacterium, Methanococcus, Methanosarcina, Methanopyrus, Pyrococcus, Picrophilus, Thermoplasma, Corynebacterium, Mycobacterium, Streptomyces, Aquifex, and the like. fex), Porphyromonas, Chlorobium, Thermus, Bacillus, Listeria, Staphylococcus, Clostridium, Thermoanaerobacter, Mycoplasma, Fusobacterium, Azarcus, Chromobacterium, Neisseria, Nitrosomonas, Desulfovibrio, Geobacter, Myxococcus, Campylobacter,The genera Wolinella, Acinetobacter, Erwinia, Escherichia, Legionella, Methylococcus, Pasteurella, Photobacterium, Salmonella, Xanthomonas, Yersinia, Treponema, and Thermotoga.

[0050] Generally, a "CRISPR system" collectively refers to transcripts and other elements involved in directing the expression or activity of CRISPR-associated ("Cas") genes, e.g., sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr mate sequences (including "direct repeats" and partial direct repeats processed by tracrRNA in the context of endogenous CRISPR systems), guide sequences (also referred to as "spacers" in the context of endogenous CRISPR systems), or other sequences and transcripts from a CRISPR locus. In some embodiments, one or more elements of a CRISPR system are derived from a type I, type II, or type III CRISPR system. In some embodiments, one or more elements of a CRISPR system are derived from a particular organism that contains an endogenous CRISPR system, e.g., Streptococcus pyogenes. Generally, CRISPR systems are characterized by elements that promote the formation of CRISPR complexes at target sequences (also referred to as protospacers in endogenous CRISPR systems). With regard to the formation of CRISPR complexes, "target sequence" refers to a sequence to which a guide sequence is designed to have complementarity, and hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. Full complementarity is not necessarily required, provided that there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR complex. The target sequence can comprise any polynucleotide, for example, a DNA or RNA polynucleotide. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence can be present in an organelle of a eukaryotic cell, such as a mitochondria or chloroplast. A sequence or template that can be used for recombination into a targeted locus containing a target sequence is referred to as an "editing template," "editing polynucleotide," or "editing sequence." In an embodiment of the present invention, an exogenous template polynucleotide can be referred to as an editing template. In one aspect of the invention, the recombination is homologous recombination.

[0051] Typically, with respect to endogenous CRISPR systems, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs therefrom). Without being bound by theory, the tracr sequence may comprise or consist of all or a portion of the wild-type tracr sequence (e.g., more than about or about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracr sequence), and may also form part of a CRISPR complex, for example, by hybridization along at least a portion of the tracr sequence with all or a portion of a tracr mate sequence operably linked to the guide sequence. In some embodiments, the tracr sequence has sufficient complementarity to the tracr mate sequence to hybridize and participate in the formation of a CRISPR complex. As with the target sequence, perfect complementarity is not required, provided that sufficient complementarity exists to be functional. In some embodiments, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr mate sequence when optimally aligned. In some embodiments, one or more vectors driving the expression of one or more elements of a CRISPR system are introduced into a host cell, such that the expression of the elements of the CRISPR system directs the formation of a CRISPR complex at one or more target sites. For example, the Cas enzyme, the guide sequence linked to the tracr mate sequence, and the tracr sequence can each be operably linked to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements can be combined in a single vector, and one or more additional vectors providing any component of the CRISPR system are not included in the first vector.The CRISPR system elements combined in a single vector can be arranged in any suitable orientation; for example, an element can be located 5' (upstream) or 3' (downstream) relative to a second element. The coding sequence of an element can be located on the same or opposite strand as the coding sequence of a second element and can be oriented in the same or opposite direction. In some embodiments, a single promoter drives the expression of a transcript encoding a CRISPR enzyme and one or more of a guide sequence, a tracr mate sequence (optionally operably linked to a guide sequence), and a tracr sequence embedded in one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron). In some embodiments, the CRISPR enzyme, guide sequence, tracr mate sequence, and tracr sequence are operably linked to and expressed from the same promoter.

[0052] In some embodiments, a vector comprises one or more insertion sites, such as restriction endonuclease recognition sequences (also referred to as "cloning sites"). In some embodiments, one or more insertion sites (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. In some embodiments, a vector comprises an insertion site upstream of a tracr mate sequence and optionally downstream of a regulatory element operably linked to the tracr mate sequence, such that after insertion of the guide sequence into the insertion site and upon expression, the guide sequence directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell. In some embodiments, a vector comprises two or more insertion sites, each located between two tracr mate sequences to allow insertion of a guide sequence at the respective site. In such an arrangement, the two or more guide sequences may comprise two or more copies of a single guide sequence, two or more different guide sequences, or a combination thereof. When using multiple different guide sequences, can use a single expression construct to target the CRISPR activity to multiple different corresponding target sequences in cells.For example, a single vector can comprise about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 or more guide sequences.In some embodiments, about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more of these guide sequence-containing vectors can be provided and optionally delivered into cells.

[0053] In some embodiments, the vector comprises regulatory elements operably linked to an enzyme coding sequence that encodes a CRISPR enzyme, e.g., a Cas protein. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. These enzymes are known; for example, the amino acid sequence of Streptococcus pyogenes (S. pyogenes) Cas9 protein can be found in the SwissProt database under accession number Q99ZW2. In some embodiments, the unmodified CRISPR enzyme has DNA cleavage activity, for example, Cas9. In some embodiments, the CRISPR enzyme is Cas9, and can be Cas9 from Streptococcus pyogenes (S. pyogenes) or Streptococcus pneumoniae (S. pneumoniae). In some embodiments, the CRISPR enzyme directs cleavage of one or both strands at the location of the target sequence, for example, within the target sequence and / or the complementary strand of the target sequence. In some embodiments, the CRISPR enzyme directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the vector encodes a CRISPR enzyme that is mutated relative to the corresponding wild-type enzyme, such that the mutant CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing the target sequence.For example, an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from Streptococcus pyogenes (S. pyogenes) converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that convert Cas9 into a nickase include, but are not limited to, H840A, N854A, and N863A. In some embodiments, Cas9 nickase can be used in combination with guide sequences, e.g., two guide sequences that target the sense and antisense strands of a DNA target, respectively. This combination allows both strands to be nicked and used to induce NHEJ. Applicants have demonstrated the efficacy of two nickase targets (i.e., sgRNAs targeted to the same location but different strands of DNA) in inducing mutagenic NHEJ (data not shown). While a single nickase (Cas9-D10A with a single sgRNA) cannot induce NHEJ and create indels, we have shown that a dual nickase (Cas9-D10A and two sgRNAs targeted to different strands at the same location) can do so in human embryonic stem cells (hESCs), with approximately 50% of the efficiency of a nuclease (i.e., regular Cas9 without the D10 mutation) in hESCs.

[0054] As a further example, two or more catalytic domains of Cas9 (RuvC I, RuvC II, and RuvC III) can be mutated to produce a mutant Cas9 that substantially lacks all DNA cleavage activity. In some embodiments, the D10A mutation is combined with one or more of the H840A, N854A, or N863A mutations to produce a Cas9 enzyme that substantially lacks all DNA cleavage activity. In some embodiments, a CRISPR enzyme is considered to substantially lack all DNA cleavage activity if the DNA cleavage activity of the mutant enzyme is less than about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less of its non-mutated form. Other mutations may be useful; if the Cas9 or other CRISPR enzyme is from a species other than S. pyogenes, corresponding amino acid mutations can be made to achieve similar effects.

[0055] In some embodiments, the enzyme coding sequence encoding CRISPR enzyme is codon-optimized for expression in specific cells, for example, eukaryotic cells.Eukaryotic cells can be derived from or derived from specific organisms, for example, mammals, for example, but not limited to, humans, mice, rats, rabbits, dogs, or non-human primates.Generally, codon optimization refers to the process of modifying nucleic acid sequences to improve expression in target host cells by replacing at least one codon (for example, about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of native sequence with the more frequently or most frequently used codon in the gene of the host cell, while maintaining the native amino acid sequence.Different species show specific bias for certain codons of specific amino acids. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, among other things, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of tRNAs selected in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, in "codon usage databases," and these tables can be adapted in a number of ways. See Nakamura, Y., et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucl. Acids Res. 28:292 (2000). Computer algorithms for codon-optimizing specific sequences for expression in specific host cells are also available, for example, in GeneForge (Aptagen; Jacobus, PA).In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in a sequence encoding a CRISPR enzyme correspond to the most frequently used codon for a particular amino acid.

[0056] In some embodiments, the vector encodes a CRISPR enzyme that comprises one or more nuclear localization sequences (NLSs), for example, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the CRISPR enzyme comprises about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the amino terminus, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the carboxy terminus, or a combination thereof (for example, one or more NLSs at the amino terminus and one or more NLSs at the carboxy terminus). When two or more NLSs are present, each can be selected independently from the others, so that a single NLS can exist in two or more copies, and / or in combination with one or more other NLSs that exist in one or more copies. In a preferred embodiment of the present invention, the CRISPR enzyme comprises at most six NLSs. In some embodiments, an NLS is considered to be near the N- or C-terminus if the nearest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N- or C-terminus. Typically, an NLS consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface, although other types of NLSs are known.Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen, which has the amino acid sequence PKKKRKV; an NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS, which has the sequence KRPAATKKAGQAKKKK); a c-myc NLS, which has the amino acid sequence PAAKRVKLD or RQRRNELKRSP; an hRNPA1 M9 NLS, which has the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY; an IBB domain from importin alpha, which has the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV; the sequences VSRKRPRP and PPKKARED of the fibroid T protein; the sequence POPKKKPL of human p53; and a mouse c-abl IV sequence SALIKKKKKMAP; influenza virus NS1 sequences DRLRR and PKQKKRK; hepatitis virus delta antigen sequence RKLKKKIKKL; mouse Mx1 protein sequence REKKKFLKRR; human poly(ADP-ribose) polymerase sequence KRKGDEVDGVDEVAKKKSKK; and steroid hormone receptor (human) glucocorticoid sequence RKCLQAGMNLEARKTKK.

[0057] Generally, one or more NLSs are strong enough to drive the accumulation of detectable amounts of CRISPR enzyme in the nucleus of eukaryotic cells.Generally, the strength of nuclear localization activity can be derived from the number of NLSs in CRISPR enzyme, the specific NLS used, or a combination of these factors.Detection of nuclear accumulation can be carried out by any suitable technique.For example, a detectable marker can be fused to CRISPR enzyme, so that intracellular localization can be visualized, for example, by combining with a means for detecting nuclear localization (for example, nuclear-specific staining, for example, DAPI).Examples of detectable markers include fluorescent proteins (for example, green fluorescent protein, or GFP; RFP; CFP) and epitope tags (HA tag, flag tag, SNAP tag).Cell nuclei can also be isolated from cells, and then their contents can be analyzed by any suitable process for detecting protein, for example, immunohistochemical analysis, Western blot, or enzyme activity assay. Accumulation in the nucleus can also be measured indirectly, for example, by assaying for the effect of CRISPR complex formation (e.g., assaying for DNA cleavage or mutation at the target sequence, or assaying for changes in gene expression activity affected by CRISPR complex formation and / or CRISPR enzymatic activity), compared to a control exposed to neither the CRISPR enzyme nor the complex, or to a CRISPR enzyme lacking one or more NLSs.

[0058] Generally, a guide sequence is any polynucleotide sequence that has sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence is greater than about or about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more when optimally aligned using a suitable alignment algorithm. Optimal alignment can be determined using any suitable algorithm for aligning sequences, including, but not limited to, the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, CA), and others. Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, the guide sequence is about or greater than about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, the guide sequence is about 75, 50, 45, 40, 35, 30, 25 The length of the guide sequence is less than 1, 20, 15, 12 or less nucleotides.The ability of the guide sequence to direct the sequence-specific binding of CRISPR complex to target sequence can be evaluated by any suitable assay.For example, the components of the CRISPR system sufficient to form a CRISPR complex, for example, the guide sequence to be tested, can be provided to the host cell having the corresponding target sequence, for example, by transfection with the vector encoding the components of the CRISPR sequence, and then the preferential cleavage within the target sequence is evaluated, for example, by the Surveyor assay described herein.Similarly, cleavage of a target polynucleotide sequence can be assessed in a test tube by providing the target sequence, a component of a CRISPR complex, for example, a guide sequence to be tested and a control guide sequence that differs from the test guide sequence, and comparing the rate of binding or cleavage at the target sequence between the test and control guide sequence reactions. Other assays are contemplated and will be recognized by those skilled in the art.

[0059] The guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the genome of a cell. Exemplary target sequences include those that are unique in the target genome. For example, for S. pyogenes Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXGG, where NNNNNNNNNNNNXGG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. A unique target sequence in the genome can include a S. pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG, where NNNNNNNNNNNXGG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. For S. thermophilus CRISPR1Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNXXAGAAW, where NNNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; W is A or T) has a single occurrence in the genome. A unique target sequence in the genome can include a S. thermophilus CRISPR1Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXXAGAAW, where NNNNNNNNNNNNXXAGAAW (N is A, G, T, or C; X can be anything; W is A or T) has a single occurrence in the genome. For S. pyogenes Cas9, a unique target sequence in the genome can include a Cas9 target site of the form MMMMMMMMNNNNNNNNNNNNNXGGXG, where NNNNNNNNNNNNXGGXG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome.A unique target sequence in a genome can include a Streptococcus pyogenes (S. pyogenes) Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGGXG, where NNNNNNNNNNNXGGXG (N is A, G, T, or C; X can be anything) has a single occurrence in the genome. In each of these sequences, "M" can be A, G, T, or C and need not be considered in identifying the sequence as unique.

[0060] In some embodiments, the guide sequence is selected to reduce the degree of secondary structure within the guide sequence. The secondary structure can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. An example of such an algorithm is mFold, described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold, developed by the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, for example, A.R. Gruber et al., 2008, Cell 106(1):23-24; and P.A. Carr and G.M. Church, 2009, Nature Biotechnology 27(12):1151-62). Further algorithms can be found in U.S. Patent Application No. TBA (Attorney Docket No. 44790.11.2022; Broad Reference No. BI-2013 / 004A), which is incorporated herein by reference.

[0061] Generally, the tracr mate sequence includes any sequence that has sufficient complementarity with the tracr sequence to promote one or more of the following: (1) excision of the guide sequence flanked by the tracr mate sequence in cells containing the corresponding tracr sequence; and (2) formation of a CRISPR complex in the target sequence (the CRISPR complex includes the tracr mate sequence hybridized to the tracr sequence). Generally, the degree of complementarity is based on optimal alignment of the tracr mate sequence and the tracr sequence along the shorter length of the two sequences. Optimal alignment can be determined by any suitable alignment algorithm and can further account for secondary structures, such as self-complementarity within the tracr sequence or the tracr mate sequence. In some embodiments, the degree of complementarity between the tracr sequence and the tracr mate sequence along the shorter length of the two sequences is greater than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more when optimally aligned. An exemplary illustration of optimal alignment between the tracr sequence and the tracr mate sequence is provided in Figures 12B and 13B. In some embodiments, the tracr sequence is about or greater than about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and the tract mate sequence are contained within a single transcript, such that hybridization between the two produces a transcript with a secondary structure, e.g., a hairpin. A preferred loop-forming sequence used in the hairpin structure is 4 nucleotides in length and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used, as can alternative sequences. The sequence preferably includes a nucleotide triplet (e.g., AAA) and an additional nucleotide (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAG. In one embodiment of the present invention, the transcript or transcribed polynucleotide sequence has at least two or more hairpins.In preferred embodiments, the transcript has two, three, four, or five hairpins. In another further embodiment of the present invention, the transcript has at most five hairpins. In some embodiments, the single transcript further comprises a transcription termination sequence; preferably, this is a poly-T sequence, e.g., six T nucleotides. An illustrative illustration of such a hairpin structure is provided in the lower portion of Figure 13B, where the final "N" and the portion of the sequence 5' upstream of the loop correspond to the tracr mate sequence, and the portion of the sequence 3' upstream of the loop corresponds to the tracr sequence. Further non-limiting examples of single polynucleotides comprising a guide sequence, a tracr mate sequence, and a tracr sequence are as follows (listed 5' to 3'), where "N" represents the base of the guide sequence, the first block of lowercase letters represents the tracr mate sequence, the second block of lowercase letters represents the tracr sequence, and the final poly-T sequence represents the transcription terminator: [ka] In some embodiments, sequences (1) through (3) are used in combination with Cas9 from S. thermophilus CRISPR1. In some embodiments, sequences (4) through (6) are used in combination with Cas9 from S. pyogenes. In some embodiments, the tracr sequence is a separate transcript from the transcript containing the tracr mate sequence (e.g., as illustrated in the top diagram of Figure 13B).

[0062] In some embodiments, a recombination template is also provided. The recombination template is a component of another vector described herein, and can be contained in a separate vector or provided as a separate polynucleotide. In some embodiments, the recombination template is designed to function as a template for homologous recombination within or near the target sequence that is nicked or cleaved by a CRISPR enzyme as part of a CRISPR complex. The template polynucleotide can be of any suitable length, for example, about or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000 or more nucleotides in length. In some embodiments, the template polynucleotide is complementary to a portion of the polynucleotide that comprises the target sequence. When optimally aligned, the template polynucleotide can overlap with one or more nucleotides of the target sequence (for example, about or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 or more nucleotides). In some embodiments, when polynucleotides comprising a template sequence and a target sequence are optimally aligned, the nearest neighbor nucleotide of the template polynucleotide is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides of the target sequence.

[0063] In some embodiments, the CRISPR enzyme is part of a fusion protein containing one or more heterologous protein domains (e.g., about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more other domains of the CRISPR enzyme). The CRISPR enzyme fusion protein can include any additional protein sequence, and optionally a linker sequence between any two domains. Examples of protein domains that can be fused to the CRISPR enzyme include, but are not limited to, epitope tags, reporter gene sequences, and protein domains with one or more of the following activities: methylase activity, demethylase activity, transcription activation activity, transcription repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tag, V5 tag, FLAG tag, influenza hemagglutinin (HA) tag, Myc tag, VSV-G tag, and thioredoxin (Trx) tag. Examples of reporter genes include, but are not limited to, glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins, such as blue fluorescent protein (BFP). CRISPR enzymes can be fused to gene sequences encoding proteins or protein fragments that bind to DNA molecules or other cellular molecules, such as, but not limited to, maltose binding protein (MBP), S-tags, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that can form part of fusion proteins comprising CRISPR enzymes are described in U.S. Patent Application Publication No. 20110059502, which is incorporated herein by reference. In some embodiments, tagged CRISPR enzymes are used to identify the location of the target sequence.

[0064] In some embodiments, the present invention provides methods comprising delivering one or more polynucleotides, such as, for example, one or more vectors described herein, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell. In some embodiments, the present invention further provides cells produced by such cells, and organisms (e.g., animals, plants, or fungi) comprising or produced from such cells. In some embodiments, a CRISPR enzyme in combination with (and optionally complexed with) a guide sequence is delivered to the cell. Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids encoding components of a CRISPR system to cells in culture or into a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery vehicle, such as a liposome. Viral vector delivery systems include DNA and RNA viruses with episomal or integrated genomes after delivery to the cell.For an overview of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Felgner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and See Immunology, Doerfler and Boehm (eds) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).

[0065] Non-viral nucleic acid delivery methods include lipofection, nucleofection, microinjection, gene gun, virosome, liposome, immunoliposome, polycation or lipid:nucleic acid conjugate, naked DNA, artificial virion, and drug-enhanced DNA uptake.Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787; and 4,897,355, and lipofection reagents are commercially available (for example, Transfectam™ and Lipofectin™).Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those of Felgner, WO 91 / 17424; WO 91 / 16024.Delivery can be to cells (for example, in vitro or ex vivo administration) or target tissue (for example, in vivo administration).

[0066] The preparation of lipid:nucleic acid complexes, including targeted liposomes, e.g., immunolipid complexes, is well known to those skilled in the art (see, e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); see U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0067] The use of RNA or DNA virus-based systems for nucleic acid delivery takes advantage of a highly evolved process that targets viruses to specific cells in the body and transports the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, and the modified cells can then be administered to patients (ex vivo). Conventional virus-based systems include retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated virus gene transfer methods, often resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues.

[0068] The tropism of retroviruses can be altered by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors contain cis-acting long terminal repeats (LTRs) that have packaging capacity for foreign sequences up to 6-10 kb. The minimal cis-acting LTRs are sufficient for vector replication and packaging, which are then used to integrate therapeutic genes into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), or combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731-2739 (1992); Johann et al., J. Virol. 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., J. Virol. 63:2374-2378 (1989); Miller et al., J. Virol. 65:2220-2224 (1991); PCT / US94 / 05700). In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors can exhibit extremely high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained with such vectors. The vectors can be produced in large quantities using a relatively simple system.For example, adeno-associated virus ("AAV") vectors can also be used to transduce cells with target nucleic acids in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Pat. No. 4,797,368; WO 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994). The construction of recombinant AAV vectors is described in numerous publications, e.g., U.S. Pat. No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al. al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989).

[0069] Typically, packaging cells are used to form viral particles that can infect host cells. Such cells include 293 cells, which package adenovirus, and ψ2 or PA317 cells, which package retrovirus. Viral vectors used in gene therapy are usually produced by creating cell lines that package nucleic acid vectors into viral particles. The vector typically contains the minimum viral sequences required for packaging and subsequent integration into the host, with other viral sequences replaced by an expression cassette for the polynucleotide to be expressed. Defective viral functions are typically provided in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome required for packaging and integration into the host genome. Viral DNA is packaged into a cell line containing a helper plasmid encoding other AAV genes, i.e., rep and cap, but lacking the ITR sequences. The cell line can also be infected with adenovirus as a helper. Helper virus promotes the replication of AAV vector and the expression of AAV gene from helper plasmid.Helper plasmid is not packaged in significant amounts due to the lack of ITR sequence.Contamination by adenovirus can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV.Additional methods for delivering nucleic acid to cells are known to those skilled in the art.For example, see US Patent Application Publication No. 20030087817, which is incorporated herein by reference.

[0070] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected as they naturally occur in a subject. In some embodiments, the transfected cells are harvested from a subject. In some embodiments, the cells are derived from cells harvested from a subject, e.g., cell lines. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa-S3, Huh1, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panc1, PC-3, TF1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, TIB55, Jurkat, J45.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, and Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis , A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd.3, BHK-21, BR29 3, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L23 / 5010, COR-L23 / R23, COS-7, COV-434, CML T1, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL-60, HMEC, HT-29, Jurkat, JY cells K562 cells, Ku8 12, KCL22, KG1, KYO1, LNCap, Ma-Mel1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCK II, MDCK II, MOR / 0.2R, MONO-MAC6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic variants thereof. Cell lines are available from a variety of sources known to those of skill in the art (see, e.g., American Type Culture Collection (ATCC), Manassus, Va.). In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines comprising one or more vector-derived sequences. In some embodiments, cells transiently transfected with components of a CRISPR system described herein (e.g., by transient transfection of one or more vectors or transfection with RNA) and modified through the activity of a CRISPR complex are used to establish new cell lines comprising cells containing the modification but lacking any other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used in the evaluation of one or more test compounds.

[0071] In some embodiments, one or more vectors described herein are used to produce non-human transgenic animals or transgenic plants.In some embodiments, the transgenic animal is a mammal, for example, a mouse, a rat, or a rabbit.In some embodiments, the organism or subject is a plant.In some embodiments, the organism or subject or plant is an algae.Methods for producing transgenic plants and animals are known in the art and generally start from, for example, the cell transfection method described herein.

[0072] In one aspect, the present invention provides a method for modifying target polynucleotide in eukaryotic cells.In some embodiments, the method comprises: allowing CRISPR complex to bind to target polynucleotide, causing the cleavage of said target polynucleotide, thereby modifying said target polynucleotide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with the guide sequence that hybridizes with the target sequence in said target polynucleotide, and said guide sequence is then bound to the tracr mate sequence that hybridizes with the tracr sequence.

[0073] In one aspect, the present invention provides a method for modifying the expression of polynucleotide in eukaryotic cells.In some embodiments, the method comprises: allowing CRISPR complex to bind to polynucleotide, and thereby causing the binding to increase or decrease the expression of said polynucleotide; CRISPR complex comprises CRISPR enzyme complexed with guide sequence that hybridizes with target sequence in said target polynucleotide, and said guide sequence is then bound to tracr mate sequence that hybridizes with tracr sequence.

[0074] With recent advances in crop genomics, the ability to use CRISPR-Cas systems to perform efficient and cost-effective gene editing and manipulation allows for the rapid selection and comparison of single and multiplexed genetic manipulations to transform such genomes for improved production and trait enhancement. In this regard, reference is made to U.S. patents and publications: U.S. Patent No. 6,603,061 - Agrobacterium-Mediated Plant Transformation Method; U.S. Patent No. 7,868,149 - Plant Genome Sequences and Uses Thereof and U.S. Patent Application Publication No. 2009 / 0100536 - Transgenic Plants with Enhanced Agronomic Traits, the entire contents and disclosures of each of which are incorporated herein by reference in their entirety. In implementing the present invention, the contents and disclosures of Morrell et al. "Crop genomics: advances and applications" Nat Rev Genet. 2011 Dec 29; 13(2): 85-96 are also incorporated herein by reference in their entirety. In an advantageous embodiment of the present invention, the CRISPR / Cas9 system is used to engineer microalgae (Example 15). Thus, references herein to animal cells may, mutatis mutandis, also apply to plant cells unless otherwise indicated.

[0075] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method includes sampling a cell or a population of cells from a human or non-human animal or plant (including microalgae) and modifying one or more cells. Culturing can be performed ex vivo at any stage. One or more cells can also be reintroduced into a non-human animal or plant (including microalgae).

[0076] In plants, pathogens are often host-specific. For example, Fusarium oxysporum f.sp. lycopersici, which causes tomato wilt, attacks only tomatoes, while F. oxysporum f. dianthiii and Puccinia graminis f.sp. tritici, which cause carnation wilt, attack only wheat. Plants have pre-existing and inducible defenses to resist most pathogens. Mutation and recombination events throughout plant development result in genetic variations that cause susceptibility, especially when the pathogen outgrows the plant. In plants, non-host resistance can exist, e.g., the host and pathogen are incompatible. Horizontal resistance, e.g., partial resistance to all species of pathogens, typically controlled by many genes, and vertical resistance, e.g., complete resistance to some species of pathogens but not others, typically controlled by a small number of genes, can also exist. At the gene-by-gene level, plants and pathogens evolve together, with genetic changes in some balanced by others. Thus, using natural variation, breeders combine genes that are most useful for yield, quality, uniformity, cold tolerance, and resistance. Sources of resistance genes include natural or exotic species, landraces, wild plant relatives, and induced mutations, such as treating plant material with mutagens. The present invention provides plant breeders with new tools for inducing mutations. Thus, those skilled in the art can analyze the genomes of sources of resistance genes and use the present invention to induce the appearance of resistance genes in varieties with desired characteristics or traits more precisely than traditional mutagens, thus accelerating and improving plant breeding programs.

[0077] In one aspect, the present invention provides a kit containing any one or more of the elements disclosed in the above methods and compositions. In some embodiments, the kit includes a vector system and instructions for use of the kit. In some embodiments, the vector system includes: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting a guide sequence upstream of the tracr mate sequence (the guide sequence, when expressed, directs the sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with the tracr mate sequence hybridized to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme comprising a nuclear localization sequence. The elements can be provided individually or in combination, and can be provided in any suitable container, for example, a vial, a bottle, or a tube. In some embodiments, the kit includes instructions in one or more languages, for example, in two or more languages.

[0078] In some embodiments, the kit includes one or more reagents used in a process utilizing one or more of the elements described herein. The reagents can be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents can be provided in a form usable in a particular assay or in a form that requires the addition of one or more other components prior to use (e.g., a concentrate or lyophilized form). The buffer can be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to the guide sequence for insertion into a vector to operably link the guide sequence and regulatory elements. In some embodiments, the kit includes a homologous recombination template polynucleotide.

[0079] In one embodiment, the present invention provides a method for using one or more elements of the CRISPR system.The CRISPR complex of the present invention provides an effective means for modifying target polynucleotide.The CRISPR complex of the present invention has a wide range of uses, for example, modifying (for example, deletion, insertion, translocation, inactivation, activation) target polynucleotide in a large number of cell types.Therefore, the CRISPR complex of the present invention has a wide range of applications, for example, in gene therapy, drug screening, disease diagnosis and prognosis.An exemplary CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence in a target polynucleotide.The guide sequence is then linked to a tracr mate sequence that hybridizes to a tract sequence.

[0080] The target polynucleotide of a CRISPR complex can be any polynucleotide that is endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide that remains in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Without being bound by theory, it is believed that the target sequence must associate with a PAM (protospacer adjacent motif); i.e., a short sequence recognized by the CRISPR complex. While the exact sequence and length requirements for the PAM vary depending on the CRISPR enzyme used, the PAM is typically a 2-5 base pair sequence adjacent to the protospacer (i.e., the target sequence). Exemplary PAM sequences are provided in the Examples section below, and those skilled in the art can identify additional PAM sequences for use with a given CRISPR enzyme.

[0081] Target polynucleotides for CRISPR complexes can include many of the disease-associated genes and polynucleotides and signaling biochemical pathway-associated genes and polynucleotides listed in U.S. Provisional Patent Applications Nos. 61 / 736,527 and 61 / 748,427 (both entitled SYSTEMS METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION, filed December 12, 2012 and January 2, 2013, respectively), having Broad Reference Nos. BI-2011 / 008 / WSGR Docket No. 44063-701.101 and BI-2011 / 008 / WSGR Docket No. 44063-701.102, respectively, the entire contents of which are incorporated herein by reference in their entireties.

[0082] Examples of target polynucleotides include sequences related to signaling biochemical pathways, such as signaling biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-related genes or polynucleotides. "Disease-related" genes or polynucleotides refer to any gene or polynucleotide that produces transcription or translation products at abnormal levels or in abnormal forms in cells derived from diseased tissues compared with tissues or cells of non-disease control. This can be a gene that is expressed at abnormally high levels; or a gene that is expressed at abnormally low levels, and expression changes correlate with the occurrence and / or progression of disease. Disease-related genes also refer to genes that are directly responsible for the pathogenesis of disease or have mutations or genetic variations that are in linkage disequilibrium with genes that are responsible for the pathogenesis of disease. The transcription or translation products can be known or unknown, and can be at normal or abnormal levels.

[0083] Examples of disease-associated genes and polynucleotides are available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and are available on the World Wide Web.

[0084] Examples of disease-associated genes and polynucleotides are listed in Tables A and B. Disease-specific information is available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.), and is available on the World Wide Web. Examples of signaling biochemical pathway-associated genes and polynucleotides are listed in Table C.

[0085] Mutations in these genes and pathways can cause inappropriate protein production or inappropriate amount of protein, which affects function.Additional examples of genes, diseases and proteins are incorporated herein by reference from U.S. Provisional Patent Application No. 61 / 736,527, filed December 12, 2012, and U.S. Provisional Patent Application No. 61 / 748,427, filed February 2, 2013.Such genes, proteins and pathways can be the target polynucleotide of CRISPR complex.

[0086] [Table 1]

[0087] [Table 2]

[0088] [Table 3]

[0089] [Table 4]

[0090] [Table 5]

[0091]

Table 6

[0092]

Table 7

[0093]

Table 8

[0094]

Table 9

[0095]

Table 10

[0096]

Table 11

[0097]

Table 12

[0098]

Table 13

[0099]

Table 14

[0100]

Table 15

[0101] [Table 16]

[0102] [Table 17]

[0103] Embodiments of the present invention also relate to methods and compositions related to gene knockout, gene amplification, and repair of specific mutations associated with DNA repeat instability and neurological diseases (Robert D. Wells, Tetsuo Ashizawa, Genetic Instabilities and Neurological Diseases, Second Edition, Academic Press, October 13, 2011 - Medical). Tandem repeat sequences of defined nature have been found to be responsible for more than 20 human diseases (New insights into repeat instability: role of RNA-DNA hybrids. McIvor EI, Polak U, Napierala M. RNA Biol. 2010 September-Oct;7(5):551-8). The CRISPR-Cas system can be used to correct these abnormalities of genomic instability.

[0104] A further aspect of the present invention relates to the use of the CRISPR-Cas system for the correction of abnormalities in the EMP2A and EMP2B genes that have been identified as being associated with Lafora disease. Malignant encephalopathy (MRSA) is an autosomal recessive condition characterized by progressive myoclonic epilepsy that may begin in adolescence as epileptic seizures. Some cases of the disease may be caused by mutations in an as-yet-unidentified gene. The disease causes seizures, muscle spasms, difficulty walking, dementia, and ultimately death. Currently, no treatments have been proven effective against disease progression. Other genetic abnormalities associated with epilepsy can also be targeted using the CRISPR-Cas system, and the underlying genetics are further described in Genetics of Epilepsy and Genetic Epilepsies, edited by Giuliano Avanzini and Jeffrey L. Noebels, Mariani Foundation Paediatric Neurology: 20; 2009).

[0105] In yet another embodiment of the present invention, the CRISPR-Cas system is used to generate a gene encoding a gene encoding a gene for the eye, as described in Genetic Diseases of the Eye, Second Edition, Eliah. It can correct eye defects resulting from several gene mutations, which are further described in "Eye defects resulting from several gene mutations," edited by I. Traboulsi, Oxford University Press, 2012.

[0106] Some further aspects of the present invention relate to the correction of abnormalities associated with a wide range of genetic diseases, which are further described on the National Institutes of Health website (website at health.nih.gov / topic / GeneticDisorders) under the topic subsection Genetic Disorders. Genetic brain diseases include, but are not limited to, adrenoleukodystrophy, agenesis of the corpus callosum, Aicardi syndrome, Alpers disease, Alzheimer's disease, Barth syndrome, Batten disease, CADASIL, cerebellar degeneration, Fabry disease, Gerstmann-Straussler-Scheinker disease, Huntington's disease and other triplet repeat diseases, Leigh disease, Lesch-Nyhan syndrome, Menkes disease, mitochondrial myopathy, and NINDS colpocephaly. These diseases are further described on the National Institutes of Health website under the subsection Genetic Brain Disorders.

[0107] In some embodiments, the condition is neoplasia. In some embodiments, where the condition may be neoplasia, the gene to be targeted may be any of those listed in Table A (such as PTEN in this case). In some embodiments, the condition may be age-related macular degeneration. In some embodiments, the condition may be schizophrenia. In some embodiments, the condition may be a trinucleotide repeat disorder. In some embodiments, the condition may be fragile X syndrome. In some embodiments, the condition may be a secretase-associated disorder. In some embodiments, the condition may be a prion-associated disorder. In some embodiments, the condition may be ALS. In some embodiments, the condition may be drug addiction. In some embodiments, the condition may be autism. In some embodiments, the condition may be Alzheimer's disease. In some embodiments, the condition may be inflammation. In some embodiments, the condition may be Parkinson's disease.

[0108] Examples of proteins associated with Parkinson's disease include, but are not limited to, alpha-synuclein, DJ-1, LRRK2, PINK1, parkin, UCHL1, synphilin-1, and NURR1.

[0109] An example of a preference-related protein is ABAT.

[0110] Examples of inflammation-related proteins include monocyte chemoattractant protein-1 (MCP1) encoded by the Ccr2 gene, CC chemokine receptor type 5 (CCR5) encoded by the Ccr5 gene, IgG receptor IIB (FCGR2b, also known as CD32) encoded by the Fcgr2b gene, or Fc epsilon R1g (FCER1g) protein encoded by the Fcer1g gene.

[0111] Examples of cardiovascular disease-related proteins include, for example, IL1B (interleukin 1, beta), XDH (xanthine dehydrogenase), TP53 (tumor protein p53), PTGIS (prostaglandin I2 (prostacyclin) synthase), MB (myoglobin), IL4 (interleukin 4), ANGPT1 (angiopoietin 1), ABCG8 (ATP-binding cassette, subfamily G (WHITE), member 8), or CTSK (cathepsin K).

[0112] Examples of Alzheimer's disease-related proteins include, for example, the very low density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, the ubiquitin-like modifier activating enzyme 1 (UBA1) encoded by the UBA1 gene, or the NEDD8 activating enzyme E1 catalytic subunit protein (UBE1C) encoded by the UBA3 gene.

[0113] Examples of proteins associated with autism spectrum disorders include, for example, benzodiazepine receptor (peripheral)-associated protein 1 (BZRAP1) encoded by the BZRAP1 gene, AF4 / FMR2 family member 2 protein (AFF2) encoded by the AFF2 gene (also known as MFR2), fragile X mental retardation autosomal homolog 1 protein (FXR1) encoded by the FXR1 gene, or fragile X mental retardation autosomal homolog 2 protein (FXR2) encoded by the FXR2 gene.

[0114] Examples of proteins associated with macular degeneration include the ATP-binding cassette subfamily A (ABC1) member 4 protein (ABCA4) encoded by the ABCR gene, the apolipoprotein E protein (APOE) encoded by the APOE gene, or the chemokine (CC motif) ligand 2 protein (CCL2) encoded by the CCL2 gene.

[0115] Examples of proteins associated with schizophrenia include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BDNF, DISC1, GSK3B, and combinations thereof.

[0116] Examples of proteins involved in tumor suppression include ATM (ataxia telangiectasia mutated), ATR (ataxia telangiectasia and Rad3 related), EGFR (epidermal growth factor receptor), ERBB2 (v-erb-b2 erythroblastic leukemia viral oncogene homolog 2), ERBB3 (v-erb-b2 erythroblastic leukemia viral oncogene homolog 3), ERBB4 (v-erb-b2 erythroblastic leukemia viral oncogene homolog 4), Notch1, Notch2, Notch3, or Notch4.

[0117] Examples of proteins associated with secretase disorders can include, for example, PSENEN (presenilin enhancer 2 homolog (C. elegans)), CTSB (cathepsin B), PSEN1 (presenilin 1), APP (amyloid beta (A4) precursor protein), APH1B (anterior pharyngeal defect 1 homolog B (C. elegans)), PSEN2 (presenilin 2 (Alzheimer's disease 4)), or BACE1 (beta-site APP cleaving enzyme 1).

[0118] Examples of proteins associated with amyotrophic lateral sclerosis include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), and any combination thereof.

[0119] Examples of proteins associated with prion diseases can include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C), and any combination thereof.

[0120] Examples of proteins associated with neurodegenerative pathology in prion disorders include, for example, A2M (alpha-2-macroglobulin), AATF (apoptosis antagonistic transcription factor), ACPP (prostatic acid phosphatase), ACTA2 (aortic smooth muscle actin alpha 2), ADAM22 (ADAM metallopeptidase domain), ADORA3 (adenosine A3 receptor), or ADRA1D (alpha-1D adrenergic receptor for alpha-1D adrenoceptor).

[0121] Examples of proteins associated with immunodeficiency include, for example, A2M [alpha-2-macroglobulin]; AANAT [arylalkylamine N-acetyltransferase]; ABCA1 [ATP-binding cassette subfamily A (ABC1), member 1]; ABCA2 [ATP-binding cassette subfamily A (ABC1), member 2]; or ABCA3 [ATP-binding cassette subfamily A (ABC1), member 3].

[0122] Examples of proteins associated with trinucleotide repeat disorders include, for example, AR (androgen receptor), FMR1 (Fragile X Mental Retardation 1), HTT (Huntington's disease), and DMPK (myotonic dystrophy protein kinase), FXN (frataxin), ATX N2 (ataxin 2) is an example.

[0123] Examples of proteins associated with impaired neurotransmission include, for example, SST (somatostatin), NOS1 (nitric oxide synthase 1 (neuronal type)), ADRA2A (adrenergic alpha-2A receptor), ADRA2C (adrenergic alpha-2C receptor), TACR1 (tachykinin receptor 1), or HTR2c (5-hydroxytryptamine (serotonin) receptor 2C).

[0124] Examples of neurodevelopment-related sequences include, for example, A2BP1 [ataxin 2-binding protein 1], AADAT [aminoadipate aminotransferase], AANAT [arylalkylamine N-acetyltransferase], ABAT [4-aminobutyrate aminotransferase], ABCA1 [ATP-binding cassette subfamily A (ABC1) member 1], or ABCA13 [ATP-binding cassette subfamily A (ABC1) member 13].

[0125] Further examples of preferred conditions treatable by the system of the present invention can be selected from the following: Alcardi-Goutières syndrome; Alexander disease; Allan-Herndon-Dudley syndrome; POLG-related disorders; alpha-mannosidosis (types II and III); Alström syndrome; Angelman syndrome; ataxia-telangiectasia; neuronal ceroid lipofuscinosis; beta-thellasamia; bilateral optic atrophy and (infantile) optic atrophy type 1; retinoblastoma (bilateral); Canavan disease; cerebro-ocular-facial-skeletal syndrome 1 [COFS1]; cerebrotendinous xanthomas Cornelia de Lange syndrome; MAPT-related disorders; inherited prion diseases; Dravet syndrome; early-onset familial Alzheimer's disease; Friedreich ataxia [FRDA]; Flins syndrome; fucosidosis; Fukuyama-type congenital muscular dystrophy; galactosialidosis; Gaucher disease; organic acidemia; hemophagocytic lymphohistiocytosis; Hutchinson-Gilford progeria syndrome; mucolipidosis type II; infantile free sialic acid storage disease; PLA2G6-associated neurodegeneration; Jervell-Lange-Nielsen syndrome; junctional epidermolysis bullosa; Huntington's disease; Krabbe disease (infantile form) ; Mitochondrial DNA-related Leigh syndrome and NARP; Lesch-Nyhan syndrome; LIS1-related lissencephaly; Lowe syndrome; Maple syrup urine disease; MECP2 duplication syndrome; ATP7A-related copper transport disorders; LAMA2-related muscular dystrophy; Arylsulfatase A deficiency; Mucopolysaccharidosis type I, II, or III; Peroxisomal biogenesis disorders, Zellweger syndrome spectrum; Neurodegeneration with cerebral iron storage; Acid sphingomyelinase deficiency; Niemann-Pick disease type C; Glycine encephalopathy; ARX-related disorders; Urea cycle disorders; COL1A1 / 2-related diaphyseal dysplasia; mitochondrial DNA deletion syndrome; PLP1-related disorders; Perry syndrome; Phelan-McDermott syndrome; glycogen storage disease type II (Pompe disease) (infantile form); MAPT-related disorders; MECP2-related disorders; rhizomelic chondrodysplasia punctata type 1; Roberts syndrome; Sandhoff disease; Schindler disease type 1; adenosine deaminase deficiency; Smith-Lemli-Opitz syndrome; spinal muscular atrophy; childhood-onset spinocerebellar ataxia; hexosaminidase A deficiency; lethal dysplasia type 1; collagen type VI-related disorders; Usher syndrome type 1; congenital muscular dystrophy;Wolf-Hirschhorn syndrome; lysosomal acid lipase deficiency; and xeroderma pigmentosum.

[0126] It is clear that it is envisioned that the system of the present invention can be used to target any target polynucleotide sequence.Some examples of pathological conditions or diseases that can be usefully treated using the system of the present invention are included in the table above, and examples of genes currently associated with these pathological conditions are also provided in the table.However, the genes exemplified are not exclusive. [Example]

[0127] The following examples are given for the purpose of illustrating various embodiments of the present invention and are not meant to limit the invention in any manner. The examples, together with the methods described herein, are representative and exemplary of presently preferred embodiments and are not intended to limit the scope of the invention. Modifications thereof and other uses encompassed within the spirit of the invention as defined by the scope of the claims will occur to those skilled in the art.

[0128] Example 1: CRISPR complex activity in the nucleus of a eukaryotic cell An exemplary type II CRISPR system is the type II CRISPR locus from Streptococcus pyogenes SF370, which contains a cluster of four genes, Cas9, Cas1, Cas2, and Csn1, and two non-coding RNA elements, tracrRNA and a characteristic array of repeat sequences (direct repeats) spaced by short stretches of non-repetitive sequences (spacers, each approximately 30 bp). In this system, targeted DNA double-strand breaks (DSBs) are generated in four sequential steps (Figure 2A). First, two non-coding RNAs, the pre-crRNA array and tracrRNA, are transcribed from the CRISPR locus. Second, tracrRNA hybridizes to the direct repeats of the pre-crRNA, which are then processed into mature crRNAs containing individual spacer sequences. Third, the mature crRNA:tracrRNA complex directs Cas9 to the DNA target consisting of the protospacer and the corresponding PAM via heteroduplex formation between the spacer region of the crRNA and the protospacer DNA. Finally, Cas9 mediates cleavage of the target DNA upstream of the PAM to create a DSB within the protospacer (Figure 2A). This example describes an exemplary process for adapting this RNA-programmable nuclease system to direct CRISPR complex activity in the nucleus of a eukaryotic cell.

[0129] Cell culture and transfection The human embryonic kidney (HEK) cell line, HEK293FT (Life Technologies), was maintained in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Life Technologies), 100 U / mL penicillin, and 100 μg / mL streptomycin at 37°C with 5% CO2 incubation. The mouse neuro2A (N2A) cell line (ATCC) was maintained in DMEM supplemented with 5% fetal bovine serum (HyClone), 2 mM GlutaMAX (Life Technologies), 100 U / mL penicillin, and 100 μg / mL streptomycin at 37°C with 5% CO2 incubation.

[0130] HEK293FT or N2A cells were seeded into 24-well plates (Corning) at a density of 200,000 cells per well one day before transfection. Cells were transfected using Lipofectamine 2000 (Life Technologies) according to the manufacturer's recommended protocol. A total of 800 ng of plasmid was used for each well of the 24-well plate.

[0131] Surveyor assay and sequencing analysis of genome modifications HEK293FT or N2A cells were transfected with the above-mentioned plasmid DNA. After transfection, the cells were incubated at 37°C for 72 hours before genomic DNA extraction. Genomic DNA was extracted using a QuickExtract DNA extraction kit (Epicentre) according to the manufacturer's protocol. Briefly, cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and 98°C for 10 minutes. The extracted genomic DNA was processed immediately or stored at -20°C.

[0132] Genomic regions surrounding the CRISPR target sites for each gene were PCR-amplified, and the products were purified using QiaQuick Spin Columns (Qiagen) according to the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10× Taq polymerase PCR buffer (Enzymatics) and brought to a final volume of 20 μl with ultrapure water. The product was then subjected to a reannealing process to allow heteroduplex formation: 95°C for 10 minutes, ramped from 95°C to 85°C at -2°C / s, ramped from 85°C to 25°C at -0.25°C / s, and held at 25°C for 1 minute. After reannealing, the products were treated with Surveyor Nuclease and Surveyor Enhancer S (Transgenomics) according to the manufacturer's recommended protocol and analyzed on a 4-20% Novex TBE polyacrylamide gel (Life Technologies). Gels were stained with SYBR Gold DNA stain (Life Technologies) for 30 minutes and imaged using a Gel Doc gel imaging system (Bio-Rad). Quantitation was based on relative band intensity as a measure of the percentage of cleaved DNA. Figure 8 provides a schematic illustration of the Surveyor assay.

[0133] Restriction fragment length polymorphism assay for detection of homologous recombination HEK293FT and N2A cells were transfected with the plasmid DNA and incubated at 37°C for 72 hours, after which genomic DNA was extracted as described above. The target genomic region was amplified by PCR using primers outside the homology arms of the homologous recombination (HR) template. PCR products were separated on a 1% agarose gel and extracted using a MinElute Gel Extraction Kit (Qiagen). Purified products were digested with HindIII (Fermentas) and analyzed on a 6% Novex TBE polyacrylamide gel (Life Technologies).

[0134] RNA secondary structure prediction and analysis RNA secondary structure prediction was performed using the online web server RNAfold, developed at the Institute for Theoretical Chemistry at the University of Vienna, using a centroid structure prediction algorithm (see, e.g., A.R. Gruber et al., 2008, Cell 106(1):23-24; and P.A. Carr and G.M. Church, 2009, Nature Biotechnology 27(12):1151-62).

[0135] Bacterial plasmid transformation interference assay Elements of the S. pyogenes CRISPR locus 1 sufficient for CRISPR activity were reconstituted in Escherichia coli (E. coli) using the pCRISPR plasmid (schematically illustrated in Figure 10A). pCRISPR contained tracrRNA, SpCas9, and a leader sequence driving the crRNA array. A spacer (also referred to as the "guide sequence") was inserted between the BsaI sites into the crRNA array using annealed oligonucleotides as described. The challenge plasmid used in the interference assay was constructed by inserting the protospacer (also referred to as the "target sequence") sequence, along with flanking CRISPR motif sequences (PAM), into pUC19 (see Figure 10B). The challenge plasmid contained ampicillin resistance. Figure 10C provides a schematic representation of the interference assay. Chemically competent E. coli strains already carrying pCRISPR and the appropriate spacer were transformed with a challenge plasmid containing the corresponding protospacer-PAM sequence. pUC19 was used to assess the transformation efficiency of each pCRISPR-carrying competent strain. CRISPR activity resulted in cleavage of the pPSP plasmid carrying the protospacer, eliminating the ampicillin resistance otherwise conferred by pUC19, which lacks the protospacer. Figure 10D illustrates the competence of each pCRISPR-carrying E. coli strain used in the assay described in Figure 4C.

[0136] RNA purification HEK293FT cells were maintained and transfected as described above. Cells were harvested by trypsinization and then washed in phosphate-buffered saline (PBS). Total cellular RNA was extracted with TRI Reagent (Sigma) according to the manufacturer's protocol. Extracted total RNA was quantified using Naonodrop (Thermo Scientific) and normalized to the same concentration.

[0137] Northern blot analysis of crRNA and tracrRNA expression in mammalian cells RNA was mixed with an equal volume of 2x loading buffer (Ambion), heated to 95°C for 5 minutes, chilled on ice for 1 minute, and then loaded onto an 8% denaturing polyacrylamide gel (SequaGel, National Diagnostics) after pre-running the gel for at least 30 minutes. Samples were electrophoresed for 1.5 hours at 40 W limit. RNA was then transferred to Hybond N+ membranes (GE Healthcare) at 300 mA in a semi-dry transfer apparatus (Bio-Rad) for 1.5 hours at room temperature. RNA was crosslinked to the membrane using the autocrosslink button on a Stratagene UV Crosslinker (Stratagene). The membrane was prehybridized in ULTRAhyb-Oligo Hybridization Buffer (Ambion) for 30 minutes at 42°C with rotation, followed by the addition of probe and overnight hybridization. Probes were ordered from IDT and labeled with [γ-P]ATP (Perkin Elmer) using T4 polynucleotide kinase (New England Biolabs). The membranes were washed once for 1 minute with prewarmed (42°C) 2× SSC, 0.5% SDS, followed by two 30-minute washes at 42°C. The membranes were exposed to a phosphor screen at room temperature for 1 hour or overnight and then scanned using a phosphorimager (Typhoon).

[0138] Bacterial CRISPR system construction and evaluation CRISPR locus elements, including tracrRNA, Cas9, and the leader, were PCR-amplified from Streptococcus pyogenes SF370 genomic DNA using flanking homology arms for Gibson assembly. Two BsaI type IIS sites were introduced between the two direct repeats to facilitate easy insertion of the spacer (Figure 9). The PCR product was cloned downstream of the tet promoter into EcoRV-digested pACYC184 using Gibson Assembly Master Mix (NEB). The last 50 bp of Csn2 was removed, excluding other endogenous CRISPR system elements. Oligos encoding the spacer with complementary overhangs (Integrated DNA Technology) were cloned into BsaI-digested vector pDC000 (NEB) and then ligated with T7 ligase (Enzymatics) to generate the pCRISPR plasmid. A challenge plasmid containing a spacer with a PAM sequence (also referred to herein as a "CRISPR motif sequence") was created by ligating hybridized oligos carrying the equivalent overhangs (Integrated DNA Technology) into BamHI-digested pUC19. Cloning for all constructs was performed in E. coli strain JM109 (Zymo Research).

[0139] pCRISPR-carrying cells were made competent using the Z-Competent E. coli Transformation Kit and Buffer Set (Zymo Research, T3001) according to the manufacturer's instructions. For the transformation assay, a 50 μL aliquot of pCRISPR-carrying competent cells was thawed on ice and transformed with 1 ng of spacer plasmid or pUC19 on ice for 30 minutes, then heat-shocked at 42°C for 45 seconds and kept on ice for 2 minutes. Subsequently, 250 μL of SOC (Invitrogen) was added, followed by incubation at 37°C with shaking for 1 hour. 100 μL of the post-SOC preculture was plated onto a double-selection plate (12.5 μg / ml chloramphenicol, 100 μg / ml ampicillin). The total colony count was multiplied by 3 to obtain cfu / 1 ng of DNA.

[0140] To improve expression of CRISPR components in mammalian cells, two genes from Streptococcus pyogenes (S. pyogenes) SF370 locus 1, Cas9 (SpCas9) and RNase III (SpRNase III), were codon-optimized. To facilitate nuclear localization, nuclear localization signals (NLSs) were included at the amino (N)- or carboxyl (C)-termini of both SpCas9 and SpRNase III (Figure 2B). To facilitate visualization of protein expression, fluorescent protein markers were also included at the N- or C-termini of both proteins (Figure 2B). A version of SpCas9 with NLSs attached to both the N- and C-termini (2xNLS-SpCas9) was also generated. Constructs containing NLS-fused SpCas9 and SpRNase III were transfected into 293FT human embryonic kidney (HEK) cells, and it was found that the relative positioning of the NLS to SpCas9 and SpRNase III affected their nuclear localization efficiency. While a C-terminal NLS was sufficient to target SpRNase III to the nucleus, attachment of a single copy of these specific NLSs to either the N- or C-terminus of SpCas9 failed to achieve proper nuclear localization in this system. In this example, the C-terminal NLS was that of nucleoplasmin (KRPAATKKAGQAKKKK), and the C-terminal NLS was that of SV40 large T antigen (PKKKRKV). Of the SpCas9 versions tested, only 2xNLS-SpCas9 exhibited nuclear localization (Figure 2B).

[0141] The tracrRNA from the CRISPR locus of Streptococcus pyogenes (S. pyogenes) SF370 has two transcription start sites, generating two transcripts of 89 nucleotides (nt) and 171 nt, which are subsequently processed into identical 75-nt mature tracrRNAs. The shorter 89-nt tracrRNA was selected for expression in mammalian cells (expression construct illustrated in Figure 7A, with functionality determined by the results of the Surveyor assay shown in Figure 7B). The transcription start site is labeled +1, and the transcription terminator and sequence probed by Northern blot are also shown. Expression of the processed tracrRNA was also confirmed by Northern blot. Figure 7C shows the results of Northern blot analysis of total RNA extracted from 293FT cells transfected with long or short tracrRNA and U6 expression constructs carrying SpCas9 and DR-EMX1(1)-DR. The left and right panels are from 293FT cells transfected without or with SpRNase III, respectively. U6 represents a loading control blotted with a probe targeting human U6 snRNA. Transfection of the short tracrRNA expression construct resulted in sufficient levels of the processed form of tracrRNA (approximately 75 bp). Very little long tracrRNA is detected on Northern blots.

[0142] To promote accurate transcription initiation, an RNA polymerase III-based U6 promoter was selected to drive tracrRNA expression (Figure 2C). Similarly, a U6 promoter-based construct was developed to express a pre-crRNA array consisting of a single spacer flanked by two direct repeats (DRs, also encompassed by the term "tracr-mate sequence"; Figure 2C). The first spacer was designed to target a 33-base pair (bp) target site (a 30-bp protospacer and a 3-bp CRISPR motif (PAM) sequence that fulfills the NGG recognition motif of Cas9) in the human EMX1 locus, a key gene in cerebral cortical development (Figure 2C).

[0143] To test whether heterologous expression of a CRISPR system (SpCas9, SpRNase III, tracrRNA, and pre-crRNA) in mammalian cells can achieve targeted mammalian chromosome cleavage, HEK293FT cells were transfected with a combination of CRISPR components. Because DSBs in mammalian nuclei are repaired, in part, by the non-homologous end joining (NHEJ) pathway, which results in the formation of indels, we used the Surveyor assay to detect potential cleavage activity at the target EMX1 locus (Figure 8) (see, e.g., Guschin et al., 2010, Methods Mol Biol 649:247). Co-transfection of all four CRISPR components could induce cleavage of up to 5.0% of the protospacer (see Figure 2D). Cotransfection of all CRISPR components except SpRNase III also induced indels in up to 4.7% of the protospacer, suggesting the presence of endogenous mammalian RNases, such as the related Dicer and Drosha enzymes, that may assist crRNA maturation. Removal of any of the remaining three components abolished the genome cleavage activity of the CRISPR system (Figure 2D). Sanger sequencing of amplicons containing the target locus confirmed cleavage activity; five mutant alleles (11.6%) were found among 43 sequenced clones. Similar experiments using various guide sequences yielded indel rates as high as 29% (see Figures 4-7, 12, and 13). These results define a three-component system for efficient CRISPR-mediated genome modification in mammalian cells. To optimize cleavage efficiency, we also tested whether different isoforms of tracrRNA affect cleavage efficiency and found that in this exemplary system, only the short (89 bp) transcript form could mediate cleavage of the human EMX1 genomic locus (Figure 7B).

[0144] Figure 14 provides additional Northern blot analysis of crRNA processing in mammalian cells. Figure 14A illustrates a schematic diagram showing an expression vector for a single spacer flanked by two direct repeats (DR-EMX1(1)-DR). The 30-bp spacer targeting the human EMX1 locus protospacer 1 (see Figure 6) and the direct repeat sequence are shown in the lower sequence of Figure 14A. The line indicates the region where the reverse complement sequence was used to generate a Northern blot probe for EMX1(1) crRNA detection. Figure 14B shows Northern blot analysis of total RNA extracted from 293FT cells transfected with a U6 expression construct carrying DR-EMX1(1)-DR. The left and right panels are from 293FT cells transfected without or with SpRNase III, respectively. DR-EMX1(1)-DR was processed to mature crRNA only in the presence of SpCas9, whereas the short tracrRNA was not dependent on the presence of SpRNase III. The mature crRNA detected from transfected 293FT total RNA was approximately 33 bp, shorter than the 39-42 bp mature crRNA from S. pyogenes. These results demonstrate that the CRISPR system can be transplanted into eukaryotic cells and reprogrammed to promote cleavage of endogenous mammalian target polynucleotides.

[0145] Figure 2 illustrates the bacterial CRISPR system described in this example. Figure 2A illustrates a schematic diagram showing CRISPR locus 1 from Streptococcus pyogenes SF370 and the proposed mechanism of CRISPR-mediated DNA cleavage by this system. Mature crRNA processed from the direct repeat-spacer array directs Cas9 to a genomic target consisting of a complementary protospacer and protospacer-adjacent motif (PAM). Upon target-spacer base pairing, Cas9 mediates a double-strand break in the target DNA. Figure 2B illustrates the engineering of S. pyogenes Cas9 (SpCas9) and RNase III (SpRNase III) with a nuclear localization signal (NLS) to enable transport into mammalian nuclei. Figure 2C illustrates mammalian expression of SpCas9 and SpRNase III driven by the constitutive EF1a promoter and tracrRNA and pre-crRNA arrays (DR-spacer-DR) driven by the RNAPol3 promoter U6 to promote accurate transcription initiation and termination. A protospacer from the human EMX1 locus with a sufficient PAM sequence is used as the spacer in the pre-crRNA array. Figure 2D illustrates surveyor nuclease assays for SpCas9-mediated small insertions and deletions. SpCas9 was expressed with or without SpRNase III, tracrRNA, and a pre-crRNA array carrying the EMX1-targeting spacer. Figure 2E illustrates a schematic representation of base pairing between the target locus and the EMX1-targeting crRNA, as well as an example chromatogram showing the microdeletion adjacent to the SpCas9 cleavage site. Figure 2F illustrates mutant alleles identified from sequencing analysis of 43 clonal amplicons exhibiting various microinsertions and deletions. Dotted lines indicate deleted bases, unaligned or mismatched bases indicate insertions or mutations. Scale bar = 10 μm.

[0146] To further simplify the three-component system, we adapted a chimeric crRNA-tracrRNA hybrid design in which the mature crRNA (including the guide sequence) is fused to a partial tracrRNA via a stem-loop, mimicking the natural crRNA:tracrRNA duplex (Figure 3A). To increase co-delivery efficiency, we created a bicistronic expression vector that drives co-expression of the chimeric RNA and SpCas9 in transfected cells (Figures 3A and 8). In parallel, we used a bicistronic vector to express pre-crRNA (DR-guide sequence-DR) ​​together with SpCas9 to direct its processing into crRNA, and to express tracrRNA separately (compare the top and bottom panels in Figure 13B). Figure 9 provides a schematic illustration of a bicistronic expression vector for a pre-crRNA array (Figure 9A) or chimeric crRNA (represented by a short line downstream of the guide sequence insertion site and upstream of the EF1α promoter in Figure 9B) with hSpCas9, showing the location of various elements and the location of guide sequence insertion. The expanded sequence around the location of the guide sequence insertion site in Figure 9B also shows the partial DR sequence (GTTTAGAGCTA) and partial tracrRNA sequence (TAGCAAGTTAAAATAAGGCTAGTCCGTTTTT). Guide sequences can be inserted between the BbsI sites using annealed oligonucleotides. The sequence design for the oligonucleotides is shown below the schematic illustration in Figure 9, and appropriate ligation adapters are indicated. WPRE stands for woodchuck hepatitis virus post-transcriptional regulatory element. The efficiency of chimeric RNA-mediated cleavage was tested by targeting the same EMX1 locus described above. Using both the Surveyor assay and Sanger sequencing of the amplicons, Applicants confirmed that the chimeric RNA design promoted cleavage of the human EMX1 locus with a modification rate of approximately 4.7% (Figure 4).

[0147] The generalizability of CRISPR-mediated cleavage in eukaryotic cells was tested by targeting additional genomic loci in both human and mouse cells by designing chimeric RNAs targeting multiple sites in the human EMX1 and PVALB and mouse Th loci. Figure 15 illustrates the selection of protospacers in several additional targeted human PVALB (Figure 15A) and mouse Th (Figure 15B) loci. A schematic diagram of the location of the three protospacers within the loci and their respective final exons is provided. The underlined sequences include the 30-bp protospacer sequence and 3 bp at the 3' end corresponding to the PAM sequence. The protospacers on the sense and antisense strands are shown above and below the DNA sequences, respectively. Modification rates of 6.3% and 0.75% were achieved for the human PVALB and mouse Th loci, respectively, demonstrating the broad applicability of the CRISPR system in modifying different loci across multiple organisms (Figures 3B and 6). While cleavage was detected for only one of the three spacers for each locus using the chimeric construct, when using the co-expressed pre-crRNA configuration, all target sequences were cleaved with an efficiency of indel generation reaching 27% (Figure 6).

[0148] Figure 13 provides further illustration of the ability of SpCas9 to reprogram and target multiple genomic loci in mammalian cells. Figure 13A provides a schematic diagram of the human EMX1 locus, showing the location of the five protospacers indicated by the underlined sequences. Figure 13B provides a schematic diagram of the pre-crRNA / trcrRNA complex (top panel) showing hybridization between the direct repeat regions of the pre-crRNA and tracrRNA, and a schematic diagram of the chimeric RNA design (bottom panel) containing a 20-bp guide sequence and a tracrmate and tracr sequence consisting of a partial direct repeat and tracrRNA sequence hybridized to a hairpin structure. Figure 13C illustrates the results of a Surveyor assay comparing the efficacy of Cas9-mediated cleavage at five protospacers in the human EMX1 locus. Either the processed pre-crRNA / tracrRNA complex (crRNA) or a chimeric RNA (chiRNA) is used to target each protospacer.

[0149] Because RNA secondary structure can be important for intermolecular interactions, we used a structure prediction algorithm based on minimum free energy and Boltzmann-weighted structural ensembles to compare the predicted secondary structures of all guide sequences used in our genome targeting experiments (Figure 3B) (see, e.g., Gruber et al., 2008, Nucleic Acids Research, 36:W70). The analysis revealed that, in most cases, effective guide sequences in the chimeric crRNA context were substantially free of secondary structure motifs, while ineffective guide sequences were more likely to form internal secondary structures that could interfere with base pairing with the target protospacer DNA. Therefore, it is conceivable that variability in spacer secondary structure may affect the efficiency of CRISPR-mediated interference when using chimeric crRNAs.

[0150] Figure 3 illustrates exemplary expression vectors. Figure 3A provides a schematic diagram of a synthetic crRNA-tracrRNA chimera (chimeric RNA) and a bicistronic vector for driving expression of SpCas9. The chimeric guide RNA contains a 20-bp guide sequence corresponding to the protospacer in the genomic target site. Figure 3B provides a schematic diagram showing guide sequences targeting the human EMX1, PVALB, and mouse Th loci, as well as their predicted secondary structures. The modification efficiency at each target site is shown below the RNA secondary structure diagram (EMX1, n = 216 amplicon sequencing reads; PVALB, n = 224 reads; Th, n = 265 reads). The folding algorithm produced output colored according to the probability that each base would assume the predicted secondary structure, as indicated by the rainbow scale reproduced in grayscale in Figure 3B. Further vector designs for SpCas9 are shown in Figure 44, which describes a single expression vector incorporating a U6 promoter linked to an insertion site for a guide oligo and a Cbh promoter linked to the SpCas9 coding sequence. The vector shown in Figure 44b contains the tracrRNA coding sequence linked to the H1 promoter.

[0151] To test whether spacers containing secondary structures can function in prokaryotic cells where CRISPR naturally operates, we tested the transformation interference of protospacer-bearing plasmids in an E. coli strain heterologously expressing the S. pyogenes SF370 CRISPR locus 1 (Figure 10). The CRISPR locus was cloned into a low-copy E. coli expression vector, and the crRNA array was replaced with a single spacer flanked by a pair of DRs (pCRISPR). E. coli strains carrying different pCRISPR plasmids were transformed with challenge plasmids containing the corresponding protospacer and PAM sequences (Figure 10C). In bacterial assays, all spacers promoted efficient CRISPR interference (Figure 4C). These results suggest that there may be additional factors that affect the efficiency of CRISPR activity in mammalian cells.

[0152] To investigate the specificity of CRISPR-mediated cleavage, we analyzed the effect of single-nucleotide mutations in the guide sequence on protospacer cleavage in mammalian genomes using a series of EMX1-targeting chimeric crRNAs with single point mutations (Figure 4A). Figure 4B illustrates the results of a Surveyor nuclease assay comparing the cleavage efficiency of Cas9 when paired with different mutant chimeric RNAs. A single-base mismatch up to 12 bp 5′ to the PAM essentially abolished genome cleavage by SpCas9, whereas spacers with mutations further upstream retained activity against the original protospacer target (Figure 4B). In addition to the PAM, SpCas9 has single-base specificity within the last 12 bp of the spacer. Furthermore, CRISPR can mediate genome cleavage as efficiently as a pair of TALE nucleases (TALENs) targeting the same EMX1 protospacer. Figure 4C provides a schematic showing the design of TALENs targeting EMX1, and Figure 4D shows a Surveyor gel comparing the efficiency of TALENs and Cas9 (n=3).

[0153] Having established a set of components for achieving CRISPR-mediated gene editing in mammalian cells through the error-prone NHEJ mechanism, we tested the ability of CRISPR to stimulate homologous recombination (HR), a high-fidelity gene repair pathway, to create precise edits in the genome. Wild-type SpCas9 can mediate site-specific DSBs that can be repaired through both NHEJ and HR. Furthermore, we engineered an aspartate-to-alanine substitution (D10A) in the RuvC I catalytic domain of SpCas9 to convert the nuclease into a nickase (SpCas9n; illustrated in Figure 5A) (see, e.g., Sapranauskas et al., 2011, Nucleic Acids Research, 39:9275; Gasiunas et al., 2012, Proc. Natl. Acad. Sci. USA, 109:E2579), allowing nicked genomic DNA to undergo high-fidelity homology-directed repair (HDR). Surveyor assays confirmed that SpCas9n did not generate indels in the EMX1 protospacer target. As illustrated in Figure 5B, coexpression of an EMX1-targeting chimeric crRNA with SpCas9 generated indels in the target site, whereas coexpression with SpCas9n did not (n = 3). Furthermore, sequencing of 327 amplicons did not detect any indels induced by SpCas9n. The same locus was selected and CRISPR-mediated HR was tested by cotransfecting HEK293FT cells with a chimeric RNA targeting EMX1, hSpCas9, or hSpCas9n, and an HR template to introduce a pair of restriction sites (HindIII and NheI) near the protospacer. Figure 5C provides a schematic illustration of the HR strategy, along with the relative locations of the recombination sites and primer annealing sequences (arrows). SpCas9 and SpCas9n indeed catalyzed the integration of the HR template into the EMX1 gene.PCR amplification of the target region followed by restriction digestion with HindIII revealed cleavage products corresponding to the predicted fragment sizes (arrows in the restriction fragment length polymorphism gel analysis shown in Figure 5D), indicating that SpCas9 and SpCas9n mediated similar levels of HR efficiency. Applicants further confirmed HR using Sanger sequencing of the genomic amplicon (Figure 5E). These results demonstrate the utility of CRISPR for facilitating targeted gene insertion in mammalian genomes. Given the 14-bp target specificity of wild-type SpCas9 (12 bp from the spacer and 2 bp from the PAM), the availability of a nickase may significantly reduce the potential for off-target modifications, as single-stranded fragments are not substrates for the error-prone NHEJ pathway.

[0154] We constructed an expression construct (Figure 2A) that mimics the natural architecture of CRISPR loci with array spacers to test the feasibility of multiplexed sequence targeting. Using a single CRISPR array encoding a pair of EMX1 and PVALB targeting spacers, we detected efficient cleavage at both loci (Figure 4F, which shows both the schematic design of the crRNA array and a Surveyor blot demonstrating efficient mediation of cleavage). We also tested targeted deletion of a larger genomic region through simultaneous DSBs using spacers for two targets in EMX1, spaced 119 bp apart, and detected a deletion efficiency of 1.6% (3 out of 182 amplicons; Figure 4G). This demonstrates that the CRISPR system can mediate multiplexed editing within a single genome.

[0155] Example 2: CRISPR-based modifications and alternatives The ability to use RNA to program sequence-specific DNA cleavage defines a new class of genome engineering tools for a variety of research and industrial applications. Some aspects of the CRISPR system can be further improved to increase the efficiency and versatility of CRISPR targeting. Optimal Cas9 activity may depend on the availability of free Mg2+ at levels higher than those present in mammalian nuclei (see, e.g., Jinek et al., 2012, Science, 337:816), and the preference for NGG motifs immediately downstream of the protospacer limits targeting ability to an average of every 12 bp in the human genome (Figure 11, evaluating both the plus and minus strands of human chromosomal sequences). Some of these constraints can be overcome by taking advantage of the diversity of CRISPR loci across microbial metagenomes (see, e.g., Makarova et al., 2011, Nat Rev Microbiol, 9:467). Other CRISPR loci can be transplanted into mammalian cell environments using methods similar to those described in Example 1. For example, Figure 12 illustrates the adaptation of a type II CRISPR system from CRISPR1 of Streptococcus thermophilus LMD-9 for heterologous expression in mammalian cells to achieve CRISPR-mediated genome editing. Figure 12A provides a schematic illustration of CRISPR1 of S. thermophilus LMD-9. Figure 12B illustrates the design of the expression system for the S. thermophilus CRISPR system. Human codon-optimized hStCas9 is expressed using a constitutive EF1α promoter. Mature versions of tracrRNA and crRNA are expressed using a U6 promoter to promote accurate transcription initiation. Sequences from the mature crRNA and tracrRNA are illustrated. A single base, indicated by a lowercase "a" in the crRNA sequence, is used to remove the polyU sequence that functions as an RNApol III transcription terminator. Figure 12C provides a schematic illustrating guide sequences targeting the human EMX1 locus and their predicted secondary structures.The modification efficiency at each target site is shown below the RNA secondary structure. The algorithm that generated this structure colors each base according to its probability of assuming the predicted secondary structure, shown in rainbow scales reproduced in grayscale in Figure 12C. Figure 12D shows the results of hStCas9-mediated cleavage in the target locus using the Surveyor assay. RNA-guided spacers 1 and 2 induced 14% and 6.4%, respectively. Statistical analysis of cleavage activity across biological replicates at these two protospacer sites is also provided in Figure 6. Figure 16 provides a schematic diagram of additional protospacers and corresponding PAM sequence targets of the S. thermophilus CRISPR system in the human EMX1 locus. Two protospacer sequences are highlighted, and their corresponding PAM sequences, which fulfill the NNAGAAW motif, are underlined 3' to the corresponding highlighted sequence. Both protospacers target the antisense strand.

[0156] Example 3: Sample target sequence selection algorithm Design a software program to identify candidate CRISPR target sequences on both strands of the input DNA sequence based on the desired guide sequence length and CRISPR motif sequence (PAM) for a given CRISPR enzyme. For example, target sites for Cas9 from Streptococcus pyogenes (S. pyogenes) can be identified by using the PAM sequence NGG to probe for 5'-Nx-NGG-3' on both the input sequence and the reverse complementary strand of the input. Similarly, target sites for Cas9 from S. thermophilus (S. thermophilus) can be identified by using the PAM sequence NNAGAAW to probe for 5'-Nx-NNAGAAW-3' on both the input sequence and the reverse complementary strand of the input. Similarly, target sites for Cas9 in S. thermophilus CRISPR3 can be identified by searching for 5'-Nx-NGGNG-3' on both the input sequence and the reverse complement of the input using the PAM sequence NGGNG. The value "x" in Nx can be fixed by the program or user-defined, e.g., 20.

[0157] Because multiple occurrences of a DNA target site in a genome can result in nonspecific genome editing, after identifying all potential sites, the program filters out sequences based on the number of times the sequence appears in the relevant reference genome. For those CRISPR enzymes whose sequence specificity is determined by a "seed" sequence, e.g., the 11-12 bp 5' from the PAM sequence, including the PAM sequence itself, the filtering step can be based on the seed sequence. Thus, to avoid editing at additional genomic loci, results are filtered based on the number of occurrences of the seed:PAM sequence in the relevant genome. The user can select the length of the seed sequence. The user can also specify the number of occurrences of the seed:PAM sequence in the genome for filtering purposes. The default is to screen for unique sequences. The filtration level can be varied by changing both the length of the seed sequence and the number of occurrences of the sequence in the genome. The program can also, or alternatively, provide a guide sequence complementary to the reported target sequence by providing the reverse complement of the identified target sequence.

[0158] Further details of methods and algorithms for optimizing sequence selection can be found in US Patent Application No. 61 / 836,080 (Attorney Docket No. 44790.11.2022), which is incorporated herein by reference.

[0159] Example 4: Evaluation of multiple chimeric crRNA-tracrRNA hybrids This example describes results obtained with chimeric RNAs (chiRNAs; containing guide, tracr mate, and tracr sequences in a single transcript) incorporating different lengths of wild-type tracrRNA sequences. Figure 18a illustrates a schematic diagram of the bicistronic expression vector for chimeric RNA and Cas9. Cas9 is driven by the CBh promoter, and the chimeric RNA is driven by the U6 promoter. The chimeric guide RNAs consist of a 20-bp guide sequence (N) linked to truncated tracr sequences (spanning from the first "U" on the lower strand to the end of the transcript) at various positions as indicated. The guide and tracr sequences are separated by the tracr mate sequence GUUUUAGAGCUA, followed by the loop sequence GAAA. The results of the SURVEYOR assay for Cas9-mediated indels at the human EMX1 and PVALB loci are illustrated in Figures 18b and 18c, respectively. Arrows indicate predicted SURVEYOR fragments. ChiRNAs are indicated by their "+n" designation, and crRNA refers to a hybrid RNA in which the guide and tracr sequences are expressed as separate transcripts. Quantification of these results, performed in triplicate, is shown by histograms in Figures 19a and 19b, corresponding to Figures 18b and 18c, respectively ("ND" indicates no indels detected). Protospacer IDs and their corresponding genomic targets, protospacer sequences, PAM sequences, and strand localizations are provided in Table D. Guide sequences were designed to be complementary to the entire protospacer sequence in the case of separate transcripts in the hybrid system, or to only the underlined portion in the case of chimeric RNAs.

[0160] [Table 18]

[0161] Cell culture and transfection Human embryonic kidney (HEK) cell line 293FT (Life Technologies) was maintained in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Life Technologies), 100 U / mL penicillin, and 100 μg / mL streptomycin at 37°C with 5% CO2 incubation. 293FT cells were seeded onto 24-well plates (Corning) at a density of 150,000 cells per well 24 hours prior to transfection. Cells were transfected using Lipofectamine 2000 (Life Technologies) according to the manufacturer's recommended protocol. A total of 500 ng of plasmid was used for each well of the 24-well plate.

[0162] SURVEYOR assay for genome modifications 293FT cells were transfected with the above plasmid DNA. Cells were incubated at 37°C for 72 hours after transfection before genomic DNA extraction. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. Briefly, pelleted cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and 98°C for 10 minutes. Genomic regions flanking the CRISPR target site for each gene were PCR amplified (primers listed in Table E), and the products were purified using QiaQuick Spin Columns (Qiagen) according to the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10× Taq DNA Polymerase PCR buffer (Enzymatics), brought to a final volume of 20 μl with ultrapure water, and subjected to a reannealing process to allow heteroduplex formation: 95°C for 10 min, ramping from 95°C to 85°C at -2°C / sec, from 85°C to 25°C at -0.25°C / sec, and holding at 25°C for 1 min. After reannealing, the product was treated with SURVEYOR nuclease and SURVEYOR enhancer S (Transgenomics) according to the manufacturer's recommended protocol. Analysis was performed on a 4-20% Novex TBE polyacrylamide gel (Life Technologies). Gels were stained with SYBR Gold DNA stain (Life Technologies) for 30 minutes and imaged using a Gel Doc gel imaging system (Bio-Rad). Quantitation was based on relative band intensity.

[0163] [Table 19]

[0164] Computational identification of unique CRISPR target sites To identify unique target sites for the Streptococcus pyogenes (S. pyogenes) SF370Cas9 (SpCas9) enzyme in the genomes of humans, mice, rats, zebrafish, fruit flies, and Caenorhabditis elegans (C. elegans), we developed a software package to scan both strands of a DNA sequence and identify all possible SpCas9 target sites. For this example, we operationally defined each SpCas9 target site as a 20-bp sequence followed by an NGG protospacer adjacent motif (PAM) sequence, and we identified all sequences that met this 5'-N20-NGG-3' definition on all chromosomes. To prevent nonspecific genome editing, after identifying all potential sites, we filtered all target sites based on the number of times they appeared in the relevant reference genome. To take advantage of the sequence specificity of Cas9 activity conferred by a "seed" sequence, which can be approximately 11-12 bp 5' from the PAM sequence, for example, we selected the 5'-NNNNNNNNNN-NGG-3' sequence as unique within the relevant genome. All genome sequences were downloaded from the UCSC Genome Browser (human genome hg19, mouse genome mm9, rat genome rn5, zebrafish genome danRer7, D. melanogaster genome dm4, and C. elegans genome ce10). All search results are available for viewing using the UCSC Genome Browser. An exemplary visualization of some target sites in the human genome is provided in Figure 21.

[0165] We first targeted three sites within the EMX1 locus in human HEK293FT cells. The genome modification efficiency of each chiRNA was assessed using the SURVEYOR nuclease assay, which detects mutations resulting from DNA double-strand breaks (DSBs) and their subsequent repair by the non-homologous end joining (NHEJ) DNA damage repair pathway. Constructs designated chiRNA(+n) indicate that up to +n nucleotides of wild-type tracrRNA are included in the chimeric RNA construct, with values ​​of 48, 54, 67, and 85 used for n. Chimeric RNAs containing longer fragments of wild-type tracrRNA (chiRNA(+67) and chiRNA(+85)) mediated DNA cleavage at all three EMX1 target sites, with chiRNA(+85) in particular demonstrating significantly higher levels of DNA cleavage than the corresponding crRNA / tracrRNA hybrids expressing the guide and tracr sequences in separate transcripts (Figures 18b and 19a). Two sites in the PVALB locus that did not produce detectable cleavage in the hybrid system (guide and tracr sequences expressed as separate transcripts) were also targeted using chiRNAs. chiRNA(+67) and chiRNA(+85) were able to mediate significant cleavage in the two PVALB protospacers (Figures 18c and 19b).

[0166] For all five targets in the EMX1 and PVALB loci, a consistent increase in genome modification efficiency was observed with increasing tracr sequence length. Without being bound by any theory, the secondary structure formed by the 3' end of tracrRNA may play a role in improving the rate of CRISPR complex formation. Figure 21 provides an illustration of the predicted secondary structure for each chimeric RNA used in this example. The secondary structure was predicted using RNAfold (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAfold.cgi), which uses a minimum free energy and partition function algorithm. The pseudocolor (reproduced in grayscale) for each base indicates the probability of pairing. It is thought that chimeric RNAs with longer tracr sequences could cleave targets that are not cleaved by the natural CRISPRcrRNA / tracrRNA hybrid, allowing chimeric RNAs to be loaded onto Cas9 more efficiently than their natural hybrid counterparts. To facilitate the application of Cas9 for site-specific genome editing in eukaryotic cells and organisms, all predicted unique target sites for Streptococcus pyogenes (S. pyogenes) Cas9 were computationally identified in the human, mouse, rat, zebrafish, C. elegans, and D. melanogaster genomes. Chimeric RNAs can be designed for Cas9 enzymes from other microorganisms to expand the target space of CRISPR RNA-programmable nucleases.

[0167] Figure 22 shows the wild-type tracrRNA sequence up to +85 nucleotides and the nuclear localization sequence. Illustrative bicistronic sequences for the expression of chimeric RNAs containing SpCas9 The expression vector is described. SpCas9 is expressed from the CBh promoter and the bGH promoter. The expanded sequence, illustrated directly below the schematic, is It corresponds to the region surrounding the guide sequence insertion site and includes, from 5' to 3', the 3' portion of the U6 promoter (first shaded area), a BbsI cleavage site (arrow), a partial direct repeat (tracr mate sequence GTTTTAGAGCTA, underlined), a loop sequence GAAA, and the +85 tracr sequence (underlined sequence after the loop sequence). Exemplary guide sequence inserts are illustrated below the guide sequence insertion site, with the guide sequence nucleotide for the selected target represented by "N."

[0168] The sequences described in the examples above are as follows (polynucleotide sequences are from 5' to 3'): U6-short tracrRNA (Streptococcus pyogenes SF370): [ka] U6-long tracrRNA (Streptococcus pyogenes SF370): [ka] U6-DR-BbsI backbone-DR (Streptococcus pyogenes SF370): [ka] U6-chimeric RNA-BbsI backbone (Streptococcus pyogenes SF370) [ka] NLS-SpCas9-EGFP: [ka] SpCas9-EGFP-NLS: [ka] NLS-SpCas9-EGFP-NLS: [ka] NLS-SpCas9-NLS: [ka] NLS-mCherry-SpRNase3: [ka] SpRNase3-mCherry-NLS: [ka] NLS-SpCas9n-NLS (D10A nickase mutation is in lowercase): [ka] hEMX1-HR template-HindII-NheI: [ka] [ka] NLS-StCsn1-NLS: [ka] U6-St_tracrRNA(7-97): [ka] U6-DR-spacer-DR (Streptococcus pyogenes SF370) [ka] Chimeric RNA containing +48 tracrRNA (S. pyogenes SF370) [ka] Chimeric RNA containing +54 tracrRNA (S. pyogenes SF370) [ka] Chimeric RNA containing +67 tracrRNA (S. pyogenes SF370) [ka] Chimeric RNA containing +85 tracrRNA (S. pyogenes SF370) [ka] CBh-NLS-SpCas9-NLS [ka] [ka] [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR1Cas9 (with PAM of NNAGAAW) [ka] Exemplary chimeric RNA for S. thermophilus LMD-9 CRISPR3Cas9 (with PAM of NGGNG) [ka] A codon-optimized version of Cas9 from the S. thermophilus LMD-9 CRISPR3 locus (with NLS at both the 5' and 3' ends) [ka] [ka] [ka] [ka]

[0169] Example 5: RNA-guided editing of bacterial genomes using CRISPR-Cas systems The present applicants used the CRISPR-associated endonuclease Cas9 to introduce precise mutations into the genomes of Streptococcus pneumoniae and Escherichia coli. This approach relied on Cas9-directed cleavage at targeted sites to kill non-mutated cells, avoiding the need for selectable markers or counterselection systems. Cas9 specificity was reprogrammed by altering the short CRISPR RNA (crRNA) sequence to generate single and multi-nucleotide changes carried on the editing template. The simultaneous use of two crRNAs enabled the introduction of multiple mutations. In S. pneumoniae, nearly 100% of cells surviving Cas9 cleavage contained the desired mutations, and when used in combination with recombineering in E. coli, 65% contained the desired mutations. Applicants have thoroughly analyzed the Cas9 target requirements to define the range of targetable sequences and provided strategies for editing sites that do not meet those requirements, suggesting the versatility of this technology for bacterial genome engineering.

[0170] Understanding gene function relies on the ability to alter DNA sequences within cells in a controlled manner. Site-specific mutagenesis in eukaryotes is achieved through the use of sequence-specific nucleases that promote homologous recombination of template DNA containing the desired mutation. Zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and homing meganucleases can be programmed to cleave the genome at defined locations, but these approaches require the engineering of new enzymes for each target sequence. In prokaryotes, mutagenesis requires a two-step process that involves introducing a selectable marker into the edited locus or a counterselection system. Recently, phage recombination proteins, a technology that promotes homologous recombination of linear DNA or oligonucleotides, have been used for recombineering. However, due to the lack of selection for mutations, recombineering efficiency can be relatively low (0.1–10% for point mutations, down to 10–5–10–6 for larger modifications) and often requires the screening of a large number of colonies. Therefore, new techniques that are inexpensive, easy to use, and efficient remain needed for genetic engineering of both eukaryotic and prokaryotic organisms.

[0171] Recent studies of the prokaryotic CRISPR (clustered regularly interspaced short batch repeats) adaptive immune system have led to the identification of nucleases whose sequence specificity is programmed by small RNAs. CRISPR loci consist of a series of repeats separated by "spacer" sequences that match the genomes of bacteriophages and other mobile genetic elements. The repeat-spacer array is transcribed as a long precursor and processed within the repeat sequence to generate small crRNAs that define the target sequence (also known as the protospacer) that is cleaved by the CRISPR system. The presence of a sequence motif immediately downstream of the target region, known as the protospacer adjacent motif (PAM), is essential for cleavage. CRISPR-associated (cas) genes typically flank the repeat-spacer array and encode the enzymatic machinery responsible for crRNA biogenesis and targeting. Cas9 is a dsDNA endonuclease that uses the crRNA guide to define the site of cleavage. Loading of the crRNA guide onto Cas9 occurs during processing of the crRNA precursor and requires a small RNA antisense to the precursor, tracrRNA, and RNase III. In contrast to genome editing with ZFNs or TALENs, altering Cas9 target specificity does not require protein engineering, but rather only the design of a short crRNA guide.

[0172] Applicants recently demonstrated that introduction of a CRISPR system targeting a chromosomal locus in Streptococcus pneumoniae (S. pneumoniae) resulted in the killing of transformed cells. Occasional survivors were observed to contain mutations in the targeted region, suggesting that the Cas9 dsDNA endonuclease activity against endogenous targets could be used for genome editing. Applicants demonstrated that markerless mutations can be introduced through transformation of a template DNA fragment that recombines into the genome and eliminates Cas9 target recognition. Directing the specificity of Cas9 with several different crRNAs allows for the simultaneous introduction of multiple mutations. Applicants also characterized the sequence requirements for Cas9 targeting in detail and demonstrated that this approach can be combined with recombineering for genome editing in Escherichia coli (E. coli).

[0173] Results: Genome editing by Cas9 cleavage of chromosomal targets The S. pneumoniae strain crR6 contains a Cas9-based CRISPR system that cleaves a target sequence present in the bacteriophage φ8232.5. This target was integrated into the srtA chromosomal locus of a second strain, R68232.5. An altered target sequence containing a mutation in the PAM region was integrated into the srtA locus of a third strain, R6370.1, rendering this strain "immune" to CRISPR cleavage (Figure 28a). Applicants transformed R68232.5 and R6370.1 cells with genomic DNA from crR6 cells and predicted that successful transformation of R68232.5 cells should result in cleavage of the target locus and cell death. Contrary to this prediction, applicants isolated R68232.5 transformants, albeit with approximately 10-fold lower efficiency than R6370.1 transformants (Figure 28b). Genetic analysis of eight R68232.5 transformants (Figure 28) revealed that the majority were the product of a double recombination event that eliminated the toxicity of Cas9 targeting by replacing the φ8232.5 target with the crR6 genomic wild-type srtA locus, which does not contain the protospacer required for Cas9 recognition. These results demonstrated that co-introduction of a CRISPR system targeting a genomic locus (targeting construct) together with a template for recombination into the targeted locus (editing template) resulted in targeted genome editing (Figure 23a).

[0174] To create a simplified system for genome editing, we modified the CRISPR locus in strain crR6 by deleting the cas1, cas2, and csn2 genes, which have been shown to be non-essential for CRISPR targeting, resulting in strain crR6M (Figure 28a). This strain retained the same properties of crR6 (Figure 28b). To demonstrate that the efficiency of Cas9-based editing can be increased and the introduced mutations can be controlled using an optimal DNA template, we cotransformed R68232.5 cells with PCR products of the wild-type srtA gene or the mutant R6370.1 target, either of which should be resistant to cleavage by Cas9. This resulted in a 5- to 10-fold increase in transformation frequency compared to genomic crR6 DNA alone (Figure 23b). The efficiency of editing also increased substantially, with 8 of 8 transformants tested containing wild-type srtA copies and 7 of 8 containing the PAM mutation present in the R6370.1 target (Figures 23b and 29a). Together, these results demonstrated the potential of Cas9-assisted genome editing.

[0175] Analysis of Cas9 target requirements: To introduce a specific change in the genome, an editing template must be used that carries a mutation that blocks Cas9-mediated cleavage, thereby preventing cell death. This is easy to achieve when target deletion or its replacement with another sequence (gene insertion) is desired. When the goal is to produce a gene fusion or a single nucleotide mutation, Cas9 nuclease activity can only be stopped by introducing a mutation in the editing template that changes either the PAM or protospacer sequence. To determine the constraints of CRISPR-mediated editing, the applicants performed a thorough analysis of PAM and protospacer mutations that block CRISPR targeting.

[0176] Previous studies have proposed that S. pyogenes Cas9 requires an NGG PAM immediately downstream of the protospacer. However, because only a very limited number of PAM-inactivating mutations have been described, we conducted a systematic analysis to find all five-nucleotide sequences after the protospacer that would eliminate CRISPR cleavage. We used randomized oligonucleotides to generate all 1,024 possible PAM sequences in heterologous PCR products transformed into crR6 or R6 cells. Constructs carrying a functional PAM were predicted to be recognized and destroyed by Cas9 in crR6 cells but not in R6 cells (Figure 24a). More than 2 x 10 colonies were pooled together to extract DNA, which was used as a template for simultaneous amplification of all targets. The PCR products were deep sequenced and found to contain all 1,024 sequences, with coverage ranging from 5 to 42,472 reads (see the section "Analysis of Deep Sequencing Data"). The functionality of each PAM was estimated by the relative proportion of its reads in the crR6 sample compared to the R6 sample. Analysis of the first three bases of the PAM, averaging the two final bases, clearly showed that the NGG pattern was underrepresented in the crR6 transformants (Figure 24b). Furthermore, the next two bases had no detectable effect on the NGG PAM (see section "Analysis of Deep Sequencing Data"), demonstrating that the NGGNN sequence was sufficient to permit Cas9 activity. Partial targeting was observed for the NAG PAM sequence (Figure 24b). The NNGGN pattern also partially inactivated CRISPR targeting (Table G), indicating that the NGG motif could still be recognized by Cas9 with reduced efficiency when shifted by 1 bp. These data shed light on the molecular mechanism of Cas9 target recognition and revealed that an NGG (or CCN on the complementary strand) sequence is sufficient for Cas9 targeting and that mutations from NGG to NAG or NNGGN in the editing template should be avoided.Due to the high frequency of these trinucleotide sequences (once every 8 bp), this means that almost any location in the genome can be edited. Indeed, we tested 10 randomly selected targets bearing different PAMs and found all to be functional (Figure 30).

[0177] Another approach to disrupting Cas9-mediated cleavage is to introduce mutations in the protospacer region of the editing template. It is known that point mutations within the "seed sequence" (the 8 to 10 protospacer nucleotides immediately adjacent to the PAM) can abolish cleavage by CRISPR nucleases. However, the exact length of this region is unknown, making it unclear which nucleotide mutations in the seed could disrupt Cas9 target recognition. Following the same deep sequencing approach described above, we randomized the entire protospacer sequence involved in base-pairing contacts with crRNA and determined all sequences that disrupted targeting. We randomized the position of each of the 20 matching nucleotides (14) in the spc1 target present in R68232.5 cells (Figure 23a) and transformed them into crR6 and R6 cells (Figure 24a). Consistent with the presence of a seed sequence, only mutations of the 12 nucleotides immediately upstream of the PAM abolished Cas9 cleavage (Figure 24c). However, different mutations had significantly different effects. The distal (from PAM) positions of the seed (positions 12 to 7) tolerated most mutations, with only one specific base substitution abolishing targeting. In contrast, mutation of any nucleotide at proximal positions (positions 6 to 1, excluding position 3) eliminated Cas9 activity, albeit at different levels for each specific substitution. At position 3, only two substitutions affected CRISPR activity, with different strengths. Applicants concluded that seed sequence mutations can interfere with CRISPR targeting, but there are limitations regarding the nucleotide changes that can be made at each position of the seed. Furthermore, these limitations are most likely to vary with different spacer sequences. Therefore, applicants believe that mutations in the PAM sequence should be the preferred editing strategy, if possible. Alternatively, multiple mutations in the seed sequence can be introduced to interfere with Cas9 nuclease activity.

[0178] Cas9-Mediated Genome Editing in S. pneumoniae: To develop a rapid and efficient method for targeted genome editing, we engineered strain crR6Rk, a strain into which spacers can be easily introduced by PCR (Figure 33). We decided to edit the β-galactosidase (bgaA) gene of S. pneumoniae, whose activity can be easily measured. We introduced alanine substitutions of amino acids in the active site of this enzyme: R481A (R→A) and N563A, E564A (NE→AA) mutations. To account for different editing strategies, we designed mutations in both the PAM sequence and the protospacer seed. In both cases, we used the same targeting construct with a crRNA complementary to the region of the β-galactosidase gene adjacent to the TGG PAM sequence (CCA in the complementary strand, Figure 26). The R→A editing template created a three-nucleotide mismatch in the protospacer seed sequence (CGT to GCA, also introducing a BtgZI restriction site). In the NE→AA editing template, we simultaneously introduced a synonymous mutation (TGG to TTG) creating an inactive PAM, along with a mutation 218 nt downstream of the protospacer region (AAT GAA to GCT GCA, also generating a TseI restriction site). This last editing strategy demonstrated the feasibility of using distant PAMs to create mutations in locations where it may be difficult to select appropriate targets. For example, the S. pneumoniae R6 genome has a GC content of 39.7% and contains an average of one PAM motif every 12 bp, although some PAM motifs are separated by up to 194 bp (Figure 33). Additionally, we engineered a 6,664-bp ΔbgaA in-frame deletion. In all three cases, co-transformation of the targeting and editing templates produced cells that were 10-fold more kanamycin-resistant than co-transformation with a control editing template containing the wild-type bgaA sequence (Figure 25b). Applicants genotyped 24 transformants (8 for each editing experiment) and found that all but one incorporated the desired changes (Figure 25c).DNA sequencing also confirmed the presence of the introduced mutations as well as the absence of secondary mutations in the target region (Figure 29b, c). Finally, we confirmed that all edited cells exhibited the expected phenotype by measuring β-galactosidase activity (Figure 25d).

[0179] Cas9-mediated editing can also be used to generate multiple mutations for biological pathway studies. We decided to demonstrate this for a sortase-dependent pathway that anchors surface proteins to the envelope of Gram-positive bacteria. We introduced a sortase deletion by cotransformation of a chloramphenicol-resistance targeting construct and a ΔsrtA editing template (Figures 33a, b), followed by a ΔbgaA deletion using a kanamycin-resistance targeting construct to replace ΔbgaA. In S. pneumoniae, β-galactosidase is covalently linked to the cell wall by sortase. Thus, deletion of srtA results in the release of surface proteins into the supernatant, while the double deletion has no detectable β-galactosidase activity (Figure 34c). Such sequential selections can be repeated as many times as necessary to generate multiple mutations.

[0180] These two mutations can also be introduced simultaneously. We designed a targeting construct containing two spacers, one matching srtA and the other matching bgaA, and co-transformed it with both editing templates (Fig. 25e). Genetic analysis of the transformants showed that editing occurred in six out of eight cases (Fig. 25f). Notably, the remaining two clones contained either the ΔsrtA or ΔbgaA deletion, respectively, suggesting the possibility of performing combinatorial mutagenesis using Cas9. Finally, to eliminate the CRISPR sequence, we introduced a plasmid containing the bgaA target and a spectinomycin resistance gene together with genomic DNA from the wild-type strain R6. Spectinomycin-resistant transformants harboring the plasmid eliminated the CRISPR sequence (Fig. 34a, d).

[0181] Editing mechanism and efficiency: To understand the mechanism underlying genome editing by Cas9, we designed an experiment to measure editing efficiency independently of Cas9 cleavage. We integrated the ermAM erythromycin resistance gene into the srtA locus and used Cas9-mediated editing to introduce a premature stop codon (Figure 33). The resulting strain (JEN53) contains the ermAM(stop) allele and is sensitive to erythromycin. This strain can be used to assess the efficiency with which the ermAM gene is repaired by measuring the rate of cells that revert antibiotic resistance with or without Cas9 cleavage. JEN53 was transformed with an editing template that reverts the wild-type allele, along with either a kanamycin-resistance CRISPR construct targeting the ermAM(stop) allele (CRISPR::ermAM(stop)) or a control construct without a spacer (CRISPR::φ) (Figures 26a, b). In the absence of kanamycin selection, the rate of edited colonies was on the order of 10 (erythromycin-resistant cfu / total cfu) (Figure 26c), representing the baseline frequency of recombination without Cas9-mediated selection for non-edited cells. However, when kanamycin selection was applied and a control CRISPR construct was co-transformed, the rate of edited colonies increased to approximately 10 (kanamycin- and erythromycin-resistant cfu / kanamycin-resistant cfu) (Figure 26c). This result indicates that selection for recombination at the CRISPR locus co-selected for recombination at the ermAM locus independently of Cas9 cleavage of the genome, suggesting that a subpopulation of cells is prone to transformation and / or recombination. Transformation of the CRISPR::ermAM(termination) construct followed by kanamycin selection resulted in an increase in the rate of erythromycin-resistant edited cells to 99% (Figure 26c). To determine whether this increase was caused by killing of non-edited cells, we compared the kanamycin-resistant colony-forming units (cfu) obtained after co-transformation of JEN53 cells with CRISPR::ermAM(terminate) or CRISPR::φ constructs.

[0182] We counted 5.3-fold fewer kanamycin-resistant colonies after transformation with the ermAM(stop) construct (2.5 × 10 / 4.7 × 10 , Figure 35a), a result suggesting that targeting of chromosomal loci with Cas9 actually results in killing of non-edited cells. Finally, because the introduction of dsDNA breaks in bacterial chromosomes is known to trigger repair mechanisms that increase the rate of recombination of damaged DNA, we investigated whether cleavage by Cas9 induces recombination of the edited template. We counted 2.2-fold more colonies after co-transformation with the CRISPR::erm(stop) construct than with the CRISPR::φ construct (Figure 26d), indicating that there was modest induction of recombination. Collectively, these results demonstrated that simultaneous selection of transformable cells, induction of recombination by Cas9-mediated cleavage, and selection against non-edited cells each contribute to highly efficient genome editing in S. pneumoniae.

[0183] Because genomic cleavage by Cas9 kills non-edited cells, it is expected that no cells that received the kanamycin resistance-containing Cas9 cassette will be recovered, except for those that did not receive the editing template. However, in the absence of an editing template, we recovered many kanamycin-resistant colonies after transformation with the CRISPR::ermAM(stop) construct (Figure 35a). These cells that "escape" from CRISPR-induced death generated background that determined the limitations of this method. This background frequency can be calculated as the ratio of CRISPR::ermAM(stop) / CRISPR::φcfu, which in this experiment was 2.6 x 10 (7.1 x 10 / 2.7 x 10). This means that if the recombination frequency of the editing template is below this value, CRISPR selection cannot efficiently recover the desired mutant above the background. To understand the origin of these cells, we genotyped eight background colonies and found that seven lacked the targeting spacer (Figure 35b) and one harbored a putative inactivating mutation in Cas9 (Figure 35c).

[0184] Genome editing with Cas9 in E. coli: Activation of Cas9 targeting through chromosomal integration of the CRISPR-Cas system is only possible in highly recombinogenic organisms. To develop a more general method applicable to other microorganisms, we decided to perform genome editing in E. coli using a plasmid-based CRISPR-Cas system. Two plasmids were constructed: the pCas9 plasmid (Figure 36), which carries tracrRNA, Cas9, and a chloramphenicol resistance cassette, and the pCRISPR kanamycin resistance plasmid, which carries an array of CRISPR spacers. To measure the efficiency of editing independent of CRISPR selection, we sought to introduce an A-to-C transversion in the rpsL gene, which confers streptomycin resistance. We constructed the pCRISPR::rpsL plasmid, which carries a spacer that guides Cas9 cleavage of the wild-type rpsL allele but not the mutant rpsL allele (Figure 27b). The pCas9 plasmid was first introduced into E. coli MG1655, and the resulting strain was cotransformed with the pCRISPR::rpsL plasmid and the editing oligonucleotide W542, which contains an A to C mutation. Only streptomycin-resistant colonies were recovered after transformation of the pCRISPR::rpsL plasmid, suggesting that Cas9 cleavage induces recombination of the oligonucleotide (Figure 37). However, the number of streptomycin-resistant colonies was two orders of magnitude lower than the number of kanamycin-resistant colonies, which are presumably cells that escape Cas9 cleavage. Thus, under these conditions, Cas9 cleavage promoted the introduction of mutations, but the efficiency was not sufficient to select mutant cells over a background of "escapers."

[0185] To improve the efficiency of genome editing in Escherichia coli (E. coli), we applied the CRISPR system of the present invention to select for desired mutations by recombineering using Cas9-induced cell death. The pCas9 plasmid was introduced into the recombineering strain HME63 (31), which contains the Gam, Exo, and Beta functions of the □-red phage. The resulting strain was cotransformed with the pCRISPR::rpsL plasmid (or pCRISPR::φ control) and the W542 oligonucleotide (Figure 27a). The recombineering efficiency, calculated as the percentage of the total number of cells that became streptomycin-resistant when the control plasmid was used, was 5.3 × 10-5 (Figure 27c). In contrast, transformation with the pCRISPR::rpsL plasmid increased the percentage of mutant cells to 65 ± 14% (Figures 27c and 29f). We observed that the number of cfu was reduced by approximately three orders of magnitude after transformation with the pCRISPR::rpsL plasmid compared to the control plasmid (4.8 × 10 / 5.3 × 10 , Figure 38a), suggesting that selection resulted from CRISPR-induced death of non-edited cells. To measure the rate at which Cas9 cleavage was inactivated, a key parameter of our method, we transformed cells with either the pCRISPR::rpsL or control plasmid without the W542 editing oligonucleotide (Figure 38a). The CRISPR "escaper" population in this background was 2.5 × 10 (1.2 × 10 / 4.8 × 10 ), measured as the ratio of pCRISPR::rpsL / pCRISPR::φ cfu. Genotyping of eight of these escapers revealed the presence of targeting spacer deletions in all cases (Figure 38b). This background was higher than the recombineering efficiency of the rpsL mutation, 5.3 x 10-5, suggesting that Cas9 cleavage must induce oligonucleotide recombination to yield 65% edited cells. To confirm this, we compared the numbers of kanamycin- and streptomycin-resistant cfu after transformation with pCRISPR::rpsL or pCRISPR::φ (Figure 27d).As with S. pneumoniae, we observed a modest induction of recombination of approximately 6.7-fold (2.0 x 10 / 3.0 x 10). Collectively, these results demonstrate that the CRISPR system provides a method for selecting mutations introduced by recombineering.

[0186] The present applicants have demonstrated that the CRISPR-Cas system can be used for targeted genome editing in bacteria by cotransformation of a targeting construct that kills wild-type cells and an editing template that eliminates CRISPR cleavage and induces the desired mutation. Different types of mutations (insertion, deletion, or scarless single-nucleotide substitution) can be generated. Multiple mutations can be introduced simultaneously. The specificity and versatility of editing using the CRISPR system depend on several unique properties of the Cas9 endonuclease: (i) its target specificity can be programmed by small RNA molecules without the need for enzyme engineering; (ii) its target specificity is extremely high, determined by a 20-bp RNA-DNA interaction, with a low probability of non-target recognition; (iii) it can target almost any sequence; the only requirement is the presence of a flanking NGG sequence; and (iv) almost any mutation in the NGG sequence and mutations in the protospacer seed sequence eliminate targeting.

[0187] The present applicants have demonstrated that genome engineering using the CRISPR system works not only in highly recombinogenic bacteria, such as Streptococcus pneumoniae, but also in Escherichia coli (E. coli). The results in E. coli suggested that this method may be applicable to other microorganisms into which plasmids can be introduced. In E. coli, this approach complements recombineering of mutagenic oligonucleotides. To use this methodology in microorganisms in which recombineering is not possible, the host's homologous recombination machinery can be utilized by providing an editing template on a plasmid. Furthermore, accumulating evidence indicates that CRISPR-mediated chromosomal cleavage leads to cell death in many bacteria and archaea, so it is possible to envision the use of endogenous CRISPR-Cas systems for editing purposes.

[0188] In both Streptococcus pneumoniae (S. pneumoniae) and Escherichia coli (E. coli), we observed that editing is facilitated by the simultaneous selection of transformable cells and the small induction of recombination at the target site by Cas9 cleavage, but that the mechanism largely contributing to editing is selection against non-editing cells. Thus, a major limitation of this method was the presence of a background of cells that escape CRISPR-induced cell death and lack the desired mutation. We showed that these "escapers" arise primarily through deletion of the targeting spacer, presumably after recombination of repeat sequences flanking the targeting spacer. Further improvements could focus on engineering flanking sequences that are sufficiently different from each other to preclude recombination while still supporting the biogenesis of a functional crRNA. Alternatively, direct transformation of chimeric crRNAs could be explored. In the specific case of E. coli, construction of a CRISPR-Cas system was not possible when this organism was also used as a cloning host. We solved this problem by placing Cas9 and tracrRNA on a different plasmid than the CRISPR array. Engineering an inducible system may also circumvent this limitation.

[0189] While new DNA synthesis technologies offer the ability to cost-effectively create arbitrary sequences at high throughput, integrating synthetic DNA into living cells to create functional genomes remains challenging. Recently, the co-selection MAGE strategy has been shown to improve the mutational efficiency of recombineering by selecting a subpopulation of cells with an increased likelihood of achieving recombination at or around a given locus. In this method, the introduction of a selectable mutation increases the chance of generating nearby unselectable mutations. In contrast to the indirect selection offered by this strategy, the use of CRISPR systems allows for the direct selection of desired mutations and their recovery with high efficiency. These techniques expand the genetic engineer's toolbox, and together with DNA synthesis, they could substantially advance the ability to decipher gene function and manipulate organisms for biotechnological purposes. Two other studies are also ongoing involving CRISPR-assisted engineering of mammalian genomes. It is predicted that these crRNA-directed genome editing techniques could be widely useful in basic and medical sciences.

[0190] Strains and Culture Conditions. S. pneumoniae strain R6 was provided by Dr. Alexander Tomasz. Strain crR6 was generated in the above study. Liquid cultures of S. pneumoniae were grown in THYE medium (30 g / L Todd-Hewitt agar, 5 g / L yeast extract). Cells were plated on tryptic soy agar (TSA) supplemented with 5% defibrinated sheep blood. Antibiotics were added as appropriate: kanamycin (400 μg / ml), chloramphenicol (5 μg / ml), erythromycin (1 μg / ml), streptomycin (100 μg / ml), or spectinomycin (100 μg / ml). β-Galactosidase activity was measured using the Miller assay as described above.

[0191] E. coli strains MG1655 and HME63 (MG1655-derived, Δ(argF-lac)U169λcI857Δcro-bioA galK tyr145UAG mutS<>amp) (31) were kindly provided by Jeff Roberts and Donald Court, respectively. Liquid cultures of E. coli were grown in LB medium (Difco). Antibiotics were added as appropriate: chloramphenicol (25 μg / ml), kanamycin (25 μg / ml), and streptomycin (50 μg / ml).

[0192] S. pneumoniae transformation. Competent cells were prepared as described previously (23). For all genome editing transformations, cells were gently thawed on ice and resuspended in 10 volumes of M2 medium supplemented with 100 ng / ml of the competence-stimulating peptide CSP1 (40), followed by the addition of the editing construct (the editing construct was added to the cells at a final concentration of 0.7 ng / μl to 2.5 μg / ul). Cells were incubated at 37°C for 20 min, after which 2 μl of the targeting construct was added, followed by incubation at 37°C for 40 min. Serial dilutions of the cells were plated on the appropriate medium to determine colony-forming unit (cfu) counts.

[0193] E. coli Lambda-red Recombineering. Strain HME63 was used for all recombineering experiments. Recombineering cells were prepared and handled according to a previously published protocol (6). Briefly, a 2 ml overnight culture (LB medium) inoculated from a single colony obtained from a plate was grown at 30°C. The overnight culture was diluted 100-fold and grown with shaking (200 rpm) at 30°C until the OD600 reached 0.4–0.5 (approximately 3 h). For lambda-red induction, the culture was transferred to a 42°C water bath and shaken at 200 rpm for 15 min. Immediately after induction, the culture was spun in an ice-water slurry and refrigerated on ice for 5–10 min. The cells were then washed and aliquoted according to the protocol. For electrotransformation, 50 μl of cells were mixed with 1 mM salt-free oligos (IDT) or 100–150 ng of plasmid DNA (prepared by QIAprep Spin Miniprep Kit, Qiagen). Cells were electroporated using a 1 mm Gene Pulser cuvette (Bio-Rad) at 1.8 kV and immediately resuspended in 1 ml of room temperature LB medium. Cells were allowed to recover for 1–2 h at 30°C before being plated on LB agar containing the appropriate antibiotic resistance and incubated overnight at 32°C.

[0194] Preparation of S. pneumoniae genomic DNA. For transformation purposes, S. pneumoniae genomic DNA was extracted using the Wizard Genomic DNA Purification Kit according to the instructions provided by the manufacturer (Promega). For genotyping purposes, 700 ul of overnight S. pneumoniae culture was pelleted, resuspended in 60 ul of lysozyme solution (2 mg / ml), and incubated at 37°C for 30 minutes. Genomic DNA was extracted using a QIAprep Spin Miniprep Kit (Qiagen).

[0195] Strain construction. All primers used in this study are provided in Table G. To generate S. pneumoniae crR6M, an intermediate strain, LAM226, was created. In this strain, the aphA-3 gene (providing kanamycin resistance) flanking the CRISPR array of S. pneumoniae crR6 strain was replaced with the cat gene (providing chloramphenicol resistance). Briefly, crR6 genomic DNA was amplified using primers L448 / L444 and L447 / L481, respectively. The cat gene was amplified from plasmid pC194 using primers L445 / L446. Each PCR product was gel purified, and all three were fused by SOEing PCR using primers L448 / L481. The resulting PCR products were transformed into competent S. pneumoniae crR6 cells, and chloramphenicol-resistant transformants were selected. To generate S. pneumoniae crR6M, S. pneumoniae crR6 genomic DNA was amplified by PCR using primers L409 / L488 and L448 / L481, respectively. Each PCR product was gel purified and fused by continuous overlay PCR using primers L409 / L481. The resulting PCR product was transformed into competent S. pneumoniae LAM226 cells, and kanamycin-resistant transformants were selected.

[0196] To generate S. pneumoniae crR6Rc, S. pneumoniae crR6M genomic DNA was amplified by PCR using primers L430 / W286, and S. pneumoniae LAM226 genomic DNA was amplified by PCR using primers W288 / L481. The PCR products were gel purified and fused by single-stranded fusion PCR using primers L430 / L481. The resulting PCR products were transformed into competent S. pneumoniae crR6M cells, and chloramphenicol-resistant transformants were selected.

[0197] To generate S. pneumoniae crR6Rk, S. pneumoniae crR6M genomic DNA was amplified by PCR using primers L430 / W286 and W287 / L481, respectively. Each PCR product was gel-purified and fused by SOEing PCR using primers L430 / L481. The resulting PCR product was transformed into competent S. pneumoniae crR6Rc cells, and kanamycin-resistant transformants were selected.

[0198] To generate JEN37, S. pneumoniae crR6Rk genomic DNA was amplified by PCR using primers L430 / W356 and W357 / L481, respectively. Each PCR product was gel-purified and fused by continuous overlay PCR using primers L430 / L481. The resulting PCR product was transformed into competent S. pneumoniae crR6Rc cells, and kanamycin-resistant transformants were selected.

[0199] To generate JEN38, R6 genomic DNA was amplified using primers L422 / L461 and L459 / L426, respectively. Primers L457 / L458 were used to amplify the ermAM gene (which confers erythromycin resistance) from plasmid pFW1543. Each PCR product was gel-purified, and all three were fused together by SOEing PCR using primers L422 / L426. The resulting PCR products were transformed into competent S. pneumoniae crR6Rc cells, and erythromycin-resistant transformants were selected.

[0200] S. pneumoniae JEN53 was generated in two steps. First, JEN43 was constructed as illustrated in Figure 33. JEN53 was generated by transforming genomic DNA of JEN25 into competent JEN43 cells and selecting on both chloramphenicol and erythromycin.

[0201] To generate S. pneumoniae JEN62, S. pneumoniae crR6Rk genomic DNA was amplified by PCR using primers W256 / W365 and W366 / L403, respectively. Each PCR product was purified and ligated by Gibson assembly. The assembly product was transformed into competent S. pneumoniae crR6Rc cells, and kanamycin-resistant transformants were selected.

[0202] Plasmid construction. pDB97 was constructed through phosphorylation and annealing of oligonucleotides B296 / B297, followed by ligation into EcoRI / BamHI-digested pLZ12spec. Applicants have fully sequenced pLZ12spec and deposited the sequence in genebank (accession number: KC112384).

[0203] pDB98 was obtained after cloning the CRISPR leader sequence together with the repeat-spacer-repeat unit into pLZ12spec. This was achieved through amplification of the crR6R cDNA using primers B298 / B320 and B299 / B321, followed by SOEing PCR of both products and cloning into pLZ12spec with the restriction sites BamHI / EcoRI. Thus, the spacer sequence in pDB98 was engineered to contain two BsaI restriction sites in opposite orientations, allowing for scarless cloning of new spacers.

[0204] pDB99 to pDB108 were constructed by annealing oligonucleotides B300 / B301 (pDB99), B302 / B303 (pDB100), B304 / B305 (pDB101), B306 / B307 (pDB102), B308 / B309 (pDB103), B310 / B311 (pDB104), B312 / B313 (pDB105), B314 / B315 (pDB106), B315 / B317 (pDB107), and B318 / B319 (pDB108) followed by ligation into pDB98 cut with BsaI.

[0205] The pCas9 plasmid was constructed as follows: the essential CRISPR elements were inserted into the pCas9 plasmid. Streptococcus pyogenes using flanking homology arms for Gibson assembly (Streptococcos pyogenes) Amplified from SF370 genomic DNA The tracrRNA and Cas9 were amplified using oligos HC008 and HC010. The leader and CRISPR sequences were amplified using oligos HC011 / HC014 and HC015 / HC009, such that two BsaI type IIS sites were introduced between the two direct repeats to facilitate easy insertion of spacers.

[0206] pCRISPR was constructed by subcloning the pCas9CRISPR array into pZE21-MCS1 via amplification with oligos B298+B299 and restriction with EcoRI and BamHI. The rpsL targeting spacer was cloned by annealing oligos B352+B353 and cloning into BsaI-cut pCRISPR, resulting in pCRISPR::rpsL.

[0207] Generation of targeting and editing constructs. Targeting constructs used for genome editing were generated by Gibson assembly of left-handed and right-handed PCR (Table G). Editing constructs were generated by single-strand assembly PCR fusing PCR product A (PCR A), PCR product B (PCR B), and PCR product C (PCR C), where applicable (Table G). CRISPR::φ and CRISPR::ermAM(termination) targeting constructs were generated by PCR amplification of JEN62 and crR6 genomic DNA using oligos L409 and L481, respectively.

[0208] Generation of targets using randomized PAM or protospacer sequences. The five nucleotides after the spacer 1 target were randomized through amplification of R68232.5 genomic DNA using primers W377 / L426. This PCR product was then assembled with the cat gene and srtA upstream region amplified from the same template using primers L422 / W376. 80 ng of assembled DNA was used to transform strains R6 and crR6. Samples for randomization were prepared using the following primers: B280–B290 / L426 to randomize bases 1–10 of the target and B269–B278 / L426 to randomize bases 10–20. Primers L422 / B268 and L422 / B279 were used to amplify the cat gene and srtA upstream region to be assembled with the first and last 10 PCR products, respectively. The assembled constructs were pooled together and 30 ng was transformed into R6 and crR6. After transformation, cells were plated under chloramphenicol selection. For each sample, more than 2 x 10 cells were pooled together in 1 ml of THYE and genomic DNA fragments were extracted. DNA was extracted using the Promega Wizard kit. Primers B250 / B251 were used to amplify the target region. PCR products were tagged and run on one Illumina MiSeq paired-end lane for 300 cycles.

[0209] Analysis of deep sequencing data Randomized PAM: For the randomized PAM experiment, 3,429,406 reads were obtained for crR6 and 3,253,998 reads for R6. Only half of these correspond to the PAM target, while the other half are predicted to sequence the other end of the PCR product. 1,623,008 crR6 reads and 1,537,131 R6 reads carried error-free target sequences. The incidence of each possible PAM among these reads is shown in the supplemental file. To estimate the functionality of PAM, its relative proportion in the crR6 sample relative to the R6 sample was calculated, denoted by rijklm (where l, j, k, l, m is one of the four possible bases). The following statistical model was constructed: log(rijklm)=μ+b2i+b3j+b4k+b2b3i,j+b3b4j, k+εijklm (where ε is the residual, b2 is the effect of the second base of the PAM, b3 is the effect of the third base, b4 is the effect of the fourth base, b2b3 is the interaction between the second and third bases, and b3b4 is the interaction between the third and fourth bases.) Analysis of variance was performed.

[0210] [Table 20]

[0211] If added to this model, b1 or b5 are not considered significant, and other interactions can be discarded except for the one included. Model selection was performed through successive comparisons of virtually complete models using the anova method in R. Tukey's honest significance test to determine whether pairwise differences between effects were significant.

[0212] The NGGNN pattern is significantly different from all other patterns and has the strongest effect (see table below).

[0213] To show that positions 1, 4, or 5 do not influence the NGGNN pattern, we considered only these sequences. These effects appear to be normally distributed (see QQ plot in Figure 71), and model comparison using anova in R shows that the null model is the best one, i.e., there is no significant role for b1, b4, and b5.

[0214] [Table 21]

[0215] Partial interference of NAGNN and NNGGN patterns The NAGNN pattern is significantly different from all other patterns, but has a much smaller effect than NGGNN (see Tukey's HSD test below).

[0216] Finally, the NTGGN and NCGGN patterns also show significantly greater CRISPR interference than the NTGHN and NCGHN patterns (where H is A, T, or C) as shown by Bonferroni-adjusted pairwise Student's test.

[0217] [Table 22]

[0218] Taken together, these results allow us to conclude that NNGGN patterns generally give rise to full interference in the case of NGGGN, or partial interference in the case of NAGGN, NTGGN or NCGGN.

[0219] [Table 23]

[0220] Randomized Target For the randomized target experiment, 540,726 reads were obtained for crR6 and 753,570 reads for R6. As mentioned above, only half of the reads are expected to sequence the intended end of the PCR product. After filtering out reads carrying error-free or single-point mutation targets, 217,656 and 353,141 reads remained for crR6 and R6, respectively. The relative proportion of each mutation in the crR6 sample relative to the R6 sample was calculated (Figure 24c). All mutations outside the seed sequence (13–20 bases away from the PAM) exhibited complete interference. These sequences were used as references to determine whether other mutations inside the seed sequence could significantly disrupt interference. A normal distribution was fitted to these sequences using the fitdistr function in the MASS R package. The 0.99 quantile of the fitted distribution is shown as a dotted line in Figure 24c. FIG. 72 shows a histogram of data density with a fitted normal distribution (black line) and the 0.99 quantile (dotted line).

[0221] [Table 24]

[0222] [Table 25]

[0223] [Table 26]

[0224] [Table 27]

[0225] [Table 28]

[0226] Example 6: Optimization of guide RNA for Streptococcus pyogenes Cas9 (referred to as SpCas9) The applicants mutated the tracrRNA and direct repeat sequences or mutated the chimeric guide RNA to improve RNA in cells.

[0227] The optimization was based on the observation that there were stretches of thymines (T) in the tracrRNA and guide RNA, which could result in premature transcription termination by the pol3 promoter. Therefore, we generated the following optimized sequences. The optimized tracrRNA and the corresponding optimized direct repeat are represented as a pair. Optimized tracrRNA1 (mutations underlined): [ka] Optimized direct repeat 1 (mutations are underlined): [ka] Optimized tracrRNA2 (mutations underlined): [ka] Optimized direct repeat 2 (mutations are underlined): [ka] Applicants also optimized the chimeric guide RNAs for optimal activity in eukaryotic cells. Original guide RNA: [ka] Optimized chimeric guide RNA sequence 1: [ka] Optimized chimeric guide RNA sequence 2: [ka] Optimized chimeric guide RNA sequence 3: [ka]

[0228] Applicants have shown that the optimized chimeric guide RNA performs better, as shown in Figure 3. This experiment was performed by co-transfecting 293FT cells with Cas9 and U6 guide RNA DNA cassettes to express one of the four RNA forms described above. The guide RNAs target the same target site in the human Emx1 locus: "GTCACCTCCAATGACTAGGG".

[0229] Example 7: Optimization of Streptococcus thermophiles LMD-9 CRISPR1 Cas9 (referred to as St1Cas9) The applicants designed the guide chimeric RNA shown in FIG.

[0230] St1Cas9 guide RNAs can undergo the same type of optimization as SpCas9 guide RNAs by degrading stretches of polythymine (T).

[0231] Example 8: Cas9 diversity and mutations The CRISPR-Cas system is an adaptive immune mechanism against invading exogenous DNA used by a wide variety of species, from bacteria to archaea. Type II CRISPR-Cas9 systems consist of a set of genes encoding proteins responsible for "getting" foreign DNA into the CRISPR locus and a set of genes encoding the "execution" DNA cleavage mechanism; these include a DNA nuclease (Cas9), a non-coding transactivating crRNA (tracrRNA), and an array of spacers (crRNAs) derived from the foreign DNA flanked by direct repeats. Upon maturation by Cas9, the tracrRNA and crRNA duplex guide the Cas9 nuclease to the target DNA sequence specified by the spacer guide sequence, where it mediates a double-stranded break in DNA near a short sequence motif in the target DNA required for cleavage and specific to each CRISPR-Cas system. Type II CRISPR-Cas systems are found throughout the bacterial kingdom and are highly diverse in Cas9 protein sequence and size, tracrRNA and crRNA direct repeat sequences, genomic organization of these elements, and motif requirements for targeted cleavage. A given species may have multiple distinct CRISPR-Cas systems.

[0232] Applicants evaluated 207 putative Cas9s from bacterial species that were identified based on sequence homology to known Cas9s and orthologous structures to known subdomains, e.g., the HNH endonuclease domain and the RuvC endonuclease domain (information from Eugene Koonin and Kira Makarova). Phylogenetic analysis of this set based on protein sequence conservation revealed five families of Cas9s, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) (Figures 39 and 40A-F).

[0233] In this example, Applicants show that the following mutations can convert SpCas9 into a nicking enzyme: D10A, E762A, H840A, N854A, N863A, D986A.

[0234] We provide sequences showing where the mutations are located within the SpCas9 gene (Figure 41). We also show that the nickase can still mediate homologous recombination (assay shown in Figure 2). Furthermore, we show that SpCas9 with these mutations (individually) does not induce double-strand breaks (Figure 47).

[0235] Example 9: Further investigation of the DNA target specificity of RNA-guided Cas9 nuclease Cell culture and transfection The human embryonic kidney (HEK) cell line 293FT (Life Technologies) was maintained in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Life Technologies), 100 U / mL penicillin, and 100 μg / mL streptomycin at 37°C with 5% CO incubation.

[0236] 293FT cells were seeded onto 6-, 24-, or 96-well plates (Corning) 24 hours prior to transfection. Cells were transfected at 80-90% confluence using Lipofectamine 2000 (Life Technologies) according to the manufacturer's recommended protocol. A total of 1 μg of Cas9 + sgRNA plasmid was used for each well of a 6-well plate. Unless otherwise noted, a total of 500 ng of Cas9 + sgRNA plasmid was used for each well of a 24-well plate. For each well of a 96-well plate, 65 ng of Cas9 plasmid was used at a 1:1 molar ratio with the U6-sgRNA PCR product.

[0237] Human embryonic stem cell line HUES9 (Harvard Stem Cell Institute core) was maintained in feeder-free conditions on GelTrex (Life Technologies) in mTesR medium (Stemcell Technologies) supplemented with 100 μg / ml Normocin (InvivoGen). HUES9 cells were transfected using the Amaxa P3 Primary Cell 4-D Nucleofector Kit (Lonza) according to the manufacturer's protocol.

[0238] SURVEYOR Nuclease Assay for Genomic Modifications 293FT cells were transfected with plasmid DNA as described above. Cells were incubated at 37°C for 72 hours after transfection before genomic DNA extraction. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. Briefly, pelleted cells were suspended in QuickExtract solution and incubated at 65°C for 15 minutes and 98°C for 10 minutes.

[0239] Genomic regions flanking the CRISPR target sites for each gene were PCR-amplified (primers listed in Tables J and K), and the products were purified using QiaQuick Spin Columns (Qiagen) according to the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10× Taq DNA Polymerase PCR buffer (Enzymatics), brought to a final volume of 20 μl with ultrapure water, and subjected to a reannealing process to allow heteroduplex formation: 95°C for 10 minutes, ramped from 95°C to 85°C at -2°C / s, ramped from 85°C to 25°C at -0.25°C / s, and held at 25°C for 1 minute. After reannealing, the products were treated with SURVEYOR Nuclease and SURVEYOR Enhancer S (Transgenomics) according to the manufacturer's recommended protocol and analyzed on a 4-20% Novex TBE polyacrylamide gel (Life Technologies). Gels were stained with SYBR Gold DNA stain (Life Technologies) for 30 minutes and imaged using a Gel Doc gel imaging system (Bio-Rad). Quantitation was based on relative band intensity.

[0240] Northern blot analysis of tracrRNA expression in human cells Northern blots were performed as described above. Briefly, RNA was heated to 95°C for 5 minutes and then loaded onto an 8% denaturing polyacrylamide gel (SequaGel, National Diagnostics). The RNA was then transferred to a prehybridized Hybond N+ membrane (GE Healthcare) and crosslinked using a Stratagene UV Crosslinker (Stratagene). The probe was labeled with [gamma-32P]ATP (Perkin Elmer) using T4 polynucleotide kinase (New England Biolabs). After washing, the membrane was exposed to a phosphor screen for 1 hour and scanned using a phosphorimager (Typhoon).

[0241] Bisulfite sequencing to assess DNA methylation status HEK293FT cells were transfected with Cas9 as described above. Genomic DNA was isolated using the DNeasy Blood & Tissue Kit (Qiagen) and bisulfite converted using the EZ DNA Methylation-Lightning Kit (Zymo Research). Bisulfite PCR was performed using KAPA2G Robust HotStart DNA Polymerase (KAPA Biosystems) with primers designed using Bisulfite Primer Seeker (Zymo Research, Tables J and K). The resulting PCR amplicons were gel-purified, digested with EcoRI and HindIII, and ligated into a pUC19 backbone before transformation. Individual clones were then Sanger sequenced to assess DNA methylation status.

[0242] In vitro transcription and cleavage assays HEK293FT cells were transfected with Cas9 as described above. Whole-cell lysates were then prepared using lysis buffer (20 mM HEPES, 100 mM KCl, 5 mM MgCl, 1 mM DTT, 5% glycerol, 0.1% Triton X-100) supplemented with Protease Inhibitor Cocktail (Roche). T7-driven sgRNAs were transcribed in vitro using custom oligos (Example 10) and the HiScribe T7 In Vitro Transcription Kit (NEB) according to the manufacturer's recommended protocol. To prepare methylation target sites, the pUC19 plasmid was methylated with M.SssI and then linearized with NheI. In vitro cleavage assays were performed as follows: for a 20 uL cleavage reaction, 10 uL of cell lysate was incubated with 2 uL of cleavage buffer (100 mM HEPES, 500 mM KCl, 25 mM MgCl, 5 mM DTT, 25% glycerol), in vitro transcribed RNA, and 300 ng of pUC19 plasmid DNA.

[0243] Deep sequencing to assess targeting specificity HEK293FT cells plated in 96-well plates were transfected with Cas9 plasmid DNA and single guide RNA (sgRNA) PCR cassettes for 72 hours, and then genomic DNA was extracted (Figure 72). Genomic regions flanking the CRISPR target sites for each gene were amplified by fusion PCR to generate Illum A P5 adapter and a unique sample-specific barcode were attached to the target amplicon (schematic diagram shown in Figure 73) (Figure 74, Figure 80 (Example 10)). PCR products were purified using EconoSpin 96-well Filter Plates (Epoch Life Sciences) according to the manufacturer's recommended protocol.

[0244] The barcoded and purified DNA samples were quantified using the Quant-iT PicoGreen dsDNA Assay Kit or Qubit 2.0 Fluorometer (Life Technologies) and pooled in equimolar ratios. The sequencing libraries were then deep sequenced using an Illumina MiSeq Personal Sequencer (Life Technologies).

[0245] Sequencing data analysis and indel detection MiSeq reads were filtered by requiring an average Phred quality (Q score) of at least 23 and a complete sequence match with the barcode and amplicon forward primer. Reads from on- and off-target loci were analyzed by first performing Smith-Waterman alignments to the amplicon sequence, including 50 nucleotides (120 bp total) upstream and downstream of the target site. Meanwhile, alignments were analyzed for indels from 5 nucleotides upstream to 5 nucleotides downstream of the target site (30 bp total). Analyzed target regions were discarded if any part of these alignments fell outside the MiSeq read itself or if the matched base pairs comprised less than 85% of their total length.

[0246] The negative controls for each sample provided a metric for the inclusion or exclusion of indels as putative cleavage events. For each sample, indels were counted only if their quality score exceeded μ-σ (μ denotes the mean quality score of the negative controls corresponding to that sample, and σ is its standard deviation). This generated a total target region indel rate for both the negative controls and their corresponding samples. Using the per-read error rate per target region of the negative controls, q, the observed indel count of the sample, n, and its read count, R, a maximum likelihood estimator, p, for the proportion of reads with target regions containing true indels was obtained by applying a binomial error model as follows:

[0247] Let E be the (unknown) number of reads in a sample with a target region that was incorrectly counted as having at least one indel.

number

number

[0248] If all values ​​of the frequency of a target region with a true indel p are considered a priori equally likely, then Prob(n|p) ∝ Prob(p|n). Therefore, the maximum likelihood estimator (MLE) for the frequency of a target region with a true indel was set as the value of p that maximizes Prob(n|p). This was evaluated numerically.

[0249] To place error bounds on the true indel read frequency within the sequencing library itself, Wilson score intervals (2) were calculated for each sample to obtain MLE estimates for the true indel target region Rp and the number of reads R. Specifically, the lower bound l and upper bound μ are given by

number

[0250] qRT-PCR analysis of relative Cas9 and sgRNA expression 293FT cells plated in 24-well plates were transfected as described above. 72 hours after transfection, total RNA was collected using the miRNeasy Micro Kit (Qiagen). Reverse strand synthesis for sgRNA was performed using the qScript Flex cDNA Kit (VWR) and custom first-strand synthesis primers (Tables J and K). qPCR analysis was performed using Fast SYBR Green Master Mix (Life Technologies) and custom primers (Tables J and K) with GAPDH as an endogenous control. Relative quantification was calculated using the ΔΔCT method.

[0251] Table I | Target site sequences. Test target sites and required PAMs for the S. pyogenes Type II CRISPR system. Cells were transfected with Cas9 and either crRNA-tracrRNA or chimeric sgRNA for each target.

[0252] [Table 29]

[0253] [Table 30]

[0254] [Table 31]

[0255] [Table 32]

[0256] Table K | Sequences for primers for testing sgRNA architectures. Primers hybridize to opposite strands of the U6 promoter unless otherwise noted. The U6 priming site is shown in italics, the guide sequence is shown as a stretch of Ns, the direct repeat sequence is highlighted in bold, and the tracrRNA sequence is underlined. The secondary structure of each sgRNA architecture is shown in Figure 43.

[0257] [Table 33]

[0258] Table L | Target sites and alternative PAMs for testing Cas9 PAM specificity. All target sites for PAM specificity testing are found within the human EMX1 locus.

[0259] [Table 34]

[0260] Example 10: Complementary Sequences All sequences are in the 5' to 3' direction. For U6 transcription, the string of underlined T's serves as the transcription terminator. >U6-short tracrRNA (Streptococcus pyogenes SF370) [ka] >U6-DR-guide sequence-DR (Streptococcus pyogenes SF370) [ka] sgRNA containing >+48 tracrRNA (Streptococcus pyogenes SF370) [ka] sgRNA containing >+54 tracrRNA (Streptococcus pyogenes SF370) [ka] sgRNA containing >+67 tracrRNA (Streptococcus pyogenes SF370) [ka] sgRNA containing >+85 tracrRNA (Streptococcus pyogenes SF370) [ka] >CBh-NLS-SpCas9-NLS [ka] [ka] [ka] [ka] >Amplicon sequencing for EMX1 guides 1.1, 1.14, and 1.17 [ka] >Amplicon sequencing for EMX1 guides 1.2 and 1.16 [ka] >Amplicon sequencing for EMX1 guides 1.3, 1.13, and 1.15 [ka] Amplicon sequencing for EMX1 guide 1.6 [ka] Amplicon sequencing for EMX1 guide 1.10 [ka] >Amplicon sequencing for EMX1 guides 1.11 and 1.12 [ka] >Amplicon sequencing for EMX1 guides 1.18 and 1.19 [ka] >Amplicon sequencing for EMX1 Guide 1.20 [ka] >T7 promoter F primer for annealing with the target strand [ka] >Oligo containing pUC19 target site 1 for methylation (T7 reverse) [ka] >Oligo containing pUC19 target site 2 for methylation (T7 reverse) [ka]

[0261] Example 11: Oligo-mediated Cas9-induced homologous recombination The oligo-homologous recombination test is a comparison of efficiency across different Cas9 variants and different HR templates (oligo vs. plasmid).

[0262] 293FT cells were used. SpCas9 = wild-type Cas9, SpCas9n = nickase Cas9 (D10A). The chimeric RNA target was the same EMX1 protospacer target 1 as in Examples 5, 9, and 10, and oligos were synthesized by IDT using PAGE purification.

[0263] Figure 44 shows the design of the oligo DNA used as the homologous recombination (HR) template in this experiment. The long oligo contains 100 bp of homology with the EMX1 locus and a HindIII restriction site. 293FT cells were co-transfected with a plasmid containing a chimeric RNA targeting the human EMX1 locus and a wild-type cas9 protein, and with the oligo DNA as the HR template. Samples were obtained from 293FT cells harvested 96 hours after transfection with Lipofectamine 2000. All products were amplified using EMX1HR primers, gel-purified, and then digested with HindIII to detect the efficiency of integration of the HR template into the human genome.

[0264] Figures 45 and 46 show a comparison of HR efficiency induced by different combinations of Cas9 protein and HR template. The Cas9 constructs used were either wild-type Cas9 or the nickase version of Cas9 (Cas9n). The HR templates used were either antisense oligo DNA (antisense oligo in the upper figure) or sense oligo DNA (sense oligo in the upper figure), or a plasmid HR template (HR template in the upper figure). The sense / antisense definition is that the actively transcribed strand, which has a sequence corresponding to the transcribed mRNA, is defined as the genomic sense strand. HR efficiency is shown as the ratio of the HindIII-digested band to the total genomic PCR amplification product (lower number).

[0265] Example 12: Autistic mice Recent large-scale sequencing initiatives have generated a large number of genes associated with disease. Gene discovery is only the beginning of understanding what that gene is and how it causes disease phenotypes. Current technologies and approaches for studying candidate genes are slow and cumbersome. The gold standard, gene targeting and gene knockout, require significant investments of time and resources, both in terms of money and research personnel. The applicants have designed hSpCas9 nuclease to target many target genes, with high efficiency and low turnaround time compared to any other technology. Due to the high efficiency of hSpCas9, the applicants can perform RNA injection into mouse zygotes and immediately obtain genome-modified animals without the need for any prior gene targeting in mESCs.

[0266] Chromodomain helicase DNA-binding protein 8 (CHD8) is a crucial gene involved in early vertebrate development and morphogenesis. Mice lacking CHD8 die during embryonic development. Mutations in the CHD8 gene have been linked to autism spectrum disorder in humans. This association was confirmed in three separate papers published simultaneously in Nature. The three identical studies identified a large number of genes associated with autism spectrum disorder. Our goal was to create knockout mice for four genes found in all papers: Chd8, Katnal2, Kctd13, and Scn2a. Additionally, we selected two other genes associated with autism spectrum disorder, schizophrenia, and ADHD: GIT1, CACNA1C, and CACNB2. Finally, as a positive control, we decided to target MeCP2.

[0267] For each gene, we selected three genes likely to knock out the gene. Two gRNAs were designed. The knockout was achieved by the hSpCas9 nuclease generating a double-strand break, and error-prone DNA repair pathways, which create non-homologous end joining, correct breaks, and repair mutations. The most likely outcome is a frame-locked gene that knocks out the gene. The targeting strategy is to target a gene with a PAM sequence NGG, which is unique in the genome. This involved finding the protospacer in the exon of a key gene. Priority was given to the protospacer in the first exon, which is harmful.

[0268] Each gRNA was validated in the mouse cell line Neuro-N2a by liposome transient co-transfection with hSpCas9. Seventy-two hours after transfection, genomic DNA was purified using QuickExtract DNA from Epicentre. PCR was performed to amplify the locus of interest. This was followed by the SURVEYOR Mutation Detection Kit from Transgenomics. The SURVEYOR results for each gRNA and their respective controls are shown in Figure A1. A positive SURVEYOR result is one large band corresponding to genomic PCR and two smaller bands representing the product of SURVEYOR nuclease, which creates a double-stranded break at the site of the mutation. The average cleavage efficiency of each gRNA was also measured for each gRNA. The gRNA selected for injection was the most unique and most efficient gRNA within the genome.

[0269] RNA (hSpCas9+gRNA RNA) was injected into the pronuclei of zygotes, which were then implanted into surrogate mothers. The surrogate mothers were allowed to carry the pregnancy to term, and the pups were sampled by tail snip 10 days after birth. DNA was extracted and used as a template for PCR, which was then processed by SURVEYOR. The PCR products were then sent for sequencing. Animals detected as positive in either the SURVEYOR assay or PCR sequencing had their genomic PCR products cloned into a pUC19 vector and sequenced to determine the putative mutations from each allele.

[0270] To date, pups from Chd8 targeting experiments have been fully processed up to the time of allele sequencing. Surveyor results for 38 surviving pups (lanes 1–38), one dead pup (lane 39), and one wild-type control pup (lane 40) are shown in Figure A2. Pups 1–19 were injected with gRNA Chd8.2, and pups 20–38 were injected with gRNA Chd8.3. Of the 38 surviving pups, 13 tested positive for the mutation. The one dead pup also had a mutation. No mutations were detected in the wild-type sample. Genomic PCR sequencing was consistent with the findings of the SURVEYOR assay.

[0271] Example 13: CRISPR / Cas-mediated transcriptional modulation Figure 67 shows the design of CRISPR-TF (transcription factor) with transcription activation activity. The chimeric RNA is expressed by the U6 promoter, while the human codon-optimized double mutant version of the Cas9 protein (hSpCas9m), which is operably linked to three NLSs and a VP64 functional domain, is expressed by the EF1a promoter. The double mutations D10A and H840A render the Cas9 protein unable to introduce any cleavage but maintain its ability to bind to target DNA when guided by the chimeric RNA.

[0272] Figure 68 shows transcriptional activation of the human SOX2 gene by the CRISPR-TF system (chimeric RNA and Cas9-NLS-VP64 fusion protein). 293FT cells were transfected with plasmids carrying two components: (1) different chimeric RNAs driven by U6 targeting 20-bp sequences within or surrounding the human SOX2 genomic locus, and (2) an hSpCas9m (double mutant)-NLS-VP64 fusion protein driven by EF1a. Ninety-six hours after transfection, 293FT cells were harvested, and the level of activation due to the induction of mRNA expression was measured using a qRT-PCR assay. All expression levels are normalized to the control group (gray bar), which represents results from cells transfected with the CRISPR-TF backbone plasmid without the chimeric RNA. The qRT-PCR probe used to detect SOX2 mRNA is the Taqman Human Gene Expression Assay (Life Technologies). All experiments represent data from three biological replicates, n=3, and error bars indicate the standard error (sem).

[0273] Example 14: NLS:Cas9NLS 293FT cells were transfected with a plasmid containing two components: (1) the EF1a promoter driving the expression of Cas9 (wild-type human codon-optimized SpCas9) with a different NLS design, and (2) the U6 promoter driving the same chimeric RNA targeting the human EMX1 locus.

[0274] Cells were harvested 72 hours posttransfection and extracted with 50 μl of QuickExtract genomic DNA extraction solution according to the manufacturer's protocol. Target EMX1 genomic DNA was PCR amplified and gel-purified on a 1% agarose gel. The genomic PCR product was reannealed and subjected to the Surveyor assay according to the manufacturer's protocol. The genomic cleavage efficiency of the different constructs was measured using SDS-PAGE on 4-12% TBE-PAGE gels (Life Technologies) and analyzed and quantified using ImageLab (Bio-Rad) software, all according to the manufacturer's protocol.

[0275] Figure 69 shows the design of different Cas9NLS constructs. All Cas9s were human codon-optimized versions of SpCas9. The NLS sequence was attached to the cas9 gene at either the N- or C-terminus. All Cas9 variants with different NLS designs were cloned into a backbone vector containing the EF1a promoter-driven chimeric RNA targeting the human EMX1 locus, driven by the U6 promoter, on the same vector, forming a two-component system.

[0276] Table M. Cas9 NLS design test results. Quantification of genome cleavage of different cas9-nls constructs by surveyor assay.

[0277] [Table 35]

[0278] Figure 70 shows the efficiency of genome cleavage induced by Cas9 variants carrying different NLS designs. The percentages indicate the fraction of human EMX1 genomic DNA cleaved by each construct. All experiments represent data from three biological replicates, n=3, and error bars indicate the standard error of the mean (SEM).

[0279] Example 15: Engineering microalgae using Cas9 Methods for delivering Cas9

[0280] Method 1: Applicants deliver Cas9 and guide RNA using a vector that expresses Cas9 under the control of a constitutive promoter, e.g., Hsp70A-Rbc S2 or beta2-tubulin.

[0281] Method 2: Applicants deliver Cas9 and T7 polymerase using vectors that express Cas9 and T7 polymerase under the control of a constitutive promoter, e.g., Hsp70A-Rbc S2 or beta2-tubulin. Delivery is performed using a vector containing a T7 promoter driving the NA.

[0282] Method 3: We deliver Cas9 mRNA and in vitro transcribed guide RNA into algal cells. RNA can be transcribed in vitro. Cas9 mRNA consists of the coding region for Cas9 and the 3'UTR from Cop1 to ensure stability of the Cas9 mRNA.

[0283] For homologous recombination, applicants provide additional homology-directed repair templates.

[0284] A cassette driving expression of Cas9 under the control of the beta-2 tubulin promoter, followed by the sequence for the 3'UTR of Cop1. [ka] [ka] [ka] [ka]

[0285] A cassette driving expression of T7 polymerase under the control of the beta-2 tubulin promoter, followed by the sequence for the 3'UTR of Cop1: [ka] [ka]

[0286] Sequence of guide RNA driven by T7 promoter (T7 promoter, N represents targeting sequence): [ka]

[0287] Gene delivery: Chlamydomonas reinhardtii strains CC-124 and CC-125 from the Chlamydomonas Resource Center are used for electroporation. The electroporation protocol follows the standard recommended protocol from the GeneArt Chlamydomonas Engineering kit.

[0288] Applicants also generate strains of Chlamydomonas reinhardtii that constitutively express Cas9. This can be done by using pChlamy1 (linearized with PvuI) and selecting for hygromycin-resistant colonies. The sequence for pChlamy1 containing Cas9 is shown below. In this approach to achieve gene knockout, it is only necessary to deliver RNA for the guide RNA. For homologous recombination, Applicants deliver the guide RNA and a linearized homologous recombination template. pChlamy1-Cas9: [ka] [ka] [ka] [ka] [ka] [ka]

[0289] For all modified Chlamydomonas reinhardtii cells, Applicants confirmed successful modifications using PCR, SURVEYOR nuclease assay, and DNA sequencing.

[0290] Example 16: Use of Cas9 as a transcriptional repressor in bacteria The ability to artificially control transcription is essential for both studying gene function and constructing synthetic gene networks with desired properties. Applicants describe herein the use of the RNA-guided Cas9 protein as a programmable transcriptional repressor.

[0291] We have previously demonstrated how the Cas9 protein of Streptococcus pyogenes SF370 can be used to direct genome editing in Streptococcus pneumoniae. In this study, we engineered the crR6Rk strain, which contains a minimal CRISPR system consisting of cas9, tracrRNA, and repeats. The D10A-H840 mutation was introduced into cas9 in this strain, resulting in strain crR6Rk**. Four spacers targeting different positions in the bgaA β-galactosidase gene promoter were cloned into a CRISPR array carried by the previously described pDB98 plasmid. We observed an X- to Y-fold reduction in β-galactosidase activity depending on the targeted position, demonstrating the potential of Cas9 as a programmable repressor (Figure 73).

[0292] To achieve Cas9** suppression in Escherichia coli, a green fluorescent protein (GFP) reporter plasmid (pDB127) was constructed to express the gfpmut2 gene from a constitutive promoter. The promoter was designed to carry several NPP PAMs on both strands to measure the effect of Cas9** binding at various positions. Applicants introduced the D10A-H840 mutation into pCas9, a described plasmid carrying tracrRNA, cas9, and a minimal CRISPR array designed for easy cloning of new spacers. 22 different spacers were designed to target different regions of the gfpmut2 promoter and open reading frame. An approximately 20-fold reduction in fluorescence was observed upon targeting the -35 and -10 promoter elements and regions overlapping or adjacent to the Shine-Dalgarno sequence. Targets on both strands showed similar suppression levels. These results suggest that binding of Cas9** anywhere in the promoter region disrupts transcription initiation, presumably through steric inhibition of RNAP binding.

[0293] To determine whether Cas9** could interfere with transcription elongation, we targeted it to the gpfmut2 reading frame. A reduction in fluorescence was observed when both the coding and noncoding strands were targeted, suggesting that Cas9 binding was indeed strong enough to represent an obstacle to RNAP activation. However, a 40% reduction in expression was observed when the coding strand was targeted, whereas a 20-fold reduction was observed when the noncoding strand was targeted (Figure 21b; compare T9, T10, and T11 with B9, B10, and B11). To directly determine the effect of Cas9** binding on transcription, we extracted RNA from strains carrying either T5, T10, B10, or a control construct not targeting pDB127 and subjected it to Northern blot analysis using probes binding either before (B477) or after (B510) the B10 and T10 target sites. Consistent with our fluorescence assay, gfpmut2 transcription was not detected when Cas9** was directed to the promoter region (T5 target), but transcription was observed after targeting the T10 region. Interestingly, a smaller transcript was observed with the B477 probe. This band corresponds to the predicted size of the transcript interrupted by Cas9** and is a direct indication of transcription termination caused by dgRNA::Cas9** binding to the coding strand. Surprisingly, we did not detect any transcript when targeting the noncoding strand (B10). This result suggests that mRNA was degraded, as Cas9** binding to the B10 region is unlikely to interfere with transcription initiation. dgRNA::Cas9 has been shown to bind to ssRNA in vitro. We speculate that binding may trigger mRNA degradation by host nucleases. Indeed, ribosome stalling can induce cleavage of translated mRNA in E. coli.

[0294] Some applications require precise modulation of gene expression rather than complete inhibition. Applicants sought to achieve intermediate levels of inhibition through the introduction of mismatches that weaken the crRNA / target interaction. Applicants created a series of spacers based on the B1, T5, and B10 constructs with increasing numbers of mutations in the 5' end of the crRNA. Up to eight mutations in B1 and T5 did not affect the level of inhibition, and a gradual increase in fluorescence was observed with additional mutations.

[0295] The observed suppression only for an 8-nt match between the crRNA and its target raises the question of off-targeting effects in the use of Cas9** as a transcriptional regulator. Because a good PAM (NGG) is also required for Cas9 binding, the number of matching nucleotides needed to obtain the same level of respiration is 10. A 10-nt match occurs randomly approximately once every 1 Mbp, and therefore such sites are likely to be found even in small bacterial genomes. However, to effectively suppress transcription, such sites must reside in the promoter region of the gene, significantly reducing the likelihood of off-targeting. Applicants have also shown that gene expression can be affected when the non-coding strand of a gene is targeted. For this to occur, the random target must be in the right orientation, but such an event is relatively likely. Indeed, during the course of this study, Applicants were unable to construct one of the spacers designed on pCas9**. Applicants later found that this spacer exhibited a 12-bp match adjacent to a good PAM in the essential murC gene. Such off-targeting can be easily avoided by systematic blast of designed spacers.

[0296] Aspects of the invention are further described in the following numbered paragraphs: 1. A vector system comprising one or more vectors, a. a first regulatory element operably linked to a traer mate sequence and one or more insertion sites for inserting a guide sequence upstream of the traer mate sequence (the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with a traer mate sequence hybridized to the traer sequence); and b. A second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme that comprises a nuclear localization sequence. Includes; Components (a) and (b) are on the same or different vectors of the system Vector system.

[0297] 2. The vector system of paragraph 1, wherein component (a) further comprises a traer sequence downstream of the traer mate sequence under the control of a first regulatory element.

[0298] 3. The vector system of paragraph 1, wherein component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of the two or more guide sequences, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in the eukaryotic cell.

[0299] 4. The vector system of paragraph 1, comprising a traer sequence under the control of a third regulatory element.

[0300] 5. A vector system according to paragraph 1, wherein the traer sequence exhibits at least 50% sequence complementarity along the length of the traer mate sequence when optimally aligned.

[0301] 6. The vector system of paragraph 1, wherein the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive accumulation of said CRISPR enzyme in a detectable amount in the nucleus of a eukaryotic cell.

[0302] 7. The vector system of paragraph 1, wherein the CRISPR enzyme is a type II CRISPR enzyme.

[0303] 8. The vector system of paragraph 1, wherein the CRISPR enzyme is a Cas9 enzyme.

[0304] 9. The vector system of paragraph 1, wherein the CRISPR enzyme is codon-optimized for expression in eukaryotic cells.

[0305] 10. The vector system of paragraph 1, wherein the CRISPR enzyme directs cleavage of one or two strands at the location of the target sequence.

[0306] 11. The vector system of paragraph 1, wherein the CRISPR enzyme lacks DNA strand cleavage activity.

[0307] 12. The vector system of paragraph 1, wherein the first regulatory element is a polymerase III promoter.

[0308] 13. The vector system of paragraph 1, wherein the second regulatory element is a polymerase II promoter.

[0309] 14. The vector system of paragraph 4, wherein the third regulatory element is a polymerase III promoter.

[0310] 15. The vector system of paragraph 1, wherein the guide sequence is at least 15 nucleotides in length.

[0311] 16. The vector system of paragraph 1, wherein less than 50% of the nucleotides of the guide sequence participate in self-complementary base pairing when optimally folded.

[0312] 17. A vector comprising a regulatory element operably linked to an enzyme-coding sequence encoding a CRISPR enzyme comprising one or more nuclear localization sequences, wherein the regulatory element drives transcription of the CRISPR enzyme in a eukaryotic cell such that the CRISPR enzyme accumulates in detectable amounts in the nucleus of the eukaryotic cell.

[0313] 18. The vector system according to paragraph 17, wherein the regulatory element is a polymerase II promoter.

[0314] 19. The vector system of paragraph 17, wherein the CRISPR enzyme is a type II CRISPR system enzyme.

[0315] 20. The vector system of paragraph 17, wherein the CRISPR enzyme is a Cas9 enzyme.

[0316] 21. The vector system of paragraph 17, wherein the CRISPR enzyme lacks the ability to cleave one or more strands of the target sequence to which it binds.

[0317] 22. A CRISPR enzyme comprising one or more nuclear localization sequences of sufficient strength to drive accumulation of a detectable amount of the CRISPR enzyme in the nucleus of a eukaryotic cell.

[0318] 23. The CRISPR enzyme of paragraph 22, which is a type II CRISPR system enzyme.

[0319] 24. The CRISPR enzyme of paragraph 22, which is a Cas9 enzyme.

[0320] 25. The CRISPR enzyme of paragraph 22, which lacks the ability to cleave one or more strands of a target sequence to which it binds.

[0321] 26. a. a first regulatory element operably linked to a traer mate sequence and one or more insertion sites for inserting a guide sequence upstream of the traer mate sequence, wherein the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with the traer mate sequence, which is hybridized to the traer sequence; and / or b. A second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme that comprises a nuclear localization sequence. A eukaryotic host cell comprising:

[0322] 27. A eukaryotic host cell according to paragraph 26, comprising components (a) and (b).

[0323] 28. The eukaryotic host cell according to paragraph 26, wherein component (a), component (b), or components (a) and (b) are stably integrated into the genome of the host eukaryotic cell.

[0324] 29. The eukaryotic host cell according to paragraph 26, wherein component (a) further comprises a traer sequence downstream of the traer mate sequence under the control of a first regulatory element.

[0325] 30. The eukaryotic host cell of paragraph 26, wherein component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of the two or more guide sequences, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in the eukaryotic cell.

[0326] 31. The eukaryotic host cell of paragraph 26, further comprising a third regulatory element operably linked to the traer sequence.

[0327] 32. A eukaryotic host cell according to paragraph 26, wherein the traer sequence exhibits at least 50% sequence complementarity along the length of the traer mate sequence when optimally aligned.

[0328] 33. The eukaryotic host cell of paragraph 26, wherein the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive accumulation of said CRISPR enzyme in a detectable amount in the nucleus of the eukaryotic cell.

[0329] 34. The eukaryotic host cell of paragraph 26, wherein the CRISPR enzyme is a type II CRISPR system enzyme.

[0330] 35. The eukaryotic host cell of paragraph 26, wherein the CRISPR enzyme is a Cas9 enzyme.

[0331] 36. The eukaryotic host cell of paragraph 26, wherein the CRISPR enzyme is codon-optimized for expression in the eukaryotic cell.

[0332] 37. The eukaryotic host cell of paragraph 26, wherein the CRISPR enzyme directs cleavage of one or two strands at the location of the target sequence.

[0333] 38. The eukaryotic host cell of paragraph 26, wherein the CRISPR enzyme lacks DNA strand cleavage activity.

[0334] 39. The eukaryotic host cell according to paragraph 26, wherein the first regulatory element is a polymerase III promoter.

[0335] 40. The eukaryotic host cell according to paragraph 26, wherein the second regulatory element is a polymerase II promoter.

[0336] 41. The eukaryotic host cell according to paragraph 31, wherein the third regulatory element is a polymerase III promoter.

[0337] 42. The eukaryotic host cell of paragraph 26, wherein the guide sequence is at least 15 nucleotides in length.

[0338] 43. The eukaryotic host cell of paragraph 26, wherein fewer than 50% of the nucleotides of the guide sequence are involved in self-complementary base pairing when optimally folded.

[0339] 44. A non-human animal comprising a eukaryotic host cell according to any one of paragraphs 26 to 43.

[0340] 45. A kit comprising a vector system and instructions for use of the kit, wherein the vector system comprises: a. a first regulatory element operably linked to a traer mate sequence and one or more insertion sites for inserting a guide sequence upstream of the traer mate sequence, wherein the guide sequence, when expressed, directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, the CRISPR complex comprising (1) a guide sequence hybridized to the target sequence, and (2) a CRISPR enzyme complexed with the traer mate sequence, which is hybridized to the traer sequence; and / or b. A second regulatory element operably linked to an enzyme coding sequence encoding the CRISPR enzyme that comprises a nuclear localization sequence. Kit including:

[0341] 46. ​​A kit according to paragraph 45, comprising components (a) and (b) of the system on the same or different vectors.

[0342] 47. The kit of paragraph 45, wherein component (a) further comprises a traer sequence downstream of the traer mate sequence under the control of a first regulatory element.

[0343] 48. The kit of paragraph 45, wherein component (a) further comprises two or more guide sequences operably linked to the first regulatory element, each of the two or more guide sequences, when expressed, directs sequence-specific binding of the CRISPR complex to a different target sequence in the eukaryotic cell.

[0344] 49. The kit of paragraph 45, wherein the system comprises a traer sequence under the control of a third regulatory element.

[0345] 50. A kit according to paragraph 45, wherein the traer sequence exhibits at least 50% sequence complementarity along the length of the traer mate sequence when optimally aligned.

[0346] 51. The kit of paragraph 45, wherein the CRISPR enzyme comprises one or more nuclear localization sequences of sufficient strength to drive accumulation of a detectable amount of said CRISPR enzyme in the nucleus of a eukaryotic cell.

[0347] 52. The kit of paragraph 45, wherein the CRISPR enzyme is a type II CRISPR system enzyme.

[0348] 53. The kit of paragraph 45, wherein the CRISPR enzyme is a Cas9 enzyme.

[0349] 54. The kit of paragraph 45, wherein the CRISPR enzyme is codon-optimized for expression in eukaryotic cells.

[0350] 55. The kit of paragraph 45, wherein the CRISPR enzyme directs cleavage of one or two strands at the location of the target sequence.

[0351] 56. The kit of paragraph 45, wherein the CRISPR enzyme lacks DNA strand cleavage activity.

[0352] 57. The kit according to paragraph 45, wherein the first regulatory element is a polymerase III promoter.

[0353] 58. The kit according to paragraph 45, wherein the second regulatory element is a polymerase II promoter.

[0354] 59. The kit of paragraph 49, wherein the third regulatory element is a polymerase III promoter.

[0355] 60. The kit of paragraph 45, wherein the guide sequence is at least 15 nucleotides in length.

[0356] 61. The kit of paragraph 45, wherein fewer than 50% of the nucleotides of the guide sequence participate in self-complementary base pairing when optimally folded.

[0357] 62. A computer system for selecting candidate target sequences within a nucleic acid sequence in a eukaryotic cell for targeting by a CRISPR complex, comprising: a. a storage device configured to receive and / or store the nucleic acid sequence; and b. one or more processors, alone or in combination, programmed to: (i) localize a CRISPR motif sequence within the nucleic acid sequence; and (ii) select sequences adjacent to the localized CRISPR motif sequence as candidate target sequences for binding by a CRISPR complex. A system including:

[0358] 63. The computer system of paragraph 62, wherein the localizing step comprises identifying a CRISPR motif sequence that is located less than about 500 nucleotides away from the target sequence.

[0359] 64. The computer system of paragraph 62, wherein the candidate target sequences are at least 10 nucleotides in length.

[0360] 65. The computer system of paragraph 62, wherein the nucleotide at the 3' end of the candidate target sequence is located no more than about 10 nucleotides upstream of the CRISPR motif sequence.

[0361] 66. The computer system of paragraph 62, wherein the nucleic acid sequence in the eukaryotic cell is endogenous to the eukaryotic genome.

[0362] 67. The method of claim 62, wherein the nucleic acid sequence in the eukaryotic cell is exogenous to the eukaryotic genome. computer systems.

[0363] 68. A computer-readable medium comprising code that, when executed by one or more processors, implements a method for selecting candidate target sequences within a nucleic acid sequence in a eukaryotic cell for targeting by a CRISPR complex, the method comprising: (a) localizing a CRISPR motif sequence within the nucleic acid sequence; and (b) selecting sequences adjacent to the localized CRISPR motif sequence as candidate target sequences to which the CRISPR complex binds.

[0364] 69. The computer-readable medium of paragraph 68, wherein the localizing step comprises localizing a CRISPR motif sequence that is less than about 500 nucleotides away from the target sequence.

[0365] 70. The computer-readable medium of paragraph 68, wherein the candidate target sequences are at least 10 nucleotides in length.

[0366] 71. The computer-readable medium of paragraph 68, wherein the nucleotide at the 3' end of the candidate target sequence is located no more than about 10 nucleotides upstream of the CRISPR motif sequence.

[0367] 72. The computer-readable medium of paragraph 68, wherein the nucleic acid sequence in the eukaryotic cell is endogenous to the eukaryotic genome.

[0368] 73. The computer-readable medium of paragraph 68, wherein the nucleic acid sequence in the eukaryotic cell is exogenous to the eukaryotic genome.

[0369] 74. A method for modifying a target polynucleotide in a eukaryote, comprising binding a CRISPR complex to the target polynucleotide to cause cleavage of the target polynucleotide, thereby modifying the target polynucleotide, wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that is hybridized to a target sequence within the target polynucleotide, the guide sequence being bound to a traer mate sequence that in turn hybridizes to a traer sequence.

[0370] 75. The method of paragraph 74, wherein the cleavage comprises cleaving one or two strands at the location of the target sequence with the CRISPR enzyme.

[0371] 76. The method of paragraph 74, wherein said cleavage results in a decrease in transcription of the target gene.

[0372] 77. The method of paragraph 74, further comprising repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides of the target polynucleotide.

[0373] 78. The method of paragraph 77, wherein the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.

[0374] 79. The method of paragraph 74, further comprising delivering one or more vectors to said eukaryotic cell, wherein the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a traer mate sequence, and a traer sequence.

[0375] 80. The method of paragraph 79, wherein the vector is delivered into a eukaryotic cell in a subject.

[0376] 81. The method of paragraph 74, wherein the modification is carried out in the eukaryotic cells in cell culture.

[0377] 82. The method of paragraph 74, further comprising isolating said eukaryotic cells from a subject prior to said modification.

[0378] 83. The method of paragraph 82, further comprising returning said eukaryotic cells and / or cells derived therefrom to said subject.

[0379] 84. A method of modifying expression of a polynucleotide in a eukaryotic cell, comprising binding a CRISPR complex to a polynucleotide, such that the binding results in increased or decreased expression of the polynucleotide; the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence within the target polynucleotide, the guide sequence then hybridizing to a target sequence. How it is joined to the mate sequence.

[0380] 85. The method of paragraph 74, further comprising delivering one or more vectors to said eukaryotic cell, wherein the one or more vectors drive expression of one or more of a CRISPR enzyme, a guide sequence linked to a traer mate sequence, and a traer sequence.

[0381] 86. A method for generating a model eukaryotic cell containing a mutant disease gene, comprising: a. introducing one or more vectors into a eukaryotic cell, the one or more vectors driving expression of one or more of a CRISPR enzyme, a guide sequence linked to a traer mate sequence, and a traer sequence; and b. Binding a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide within said disease gene, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) a guide sequence that hybridizes to a target sequence within the target polynucleotide, and (2) a traer mate sequence that hybridizes to the traer sequence, thereby generating a model eukaryotic cell containing a mutant disease gene.

[0382] 87. The method of paragraph 86, wherein the cleaving comprises cleaving one or two strands at the location of the target sequence by the CRISPR enzyme.

[0383] 88. The method of paragraph 86, wherein said cleavage results in a decrease in transcription of the target gene.

[0384] 89. The method of paragraph 86, further comprising repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising an insertion, deletion, or substitution of one or more nucleotides of the target polynucleotide.

[0385] 90. The method of paragraph 89, wherein the mutation results in one or more amino acid changes in a protein expressed from a gene comprising the target sequence.

[0386] 91. A method for developing a bioactive agent that modulates a cell signaling event associated with a disease gene, comprising: a. contacting a test compound with the model cell of any one of paragraphs 86-90; and b. detecting a change in a readout indicative of a reduction or an increase in a cellular signaling event associated with said mutation in said disease gene, thereby developing said bioactive agent that modulates said cellular signaling event associated with said disease gene.

[0387] 92. A recombinant polynucleotide comprising a guide sequence upstream of a target mate sequence, which, when expressed, directs sequence-specific binding of a CRISPR complex to a corresponding target sequence present in a eukaryotic cell.

[0388] 93. The recombinant vector according to paragraph 89, wherein the target sequence is a viral sequence present in a eukaryotic cell. e polynucleotide.

[0389] 94. The recombinant polynucleotide according to paragraph 89, wherein the target sequence is a proto-oncogene or an oncogene.

[0390] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein can be employed in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of those claims and their equivalents be covered thereby.

[0391] References: 1.Urnov,F.D.,Rebar,E.J.,Holmes,M.C.,Zhang,H.S.&Gregory,P.D.Genome editing with engineered zinc finger nucleases.Nat.Rev.Genet.11,636-646(2010). 2.Bogdanove,A.J.&Voytas,D.F.TAL effectors:customizable proteins for DNA targeting.Science 333,1843-1846(2011). 3.Stoddard,B.L.Homing endonuclease structure and function.Q.Rev.Biophys.38,49-95(2005). 4.Bae,T.&Schneewind,O.Allelic replacement in Staphylococcus aureus with inducible counter-selection.Plasmid 55,58-63(2006). 5.Sung,C.K.,Li,H.,Claverys,J.P.&Morrison,D.A.An rpsL cassette,janus,for gene replacement through negative selection in Streptococcus pneumoniae.Appl.Environ.Microbiol.67,5190-5196(2001). 6.Sharan,S.K.,Thomason,L.C.,Kuznetsov,S.G.&Court,D.L.Recombineering:a homologous recombination-based method of genetic engineering.Nat.Protoc.4,206-223(2009). 7.Jinek,M.et al.A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.Science 337,816-821(2012). 8.Deveau,H.,Garneau,J.E.&Moineau,S.CRISPR / Cas system and its role in phage-bacteria interactions.Annu.Rev.Microbiol.64,475-493(2010). 9.Horvath,P.&Barrangou,R.CRISPR / Cas,the immune system of bacteria and archaea.Science 327,167-170(2010). 10.Terns,M.P.&Terns,R.M.CRISPR-based adaptive immune systems.Curr.Opin.Microbiol.14,321-327(2011). 11.van der Oost,J.,Jore,M.M.,Westra,E.R.,Lundgren,M.&Brouns,S.J.CRISPR-based adaptive and heritable immunity in prokaryotes.Trends.Biochem.Sci.34,401-407(2009). 12.Brouns,S.J.et al...

Claims

1. (a) a Cas9 protein or a polynucleotide encoding a Cas9 protein, wherein the Cas9 protein is S. pyogenes Cas9 and is fused to two or more nuclear localization signals (NLS); (b) a CRISPR-Cas9 based chimeric RNA comprising: NNNNNNNNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA, where NNNNNNNNNNNNNNNNNNNNNNNNNN is a guide sequence capable of hybridizing to a target sequence adjacent to a protospacer adjacent motif (PAM) in a genomic locus of interest in a eukaryotic cell: wherein the chimeric RNA and the Cas9 protein are capable of forming a CRISPR complex in a eukaryotic cell, and the guide sequence is capable of directing sequence-specific binding of the CRISPR complex to a target sequence adjacent to a PAM in a genomic locus of interest in the eukaryotic cell.

2. (a) a Cas9 protein or a polynucleotide encoding a Cas9 protein, wherein the Cas9 protein is S. pyogenes Cas9 and is fused to two or more nuclear localization signals (NLS); (b) a CRISPR-Cas9 based chimeric RNA comprising: NNNNNNNNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUG, where NNNNNNNNNNNNNNNNNNNNNNNNNN is a guide sequence capable of hybridizing to a target sequence adjacent to a protospacer adjacent motif (PAM) in a genomic locus of interest in a eukaryotic cell: wherein the chimeric RNA and the Cas9 protein are capable of forming a CRISPR complex in a eukaryotic cell, and the guide sequence is capable of directing sequence-specific binding of the CRISPR complex to a target sequence adjacent to a PAM in a genomic locus of interest in the eukaryotic cell.

3. (a) a Cas9 protein or a polynucleotide encoding a Cas9 protein, wherein the Cas9 protein is S. pyogenes Cas9 and is fused to two or more nuclear localization signals (NLS); (b) a CRISPR-Cas9 based chimeric RNA comprising: NNNNNNNNNNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGC, where NNNNNNNNNNNNNNNNNNNNNNNNNN is a guide sequence capable of hybridizing to a target sequence adjacent to a protospacer adjacent motif (PAM) in a genomic locus of interest in a eukaryotic cell: wherein the chimeric RNA and the Cas9 protein are capable of forming a CRISPR complex in a eukaryotic cell, and the guide sequence is capable of directing sequence-specific binding of the CRISPR complex to a target sequence adjacent to a PAM in a genomic locus of interest in the eukaryotic cell.

4. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the PAM is NGG.

5. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the chimeric RNA further comprises a polyU sequence.

6. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the chimeric RNA comprises one or more modified nucleotides.

7. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the chimeric RNA comprises one or more methylated nucleotides or nucleotide analogues.

8. Two or more NLSs may independently be PKKKRKV, KRPAATTKKAGQAKKKK, 3. The engineered CRISPR-Cas9 system of any preceding claim, selected from the group consisting of PAAKRVKLD, RQRRNELKRSP, NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV, VSRKRPRP, PPKKARED, PQPKKKPL, SALIKKKKKKMAP, DRLRR, PKQKKRK, RKLKKKIKKL, REKKKFLKRR, KRKGDEVDGVDEVAKKKSKK and RKCLQAGMNLEARKTKK.

9. 9. The engineered CRISPR-Cas9 system of claim 8, wherein at least one of the NLSs comprises PKKKRKV.

10. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the Cas9 protein comprises a D10A, H840A, N854A, or N863A mutation.

11. 11. The engineered CRISPR-Cas9 system of claim 10, wherein the Cas9 protein is fused to at least one heterologous protein domain.

12. 12. The engineered CRISPR-Cas9 system of claim 11, wherein the heterologous protein domain is selected from the group consisting of an epitope tag, a reporter sequence, and a protein domain having one or more of the following activities: methylase activity, demethylase activity, transcription activator activity, transcription repressor activity, transcription release factor activity, histone modifying activity, RNA cleavage activity, or nucleic acid binding activity.

13. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the polynucleotide encoding the Cas9 protein is codon-optimized for expression in a eukaryotic cell.

14. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the polynucleotide encoding the Cas9 protein comprises a polyadenylation signal.

15. 2. The engineered CRISPR-Cas9 system of any preceding claim, wherein the CRISPR-Cas9 system is comprised in a liposome for delivery.

16. 2. The engineered CRISPR-Cas9 system of any preceding claim, further comprising an exogenous polynucleotide for targeted integration into the DNA break introduced by the CRISPR complex.

17. A pharmaceutical composition comprising an engineered CRISPR-Cas9 system according to any of the preceding claims for use in the treatment of a genetic disease or disorder, provided that said use does not include a step for altering the genetic identity of a human germline.