CRISPR-CAS component systems, methods, and compositions for sequence manipulation
Patent Information
- Application Number
- CN202111233527.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2013-06-17
- Filing Date
- 2013-12-12
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2033-12-12
AI Technical Summary
本发明着手解决这种需要并且提供了相关优点
Smart Images

Figure CN114634950B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 201380072868.6, filed on December 12, 2013, entitled "CRISPR-Cas component system, method and composition for sequence manipulation".
[0002] Related applications and references
[0003] This application requires reference numbers BI-2011 / 008 / WSGR, file numbers 44063-701.101 and 44063-701.102, BI-2011 / 008 / VP, file numbers 44790.01.2003 and 44790.02.2003, and BI-2011 / 008 / VP, file number 44790.02.2003, respectively. Priority claims are made to U.S. Provisional Patent Application Nos. 4790.03.2003, 61 / 736,527, 61 / 748,427, 61 / 768,959, 61 / 791,409, and 61 / 835,931, all entitled “Systems, methods, and compositions for sequence manipulation”, filed on December 12, 2012, January 2, 2013, February 25, 2013, March 15, 2013, and June 17, 2013, respectively.
[0004] Reference is made to U.S. provisional patent applications 61 / 758,468; 61 / 769,046; 61 / 802,174; 61 / 806,375; 61 / 814,263; 61 / 819,803 and 61 / 828,130, each entitled “Engineer and Optimizer of Systems, Methods and Compositions for Sequence Manipulation”, filed on January 30, 2013; February 25, 2013; March 15, 2013; March 28, 2013; April 20, 2013; May 6, 2013 and May 28, 2013, respectively. Reference is also made to U.S. Provisional Patent Applications 61 / 835,936, 61 / 836,127, 61 / 836,101, 61 / 836,080, 61 / 836,123, and 61 / 835,973, each filed on June 17, 2013. Reference is also made to U.S. Provisional Patent Applications 61 / 842,322 and 14 / 054,414, each bearing Bode Reference No. BI-2011 / 008A, entitled “CRISPR-Cas System and Method for Altering the Expression of Gene Products,” filed on July 2, 2013, and October 15, 2013, respectively.
[0005] All documents cited in or during the examination of the foregoing applications (“Application References”), all documents referenced or cited in those Application References, and all documents referenced or cited herein (“Documents Referenced herein”), together with any manufacturer’s specifications, descriptions, product specifications, and product tables for any product mentioned herein or incorporated herein by reference, are hereby incorporated by reference and may be used in the practice of the invention. More specifically, all referenced documents are incorporated herein by reference to the extent that each individual document is explicitly and individually indicated as incorporated herein by reference. Invention Field
[0006] The present invention generally relates to systems, methods, and compositions for controlling gene expression involving sequence targeting, such sequence targeting may be achieved, for example, using vector systems involving regularly spaced clustered short palindromic repeats (CRISPR) and their components, genomic perturbation, or gene editing.
[0007] Statement on Federally Funded Research
[0008] This invention was completed with government support, receiving the NIH Pioneer Award (DP1MH100706) from the National Institutes of Health (NIH). The U.S. government holds certain rights to this invention. Background of the Invention
[0009] Recent advances in genome sequencing technologies and analytical methods have significantly accelerated the ability to catalog and map genetic factors associated with a wide range of biological functions and diseases. Precise genome targeting technologies are needed to enable systematic reverse engineering of causal genetic variation by allowing selective interference with individual genetic elements, and to advance synthetic biology, biotechnology applications, and medical applications. While genome editing technologies, such as designer zinc fingers, transcription activator-like effector factors (TALEs), or homing meganucleases, are available for generating targeted genomic interference, new, affordable, easily implemented, scalable genome engineering techniques that can target multiple locations within the eukaryotic genome remain essential. Invention Overview
[0010] There is an urgent need for alternative and robust systems and techniques for sequence targeting with a wide range of applications. This invention addresses this need and provides relevant advantages. CRISPR / Cas or CRISPR-Cas systems (the two terms are used interchangeably throughout this application) do not require the production of custom proteins targeting specific sequences; instead, they recognize a specific DNA target via a short RNA molecule, allowing for the programming of a single Cas enzyme—in other words, the recruitment of a Cas enzyme to a specific DNA target using the short RNA molecule. Adding this CRISPR-Cas system to the repertoire of genome sequencing technologies and analytical methods can significantly simplify methodologies and improve the ability to catalog and map genetic factors associated with a wide range of biological functions and diseases. Understanding the engineering and optimization aspects of these genome engineering tools—which are the aspects of this claimed invention—is crucial for the effective and harmless use of CRISPR-Cas systems for genome editing.
[0011] In one aspect, the present invention provides a vector system comprising one or more vectors. In some embodiments, the system comprises: (a) a first regulatory element operatively linked to a tracr pairing sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr pairing sequence, wherein, upon expression, the guide sequences direct a CRISPR complex in eukaryotic cells to sequence-specific binding to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme conjugated to: (1) a guide sequence hybridizing to the target sequence, and (2) a tracr pairing sequence hybridizing to the tracr sequence; and (b) a second regulatory element operatively linked to an enzyme-coding sequence comprising a nuclear localization sequence encoding the CRISPR enzyme; wherein components (a) and (b) are located on the same or different vectors of the system. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr pairing sequence under the control of the first regulatory element. In some embodiments, component (a) further includes two or more guide sequences operatively linked to the first regulatory element, wherein, upon expression, each of the two or more guide sequences guides the CRISPR complex to bind sequence-specifically to a different target sequence in eukaryotic cells. In some embodiments, the system includes a tracr sequence under the control of a third regulatory element, such as a polymerase III promoter. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr pairing sequence when optimal alignment is performed. Determining optimal alignment is within the capabilities of a person of ordinary skill in the art. For example, publicly available and commercially available alignment algorithms and programs exist, such as, but not limited to, ClustalW, the Smith-Waterman algorithm in MATLAB, Bowtie, Geneious, Biopython, and SeqMan. In some embodiments, the CRISPR complex includes one or more nuclear localization sequences that have sufficient strength to drive the CRISPR complex to accumulate in a detectable amount in the nucleus of a eukaryotic cell. Without being bound by theory, it is assumed that nuclear localization sequences are not essential for the activity of CRISPR complexes in eukaryotes, but including such sequences enhances the activity of the system, particularly for targeting nucleic acid molecules in the cell nucleus. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or Streptococcus thermophilus Cas9, and may include mutated Cas9 derived from these organisms.The enzyme can be a Cas9 homologue or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands at the target sequence location. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guiding sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides in length, or between 10 and 30, or between 15 and 25, or between 15 and 20 nucleotides. Generally, and throughout this specification, the term "vector" refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked thereto. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. One type of vector is the "plasmid," which is a circular double-stranded DNA loop in which another DNA fragment can be inserted, for example, using standard molecular cloning techniques. Another type of vector is the viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package viruses (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by a virus used for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cells in which they are introduced. Other vectors (e.g., non-episodic mammalian vectors) integrate into the host cell's genome after introduction and thereby replicate along with the host genome. Moreover, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors." Common expression vectors used in recombinant DNA technology are typically in plasmid form.
[0012] Recombinant expression vectors may contain the nucleic acids of the present invention in a form suitable for nucleic acid expression in host cells. This means that these recombinant expression vectors contain one or more regulatory elements selected based on the host cell to be used for expression, said regulatory elements being operatively linked to the nucleic acid sequence to be expressed. Within the recombinant expression vector, "operatively linked" is intended to indicate that the nucleotide sequence of interest is linked to the one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0013] The term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Such regulatory sequences are described, for example, in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). Regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). Regulatory elements can also be directed to express in a time-dependent manner (e.g., in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue- or cell type-specific. In some embodiments, a vector comprises one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or combinations thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters. Examples of pol II promoters include, but are not limited to, the retro-transcribed Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer) [see, for example, Boshart et al., Cell 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the glycerol phosphokinase (PGK) promoter, and the EF1α promoter. Also covered by the term “regulatory element” are enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Molecular Cell Biology, Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proceedings of the National Academy of Sciences, Vol. 78(3), pp. 1527-31, 1981).Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including fusion proteins or peptides encoded by nucleic acids as described herein (e.g., regularly spaced clustered short palindromic repeats (CRISPR) transcripts, proteins, enzymes, their mutant forms, their fusion proteins, etc.).
[0014] Advantageous vectors include lentiviruses and adeno-associated viruses, and these types of vectors can also be selected to target specific cell types.
[0015] In one aspect, the present invention provides a vector comprising a regulatory element operatively linked to an enzyme-coding sequence encoding a CRISPR enzyme, the CRISPR enzyme comprising one or more nuclear localization sequences. In some embodiments, the regulatory element drives transcription of the CRISPR enzyme in a eukaryotic cell, causing the CRISPR enzyme to accumulate in a detectable amount in the nucleus of the eukaryotic cell. In some embodiments, the regulatory element is a polymerase II promoter. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or Streptococcus thermophilus Cas9, and may include mutated Cas9 derived from these organisms. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs the cutting of one or both strands at the target sequence location. In some embodiments, the CRISPR enzyme lacks DNA strand cutting activity.
[0016] In one aspect, the present invention provides a CRISPR enzyme comprising one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or Streptococcus thermophilus Cas9, and may include mutated Cas9 derived from these organisms. The enzyme may be a Cas9 homologue or ortholog. In some embodiments, the CRISPR enzyme lacks the ability to cleave one or more strands of a target sequence to which it is bound.
[0017] In one aspect, the present invention provides a eukaryotic host cell comprising: (a) a first regulatory element operatively linked to a tracr pairing sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr pairing sequence, wherein, when expressed, the guide sequences direct a CRISPR complex in the eukaryotic cell to bind sequence-specifically to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme complexed with: (1) a guide sequence hybridizing to the target sequence, and (2) a tracr pairing sequence hybridizing to the tracr sequence; and / or (b) a second regulatory element operatively linked to an enzyme-coding sequence including a nuclear localization sequence encoding the CRISPR enzyme. In some embodiments, the host cell comprises components (a) and (b). In some embodiments, components (a), (b), or (a) and (b) are stably integrated into a genome of the host eukaryotic cell. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr pairing sequence under the control of the first regulatory element. In some embodiments, component (a) further comprises two or more guide sequences operatively linked to the first regulatory element, wherein, upon expression, each of the two or more guide sequences guides the CRISPR complex to bind sequence-specifically to a different target sequence in a eukaryotic cell. In some embodiments, the eukaryotic host cell further comprises a third regulatory element, such as a polymerase III promoter, operatively linked to the tracr sequence. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr pairing sequence when optimal alignment is performed. In some embodiments, the CRISPR enzyme comprises one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or Streptococcus thermophilus Cas9, and may include mutated Cas9 derived from these organisms. The enzyme may be a Cas9 homologue or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands at the target sequence location. In some embodiments, the CRISPR enzyme lacks DNA strand cleavage activity. In some embodiments, the first regulatory element is a polymerase III promoter.In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides in length, or between 10 and 30, or between 15 and 25, or between 15 and 20 nucleotides. In one aspect, the invention provides a non-human eukaryote; preferably a multicellular eukaryote comprising a eukaryotic host cell according to any of the embodiments described. In other aspects, the invention provides a eukaryote; preferably a multicellular eukaryote comprising a eukaryotic host cell according to any of the embodiments described. In some embodiments of these aspects, the organism may be an animal; for example, a mammal. Additionally, the organism may be an arthropod, such as an insect. The organism may also be a plant. Further, the organism may be a fungus.
[0018] In one aspect, the present invention provides a kit comprising one or more of the components described herein. In some embodiments, the kit comprises a vector system and instructions for using the kit. In some embodiments, the vector system comprises: (a) a first regulatory element operatively linked to a tracr pairing sequence and one or more insertion sites for inserting one or more guide sequences upstream of the tracr pairing sequence, wherein, when expressed, the guide sequences direct a CRISPR complex in eukaryotic cells to bind sequence-specifically to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme conjugated to: (1) a guide sequence hybridizing to the target sequence, and (2) a tracr pairing sequence hybridizing to the tracr sequence; and / or (b) a second regulatory element operatively linked to an enzyme-coding sequence comprising a nuclear localization sequence encoding the CRISPR enzyme. In some embodiments, the kit comprises components (a) and (b) located on the same or different vectors of the system. In some embodiments, component (a) further comprises a tracr sequence downstream of the tracr pairing sequence under the control of the first regulatory element. In some embodiments, component (a) further includes two or more guide sequences operatively linked to the first regulatory element, wherein, upon expression, each of the two or more guide sequences guides the CRISPR complex to bind sequence-specifically to a different target sequence in eukaryotic cells. In some embodiments, the system further includes a third regulatory element, such as a polymerase III promoter, operatively linked to the tracr sequence. In some embodiments, the tracr sequence exhibits at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr pairing sequence when optimal alignment is performed. In some embodiments, the CRISPR enzyme includes one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of a eukaryotic cell. In some embodiments, the CRISPR enzyme is a type II CRISPR system enzyme. In some embodiments, the CRISPR enzyme is a Cas9 enzyme. In some embodiments, the Cas9 enzyme is Streptococcus pneumoniae, Streptococcus pyogenes, or Streptococcus thermophilus Cas9, and may include mutated Cas9 derived from these organisms. The enzyme may be a Cas9 homologue or ortholog. In some embodiments, the CRISPR enzyme is codon-optimized for expression in eukaryotic cells. In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands at the target sequence location. In some embodiments, the CRISPR enzyme lacks DNA strand cutting activity.In some embodiments, the first regulatory element is a polymerase III promoter. In some embodiments, the second regulatory element is a polymerase II promoter. In some embodiments, the guide sequence is at least 15, 16, 17, 18, 19, 20, or 25 nucleotides in length, or between 10 and 30, or between 15 and 25, or between 15 and 20 nucleotides.
[0019] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method includes allowing a CRISPR complex to bind to the target polynucleotide to perform a cleavage of the target polynucleotide, thereby modifying the target polynucleotide, wherein the CRISPR complex includes a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, wherein the guide sequence is linked to a tracr pairing sequence, which in turn hybridizes to a tracr sequence. In some embodiments, the cleavage includes cutting one or both strands at the target sequence location by the CRISPR enzyme. In some embodiments, the cleavage results in a reduction in transcription of the target gene. In some embodiments, the method further includes repairing the cleaved target polynucleotide by homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation, including the insertion, deletion, or substitution of one or more nucleotides of the target polynucleotide. In some embodiments, the mutation results in a change in one or more amino acids in a protein expressed from a gene containing the target sequence. In some embodiments, the method further includes delivering one or more vectors to the eukaryotic cells, wherein the one or more vectors drive the expression of one or more of the following: the CRISPR enzyme, a guide sequence linked to the tracr pairing sequence, and the tracr sequence. In some embodiments, the vector is delivered to eukaryotic cells within a subject. In some embodiments, the modification occurs in the eukaryotic cells in a cell culture. In some embodiments, the method further includes isolating the eukaryotic cells from the subject prior to the modification. In some embodiments, the method further includes returning the eukaryotic cells and / or cells derived therefrom to the subject.
[0020] In one aspect, the present invention provides a method for modifying the expression of a polynucleotide in eukaryotic cells. In some embodiments, the method includes allowing a CRISPR complex to bind to the polynucleotide, such that the binding results in an increase or decrease in the expression of the polynucleotide; wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, wherein the guide sequence is linked to a tracr pairing sequence, which in turn hybridizes to a tracr sequence. In some embodiments, the method further includes delivering one or more vectors to the eukaryotic cells, wherein the one or more vectors drive the expression of one or more of the following: the CRISPR enzyme, the guide sequence linked to the tracr pairing sequence, and the tracr sequence.
[0021] In one aspect, the present invention provides a method for generating model eukaryotic cells containing a mutated disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes (a) introducing one or more vectors into a eukaryotic cell, wherein the one or more vectors drive the expression of one or more of the following: a CRISPR enzyme, a guide sequence linked to a tracr pairing sequence, and a tracr sequence; and (b) allowing a CRISPR complex to bind to a target polynucleotide to perform a cleavage of the target polynucleotide within the disease gene, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guide sequence hybridized to the target sequence within the target polynucleotide and (2) the tracr pairing sequence hybridized to the tracr sequence, thereby generating model eukaryotic cells containing the mutated disease gene. In some embodiments, the cleavage comprises cleaving one or both strands at the target sequence location by the CRISPR enzyme. In some embodiments, the cleavage results in a reduction in transcription of the target gene. In some embodiments, the method further includes repairing the cleaved target polynucleotide via homologous recombination with an exogenous template polynucleotide, wherein the repair results in a mutation comprising the insertion, deletion, or substitution of one or more nucleotides of the target polynucleotide. In some embodiments, the mutation results in a change of one or more amino acids in a protein expressed from a gene containing the target sequence.
[0022] In one aspect, the present invention provides a method for developing a bioactive agent that modulates cell signaling events associated with a disease gene. In some embodiments, the disease gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the method includes (a) contacting a test compound with a model cell of any of the embodiments; and (b) detecting a change in reading indicating a decrease or increase in cell signaling events associated with the mutation of the disease gene, thereby developing the bioactive agent that modulates the cell signaling events associated with the disease gene.
[0023] In one aspect, the present invention provides a recombinant polynucleotide comprising a guide sequence upstream of a tracr pairing sequence, wherein, upon expression, the guide sequence directs a CRISPR complex to bind sequence-specifically to a corresponding target sequence present in a eukaryotic cell. In some embodiments, the target sequence is a viral sequence present in a eukaryotic cell. In some embodiments, the target sequence is a proto-oncogene or an oncogene.
[0024] In one aspect, the present invention provides a method for selecting one or more prokaryotic cells by introducing one or more mutations into a gene in one or more prokaryotic cells, the method comprising: introducing one or more vectors into one or more prokaryotic cells, wherein the one or more vectors drive the expression of one or more of the following: a CRISPR enzyme, a guide sequence linked to a tracr pairing sequence, a tracr sequence, and an editing template; wherein the editing template contains one or more mutations that eliminate CRISPR cleavage; allowing the editing template to homologously recombine with a target polynucleotide in the one or more cells to be selected; allowing a CRISPR complex to bind to the target polynucleotide to perform cleavage of the target polynucleotide within the gene, wherein the CRISPR complex comprises a CRISPR enzyme complexed with (1) a guide sequence hybridized to a target sequence within the target polynucleotide and (2) a tracr pairing sequence hybridized to the tracr sequence, wherein the binding of the CRISPR complex to the target polynucleotide induces cell death, thereby allowing selection of one or more prokaryotic cells in which one or more mutations have been introduced. In a preferred embodiment, the CRISPR enzyme is Cas9. In another aspect of the invention, the cells to be selected may be eukaryotic cells. Aspects of the present invention allow for the selection of specific cells without the need for selection markers or a two-step process that may include an anti-selection system.
[0025] Therefore, the object of this invention is to exclude any previously known products, processes for manufacturing such products, or methods of using such products, thereby allowing the applicant to reserve and hereby disclose the waiver of any rights to previously known products, processes, or methods. Furthermore, within the scope of this invention, it is not intended to cover any product, process, or method of manufacturing or using such product that does not meet the written description and implementability requirements of the USPTO (35 USC §112, paragraph 1) or EPO (Section 83 of the EPC), thereby allowing the applicant to reserve and hereby disclose the waiver of any previously described products, processes for manufacturing such products, or methods of using such products.
[0026] It should be noted that, in this disclosure and particularly in the claims and / or paragraphs, terms such as “comprises,” “comprised,” “comprising,” etc., may have the meanings that fall under their jurisdiction in U.S. patent law; for example, they may mean “includes,” “included,” “including,” etc.; and terms such as “consisting essentially of” and “consists essentially of” have the meanings that fall under their jurisdiction in U.S. patent law, for example, they allow elements not explicitly stated but exclude elements found in the prior art or affecting the essential or novel features of the invention. These and other embodiments are disclosed in the following detailed description, or are as obvious therein and are covered therein. Brief description of the attached figures
[0027] The novel features of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description of illustrative embodiments, in which the principles of the invention are utilized, and in these drawings:
[0028] Figure 1 A schematic model of the CRISPR system is shown. The Cas9 nuclease from *Streptococcus pyogenes* (yellow) is targeted to genomic DNA via a synthesized guide RNA (sgRNA) containing a 20 nt guide sequence (blue) and a scaffold (red). The guide sequence base pair with the DNA target (blue) is directly upstream of the necessary 5'-NGG prototype spacer adjacent motif (PAM; red-purple), and Cas9 mediates a double-strand break (DSB) approximately 3 bp upstream of this PAM (red triangle).
[0029] Figure 2A-F shows an exemplary CRISPR system, a possible mechanism of action, an exemplary adaptive alteration of expression in eukaryotic cells, and the results of an experiment assessing nuclear localization and CRISPR activity.
[0030] Figure 3 An exemplary expression cassette for expressing CRISPR system elements in eukaryotic cells, a predicted structure of an exemplary guide sequence, and CRISPR system activity as measured in eukaryotic and prokaryotic cells are shown.
[0031] Figure 4A -D shows the evaluation results of SpCas9 specificity for the exemplary target.
[0032] Figure 5A -G shows an exemplary vector system and its results in guiding homologous recombination in eukaryotic cells.
[0033] Figure 6 A table of prototype spacer sequences is provided, and modification efficiency results for prototype spacer targets designed based on exemplary *Streptococcus pyogenes* and *Streptococcus thermophilus* CRISPR systems with corresponding PAMs targeting loci in the human and mouse genomes are summarized. Cells were transfected with Cas9 as well as pre-crRNA / tracrRNA or chimeric RNA, and analyses were performed 72 hours after transfection. The percentage of indels (insertions / deletions) was calculated based on Surveyor analysis results from indicator cell lines (N=3 for all prototype spacer targets, error is SEM, ND indicates not detected by Surveyor assay, and NT indicates not detected in this study).
[0034] Figure 7A -C shows a comparison of different tracrRNA transcripts for Cas9-mediated gene targeting.
[0035] Figure 8 A schematic diagram of a manned nuclease assay for detecting microinsertions and microdeletions induced by double-strand breaks is shown.
[0036] Figure 9A -B shows an exemplary bicistronic expression vector for the expression of CRISPR system elements in eukaryotic cells.
[0037] Figure 10 This displays a bacterial plasmid transformation interference assay, the expression cassette and plasmid used therein, and the cell transformation rate used therein.
[0038] Figure 11A -C indicates the presence of 1PAM(NGG) at the adjacent Streptococcus pyogenes SF370 locus in the human genome. Figure 10 A) and Streptococcus thermophilus LMD9 locus 2PAM(NNAGAAW) ( Figure 10 The distance between B) and the distance (Chr) for each PAM relative to the chromosome (B) Figure 10 C) Histogram.
[0039] Figure 12A -C shows an exemplary CRISPR system, an exemplary adaptive alteration of expression in eukaryotic cells, and the results of an experiment evaluating CRISPR activity.
[0040] Figure 13A -C illustrates an exemplary manipulation of a CRISPR system for targeting genomic loci in mammalian cells.
[0041] Figure 14A -B shows the results of Northern blot analysis of crRNA processing in mammals.
[0042] Figure 15 Exemplary screening of prototype spacers in human PVALB and mouse Th loci is shown.
[0043] Figure 16 An exemplary prototype spacer and corresponding PAM sequence target of the thermophilic streptococcal CRISPR system in the human EMX1 locus are shown.
[0044] Figure 17 A table of primer and probe sequences for Surveyor, RFLP, genome sequencing, and Northern blotting analysis is provided.
[0045] Figure 18A -C shows the results of a SURVEYOR analysis of exemplary manipulation of a CRISPR system with chimeric RNA and the system's activity in eukaryotic cells.
[0046] Figure 19A -B shows a graph illustrating the results of a SURVEYOR analysis of CRISPR system activity in eukaryotic cells.
[0047] Figure 20 This paper presents an exemplary visualization of some Streptococcus pyogenes Cas9 target sites in the human genome using the UCSC Genome Browser.
[0048] Figure 21 This shows the predicted secondary structure of an exemplary chimeric RNA, including the guide sequence, the tracr pairing sequence, and the tracr sequence.
[0049] Figure 22An exemplary bicistronic expression vector for the expression of CRISPR system elements in eukaryotic cells is shown.
[0050] Figure 23 This demonstrates that Cas9 nuclease activity targeting endogenous targets can be used for genome editing. (a) Concept of genome editing using the CRISPR system. The CRISPR targeting construct guides the cutting of a chromosomal locus and is co-transformed with an editing template that recombines with the target to prevent cutting. Kanamycin-resistant transformants surviving CRISPR attack contain modifications introduced through the editing template: tracr, trans-activating CRISPR RNA; aphA-3, kanamycin resistance gene. (b) Cas9 nuclease activity against endogenous targets, R6 wild-type srtA, or R6370.1 editing templates in the absence of an editing template. 8232.5 Transformation of crR6 MDNA in cells. R6 srtA or R6 370.1 Recombination was prevented by Cas9 to prevent cleavage. Transformation efficiency was calculated as colony-forming units (cfu) / μg crR6M DNA; mean and standard deviation from at least three independent experiments are shown. PCR analysis was performed on 8 clones in each transformation. “Un.” indicates strain R6 8232.5 The unedited srtA seat; “Ed.” displays the edit template. R6 8232.5 and R6 370.1 The targets are distinguished by restriction enzyme cleavage with EaeI.
[0051] Figure 24 Analysis of PAM and seed sequences after Cas9 cleavage removal is shown. (a) PCR products with randomized PAM or randomized seed sequences were transformed into crR6 cells. These cells express crRNA loaded with crRNA (targeting R6). 8232.5(a) Cas9 of the chromosomal region of the cell (highlighted in pink), which is deleted from the R6 genome. More than 2 × 10⁵ chloramphenicol-resistant transformants (carrying inactive PAM or seed sequences) were combined for amplification and deep sequencing of the target region. (b) Relative proportion of reads after transformation of random PAM constructs in crR6 cells (compared to the number of reads in R6 transformants). The relative abundance of each 3-nucleotide PAM sequence is shown. Sequences with severe representative deficiency (NGG) are shown in red; sequences with partial representative deficiency are shown in orange (NAG). (c) Relative proportion of reads after transformation of random seed sequence constructs in crR6 cells (compared to the number of reads in R6 transformants). The relative abundance of each nucleotide at each position for the first 20 nucleotides of the prototype spacer is shown. High abundance indicates the absence of Cas9 cuts, i.e., a CRISPR inactivation mutation. The gray line shows the level of wild-type sequences. The dashed line represents the level above which a mutation significantly disrupts cleavage (see the "Deep Sequencing Data Analysis" section in Example 5).
[0052] Figure 25 This study demonstrates the introduction of single and multiple mutations using the CRISPR system in *Streptococcus pneumoniae*. (a) Nucleotide and amino acid sequences of wild-type and edited (green nucleotides; underlined amino acid residues) bgaA. The prototype spacer, PAM, and restriction enzyme sites are shown. (b) Transformation rates of cells transformed with the targeted construct in the presence of an edited template or control. (c) PCR analysis of eight transformants for each edit assay, followed by digestion with BtgZI (R→A) and TseI (NE→AA). The deletion of bgaA is shown as a small PCR product. (d) Miller assays were performed to measure the β-galactosidase activity of wild-type and edited strains. (e) For single-step, bidirectional deletions, the targeted construct contains two spacers (in this case matching srtA and bgaA) and is co-transformed with two different edited templates. (f) PCR analysis of eight transformants to detect deletions at the srtA and bgaA loci. Six out of eight transformants contained deletions of both genes.
[0053] Figure 26 Provide potential mechanisms for editing using the CRISPR system. (a) Introducing a stop codon into the erythromycin resistance gene ermAM to produce strain JEN53. The wild-type sequence can be repaired by targeting the stop codon with the CRISPR::ermAM (stop codon) construct and using the wild-type ermAM sequence as an editing template. (b) Mutant and wild-type ermAM sequences. (c) From total or kanamycin resistance (kan... RErythromycin resistance calculated by CFU (erm) R (d) The total cell score for both the CRISPR construct and the editing template. Co-transformation of the CRISPR construct with the target construct yielded more transformants (t-test, p = 0.011). In all cases, these values are shown as mean ± SD for three independent experiments.
[0054] Figure 27 This demonstrates the use of the CRISPR system for genome editing in *E. coli*. (a) A kanamycin resistance plasmid carrying a CRISPR array (pCRISPR) targeting the gene to be edited can be transformed into an HME63 recombinant strain containing a chloramphenicol resistance plasmid (containing cas9 and tracr (pCas9)) along with an oligonucleotide specifying the mutation. (b) Introducing a K42T mutation conferring streptomycin resistance into the rpsL gene. (c) From total or kanamycin resistance (kan... R Streptomycin resistance calculated by CFU (streptomycin resistance test) R (d) The total number of cells obtained from the pCRISPR plasmid and the edited oligonucleotide. Co-transformation of the pCRISPR targeting the plasmid produced more transformants (t-test, p = 0.004). In all cases, these values are shown as mean ± SD for three independent experiments.
[0055] Figure 28 The transformation of crR6 genomic DNA leads to editing of the targeted loci. (a) The IS1167 element of Streptococcus pneumoniae R6 is replaced by the CRISPR01 locus of Streptococcus pyogenes SF370 to produce the crR6 strain. This locus encodes the Cas9 nuclease, a CRISPR array with six spacers, the tracrRNA required for crRNA biologism, and Cas1, Cas2, and Csn2 (proteins not essential for targeting). The strain crR6M contains a minimally functional CRISPR system without Cas1, Cas2, and Csn2. The aphA-3 gene encodes kanamycin resistance. The prototype spacers of bacterial phages φ8232.5 and φ370.1 from Streptococcus are fused to a chloramphenicol resistance gene (cat) and integrated into the srtA gene of strain R6 to produce strains R68232.5 and R6370.1. (b) Left inset: genomic DNA of crR6 and crR6M in R6 8232.5 and R 6370.1 The transformation was performed. As a control for competent cells, a streptomycin resistance gene was also transformed. Right inset: 8R6 cells with crR6 genomic DNA. 8232.5PCR analysis of transformants. Primers amplifying the srtA locus were used for PCR. 7 / 8 of the genotyped colonies replaced the R68232.5 srtA locus with the WT locus from the crR6 genomic DNA.
[0056] Figure 29 Chromatograms of the DNA sequences of the edited cells obtained in this study are provided. In all cases, the wild-type and mutant prototype spacers and the PAM sequence (or its inverse complement) are indicated. The amino acid sequence encoded by the prototype spacer is provided where relevant. For each editing assay, all strains were sequenced, and PCR and restriction enzyme digestion analyses confirmed the introduction of the desired modification for these strains. Representative chromatograms are shown. (a) Introduction of a PAM mutation into the R6 8232.5 Chromatogram of the target ( Figure 24 a). (b) Chromatograms of the introduction of R>A and NE>AA mutations into β-galactosidase (bgaA) ( Figure 25 c). (c) Chromatogram of bgaA ORF with the 6664bp deletion introduced ( Figure 25 c and 25f). Dashed lines indicate the deletion limit. (d) Chromatogram of introducing a 729bp deletion into srtA ORF ( Figure 25 f). Dashed lines indicate the deletion limit. (e) Chromatogram of the generation of a premature stop codon within ermAM ( Figure 33 (f) Editing of rpsL in Escherichia coli ( Figure 27 ).
[0057] Figure 30 CRISPR immunization against random Streptococcus pneumoniae targets containing different PAMs is shown. (a) Locations of 10 random targets on the Streptococcus pneumoniae R6 genome. The selected targets have different PAMs and are on both strands. (b) Spacers corresponding to these targets were cloned in a minimal CRISPR array on plasmid pLZ12 and transformed into strain crR6Rc, which provides reverse processing and targeting mechanisms. (c) Transformation rates of different plasmids in strains R6 and crR6Rc. Transformation of pDB99-108 (T1-T10) in crR6Rc did not regenerate colonies. The dashed line indicates the limit of detection for this assay.
[0058] Figure 31A general protocol for targeted genome editing is provided. To facilitate targeted genome editing, crR6M was further engineered to contain tracrRNA, Cas9, and a single-repeat CRISPR array, followed by a kanamycin resistance marker (aphA-3), to generate the strain crR6Rk. DNA from this strain was used as a template for PCR, with primers designed to introduce a novel spacer (green box indicated by N). Left and right PCR were spliced using the Gibson method to construct the targeted construct. Both the targeted and editing constructs were then transformed into the strain crR6Rc, an equivalent of crR6Rk but with the kanamycin resistance marker replaced by a chloramphenicol resistance marker (cat). Approximately 90% of the kanamycin-resistant transformants contained the desired mutation.
[0059] Figure 32 The distance distribution between PAMs is illustrated. NGG and CCN are considered valid PAMs. Data for the Streptococcus pneumoniae R6 genome are shown along with a random sequence of the same length and GC content (39.7%). The dashed line represents the average distance between PAMs in the R6 genome (12).
[0060] Figure 33 This illustrates CRISPR-mediated editing of the ermAM locus using genomic DNA as a targeting construct. To use genomic DNA as a targeting construct, it is necessary to avoid CRISPR autoimmunity, and therefore a spacer targeting a sequence not present in the chromosome (in this case, the ermAM erythromycin resistance gene) must be used. (a) Nucleotide and amino acid sequences of the wild-type and mutant (highlighted in red) ermAM genes. The prototype spacer and PAM sequence are shown. (b) Schematic diagram of CRISPR-mediated editing of the ermAM locus using genomic DNA. A construct carrying an ermAM-targeting spacer (blue box) was prepared by PCR and Gibson assembly and transformed into strain crR6Rc to generate strain JEN37. The genomic DNA of JEN37 was then used as the targeting construct, and this editing template was co-transformed into JEN38, a strain in which the srtA gene is replaced by a wild-type copy of ermAM. The kanamycin-resistant transformants contained the edited genotype (JEN43). (c) Number of kanamycin-resistant cells obtained after co-transformation of the targeted, edited, or control templates. In the presence of a control template, 5.4 × 10⁴ cells were obtained. 3 The cfu / ml value was 4.3 × 10⁻⁶ when using the editing template. 5cfu / ml. This difference indicates an editing rate of approximately 99% [(4.3×10]]. 5 -5.4×10 3 ) / 4.3×10 5 (d) To check for the presence of edited cells, seven kanamycin-resistant clones and JEN38 were streaked onto agar plates with (erm+) or without (erm–) erythromycin. Only positive controls showed resistance to erythromycin. The ermAM mut genotype of one of these transformants was also verified by DNA sequencing. Figure 29 e).
[0061] Figure 34 The introduction of mutations via CRISPR-mediated genome editing sequence is illustrated. (a) Schematic diagram of the introduction of mutations via CRISPR-mediated genome editing sequence. First, R6 is engineered to generate crR6Rk. crR6Rk is co-transformed with an srtA-targeting construct fused to cat for chloramphenicol screening of edited cells, along with co-transformation with an editing construct targeting deletions within the ΔsrtA frame. The strain crR6ΔsrtA is generated through chloramphenicol screening. Subsequently, the ΔsrtA strain is co-transformed with a bgaA-targeting construct fused to aphA-3 for kanamycin screening of edited cells, and co-transformed with an editing construct containing deletions within the ΔbgaA frame. Finally, the engineered CRISPR locus can be removed from the chromosome by the following steps: first, co-transformation of R6 DNA containing the wild-type IS1167 locus with a plasmid carrying a bgaA prototype spacer (pDB97), and screening for spectinomycin. (b) PCR analysis of eight chloramphenicol (Cam) resistant transformants was performed to detect the deletion of the srtA locus. (c) β-galactosidase activity was measured by Miller assay. In Streptococcus pneumoniae, this enzyme is anchored to the cell wall via sorting enzyme A. The deletion of the srtA gene resulted in the release of β-galactosidase into the supernatant. The ΔbgaA mutant showed no activity. (d) PCR analysis of eight spectinomycin (Spec) resistant transformants was performed to detect the substitution of the wild-type IS1167 for the CRISPR locus.
[0062] Figure 35 The background mutation frequency of CRISPR in Streptococcus pneumoniae is shown. (a) JEN53 Alternatively, CRISPR::erm (terminator) targeting the transformation of the construct, with or without using the ermAM editing template. and the kan between CRISPR::erm (terminator) R The difference in CFU indicates that Cas9 cleavage kills unedited cells. (At 3 × 10⁻⁶) -3(a) Frequency of mutants escaping CRISPR interference in the absence of an editing template was observed. (b) PCR analysis of the CRISPR loci of the escapees showed that 7 / 8 had spacer deletions. (c) Escapee #2 carried a point mutation in cas9.
[0063] Figure 36 This illustrates recombination in *E. coli* using pCas9, an essential element of the *Streptococcus pyogenes* CRISPR locus 1. The plasmid contains tracrRNA, Cas9, and a leader sequence driving the crRNA array. These pCRISPR plasmids contain only the leader and the array. Using annealed oligonucleotides, spacers can be inserted between BsaI sites in the crRNA array. The oligonucleotide design is shown at the bottom. pCas9 carries chloramphenicol resistance (CmR) and is based on the low-copy pACYC184 plasmid backbone. pCRISPR is based on the high-copy-number pZE21 plasmid. Both plasmids are required because using this organism as a cloning host, it is impossible to construct a pCRISPR plasmid containing spacers targeting the *E. coli* chromosome if Cas9 is also present (which would kill the host).
[0064] Figure 37 This illustrates CRISPR-guided editing in *E. coli* MG1655. An oligonucleotide (W542) carrying a point mutation conferring both streptomycin resistance and CRISPR immunity, along with either a plasmid targeting rpsL (pCRISPR::rpsL) or a control plasmid, was used. The transformants were co-transformed into wild-type Escherichia coli strain MG1655 containing pCas9. Transformants were screened on media containing streptomycin or kanamycin. The dashed line indicates the limit of detection for this transformation assay.
[0065] Figure 38 The background mutation frequency of CRISPR in E. coli HME63 is shown. (a) Alternatively, the pCRISPR::rpsL plasmid can be transformed into HME63 competent cells. (The text abruptly ends here, likely due to an incomplete sentence or missing information.) -4 (a) Frequency of mutants escaping CRISPR interference was observed. (b) Amplification of the CRISPR array of the escapees showed an 8 / 8 missing spacer.
[0066] Figure 39A -D shows a circular depiction revealing the phylogenetic analysis of five Cas9 families, which include three large Cas9 families (approximately 1400 amino acids) and two small Cas9 families (approximately 1100 amino acids).
[0067] Figure 40A-F shows a linear depiction revealing the phylogenetic analysis of five Cas9 families, which include three large Cas9 families (approximately 1400 amino acids) and two small Cas9 families (approximately 1100 amino acids).
[0068] Figure 41A -M indicates the sequence when the mutation site is located within the SpCas9 gene.
[0069] Figure 42 The following is a schematic construct in which the transcriptional activation domain (VP64) is fused to Cas9 and there are two mutations (D10 and H840) in the catalytic domain.
[0070] Figure 43A -D shows genome editing via homologous recombination. (a) Schematic diagram of the SpCas9 cleavage enzyme with the D10A mutation in the RuvC I catalytic domain. (b) Schematic diagram of representative homologous recombination (HR) at the human EMX1 locus using sense or antisense single-stranded oligonucleotides as repair templates. The red arrows above indicate sgRNA cleavage sites; the PCR primers used for genotyping (Tables J and K) are shown as arrows in the right figure. (c) Sequence of the region modified by HR. d, SURVEYOR analysis (n=3) of SpCas9-mediated indels at the EMX1 target 1 locus for wild-type (wt) and cleavage enzyme (D10A). Arrows indicate the location of the expected fragment size.
[0071] Figure 44A -B shows a single carrier design for SpCas9.
[0072] Figure 45 The quantization of cleavage of the NLS-Csn1 constructs NLS-Csn1, Csn1, Csn1-NLS, NLS-Csn1-NLS, NLS-Csn1-GFP-NLS, and UnTFN is shown.
[0073] Figure 46 Displays the index frequencies of NLS-Cas9, Cas9, Cas9-NLS, and NLS-Cas9-NLS.
[0074] Figure 47 A gel is shown that SpCas9 with a nicking enzyme mutation (alone) does not induce double-strand breaks.
[0075] Figure 48 This study shows a design of oligoDNA used as a template for homologous recombination (HR) in this experiment and a comparison of HR rates induced by different combinations of Cas9 protein and HR template.
[0076] Figure 49AThe map of the conditional Cas9 and Rosa26 targeting vectors is displayed.
[0077] Figure 49B The compositional Cas9 and Rosa26 targeting vector maps are shown.
[0078] Figure 50A -H indicates that it exists. Figure 49A The sequence of each element in the carrier spectrum of -B.
[0079] Figure 51 A schematic diagram of the key components in compositional and conditional Cas9 constructs is shown.
[0080] Figure 52 This demonstrates the functional confirmation of the expression of constitutive and conditional Cas9 constructs.
[0081] Figure 53 The activity of the Cas9 nuclease was confirmed by Surveyor.
[0082] Figure 54 This demonstrates the quantification of Cas9 nuclease activity.
[0083] Figure 55 Display the architecture design and homology recombination (HR) strategy.
[0084] Figure 56 This shows the genomic PCR genotyping results of constitutive (right) and conditional (left) constructs at two different gel exposure times (3 min for the top row and 1 min for the bottom row).
[0085] Figure 57 The Cas9 activation in mESC is shown.
[0086] Figure 58 This diagram illustrates a strategy for gene knockout mediated by NHEJ using a nickase version of Cas9 along with two guide RNAs.
[0087] Figure 59 This demonstrates how DNA double-strand break (DSB) repair facilitates gene editing. In the error-prone non-homologous end joining (NHEJ) pathway, the ends of DSBs are processed and rejoined by endogenous DNA repair mechanisms, potentially leading to random indel (indel) mutations at the joining sites. Indel mutations occurring within the coding region of a gene can produce frameshifts and premature stop codons, resulting in gene knockout. Alternatively, repair templates in plasmid or single-stranded oligodeoxynucleotide (ssODN) form can be provided to utilize the homology-directed repair (HDR) pathway, which allows for high-fidelity and precise editing.
[0088] Figure 60This section shows the timeline and summary of the experiments, outlining the steps involved in reagent design, construction, validation, and cell line expansion. Custom sgRNAs (light blue bars) for each target, along with genotyping primers, were designed in a computer (in silico) using an online design tool (available at genome-engineering.org / tools). The sgRNA expression vector was then cloned into a plasmid containing Cas9 (PX330) and validated by DNA sequencing. The complete plasmid (pCRISPR) and an optional repair template promoting homology-directed repair were then transfected into cells, and the ability to mediate targeted cleavage was assessed. Finally, the transfected cells were clonally expanded to obtain syngeneic cell lines with defined mutations.
[0089] Figure 61 AC shows target screening and reagent preparation. (a) The 20-bp target (highlighted in blue) for *Streptococcus pyogenes* Cas9 must be followed by 5'-NGG, which can occur on any strand of the genomic DNA. We recommend using the online tool described in this protocol to assist in target screening (www.genome-engineering.org / tools). (b) Schematic diagram of co-transfection of the Cas9 expression plasmid (PX165) and the PCR-amplified U6-driven sgRNA expression cassette. Using a PCR template containing the U6 promoter and a fixed forward primer (U6 Fwd), the sgRNA-encoding DNA can be attached to the U6 reverse primer (U6 Rev) and synthesized as an extended DNA oligomer (Ultramer oligomer from IDT). Note that the guide sequence (blue N) in U6 Rev is the reverse complement of the 5'-NGG flanking target sequence. (c) Schematic diagram of scarless cloning of the guide sequence oligomer into a plasmid containing Cas9 and the sgRNA scaffold (PX330). The guide oligomer (blue N) contains overhangs to attach to paired BbsI sites on PS330, with the top and bottom strands oriented to match those of the genomic target (i.e., the top oligomer is the 20-bp sequence in the genomic DNA preceding the 5'-NGG). Digestion of PX330 with BbsI allows the type II restriction enzyme sites (blue outlines) to be replaced by unidirectional insertion of the annealed oligomer. Notably, an extra G is placed before the first base of this guide sequence. The applicant has found that an extra G preceding this guide sequence does not adversely affect the targeting rate. When the selected 20-nt guide sequence does not begin with guanine, the extra guanine will ensure that the sgRNA is efficiently transcribed by the U6 promoter, preferably as a guanine in the first base of the transcript.
[0090] Figure 62AD shows the expected results for multivariate NHEJ. (a) Schematic diagram of the SURVEYOR assay used to determine the percentage of indels. First, genomic DNA from a heterogeneous population of Cas9-target cells was amplified by PCR. The amplicons were then slowly re-annealed to produce heteroduplexes. These re-annealed heteroduplexes were cleaved by the SURVEYOR nuclease, while the homoduplexes remained intact. The Cas9-mediated cleavage rate (%indel) was calculated based on the fraction of DNA cleaved, as determined by the cumulative intensity of the gel bands. (b) Two sgRNAs (orange and blue bars) were programmed to target the human GRIN2B and DYRK1A loci. The SURVEYOR gel shows the modifications at the two loci in transfected cells. Colored arrows indicate the predicted fragment size for each locus. (c) A pair of sgRNAs (light blue and green bars) were programmed to excise the exon (dark blue) at the human EMX1 locus. The target sequence and PAM (red) are shown in their respective colors, and red triangles indicate the cleavage sites. The predicted junctions are shown below. Individual clones isolated from cell populations transfected with sgRNAs 3, 4, or both were determined by PCR (external Fwd, external Rev) to reflect deletions of approximately 270 bp. Representative clones with no modification (12 / 23), monoallelic modification (10 / 23), and bialelic modification (1 / 23) are shown. Internal Fwd and internal Rev primers were used for screening against inversion events. (d) Quantification of clone lines with EMX1 exon deletions. Two pairs of sgRNAs (3.1, 3.2 left-wing sgRNA; 4.1, 4.2 right-wing sgRNA) were used to mediate a variable-sized deletion around an EMX1 exon. Transfected cells were clonally isolated and amplified for genotyping analysis against deletions and inversion events. 105 clones were screened, of which 51 (49%) and 11 (10%) carried heterozygous and homozygous deletions, respectively. Since the contact points can be variable, an approximate missing size is given.
[0091] Figure 63A -C shows that the application of ssODN and targeting vectors to mediate the effects of wild-type and nickase mutant Cas9 in HR, HEK293FT and HUES9 cells has an efficiency range of 1.0%-27%.
[0092] Figure 64This diagram illustrates a PCR-based method for rapid and efficient CRISPR targeting in mammalian cells. A plasmid containing the human RNA polymerase III promoter U6 is amplified by PCR using a U6-specific forward primer, an inverse complementary pair carrying a portion of the U6 promoter, an sgRNA (+85) scaffold with a guide sequence, and a 7-T nucleotide inverse primer for transcription termination. The resulting PCR product is purified and co-delivered using a plasmid carrying a Cas9 promoter driven by the CBh promoter.
[0093] Figure 65 The results from the Transgenomics SURVEYOR mutation detection kit are shown for each gRNA and its respective control. A positive SURVEYOR result corresponds to a large band and two smaller bands in the genomic PCR, which are the products of the SURVEYOR nuclease that causes a double-strand break at a mutation site. Each gRNA was validated in the mouse cell line Neuro-N2a by transient co-transfection with hSpCas9 via liposomes. Genomic DNA was purified using QuickExtract DNA from Epicentre 72 hours post-transfection. PCR was performed to amplify the loci of interest.
[0094] Figure 66 The Surveyor results are shown for 38 live pups (Dates 1-38), 1 dead pup (Date 39), and 1 wild-type pup for comparison (Date 40). Pups 1-19 were injected with gRNA Chd8.2, and pups 20-38 were injected with gRNA Chd8.3. Thirteen of the 38 live pups were positive for a mutation. The dead pup also had a mutation. The mutation was not detected in the wild-type sample. Genomic PCR sequencing findings were consistent with the Surveyor assay.
[0095] Figure 67 This illustrates a design of different Cas9 NLS constructs. All Cas9s are human-codon-optimized versions of Sp Cas9. The NLS sequence is ligated to the Cas9 gene at either the N-terminus or C-terminus. All Cas9 variants with different NLS designs are cloned into a backbone vector containing the EF1a promoter, thus being driven by the EF1a promoter. Within the same vector, a chimeric RNA targeting the human EMX1 locus and driven by the U6 promoter is present, forming a two-component system.
[0096] Figure 68This shows the efficiency of genome cleavage induced by Cas9 variants with different NLS designs. The percentage indicates the portion of human EMX1 genomic DNA cleaved by each construct. All experiments were performed from 3 biological replicates. n=3, and the error is expressed as SEM.
[0097] Figure 69 A design shows a CRISPR-TF (transcription factor) with transcriptional activation activity. The chimeric RNA is expressed by the U6 promoter, while a human-codon-optimized double-mutant version of the Cas9 protein (hSpCas9m), operably linked to the three NLS and VP64 functional domains, is expressed by an EF1a promoter. This double mutation, D10A and H840A, results in the Cas9 protein being unable to introduce any cleavage when directed by the chimeric RNA, but maintaining its ability to bind to the target DNA.
[0098] Figure 69 B shows transcriptional activation of the human SOX2 gene using the CRISPR-TF system (chimeric RNA and Cas9-NLS-VP64 fusion protein). 293FT cells were transfected with a plasmid containing two components: (1) a U6-driven, different chimeric RNA targeting a 20-bp sequence within or around the human SOX2 genomic locus, and (2) an EF1a-driven hSpCas9m (double mutant)-NLS-VP64 fusion protein. 293FT cells were harvested 96 hours post-transfection, and activation levels were measured by introducing mRNA expression using a qRT-PCR assay. All expression levels were normalized against a control group (grey bars), representing results from cells transfected with this CRISPR-TF backbone plasmid (without chimeric RNA). The qRT-PCR probe used to detect SOX2 mRNA was the Taqman Human Gene Expression Assay Probe (Life Technologies). All experiments represent data from three biological replicates, n=3, and error bars are shown as sem.
[0099] Figure 70 An optimization of the NLS construction for SpCas9 is described.
[0100] Figure 71 Displays a QQ plot for the NGGNN sequence.
[0101] Figure 72 A histogram showing the data density with a fitted normal distribution (black line) and the .99 quantile (dashed line).
[0102] Figure 73AC shows RNA-guided repression of bgaA expressed by dgRNA::cas9**. a. The Cas9 protein binds to the tracrRNA and then to the precursor CRISPR RNA processed by RNase III to form crRNA. This crRNA guides Cas9 to bind to the bgaA promoter and represses transcription. b. The targets used to guide Cas9** to the bgaA promoter are shown. The assumed -35, -10, along with the bgaA start codon, are bolded. c. β-galactosidase activity measured by Miller assay in the absence of a target and for four different targets.
[0103] Figure 74 AE shows the characterization of Cas9**-mediated repression. a. The gfpmut2 gene and its promoter (including -35 and -10 signals) are shown alongside the locations of the different target sites used in this study. b. Relative blotting when targeting the coding strand. c. Relative blotting when targeting the non-coding strand. d. Northern blotting of RNA extracted from T5, T10, B10, or control strains without targets using probes B477 and B478. e. The effect of increasing the number of mutations at the 5' end of crRNA in B1, T5, and B10.
[0104] The figures in this article are for illustrative purposes only and are not necessarily drawn to scale.
[0105] Detailed Description of the Invention
[0106] The terms “polynucleotide,” “nucleotide,” “nucleotide sequence,” “nucleic acid,” and “oligonucleotide” are used interchangeably. They refer to polymeric forms of nucleotides of any length, which are deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of genes or gene fragments, multiple loci (one locus) as defined by ligation analysis, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, the nucleotide structure may be modified before or after polymer assembly. The sequence of nucleotides may be interrupted by non-nucleotide components. Polynucleotides may be further modified after polymerization, such as by conjugation with labeled components.
[0107] In various aspects of this invention, the terms "chimeric RNA," "chimeric guide RNA," "guide RNA," "single guide RNA," and "synthetic guide RNA" are used interchangeably and refer to a polynucleotide sequence comprising a guide sequence, a tracr sequence, and a tracr pairing sequence. The term "guide sequence" refers to a sequence of approximately 20 bp within the guide RNA at a specified target site and is used interchangeably with the terms "guide" or "spacer." The term "tracr pairing sequence" is also used interchangeably with the term "(one or more) unidirectional repeats."
[0108] As used herein, the term "wild type" is a term understood by those skilled in the art and refers to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature.
[0109] As used in this article, the term "variant" should be understood as a representation of a property that is derived from a pattern existing in nature.
[0110] The terms “non-naturally occurring” and “engineered” are used interchangeably and imply artificial involvement. These terms, when referring to nucleic acid molecules or polypeptides, indicate that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component bound to it, either naturally occurring or as found in nature.
[0111] "Complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via traditional Watson-Crick base pairing or other non-traditional types. The complementarity percentage indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary, respectively). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0112] As used herein, “strict conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence. A non-limiting example of strict conditions is described in Tijssen (1993), *Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization With Nucleic Acid Probes*, Part I, Chapter 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, New York.
[0113] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of these nucleotide residues. Hydrogen bonds can occur via Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization can be a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complement" of that given sequence.
[0114] As used herein, “expression” refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as “gene products.” If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.
[0115] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to a polymer having amino acids of any length. The polymer may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acid components. These terms also cover polymers containing modified amino acids; such modifications include disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as binding to a labeled component. As used herein, the term “amino acid” includes natural and / or non-natural or synthetic amino acids, including glycine and its D and L optical isomers, as well as amino acid analogs and peptide mimics.
[0116] The terms “subject,” “individual,” and “patient” are used interchangeably herein and refer to vertebrates, preferably mammals, and more preferably humans. Mammals include, but are not limited to, rodents, monkeys, humans, livestock, sporting animals, and pets. It also includes tissues, cells, and progeny of a biological entity obtained in vivo or cultured in vitro.
[0117] The terms "therapeutic agent," "therapeutic capable agent," or "treatment agent" are used interchangeably and refer to a molecule or compound that, when administered to a subject, imparts a beneficial effect. This beneficial effect includes the achievement of a confirmed diagnosis; improvement of a disease, symptom, disorder, or pathological condition; reduction or prevention of the onset of a disease, symptom, disorder, or condition; and overall counteraction against a disease, symptom, disorder, or pathological condition.
[0118] As used herein, “treatment,” “treating,” “relief,” or “improvement” are used interchangeably. These terms refer to a pathway used to obtain a beneficial or desired outcome, including, but not limited to, a therapeutic benefit and / or a preventive benefit. A therapeutic benefit means any treatment-related improvement or effect on one or more diseases, conditions, or symptoms during treatment. For a preventive benefit, the composition may be given to a subject at risk of developing a specific disease, condition, or symptom, or to a subject who reports one or more physiological symptoms of a disease, even if the disease, condition, or symptom may not yet be present.
[0119] The term "effective amount" or "therapeutic effective amount" refers to an amount of a pharmaceutical agent sufficient to achieve a beneficial or desired result. Therapeutic effective amounts can vary depending on one or more of the subject being treated and the condition of the disease, the subject's weight and age, the severity of the disease, the route of administration, etc., and can be readily determined by one of ordinary skill in the art. This term also applies to a dose that provides an image for detection by any of the imaging methods described herein. The specific dose can vary depending on one or more of the following: the specific pharmaceutical agent selected, the administration regimen followed, whether it is administered in combination with other compounds, the timing of administration, the tissue to be imaged, and the physical delivery system carrying it.
[0120] Unless otherwise stated, the practice of this invention employs conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of the art. See Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd ed. (1989); *Current Protocols in Molecular Biology* (edited by F.M. Ausubel et al., (1987)); *Methods in Enzymology* series (academic publishing company): *PCR 2: A PRACTICAL APPROACH* (edited by M.J. MacPherson, BD. Hames, and GR. Taylor, (1995)), Harlow and Lane, (1988); *Antibodies, A Laboratory Manual*. MANUAL, and ANIMAL CELL CULTURE (edited by R.R. Freshney (1987)).
[0121] Several aspects of this invention relate to vector systems comprising one or more vectors, or the vectors themselves. Vectors can be designed for the expression of CRISPR transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, CRISPR transcripts can be expressed in bacterial cells such as *Escherichia coli*, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are further discussed in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). Alternatively, recombinant expression vectors can be transcribed and translated in vitro, for example, using a T7 promoter regulatory sequence and T7 polymerase.
[0122] Vectors can be introduced into and proliferated in prokaryotes. In some embodiments, prokaryotes are used to amplify multiple copies of a vector to be introduced into eukaryotic cells, or as an intermediate vector in the production of a vector to be introduced into eukaryotic cells (e.g., amplifying plasmids as part of a viral vector packaging system). In some embodiments, prokaryotes are used to amplify multiple copies of a vector and express one or more nucleic acids, such as providing a source for delivery to one or more proteins into a host cell or host organism. Protein expression in prokaryotes is most often carried out in *E. coli* using vectors containing constitutive or inducible promoters that direct the expression of fusion or non-fusion proteins. Fusion vectors add multiple amino acids to the protein encoded therein, such as the amino terminus of the recombinant protein. Such fusion vectors can be used for one or more purposes, such as: (i) increasing the expression of the recombinant protein; (ii) increasing the solubility of the recombinant protein; and (iii) assisting in the purification of the recombinant protein by acting as a ligand in affinity purification. Typically, in fusion expression vectors, protein cleavage sites are introduced at the junction of the fusion moiety and the recombinant protein to allow the recombinant protein to be separated from the fusion moiety after purification of the fusion protein. These enzymes and their homologous recognition sequences include factor Xa, thrombin, and enterokinase. Exemplary fusion expression vectors include pGEX (Pharmacia Biotech Inc.; Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, Massachusetts), and pRIT5 (Pharmacia, Piscataway, New Jersey), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A, respectively, to the target recombinant protein.
[0123] Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., Gene, 69:301-315, 1988) and pET 11d (Studier et al., Gene Expression Technology: Methods in Enzymology, 185, Academic Press, San Diego, California, 1990, 60-89).
[0124] In some embodiments, the vector is a yeast expression vector. Examples of vectors used for expression in yeast brewing include pYepSec1 (Baldari et al., 1987. EMBO J 6:229-234), pMFa (Kurjan and Herskowitz, 1982. Cell 30:933-943), pJRY88 (Schultz et al., 1987. Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, Calif), and picZ (Invitrogen Corp., San Diego, Calif).
[0125] In some embodiments, baculovirus vectors are used to drive protein expression in insect cells. Baculovirus vectors that can be used to express proteins in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith et al., 1983. Molecular Cell Biology 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39).
[0126] In some embodiments, mammalian expression vectors are used, which are capable of driving the expression of one or more sequences in mammalian cells. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329:840) and pMT2PC (Kaufman et al., 1987. EMBO J. 6:187-195). When used in mammalian cells, the control function of the expression vector is typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyomavirus, adenovirus 2, cytomegalovirus, simian virus 40, and other viruses described herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of *Molecular Cloning: A Laboratory Manual* (2nd ed.), Sambrook et al., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989.
[0127] In some embodiments, the recombinant mammalian expression vector can direct the preferential expression of nucleic acids in specific cell types (e.g., using tissue-specific regulatory elements to express nucleic acids). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include albumin promoters (liver-specific; Pinkert et al., 1987. Genes Dev. 1:268-277), lymphocyte-specific promoters (Calame and Eaton, 1988. Adv. Immunol 43:235-275), and particularly T-cell receptor promoters (Winoto and Baltimore, 1989. Journal of the European Society for Molecular Biology). J.) 8:729-733) and immunoglobulins (Baneiji et al., 1983. Cell 33:729-740; Queen and Baltimore, 1983. Cell 33:741-748), neuron-specific promoters (e.g., neurofilament promoters; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86:5473-5477), pancreas-specific promoters (Edlund et al., 1985. Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoters; U.S. Patent No. 4,873,316 and European Application Publication No. 264,166). It also covers developmental regulatory promoters, such as the hox promoters of murine homologous proteins (Kessel and Gruss, 1990. Science 249:374-379) and the alpha-fetoprotein promoter (Campes and Tilghman, 1989. Genes Dev. 3:537-546).
[0128] In some embodiments, a regulatory element is operatively connected to one or more elements of a CRISPR system to drive the expression of those elements. Generally, CRISPR (regularly spaced clustered short palindromic repeats), also known as SPIDR (spacer-spaced parallel repeats), constitutes a family of DNA loci that are typically specific to a particular bacterial species. These CRISPR loci contain a distinct class of spaced short sequence repeats (SSRs) recognized in Escherichia coli (Ishino et al., J. Bacteriol., 169:5429-5433
[1987] ; and Nakata et al., J. Bacteriol., 171:3553-3556
[1989] ), as well as associated genes. Similar spaced SSRs have been identified in *Haloferax mediterranei*, *Streptococcus pyogenes*, *Anabaena*, and *Mycobacterium tuberculosis* (see Groenen et al., *Molecular Microbiology*, 10:1057-1065
[1993] ; Hoe et al., *Emerging Infectious Diseases*, 5:254-263
[1999] ; Masepohl et al., *Biochim. Biophys. Acta*, 1307:26-30
[1996] ; and Mojica et al., *Molecular Microbiology*, 17:85-93
[1995] ). These CRISPR loci are typically distinct from the repeating structures of other SSRs, which have been termed regularly spaced short repeats (SRSRs) (Janssen et al., OMICS J. Integ. Biol., 6:23-33
[2002] ; and Mojica et al., Molecular Microbiology, 36:244-246
[2000] ). Generally, these repeats are short elements present in clusters, regularly spaced by distinctive intercalation sequences of substantially constant length (Mojica et al.,
[2000] , ibid.). While the repeat sequences are highly conserved across strains, many of the spaced repeats and the sequences of these spacers generally differ between strains (van Embden et al., J. Bacteriol.).(182:2393-2401
[2000] ). CRISPR loci have been identified in more than 40 prokaryotes (see, for example, Janssen et al., Molecular Microbiology).), 43:1565-1575
[2002] ; and Mojica et al.,
[2005] ), including but not limited to: Aeropyrum, Pyrobaculum, Sulfolobus, Archaeoglobus, Halocarcula, Methanobacterium, Methanococcus, Methanosarcina, Methanopyrus, and pyrophorus. Genus: Pyrococcus, Picrophilus, Thermoplasma, Corynebacterium, Mycobacterium, Streptomyces, Aquifex, Porphyromonas, Chlorobium, Thermus, Bacillus, Listeria, Staphylococcus, Clostridium *Clostridium*, *Thermoanaerobacter*, *Mycoplasma*, *Fusobacterium*, *Azarcus*, *Chromobacterium*, *Neisseria*, *Nitrosomonas*, *Desulfovibrio*, *Geobacter*, *Myxococcus*, *Campylobacter*, and other genera. Genus *Wolinella*, *Acinetobacter*, *Erwinia*, *Escherichia*, *Legionella*, *Methylococcus*, *Pasteurella*, *Photobacterium*, *Salmonella*, *Xanthomonas*, *Yersinia*, *Treponema*, and *Thermotoga*.
[0129] Generally, the term "CRISPR system" refers to transcripts and other elements involved in the expression of CRISPR-related ("Cas") genes or directing their activity, including sequences encoding Cas genes, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or active tracrRNA), tracr pairing sequences (covering "direct repeats" and partial direct repeats processed by tracrRNA in the context of an endogenous CRISPR system), directing sequences (also referred to as "spacers" in the context of an endogenous CRISPR system), or other sequences and transcripts derived from CRISPR loci. In some embodiments, one or more elements of the CRISPR system are derived from a type I, II, or III CRISPR system. In some embodiments, one or more elements of the CRISPR system are derived from a specific organism containing an endogenous CRISPR system, such as Streptococcus pyogenes. Generally, a CRISPR system is characterized by elements that promote the formation of a CRISPR complex (also referred to as a prototypical spacer in the context of an endogenous CRISPR system) at the site of the target sequence. In the context of CRISPR complex formation, a "target sequence" refers to a guide sequence designed to be complementary to it, wherein hybridization between the target sequence and the guide sequence promotes the formation of a CRISPR complex. Perfect complementarity is not required, provided that sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR complex. A target sequence can comprise any polynucleotide, such as a DNA or RNA polynucleotide. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. A sequence or template that can be used for recombination into a target locus including the target sequence is referred to as an "edit template," "edit polynucleotide," or "edit sequence." In aspects of the invention, the exogenous template polynucleotide may be referred to as an edit template. In one aspect of the invention, the recombination is homologous recombination.
[0130] Typically, in the context of an endogenous CRISPR system, the formation of a CRISPR complex (containing a guide sequence that hybridizes to a target sequence and complexes with one or more Cas proteins) results in the cleavage of one or both strands in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs). Not wishing to be bound by theory, the tracr sequence (which may contain or constitute all or part of a wild-type tracr sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of a wild-type tracr sequence)) can also form part of a CRISPR complex, such as by hybridizing along at least a portion of the tracr sequence to all or part of a tracr pairing sequence operatively linked to the guide sequence. In some embodiments, the tracr sequence has sufficient complementarity with a tracr pairing sequence to hybridize and participate in the formation of a CRISPR complex. As for the target sequence, complete complementarity is believed to be unnecessary, only sufficient to be functional. In some embodiments, when optimal alignment is performed, the tracr sequence has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% sequence complementarity along the length of the tracr pairing sequence. In some embodiments, one or more vectors driving the expression of one or more elements of the CRISPR system are introduced into a host cell such that the expression of these elements of the CRISPR system directs the formation of the CRISPR complex at one or more target sites. For example, the Cas enzyme, the guide sequence linked to the tracr pairing sequence, and the tracr sequence may each be operatively linked to separate regulatory elements on separate vectors. Alternatively, two or more elements expressed from the same or different regulatory elements may be bound in a single vector, wherein one or more additional vectors providing any component of the CRISPR system are not included in the first vector. The CRISPR system elements bound in a single vector may be arranged in any suitable orientation, such as one element being located at 5' (“upstream”) or 3' (“downstream”) relative to a second element. The coding sequence of one element may be located on the same or opposite strand of the coding sequence of the second element, and oriented in the same or opposite directions. In some embodiments, a single promoter drives the expression of one or more of the transcript encoding the CRISPR enzyme, the guide sequence, the tracr pairing sequence (optionally operably linked to the guide sequence), and the tracr sequence embedded within one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron).In some embodiments, the CRISPR enzyme, the guide sequence, the tracr pairing sequence, and the tracr sequence are operatively linked to and expressed from the same promoter.
[0131] In some embodiments, a vector contains one or more insertion sites, such as restriction endonuclease recognition sequences (also known as “cloning sites”). In some embodiments, one or more insertion sites (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more insertion sites) are located upstream and / or downstream of one or more sequence elements of one or more vectors. In some embodiments, a vector contains an insertion site upstream of a tracr pairing sequence and optionally downstream of a regulatory element operatively linked to the tracr pairing sequence, such that, after a guide sequence is inserted into the insertion site and during expression, the guide sequence guides the CRISPR complex to sequence-specific binding to a target sequence in eukaryotic cells. In some embodiments, a vector contains one or more insertion sites, each located between two tracr pairing sequences, thereby allowing the insertion of a guide sequence at each site. In such an arrangement, two or more guide sequences may comprise two or more copies of a single guide sequence, two or more different guide sequences, or a combination thereof. When using multiple different guide sequences, a single expression construct can be used to target CRISPR activity to multiple different corresponding target sequences within the cell. For example, a single vector may contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more guide sequences. In some embodiments, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more such vectors containing target sequences may be provided and optionally delivered into cells.
[0132] In some embodiments, a vector includes a regulatory element operatively linked to an enzyme-coding sequence encoding a CRISPR enzyme (such as the Cas protein). Non-limiting examples of Cas proteins include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, their homologues, or modified forms thereof. These enzymes are known; for example, the amino acid sequence of the *Streptococcus pyogenes* Cas9 protein is available in the SwissProt database accession number Q99ZW2. In some embodiments, an unmodified CRISPR enzyme, such as Cas9, has DNA cleaving activity. In some embodiments, the CRISPR enzyme is Cas9, and may be Cas9 derived from *Streptococcus pyogenes* or *Streptococcus pneumoniae*. In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands at a target sequence location (e.g., within the target sequence and / or within the complement of the target sequence). In some embodiments, the CRISPR enzyme directs the cleavage of one or both strands within approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, a vector encodes a CRISPR enzyme mutated relative to the corresponding wild-type enzyme, such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. For example, an aspartic-to-alanine substitution (D10A) in the RuvC I catalytic domain of Cas9 from Streptococcus pyogenes transforms Cas9 from a two-strand nuclease into a nicking enzyme (cutting a single strand). Other mutant examples that make Cas9 a nicking enzyme include, but are not limited to, H840A, N854A, and N863A. In some embodiments, a Cas9 nicking enzyme may be used in combination with one or more guide sequences, for example, two guide sequences that target the sense and antisense strands of the DNA target, respectively. This combination allows both strands to be nicked and used to induce NHEJ. The applicant has demonstrated (data not shown) the efficacy of two nicking enzyme targets (i.e., sgRNAs targeting different strands of DNA at the same location) in inducing mutagenic NHEJ.A single-nickelase (Cas9-D10A with a single sgRNA) was unable to induce NHEJ and establish an indel, but the applicant has shown that a dual-nickelase (Cas9-D10A and two sgRNAs targeting different strands at the same location) can do so in human embryonic stem cells (hESCs). This efficiency in hESCs is approximately 50% that of nucleases (i.e., conventional Cas9 without the D10 mutation).
[0133] As another example, two or more catalytic domains of Cas9 (RuvC I, RuvC II, and RuvCIII) can be mutated to produce a Cas9 that substantially lacks all DNA cleavage activity. In some embodiments, the D10A mutation is combined with one or more of the H840A, N854A, or N863A mutations to produce a Cas9 enzyme that substantially lacks all DNA cleavage activity. In some embodiments, an enzyme is considered substantially lacking all DNA cleavage activity when the DNA cleavage activity of the mutated CRISPR enzyme is about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less lower than that of its non-mutated form. Other mutations may be useful; where Cas9 or other CRISPR enzymes are from species other than Streptococcus pyogenes, mutations in the corresponding amino acids can be produced to achieve a similar effect.
[0134] In some embodiments, the enzyme coding sequence encoding the CRISPR enzyme is codon-optimized for expression in specific cells, such as eukaryotic cells. These eukaryotic cells may be those of a specific organism or derived from a specific organism, such as mammals, including but not limited to humans, mice, rats, rabbits, dogs, or non-human primates. Generally, codon optimization refers to the modification of a nucleic acid sequence to enhance expression in host cells of interest by replacing at least one codon of the natural sequence with codons that are used more frequently or most frequently in the gene in the host cell (e.g., about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons while maintaining the natural amino acid sequence). Different species exhibit specific preferences for certain codons containing specific amino acids. Codon preference (differences in codon use between organisms) is often associated with the translation efficiency of messenger RNA (mRNA), which is thought to depend (among other things) on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be tailored to achieve optimal gene expression in a given organism based on codon optimization. Codon utilization tables are readily available, for example, in codon usage databases (“Codon Usage Database”), and these tables can be adapted in various ways. See, Y. Nakamura Y. et al., “Codonusage tabulated from the international DNA sequence databases: status for the year 2000”, Nucleic Acids Res. 28:292 (2000). Computer algorithms for optimizing specific sequences of codons for expression in specific host cells are also available, such as GeneForge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more, or all codons) in the sequence encoding the CRISPR enzyme correspond to the most frequently used codon for a particular amino acid.
[0135] In some embodiments, a vector encodes a CRISPR enzyme comprising one or more nuclear localization sequences (NLS), such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS. In some embodiments, the CRISPR enzyme comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS at or near the N-terminus, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLS at or near the C-terminus, or combinations thereof (e.g., one or more NLS at the N-terminus and one or more NLS at the C-terminus). When more than one NLS is present, each can be selected to be independent of the other NLS, such that a single NLS can be present in more than one copy and / or combined with one or more other NLS present in one or more copies. In a preferred embodiment of the invention, the CRISPR enzyme comprises up to 6 NLS. In some embodiments, an NLS can be considered to be close to the N-terminus or C-terminus when the nearest amino acid of the NLS is within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N-terminus or C-terminus. Typically, an NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the protein surface, but other types of NLS are known. Non-limiting examples of NLS include NLS sequences derived from: NLS of the SV40 viral large T antigen having the amino acid sequence PKKKRKV; NLS from nucleoplasmic proteins (e.g., nucleoplasmic protein dichotomous NLS having the sequence KRPAATKKAGQAKKKK); c-myc NLS having the amino acid sequence PAAKRVKLD or RQRRNELKRSP; hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY; the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV from the IBB domain of input protein-α; the sequences VSRKRPRP and PPKKARED of fibroid T protein; the sequence POPKKKPL of human p53; and mouse c-abl. The sequence of IV is SALIKKKKKMAP; the sequence of influenza virus NS1 is DRLRR and PKQKKRK; the sequence of hepatitis virus delta antigen is RKLKKKIKKL; the sequence of mouse Mx1 protein is REKKKFLKRR; the sequence of human poly(ADP-ribose) polymerase is KRKGDEVDGVDEVAKKKSKK; and the sequence of steroid hormone receptor (human) glucocorticoid is RKCLQAGMNLEARKTKK.
[0136] Generally, the one or more NLSs have sufficient strength to drive the CRISPR complex to accumulate in a detectable amount in the nucleus of a eukaryotic cell. The strength of nuclear localization activity can generally drive the number of NLSs in the CRISPR enzyme, the one or more specific NLSs used, or a combination of these factors. Accumulation in the nucleus can be detected using any suitable technique. For example, a detection tag can be fused to the CRISPR enzyme to visualize the intracellular location, such as in conjunction with means for detecting the location of the nucleus (e.g., a nucleus-specific dye, such as DAPI). Examples of detectable tags include fluorescent proteins (e.g., green fluorescent protein, or GFP; RFP; CFP) and epitope tags (HA tag, flag tag, SNAP tag). The nucleus can also be isolated from the cell, and its contents can then be analyzed using any suitable method for protein detection, such as immunohistochemistry, Western blotting, or enzyme activity assays. Accumulation in the cell nucleus can also be indirectly determined by measuring the effect of CRISPR complex formation (e.g., measuring DNA cleavage or mutation at the target sequence, or measuring altered gene expression activity due to the effects of CRISPR complex formation and / or CRISPR enzyme activity) and comparing it with a control that is not exposed to CRISPR enzyme or complex, or exposed to a CRISPR enzyme lacking one or more NLS.
[0137] Generally, a guide sequence is any polynucleotide sequence that is sufficiently complementary to a target polynucleotide sequence to hybridize with the target sequence and to guide the CRISPR complex to bind sequence-specifically to the target sequence. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the guide sequence and its corresponding target sequence is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. The best alignment can be determined using any suitable algorithm for aligning sequences. Non-limiting examples include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (such as the Burrows WheelerAligner), ClustalW, ClustalX, BLAT, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some embodiments, a guide sequence is about 5, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 75, or more nucleotides in length. In some embodiments, a guide sequence is less than about 75, 50, 45, 40, 35, 30, 25, 20, 15, 12, or fewer nucleotides in length. The ability of the guide sequence to direct the sequence-specific binding of the CRISPR complex to the target sequence can be assessed by any suitable assay. For example, components of a CRISPR system sufficient to form a CRISPR complex, including the guide sequence to be tested, can be provided, for example, by transfection with a vector encoding the CRISPR sequence into a host cell having the corresponding target sequence, and then the preferential cleavage within the target sequence can be assessed by an assay as described herein. Similarly, by providing the target sequence, components of the CRISPR complex including the guide sequence to be tested, and a control guide sequence different from the test guide sequence, and comparing the binding or cleavage rate at the target sequence between the reactions of the test guide sequence and the control guide sequence, the cleavage of the target polynucleotide sequence can be assessed in a test tube. Other assays are possible and will be conceived by those skilled in the art.
[0138] A guide sequence can be selected to target any target sequence. In some embodiments, the target sequence is a sequence within the cell's genome. Exemplary target sequences include those that are unique within the target genome. For example, for Streptococcus pyogenes Cas9, a unique target sequence within the genome could include a Cas9 target site of the form MMMMMMMNNNNNNNNNNNNXGG, where NNNNNNNNNNNNNXGG (N is A, G, T, or C; and X can be anything) occurs once in the genome. A unique target sequence within the genome could include a Streptococcus pyogenes Cas9 target site of the form MMMMMMMMMNNNNNNNNNNNXGG, where NNNNNNNNNNNNXGG (N is A, G, T, or C; and X can be anything) occurs once in the genome. For *Streptococcus thermophilus* CRISPR1 Cas9, unique target sequences in the genome can include Cas9 target sites of the form MMMMMMMMMNNNNNNNNNNNNXXAGAAW, where N is A, G, T, or C; X can be anything; and W is A or T. For *Streptococcus pyogenes* Cas9, unique target sequences in the genome can include Cas9 target sites of the form MMMMMMMNNNNNNNNNNXXAGAAW, where N is A, G, T, or C; X can be anything; and W is A or T. Unique target sequences in the genome can include Streptococcus pyogenes Cas9 target sites of the form MMMMMMMMMNNNNNNNNNNNXGGXG, where NNNNNNNNNNXGGXG (N is A, G, T, or C; and X can be anything) appears once in the genome. In each of these sequences, "M" can be A, G, T, or C, and need not be considered unique in sequence identification.
[0139] In some embodiments, a guide sequence is selected to reduce the level of secondary structure within that guide sequence. Secondary structures can be determined by any suitable polynucleotide folding algorithm. Some algorithms are based on calculating the minimum Gibbs free energy. An example of such an algorithm is mFold, as described by Zuker and Stiegler in Nucleic Acids Res. 9 (1981), 133–148. Another example of a folding algorithm is the use of the online web server RNAfold, developed by the Institute for Theoretical Chemistry at the University of Vienna (see, for example, AR Gruber et al., 2008, Cell 106(1):23–24; and PA Carr and GMChurch, 2009, Nature Biotechnology 27(12):1151–62). Another algorithm can be found in the U.S. application serial number TBA (Agent's Case No. 44790.11.2022; Bode Reference No. BI-2013 / 004A) which is incorporated herein by reference.
[0140] Generally, a tracr pairing sequence includes any sequence that is sufficiently complementary to a tracr sequence to facilitate one or more of the following: (1) excision of a guide sequence flanking the tracr pairing sequence in a cell containing the corresponding tracr sequence; and (2) formation of a CRISPR complex at a target sequence, wherein the CRISPR complex contains a tracr pairing sequence hybridized to the tracr sequence. Typically, complementarity is defined in terms of optimal alignment of the tracr pairing sequence and the tracr sequence along the shorter of the two sequences. Optimal alignment can be determined by any suitable alignment algorithm and can be further interpreted as self-complementarity within the tracr sequence or tracr pairing sequence. In some embodiments, at the time of optimal alignment, the complementarity between the tracr sequence and the tracr pairing sequence along the shorter of the two is about or more than about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or higher. Figure 12BAnd 13B provides an example illustrating the optimal alignment between a tracr sequence and a tracr pairing sequence. In some embodiments, the tracr sequence is about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more nucleotides in length. In some embodiments, the tracr sequence and the tracr pairing sequence are included in a single transcript such that hybridization between the two produces a transcript with a secondary structure (such as a hairpin). The preferred loop-forming sequence used in the hairpin structure is four nucleotides in length and most preferably has the sequence GAAA. However, longer or shorter loop sequences can be used, as can alternative sequences. These sequences preferably include triples (e.g., AAA) and additional nucleotides (e.g., C or G). Examples of loop-forming sequences include CAAA and AAAAG. In one embodiment of the invention, the transcript or the transcribed polynucleotide sequence has at least two or more hairpins. In a preferred embodiment, the transcript has two, three, four, or five hairpins. In another embodiment of the invention, the transcript has up to five hairpins. In some embodiments, the single transcript further includes a transcription termination sequence; preferably, this is a polyT sequence, such as a six-T nucleotide sequence. Figure 13BThe lower part of the text provides an example of such a hairpin structure, in which the portion of the final "N" sequence 5' and the upstream of the loop correspond to the tracr pairing sequence, and the portion of the loop sequence 3' corresponds to the tracr sequence. Other non-restrictive examples of single polynucleotides containing a guide sequence, a tracr pairing sequence, and a tracr sequence are as follows (listed as 5' to 3'), where "N" represents the base of the guide sequence, the first region in lowercase represents the tracr pairing sequence, the second region in lowercase represents the tracr sequence, and the final poly-T sequence represents a transcription terminator: (1) NNNNNNNNNNNNNNNNNNNNNNgtttttgtactctcaagatttaGAAAtaaatcttgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTTTT; (2) NNNNNNNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcattttatggcagggtgttttcgttatttaaTTTT TT; (3)NNNNNNNNNNNNNNNNNNNNgtttttgtactctcaGAAAtgcagaagctacaaagataaggcttcatgccgaaatcaacaccctgtcatttta tggcagggtgtTTTTTT; (4)NNNNNNNNNNNNNNNNNNNNgttttagagctaGAAAtagcaagttaaaataaggctagtccgttatcaacttgaaaaa gtggcaccgagtcggtgcTTTTTT; (5)NNNNNNNNNNNNNNNNNNgttttagagctaGAAATAGcaagttaaaataaggctagtccgttatcaact tgaaaaagtgTTTTTTTT; and (6)NNNNNNNNNNNNNNNNNNgttttagagctagAAATAGcaagttaaaataaggctagtccgttatcaTTTTTTTT. In some embodiments, sequences (1) to (3) are used in combination with Cas9 from Streptococcus thermophilus CRISPR1. In some embodiments, sequences (4) to (6) are used in combination with Cas9 from Streptococcus pyogenes CRISPR1.In some embodiments, the tracr sequence is a transcript separate from the transcript containing the tracr pairing sequence (e.g., in...). Figure 13B (As shown in the top section).
[0141] In some embodiments, a recombinant template is also provided. The recombinant template may be a component of another vector as described herein, contained in a separate vector, or provided as a separate polynucleotide. In some embodiments, the recombinant template is designed to be used as a template in homologous recombination, such as within or near a target sequence cleaved or cut by a CRISPR enzyme that is part of a CRISPR complex. The template polynucleotide may have any suitable length, such as about or more than about 10, 15, 20, 25, 50, 75, 100, 150, 200, 500, 1000, or more nucleotides. In some embodiments, the template polynucleotide is complementary to a portion of the polynucleotide containing the target sequence. When optimally aligned, the template polynucleotide is capable of overlapping with one or more nucleotides of the target sequence (e.g., about or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, or more nucleotides). In some embodiments, when a template sequence is optimally aligned with a polynucleotide containing a target sequence, the nearest nucleotide of the template polynucleotide is within approximately 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 5000, 10000, or more nucleotides from the target sequence.
[0142] In some embodiments, the CRISPR enzyme is part of a fusion protein comprising one or more heterologous protein domains (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more domains other than the CRISPR enzyme). The CRISPR enzyme fusion protein may comprise any other protein, and optionally a linker sequence between any two domains. Examples of protein domains that may be fused to a CRISPR enzyme include, but are not limited to, epitope tags, reporter gene sequences, and protein domains having one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza virus hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green luminescent protein (GFP), HcRed, DsRed, cyan luminescent protein (CFP), yellow luminescent protein (YFP), and spontaneous luminescent proteins including blue luminescent protein (BFP). CRISPR enzymes can be fused to a gene sequence encoding a protein or protein fragment that binds to DNA molecules or other cellular molecules, including, but not limited to, maltose-binding protein (MBP), S-tag, Lex A DNA-binding domain (DBD) fusions, GAL4 DNA-binding domain fusions, and herpes simplex virus (HSV) BP16 protein fusions. Additional domains that can form part of a fusion protein containing a CRISPR enzyme are described in US 20110059502, which is incorporated herein by reference. In some embodiments, a labeled CRISPR enzyme is used to identify the location of the target sequence.
[0143] In some aspects, the present invention provides methods for delivering one or more polynucleotides, such as or as described herein, one or more vectors, one or more transcripts thereof, and / or a protein transcribed therefrom into a host cell. In some aspects, the present invention further provides cells produced by such methods and organisms comprising or derived from such cells (e.g., animals, plants, or fungi). In some embodiments, a CRISPR enzyme combined with (and optionally in combination with) a guide sequence is delivered to the cell. Nucleic acids can be introduced into mammalian cells or target tissues using conventional viral and nonviral gene transfer methods. Nucleic acids encoding components of a CRISPR system can be administered to cells in a culture or in a host organism using such methods. Nonviral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of vectors described herein), naked nucleic acids, and nucleic acids in combination with a delivery excipient (e.g., liposomes). Viral vector delivery systems include DNA and RNA viruses that have a free or integrated genome after delivery to cells. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Felgner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience Neuroscience) 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddad et al., Doerfler and Bohm in Current Topics in Microbiology and Immunology (Edited) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).
[0144] Non-viral methods for nucleic acid delivery include lipid transfection, nuclear transfection, microinjection, gene gun, viral particles, liposomes, immunoliposomes, polycationic or lipid:nucleic acid conjugates, naked DNA, artificial viruses, and reagent-enhanced uptake of DNA. Lipid transfection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipid transfection reagents are commercially available (e.g., Transfectam). TM and Lipofectin TM Efficient receptor recognition for polynucleotides in cationic and neutral lipid transfection includes those of Felgner (WO 91 / 17424; WO 91 / 16024). Delivery can be targeted to cells (e.g., in vitro or ex vivo) or to target tissues (e.g., in vivo).
[0145] The preparation of lipid-nucleic acid complexes (including targeted liposomes, such as immunolipid complexes) is well known to those skilled in the art (see, for example, Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Therapy 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Research...). Res. 52:4817-4820 (1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787.
[0146] Systems using RNA or DNA viral bases for nucleic acid delivery utilize highly evolved processes to target viruses to specific cells in vivo and deliver the viral payload to the cell nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, and optionally, modified cells can be administered to patients (ex vivo). Conventional viral-based systems can include retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, and herpes simplex virus vectors for gene transfer. Integration into the host genome using retroviral, lentiviral, and adeno-associated virus gene transfer methods is possible, typically resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiency has been observed in various cell types and target tissues.
[0147] Retroviral tropism can be altered by incorporating exogenous envelope proteins to expand the potential target population of target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically producing high viral titers. Therefore, the choice of retroviral gene transfer system will depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats (LTRs) with the ability to package exogenous sequences up to 6-10 kb. A minimum amount of cis-acting LTRs is sufficient for vector replication and packaging, and these vectors are then used to integrate therapeutic genes into target cells to provide permanent transgenic expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibberish leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, for example, Buchscher et al., Journal of Virology 66:2731-2739 (1992); Johann et al., Journal of Virology 66:1635-1640 (1992); Sommnerfelt et al., Virol. 176:58-59 (1990); Wilson et al., Journal of Virology 63:2374-2378 (1989); Miller et al., Journal of Virology 65:2220-2224 (1991); PCT / US 94 / 05700). In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors are capable of very high transduction efficiency in many cell types without the need for cell division. High titers and expression levels have been achieved with such vectors. This vector can be produced in large quantities in relatively simple systems. Adeno-associated virus (“AAV”) vectors can also be used to transduce cells with target nucleic acids, for example, to produce nucleic acids and peptides in vitro, and for use in in vivo and in vitro gene therapy procedures (see, for example, West et al., Virology 160:38-47 (1987); US Patent No. 4,797,368; WO 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)).The construction of recombinant AAV vectors has been described in several publications, including U.S. Patent No. 5,173,414; Tratschin et al., *Molecular and Cell Biology* 5:3251-3260 (1985); Tratschin et al., *Molecular and Cell Biology* 4:2072-2081 (1984); Hermonat & Muzyczka, *Proceedings of the National Academy of Sciences* 81:6466-6470 (1984); and Samulski et al., *Journal of Virology* 63:03822-3828 (1989).
[0148] Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells, which package adenoviruses, and ψ2 or PA317 cells, which contain retroviruses. Viral vectors used in gene therapy are typically produced by cell lines that package nucleic acid vectors into viral particles. These vectors typically contain the minimum amount of sequences required for packaging and subsequent integration into the host, with other viral sequences replaced by expression cassettes for one or more polynucleotides to be expressed. Lost viral function is typically provided trans-formally by the packaging cell lines. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome, which are required for packaging and integration into the host genome. Viral DNA is packaged into cell lines containing additional AAV genes encoding helper plasmids, namely rep and cap, but lacking the ITR sequences. These cell lines can also be infected with adenoviruses acting as helpers. The helper virus promotes the replication of the AAV vector and the expression of AAV genes from the helper plasmid. Due to the lack of the ITR sequences, the helper plasmid is not packaged in significant quantities. Adenovirus contamination can be reduced, for example, by heat treatment, which is more sensitive to adenoviruses than to AAVs. Other methods for delivering nucleic acids into cells are known to those skilled in the art. See, for example, US 20030087817, which is incorporated herein by reference.
[0149] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, transfection is performed when cells are naturally present in the subject. In some embodiments, the transfected cells are taken from the subject. In some embodiments, the cells are derived from cells taken from the subject, such as cell lines. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLa-S3, Huh1, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panc1, PC-3, TF1, CTLL-2, C1R, Rat6, CV1, RPTE, A10, T24, J82, A375, ARH-77, Calu1, SW480, SW620, SKOV3, SK-UT, CaCo2, P388D1, SEM-K2, WEHI-231, HB56, TIB55, Jurkat, J45.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial cells, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-L1, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-1 cells, BEAS-2B, bEnd.3, BHK-21, BR293, BxPC3, C3H-10T1 / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-K1, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L23 / 5010, COR-L23 / R23, COS-7, COV-434, CML T1, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepa1c1c7, HL-60, HMEC, HT-29, Jurkat, JY cells, K562 cells, Ku812, KCL22, KG1, KYO1, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCK II, MDCK II, MOR / 0.2R, MONO-MAC 6. MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR and their transgenic variants. Cell lines are available from a variety of sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassus, Virginia)). In some embodiments, new cell lines are established using cells transfected with one or more vectors described herein, the new cell lines comprising sequences derived from one or more vectors. In some embodiments, new cell lines are established using cells transfected with components of a CRISPR system as described herein (e.g., by transient transfection with one or more vectors or by RNA transfection) and modified for CRISPR complex activity, the new cell lines comprising cells that are water-absorbing but lack any other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used in the evaluation of one or more test compounds.
[0150] In some embodiments, a non-human transgenic animal or transgenic plant is produced using one or more vectors described herein. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat, or rabbit. In some embodiments, the organism or subject is a plant. In some embodiments, the organism or subject or plant is algae. Methods for producing transgenic plants and animals are known in the art and typically begin with a cell transfection method as described herein.
[0151] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. In some embodiments, the method includes allowing a CRISPR complex to bind to the target polynucleotide to perform cleavage of the target polynucleotide, thereby modifying the target polynucleotide, wherein the CRISPR complex includes a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, wherein the guide sequence is linked to a tracr pairing sequence, which in turn hybridizes to a tracr sequence.
[0152] In one aspect, the present invention provides a method for modifying the expression of a polynucleotide in eukaryotic cells. In some embodiments, the method includes allowing a CRISPR complex to bind to the polynucleotide such that the binding results in an increase or decrease in the expression of the polynucleotide; wherein the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, wherein the guide sequence is linked to a tracr pairing sequence, which in turn hybridizes to a tracr sequence.
[0153] With recent advances in crop genomics, the ability to perform efficient and cost-effective gene editing and manipulation using CRISPR-Cas systems will allow for the rapid screening and comparison of single and multiple genetic operations to transform genomes in order to improve production and enhance traits. In this regard, reference is made to the following U.S. patents and publications: U.S. Patent No. 6,603,061 – Agrobacterium-Mediated Plant Transformation Method; U.S. Patent No. 7,868,149 – Plant Genome Sequences and Uses Thereof; and U.S. 2009 / 0100536 – Transgenic Plants with Enhanced Agronomic Traits, all contents and disclosures of which are incorporated herein by reference in their entirety. In the practice of this invention, the contents and disclosures of Morrell et al., “Cropgenomics: advances and applications,” *Nat Rev Genet.*, December 29, 2011; 13(2):85-96, are also incorporated herein by reference in their entirety. In an advantageous embodiment of the invention, the CRISPR / Cas9 system is used to engineer microalgae (Example 15). Therefore, with necessary modifications, references to animal cells herein may also apply to plant cells, unless otherwise apparent.
[0154] In one aspect, the present invention provides methods for modifying target polynucleotides in eukaryotic cells, which can be performed in vivo, in vitro, or outside the body. In some embodiments, the method includes sampling cells or cell populations from a human or non-human animal or plant (including microalgae) and modifying the cells or these cells. Culture can occur at any stage in vitro. The cells or these cells can even be reintroduced into the non-human animal or plant (including microalgae).
[0155] In plants, pathogens are often host-specific. For example, *Fusarium oxysporum* f.sp. lycopersici causes tomato wilt and attacks only tomatoes, while *F. oxysporum f. dianthii* Puccinia graminis f.sp. tritici attacks only wheat. Plants possess existing and induced defenses against most pathogens. Mutation and recombination events across plant generations lead to genetic variability in susceptibility, especially when pathogens reproduce at a higher frequency than the plant. Non-host resistance can exist in plants, for example, when the host and pathogen are incompatible. Horizontal resistance can also exist, such as partial resistance to all species of pathogens, typically controlled by many genes, and vertical resistance, such as complete resistance to certain species of pathogens but not others, typically controlled by a few genes. At the gene-to-gene level, plants and pathogens co-evolve, and genetic changes in one balance the changes in the other. Therefore, by utilizing natural variation, breeders combine the most useful genes for yield, quality, uniformity, tolerance, and resistance. Sources of resistance genes include natural or introduced varieties, heirloom varieties, closely related wild plants, and induced mutations, such as treatment of plant material with mutagens. This invention provides plant breeders with a new tool for inducing mutations. Therefore, those skilled in the art can analyze the genomes of resistance gene sources and, with respect to varieties possessing the desired characteristics or traits, use this invention to induce the occurrence of resistance genes with greater precision than mutagens, thus accelerating and improving plant breeding programs.
[0156] In one aspect, the present invention provides kits comprising any one or more of the elements disclosed in the methods and compositions described above. In some embodiments, the kit includes a vector system and instructions for using the kit. In some embodiments, the vector system includes: (a) a first regulatory element operatively linked to a tracr pairing sequence and one or more insertion sites for inserting a guide sequence upstream of the tracr pairing sequence, wherein, upon expression, the guide sequence directs a CRISPR complex in eukaryotic cells to sequence-specific binding to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme conjugated to: (1) a guide sequence hybridizing to the target sequence, and (2) a tracr pairing sequence hybridizing to the tracr sequence; and / or (b) a second regulatory element operatively linked to an enzyme-coding sequence including a nuclear localization sequence encoding the CRISPR enzyme. Elements may be provided individually or in combination and may be provided in any suitable container, such as vials, bottles, or tubes. In some embodiments, the kit includes instructions in one or more languages, such as instructions in more than one language.
[0157] In some embodiments, the kit includes one or more reagents for use in methods utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. Reagents may be provided in a form available in a specific assay or in a form requiring the addition of one or more other components prior to use (e.g., in concentrated or lyophilized form). The buffer may be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH from about 7 to about 10. In some embodiments, the kit includes one or more oligonucleotides corresponding to a guide sequence for insertion into a vector to operatively link the guide sequence and a regulatory element. In some embodiments, the kit includes a homologous recombinant template polynucleotide.
[0158] In one aspect, the present invention provides a method for using one or more elements of a CRISPR system. The CRISPR complex of the present invention provides an efficient means for modifying target polynucleotides. The CRISPR complex of the present invention has a wide range of applications, including modifying (e.g., deletion, insertion, translocation, inactivation, activation) target polynucleotides in various cell types. Therefore, the CRISPR complex of the present invention has a broad spectrum of applications, such as in gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary CRISPR complex includes a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence within the target polynucleotide. The guide sequence is linked to a tracr pairing sequence, which in turn hybridizes to a tracr sequence.
[0159] The target polynucleotide of the CRISPR complex can be any polynucleotide, endogenous or exogenous to the eukaryotic cell. For example, the target polynucleotide can be a polynucleotide residing in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). It is not desirable to be bound by theory; it is believed that the target sequence should be associated with a PAM (prototype spacer adjacent motif); that is, a short sequence recognized by the CRISPR complex. The precise sequence and length requirements for the PAM vary depending on the CRISPR enzyme used, but the PAM is typically a 2-5 base pair sequence adjacent to the prototype spacer (i.e., the target sequence). Examples of PAM sequences are given in the Examples section below, and skilled personnel will be able to identify additional PAM sequences used with a given CRISPR enzyme.
[0160] The target polynucleotides of the CRISPR complex may include multiple disease-related genes and polynucleotides, as well as genes and polynucleotides related to signal transduction biochemical pathways, as illustrated in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed December 12, 2012 and January 2, 2013, respectively, both entitled "Systems Methods and Compositions for Sequence Manipulation" with Bode Reference Nos. BI-2011 / 008 / WSGR 44063-701.101 and BI-2011 / 008 / WSGR 44063-701.102, the contents of which are incorporated herein by reference in their entirety.
[0161] Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as genes or polynucleotides related to signal transduction biochemical pathways. Examples of target polynucleotides include disease-related genes or polynucleotides. A “disease-related” gene or polynucleotide refers to any gene or polynucleotide that produces its transcriptional or translational product at an abnormal level or in an abnormal form in cells derived from tissues affected by a disease, compared to tissues or cells from non-disease control tissues or cells. In cases where altered expression is associated with the onset and / or progression of disease, it can be a gene expressed at an abnormally high level; it can be a gene expressed at an abnormally low level. Disease-related genes also refer to genes with one or more mutations or genetic variations that are directly responsible for or linked in disequilibrium to one or more genes responsible for the etiology of the disease. The transcribed or translated product can be known or unknown and can be at normal or abnormal levels.
[0162] Examples of disease-related genes and polynucleotides were available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Maryland) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Maryland).
[0163] Examples of disease-related genes and polynucleotides are listed in Tables A and B. Specific disease information available on the World Wide Web is from the McKusick-Nathan's Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Maryland) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Maryland). Examples of genes and polynucleotides related to signal transduction biochemical pathways are listed in Table C.
[0164] Mutations in these genes and pathways can lead to the production of inappropriate proteins or affect the function of proteins in inappropriate amounts. Further examples of genes, diseases, and proteins are cited and incorporated herein by reference in U.S. Provisional Application 61 / 736,527, filed December 12, 2012, and 61 / 748,427, filed February 2, 2013. Such genes, proteins, and pathways may be target polynucleotides of the CRISPR complex.
[0165] Table A
[0166]
[0167]
[0168]
[0169] Table B:
[0170]
[0171]
[0172]
[0173]
[0174]
[0175] Table C:
[0176]
[0177]
[0178]
[0179]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193] Embodiments of the present invention also relate to methods and compositions relating to gene knockout, gene amplification, and repair of specific mutations associated with DNA repeat instability and neurological disorders (Robert D. Wells and Tetsuo Ashizawa, Genetic Instabilities and Neurological Diseases, 2nd ed., Academic Press, October 13, 2011 - Medical). Specific aspects of tandem repeat sequences have been found to be responsible for more than twenty human diseases (New insights into repeat instability: role of RNA-DNA hybrids, McIvor EI, Polak U, Napierala M, RNA Biology, Sep-October 2010; 7(5):551-8). Genomic instability defects can be corrected using the CRISPR-Cas system.
[0194] Another aspect of this invention relates to the use of the CRISPR-Cas system to correct defects in the EMP2A and EMP2B genes, which have been identified as being associated with Lafora disease. Lafora disease is an autosomal recessive disorder characterized by progressive myoclonic epilepsy that can begin as seizures in adolescence. A minority of cases can be caused by mutations in unidentified genes. The disease causes seizures, muscle spasms, difficulty walking, dementia, and ultimately death. Currently, no therapies have been shown to be effective against disease progression. Other genetic abnormalities associated with epilepsy can also be targeted to the CRISPR-Cas system, and the underlying genetics are further described in *Genetics of Epilepsy and Genetic Epilepsies*, edited by Giuliano Avanzini and Jeffrey L. Noebels (Mariani Foundation Paediatric Neurology: 20; 2009).
[0195] In another aspect of the invention, the CRISPR-Cas system can be used to correct several eye defects caused by gene mutations, as further described in Genetic Diseases of the Eye, Second Edition, edited by Elias I. Traboulsi, Oxford University Press, 2012.
[0196] Several other aspects of this invention relate to correcting deficiencies associated with a wide range of genetic diseases, which are further described in the topic section Genetic Disorders on the National Institutes of Health website (health.nih.gov / topic / GeneticDisorders). Genetic brain diseases may include, but are not limited to, adrenoleukodystrophy, corpus callosum agenesis, Aicardi syndrome, Alpers' disease, Alzheimer's disease, Barth syndrome, Batten disease, CADASIL, cerebellar degeneration, Fabry's disease, Gerstmann-Straussler-Scheinker disease, Huntington's disease and other triadic duplication disorders, Leigh's disease, Lesch-Nearn syndrome, Menx's disease, mitochondrial myopathy, and Colpocephaly (NINDS). These diseases are further described under Genetic Brain Disorders on the National Institutes of Health website.
[0197] In some embodiments, the condition can be tumor formation. In some embodiments, where the condition is tumor formation, the gene to be targeted is any of those genes listed in Table A (in this case, PTEN, etc.). In some embodiments, the condition can be age-related macular degeneration. In some embodiments, the condition can be a schizophrenia disorder. In some embodiments, the condition can be a trinucleotide repeat disorder. In some embodiments, the condition can be Fragile X syndrome. In some embodiments, the condition can be a secretase-related disorder. In some embodiments, the condition can be a prion-related disorder. In some embodiments, the condition can be ALS. In some embodiments, the condition can be a substance addiction. In some embodiments, the condition can be autism. In some embodiments, the condition can be Alzheimer's disease. In some embodiments, the condition can be inflammation. In some embodiments, the condition can be Parkinson's disease.
[0198] Examples of proteins associated with Parkinson's disease include, but are not limited to, α-synuclein, DJ-1, LRRK2, PINK1, Parkinin, UCHL1, Synphilin-1, and NURR1.
[0199] Examples of addiction-related proteins may include, for example, ABAT.
[0200] Examples of inflammation-related proteins may include, for example, monocyte chemoattractant protein-1 (MCP1) encoded by the Ccr2 gene, CC chemokine receptor type 5 (CCR5) encoded by the Ccr5 gene, IgG receptor IIB (FCGR2b, also known as CD32) encoded by the Fcgr2b gene, or FcεR1g (FCER1g) protein encoded by the Fcer1g gene.
[0201] Examples of cardiovascular disease-related proteins may include, for example, IL1B (interleukin-1, β), XDH (xanthine dehydrogenase), TP53 (tumor protein p53), PTGIS (prostaglandin I2 (prostaglandin) synthase), MB (myoglobin), IL4 (interleukin-4), ANGPT1 (angiopoietin-1), ABCG8 (ATP-binding cassette, subfamily G (WHITE), member 8) or CTSK (cathepsin K).
[0202] Examples of Alzheimer's disease-related proteins may include, for example, the very low-density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, ubiquitin-like modifier activator 1 (UBA1) encoded by the UBA1 gene, or the NEDD8-activator E1 catalytic subunit protein (UBE1C) encoded by the UBA3 gene.
[0203] Examples of proteins associated with autism spectrum disorder may include, for example, benzodiazepine receptor (peripheral)-associated protein 1 (BZRAP1) encoded by the BZRAP1 gene, AF4 / FMR2 family member 2 protein (AFF2) (also known as MFR2) encoded by the AFF2 gene, Fragile X intellectual disability autosomal homologue 1 protein (FXR1) encoded by the FXR1 gene, or Fragile X intellectual disability autosomal homologue 2 protein (FXR2) encoded by the FXR2 gene.
[0204] Examples of macular degeneration-related proteins may include, for example, the ATP-binding cassette encoded by the ABCR gene, subfamily A (ABC1) member 4 protein (ABCA4), the apolipoprotein E protein (APOE) encoded by the APOE gene, or the chemokine (CC motif) ligand 2 protein (CCL2) encoded by the CCL2 gene.
[0205] Examples of schizophrenia-related proteins may include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BDNF, DISC1, GSK3B, and combinations thereof.
[0206] Examples of proteins involved in tumor suppression may include, for example, ATM (mutated in ataxia-telangiectasia), ATR (aftaxia-telangiectasia and Rad3-associated), EGFR (epidermal growth factor receptor), ERBB2 (v-erb-b2 erythroleukemia virus oncogene homologue 2), ERBB3 (v-erb-b2 erythroleukemia virus oncogene homologue 3), ERBB4 (v-erb-b2 erythroleukemia virus oncogene homologue 4), Notch 1, Notch 2, Notch 3, or Notch 4.
[0207] Examples of proteins associated with secretion enzyme disorders may include, for example, PSENEN (a homologue of presenilin enhancer 2 (C. elegans)), CTSB (cathepsin B), PSEN1 (presenilin 1), APP (amyloid β (A4) precursor protein), APH1B (a homologue of propharyngeal defect 1 (C. elegans)), PSEN2 (presenilin 2 (Alzheimer's disease 4)), or BACE1 (β-site APP-cleaving enzyme 1).
[0208] Examples of proteins associated with amyotrophic lateral sclerosis (ALS) may include SOD1 (superoxide dismutase 1), ALS2 (ALS 2), FUS (fused in sarcomas), TARDBP (TAR DNA-binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C) and any combination thereof.
[0209] Examples of proteins associated with prion diseases may include SOD1 (superoxide dismutase 1), ALS2 (amyotrophic lateral sclerosis 2), FUS (fused in sarcoma), TARDBP (TAR DNA-binding protein), VAGFA (vascular endothelial growth factor A), VAGFB (vascular endothelial growth factor B), and VAGFC (vascular endothelial growth factor C) and any combination thereof.
[0210] Examples of proteins associated with neurodegenerative diseases in prion diseases may include, for example, A2M (α-2-macroglobulin), AATF (apoptotic antagonistic transcription factor), ACPP (prostatic acid phosphatase), ACTA2 (actin α2 smooth muscle aorta), ADAM22 (ADAM metallopeptidase domain), ADORA3 (adenosine A3 receptor), or ADRA1D (α-1D adrenergic receptor).
[0211] Examples of proteins associated with immunodeficiency may include, for example, A2M [α-2-macroglobulin]; AANAT [aralkylamine N-acetyltransferase]; ABCA1 [ATP-binding cassette, subfamily A (ABC1), member 1]; ABCA2 [ATP-binding cassette, subfamily A (ABC1), member 2]; or ABCA3 [ATP-binding cassette, subfamily A (ABC1), member 3].
[0212] Examples of proteins associated with trinucleotide repeat disorders include, for example, AR (androgen receptor), FMR1 (fragile X intellectual disability 1), HTT (huntington protein), or DMPK (myotonic dystrophy protein kinase), FXN (frataxin), and ATXN2 (ataxia protein 2).
[0213] Examples of proteins associated with neurotransmission disorders include, for example, SST (somatostatin), NOS1 (nitric oxide synthase 1 (neuronal)), ADRA2A (adrenergic, α-2A-, receptor), ADRA2C (adrenergic, α-2C-, receptor), TACR1 (tachykinin receptor 1), or HTR2c (5-hydroxytryptamine (serotonin) receptor 2C).
[0214] Examples of neurodevelopment-related sequences include, for example, A2BP1 [ataxia protein 2-binding protein 1], AADAT [aminoadipic acid aminotransferase], AANAT [aralkylamine N-acetyltransferase], ABAT [4-aminobutyric acid aminotransferase], ABCA1 [ATP-binding cassette, subfamily A (ABC1), member 1] or ABCA13 [ATP-binding cassette, subfamily A (ABC1), member 13].
[0215] Further examples of preferred conditions that can be treated with the system of the present invention may be selected from: Aicardi-Goutières Syndrome; Alexander disease; Allan-Herndon-Dudley Syndrome; POLG-related disorders; α-mannosin storage disorders (types II and III); Alstrom syndrome ( Syndrome; Engmann syndrome; ataxia-telangiectasia; neuronal ceroid-lipofuscin deposition; β-thalassemia; bilateral optic atrophy and type 1 (infant) optic atrophy; retinoblastoma (bilateral); Canavan disease; Cerebrooculofacial-skeletal syndrome (COFS1); cerebral tendon xanthoma; Cornelia de Lange syndrome; MAPT-related disorder; hereditary prion disease; Dravet syndrome; early-onset familial Alzheimer's disease; Friedreich ataxia (FRDA); multiple malformations; fucoside storage disease; Fukuyama congenital muscular dystrophy; galactosialidosis; Gaucher disease Diseases; organic acidemia; hemophagocytic lymphohistiocytosis; progeria; mucolipidemia II; infantile free sialic acid storage disease; PLA2G6-associated neurodegeneration; Jervell and Lange-Nielsen syndrome; conjugated bullous epidermolysis; Huntington's disease; Krabbe disease (infant); mitochondrial DNA-associated Leigh syndrome and NARP; Lesch-Nair syndrome; LIS1-associated anencephaly; Lowe syndrome; maple syrup diabetes; MECP2 replication syndrome; ATP7A-associated copper transport disorder; LAMA2-associated muscular dystrophy; arylsulfatase A deficiency; type I, II, or III mucopolysaccharidosis; peroxisome biogenesis disorder; Zellweger syndrome Spectrum; Neurodegeneration associated with brain iron accumulation disorders; Acid sphingomyelinase deficiency; Niemann-Pick disease type C; Glycine encephalopathy; ARX-related disorders; Urea cycle disorders; COL1A1 / 2-related osteogenesis imperfecta; Mitochondrial DNA deletion syndrome; PLP1-related disorders; Perry syndrome; Phelan-McDermid syndrome; Type II glycogen storage disease (Pompe disease) (infants); MAPT-related disorders; MECP2-related disorders; Rhizomelic chondrodysplasia punctata type 1; Roberts syndrome;Sandhoff disease; Schindler disease type 1; adenosine deaminase deficiency; Schlein-Lonzo-Oxley syndrome; spinal muscular atrophy; infantile spinocerebellar ataxia; hexosamine A deficiency; type 1 lethal dysplasia; collagen type VI related disorder; Usher syndrome type I; congenital muscular dystrophy; Wolf-Hirschhorn syndrome; lysosomal acid lipase deficiency; and xeroderma pigmentosum.
[0216] As will be obvious, it is envisioned that the system of the present invention can be used to target any polynucleotide sequence of interest. Some examples of conditions or diseases that may be useful for treatment using the system of the present invention are included in the table above, and examples of genes currently associated with those conditions are also provided herein. However, the example genes are not exhaustive.
[0217] Example
[0218] The following examples are given for the purpose of illustrating different embodiments of the invention and are not intended to limit the invention in any way. The examples of the invention, together with the methods described herein, currently represent preferred embodiments, are exemplary, and are not intended to limit the scope of the invention. Variations and other uses covered herein within the spirit of the invention as defined by the scope of the claims are as will be apparent to those skilled in the art.
[0219] Example 1: CRISPR complex activity in the nucleus of eukaryotic cells
[0220] An exemplary type II CRISPR system is a type II CRISPR locus from *Streptococcus pyogenes* SF370, which contains a cluster of four genes Cas9, Cas1, Cas2, and Csn1, two non-coding RNA elements (tracrRNA), and a characteristic array of repetitive sequences (direct repeats) separated by short segments of non-repetitive sequences (spacers, each approximately 30 bp). In this system, targeted DNA double-strand breaks (DSBs) are generated in four consecutive steps. Figure 2AThe first step involves transcribing two non-coding RNAs, a pre-crRNA array, and a tracrRNA from a CRISPR locus. The second step involves hybridizing the tracrRNA to a direct repeat of the pre-crRNA, which is then processed into a mature crRNA containing a separate spacer sequence. The third step involves the mature crRNA:tracrRNA complex directing Cas9 towards a DNA target composed of the prototype spacer and its corresponding PAM via the formation of a heteroduplex between the crRNA's spacer region and the prototype spacer DNA. Finally, Cas9 mediates the cleavage of the target DNA upstream of the PAM to generate a DSB (Dual Substrate Buffer) within the prototype spacer. Figure 2A This example describes an exemplary process for adapting this RNA-programmable nuclease system to guide the activity of CRISPR complexes in the nucleus of eukaryotic cells.
[0221] Cell culture and transfection
[0222] Human embryonic kidney (HEK) cell line HEK293FT (Biotech) was maintained at 37°C with 5% CO2 in Durbecco's modified Eagle's Medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Biotech), 100 U / mL penicillin, and 100 μg / mL streptomycin. Mouse neuro2A (N2A) cell line (ATCC) was maintained at 37°C with 5% CO2 in DMEM supplemented with 5% fetal bovine serum (HyClone), 2 mM GlutaMAX (Biotech), 100 U / mL penicillin, and 100 μg / mL streptomycin.
[0223] One day prior to transfection, HEK 293FT or N2A cells were seeded into 24-well plates (Corning). Cells were transfected using Lipofectamine 2000 (Life Technologies) following the manufacturer's recommended protocol. A total of 800 ng of plasmid was used for each well of the 24-well plate.
[0224] Surveyor determination and sequencing analysis of genome modifications
[0225] As described above, HEK 293FT or N2A cells were transfected with plasmid DNA. Following transfection, cells were incubated at 37°C for 72 hours before genomic DNA extraction. Genomic DNA was extracted using the QuickExtract DNA Extraction Kit (Epicentre) following the manufacturer's protocol. In short, cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and then at 98°C for 10 minutes. The extracted genomic DNA was processed immediately or stored at -20°C.
[0226] For each gene, PCR amplification was performed on the genomic region surrounding the CRISPR target site, and the products were purified using a QiaQuick Spin column (Qiagen) following the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10X Taq polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 20 μl, and subjected to a re-annealing process to allow for heteroduplex formation: 95 °C for 10 min, decreasing from 95 °C to 85 °C at -2 °C / s, decreasing from 85 °C to 25 °C at -0.25 °C / s, and holding at 25 °C for 1 min. After re-annealing, the products were treated with Surveyor nuclease and Surveyor enhancer S (Transgenomics) according to the manufacturer's recommended protocol and analyzed on 4%–20% Novex TBE polyacrylamide gels (Life Technologies). The gels were stained with SYBR Gold DNA staining agent (Lifetech Corporation) for 30 minutes and imaged using a Gel Doc gel imaging system (Bio-rad Corporation). Quantification was based on the relative band intensity as a measure of the cut DNA portions. Figure 8 A schematic diagram of this Surveyor measurement is provided.
[0227] Determination of restriction fragment length polymorphism for detecting homologous recombination
[0228] HEK 293FT and N2A cells were transfected with plasmid DNA and incubated at 37°C for 72 hours prior to genomic DNA extraction, as described above. Target genomic regions were amplified by PCR in the homologous arms of the homologous recombination (HR) template using primers. PCR products were isolated onto 1% agarose gels and extracted using a MinElute Gel Extraction Kit (Kiagenes). The purified products were digested with HindIII (Fermentas) and analyzed on 6% Novex TBE polyacrylamide gels (Lifetechnologies).
[0229] RNA secondary structure prediction and analysis
[0230] The centroid structure prediction algorithm was used to predict RNA secondary structure using the online web server RNAfold, developed by the Institute for Theoretical Chemistry at the University of Vienna (see, for example, AR Gruber et al., 2008, Cell 106(1):23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12):1151-62).
[0231] Bacterial plasmid transformation interference assay
[0232] Elements of the Streptococcus pyogenes CRISPR locus 1 sufficient for CRISPR activity were recombined in Escherichia coli using the pCRISPR plasmid (illustrated schematically in...). Figure 10 (See A). pCRISPR contains tracrRNA, SpCas9, and a leader sequence for the driver crRNA array. Using annealed oligonucleotides, spacers (also called “guide sequences”) are inserted between BsaI sites in the crRNA array, as shown. The activation plasmid used in the interference assay is constructed by inserting the prototype spacer sequence (also called the “target sequence”) along with a neighboring CRISPR motif sequence (PAM) into pUC19 (see A). Figure 10 B). The activation plasmid contains ampicillin resistance. Figure 10 C provides a schematic representation of this interference assay. Chemically competent *E. coli* strains carrying pCRISPR and a suitable spacer were transformed with an excitation plasmid containing the corresponding prototype spacer-PAM sequence. pUC19 was used to evaluate the transformation rate of each pCRISPR-carrying competent strain. CRISPR activity resulted in the cleavage of the pPSP plasmid carrying the prototype spacer (excluding ampicillin resistance conferred by pUC19 lacking the prototype spacer). Figure 10 D shows that Figure 4C The competent cells of each E. coli strain carrying pCRISPR used in the assays shown are illustrated.
[0233] RNA purification
[0234] HEK 293FT cells were maintained and transfected as described above. Cells were harvested by trypsin digestion and then washed in phosphate-buffered saline (PBS). Total cellular RNA was extracted using TRI reagent (Sigma) following the manufacturer's protocol. The extracted total RNA was quantified and normalized to the same concentration using Naonodrop (Thermo Scientific).
[0235] Northern blot analysis of crRNA and tracrRNA expression in mammalian cells
[0236] RNA was mixed with an equal volume of 2X loading buffer (Ambion), heated to 95°C for 5 min, frozen on ice for 1 min, and then loaded onto an 8% denaturing polyacrylamide gel (SequaGel, National Diagnostics) after a pre-run of the gel for at least 30 min. The sample was subjected to limiting electrophoresis at 40 W for 1.5 h. Afterward, at room temperature, RNA was transferred to a 300 mA Hybond N+ membrane (GE Healthcare) in a semi-dry transfer device (Bio-rad) and allowed to stand for 1.5 h. RNA was crosslinked to the membrane using the automatic crosslinking button on the Stratalinker (Stratagene) UV Crosslinker. The membrane was pre-hybridized in ULTRAhyb-Oligo hybridization buffer (Ambion) at 42°C with rotation for 30 min, and then probes were added and hybridized overnight. The probe was ordered from IDT and used with [γ-] a polynucleotide kinase containing T4 (New England Biolabs). 32 P]ATP (PerkinElmer) labeling. The membrane was washed once with preheated (42°C) 2xSSC, 0.5% SDS for 1 min, followed by two washes at 42°C for 30 min each. The membrane was then exposed to a luminescent screen at room temperature for one hour or overnight and scanned using a Typhoon imager.
[0237] Construction and evaluation of bacterial CRISPR systems
[0238] CRISPR locus elements (including tracrRNA, Cas9, and a leader) were amplified by PCR from *Streptococcus pyogenes* SF370 genomic DNA with flanking homologous arms for Gibson Assembly. Two BsaI-type IIS sites were introduced between two homologous repeats to facilitate spacer insertion (Figure 9). The PCR product was cloned into EcoRV-digested pACYC184 downstream of the tet promoter using Gibson Assembly Master Mix (NEB). Other endogenous CRISPR system elements were omitted, except for the last 50 bp of Csn2. An oligomer encoding a spacer with complementary overhangs (Integrated DNA Technology) was cloned into the BsaI-digested vector pDC000 (NEB) and ligated using T7 ligase (Enzymatics) to generate the pCRISPR plasmid. An excitation plasmid containing a spacer (carrying a PAM sequence, also referred to here as a “CRISPR motif sequence”) was generated by ligating a hybrid oligonucleotide (Integral DNA Technologies, Inc.) carrying compatible overhangs into BamHI-digested pUC19. Cloning of all constructs was performed on E. coli strain JM109 (Zymo Research, Inc.).
[0239] According to the manufacturer's instructions, pCRISPR-carrying cells were converted to competent cells using the Z-competent E. coli Transformation Kit and buffer kit (Zymo Research, T3001). In this transformation assay, 50 μL aliquots of pCRISPR-carrying competent cells were thawed on ice and transformed with 1 ng of spacer plasmid or pUC19 on ice for 30 min, followed by heat shock at 42°C for 45 s and then on ice for 2 min. Next, 250 μL of SOC (Invitrogen) was added, followed by incubation at 37°C with shaking for 1 h, and 100 μL of the SOC post-product was plated onto a double-selection plate (12.5 μg / ml chloramphenicol, 100 μg / ml ampicillin). To obtain CFU / ng DNA, the total colony count was multiplied by 3.
[0240] To improve the expression of CRISPR components in mammalian cells, two genes from SF370 locus 1 of *Streptococcus pyogenes* (S. pyogenes) are codon-optimized Cas9 (SpCas9) and RNase III (SpRNase III). To facilitate nuclear localization, a nuclear localization signal (NLS) is included at the amino (N)- or carboxyl (C)-terminus of both SpCas9 and SpRNase III. Figure 2BTo facilitate visualization of protein expression, a fluorescent protein marker is also included at the N- or C-terminus of both proteins. Figure 2B A version of SpCas9 with NLS attached to both the N- and C-termini (2xNLS-SpCas9) was also generated. A construct containing the NLS-fused SpCas9 and spRNase III was transfected into 293FT human embryonic kidney (HEK) cells, and it was found that the relative localization of the NLS to SpCas9 and spRNase III affected its nuclear localization efficiency. The C-terminal NLS was sufficient to target spRNase III to the cell nucleus; in this system, attachment of single copies of these specific NLS to the N- or C-terminus of SpCas9 was insufficient for adequate nuclear localization. In this example, the C-terminal NLS was the C-terminal NLS of the nucleoplasmic protein (KRPAATKKAGQAKKKK) and the C-terminal NLS was the C-terminal NLS of the SV40 large T-antigen (PKKKRKV). In this version of SpCas9 tested, only 2xNLS-SpCas9 exhibited nuclear localization ( Figure 2B ).
[0241] The tracrRNA from the CRISPR locus of *Streptococcus pyogenes* SF370 has two transcription initiation sites, resulting in two transcripts, 89 nucleotides (nt) and 171 nt, which are subsequently processed into the same 75 nt mature tracrRNA. The shorter 89 nt tracrRNA was selected for expression in mammalian cells. Figure 7A The expression construct shown here, where functionality is as shown in Figure 7B (As determined by Surveyor assays). Transcription start sites were marked with +1, and transcription terminators and sequences detected by Northern blotting were also indicated. Expression of processed tracrRNA was also confirmed by Northern blotting. Figure 7C The results of Northern blot analysis of total RNA extracted from 293FT cells, transfected with U6 expression constructs carrying long or short tracrRNA as well as SpCas9 and DR-EMX1(1)-DR, are shown. The left and right panels are from 293FT cells transfected with and without SpRNase III, respectively. U6 indicates a control sample added using a probe targeting human U6 snRNA. Transfection with the short tracrRNA expression construct produced abundant levels of the processed form of tracrRNA (approximately 75 bp). Very low levels of long tracrRNA were detected on the Northern blot.
[0242] To promote precise transcription initiation, the U6 promoter of RNA polymerase III was selected to drive the expression of tracrRNA. Figure 2C Similarly, constructs of the U6 promoter base were developed to express a sequence containing two inverse repeats (DRs) flanking the DRs, also covered by the term "tracr pairing sequence". Figure 2C A pre-crRNA array of single spacers. The starting spacers were designed to target the human EMX1 locus ( Figure 2C The 33-base pair (bp) target site in Cas9 (a 30-bp prototype spacer plus a 3-bp CRISPR motif (PAM) sequence that satisfies the NGG recognition motif of Cas9) is a key gene in the development of the cerebral cortex.
[0243] To test whether heterologous expression of the CRISPR system (SpCas9, SpRNase III, tracrRNA, and pre-crRNA) in mammalian cells could achieve targeted cleavage of mammalian chromosomes, HEK 293FT cells were transfected with a combination of CRISPR components. Since DSBs in mammalian cell nuclei are partially repaired via the non-homologous end joining (NHEJ) pathway leading to indel formation, potential cleavage activity at the target EMX1 locus was assessed using Surveyor assays. Figure 8 (See, for example, Guschin et al., 2010, Methods in Molecular Biology, 649:247). Co-transfection of all four CRISPR components could induce up to 5.0% cleavage in the prototype spacer (see...) Figure 2D Co-transfection of all CRISPR components (minus spRNA enzyme III) also induced up to 4.7% indels in the prototype spacer, suggesting the possible presence of endogenous mammalian RNases that assist crRNA maturation, such as the associated dicer and Drosha enzymes. Removing any of the remaining three components eliminates the genome-cutting activity of the CRISPR system. Figure 2D Sanger sequencing of amplicones containing confirmed cleavage activity at target loci: 5 mutated alleles (11.6%) were found in 43 sequenced clones. Similar experiments using multiple guide sequences yielded indel percentages as high as 29% (see Figures 4–7, 12, and 13). These results define a three-component system for efficient CRISPR-mediated genome modification in mammalian cells. To optimize cleavage efficiency, the applicant also tested whether different isoforms of the tracrRNA affected cleavage efficiency and found that, in the example system, only the short (89-bp) transcript form was able to mediate cleavage of the human EMX1 genomic locus. Figure 7B ).
[0244] Figure 14 provides additional Northern blot analysis of crRNA processing in mammalian cells. Figure 14A A schematic diagram is shown illustrating an expression vector flanked by two single spacers with in-direction repeats (DR-EMX1(1)-DR). Targeting human EMX1 seat prototype spacer 1 (see...) Figure 6 The 30bp spacer and the same-direction repeat sequence are shown in Figure 14A In the sequence below, the straight line indicates the region where its inverse complementary sequence is used to generate the Northern blot probe for EMX1(1)crRNA detection. Figure 14B Northern blot analysis of total RNA extracted from 293FT cells transfected with a U6 expression construct carrying DR-EMX1(1)-DR is shown. The left and right panels are from 293FT cells transfected with and without spRNA enzyme III, respectively. DR-EMX1(1)-DR is processed into mature crRNA only in the presence of SpCas9 and short tracrRNA and is independent of the presence of spRNA enzyme III. The mature crRNA detected from the transfected 293FT total RNA was approximately 33 bp, shorter than the 39–42 bp mature crRNA from Streptococcus pyogenes. These results demonstrate that the CRISPR system can be transplanted into eukaryotic cells and reprogrammed to facilitate the cleavage of endogenous mammalian target polynucleotides.
[0245] Figure 2 illustrates the bacterial CRISPR system described in this example. Figure 2A A schematic diagram is shown illustrating the CRISPR locus from *Streptococcus pyogenes* SF370 and the proposed mechanism for CRISPR-mediated DNA cleavage via this system. Mature crRNA processed from an array of co-repetitive spacers directs Cas9 to genomic targets composed of complementary proto-spacer motifs and proto-spacer adjacent motifs (PAMs). Upon target-spacer base pairing, Cas9 mediates double-strand breaks in the target DNA. Figure 2B The engineered Streptococcus pyogenes Cas9 (SpCas9) and RNase III (SpRNase III) with nuclear localization signals (NLS) that enable their introduction into the nucleus of mammalian cells are shown. Figure 2C Mammalian expression of SpCas9 and SpRNA enzyme III driven by the constitutive EF1a promoter, and tracrRNA and pre-crRNA arrays (DR-spacer-DR) driven by the RNA Pol3 promoter U6, is shown. These promoters facilitate precise transcription initiation and termination. A prototype spacer from the human EMX1 locus with a satisfactory PAM sequence was used as the spacer in the pre-crRNA array. Figure 2D The Surveyor nuclease assays for minor insertions and deletions mediated by SpCas9 are shown. SpCas9 was expressed with and without SpRNase III, tracrRNA, and pre-crRNA arrays carrying EMX1-target spacers. Figure 2E A schematic representation of the base pairing between the target site and the EMX1-targeting crRNA is shown, along with an exemplary chromatogram showing a small deletion adjacent to the SpCas9 cleavage site. Figure 2F The mutated alleles identified from sequencing analysis of 43 clonal amplicones are shown, exhibiting a variety of small insertions and deletions. Dashes indicate deleted bases, and unaligned or mismatched bases indicate insertions or mutations. Scale bar = 10 μm.
[0246] To further simplify this three-component system, a chimeric crRNA-tracrRNA hybrid design was adapted, in which mature crRNA (including a guide sequence) is fused to a portion of tracrRNA via a stem-loop configuration to mimic the natural crRNA:tracrRNA duplex. Figure 3 A). To increase co-delivery efficiency, a bicistronic expression vector was generated to drive the co-expression of chimeric RNA and SpCas9 in transfected cells. Figure 3 A and 8). In parallel, using this bicistronic vector, pre-crRNA (DR-director sequence-DR) was expressed with SpCas9 to induce processing of separately expressed tracrRNA into crRNA (comparative). Figure 13B Top and bottom images). Figure 9 shows the pre-crRNA array ( Figure 9A ) or chimeric crRNA (from Figure 9B The diagram shows the positions of various elements and the insertion sites of the guide sequence downstream of the insertion site and upstream of the EF1α promoter (short lines representing these elements) in the bicistronic expression vector of hSpCas9, illustrating the locations of the various elements and the insertion sites of the guide sequence. Figure 9B The unfolded sequence around the insertion site of the guide sequence also shows a partial DR sequence (GTTTAGAGCTA) and a partial tracrRNA sequence (TAGCAAGTTAAAATAAGGCTAGTCCGTTTTT). The guide sequence can be inserted between BbsI sites using annealed oligonucleotides. The oligonucleotide sequence design is shown below the schematic diagram in Figure 9, indicating the appropriate linker. WPRE represents a posttranscriptional regulatory element of marmot hepatitis virus. The chimeric RNA-mediated cleavage efficiency was tested by targeting the same EMX1 locus as described above. Using both amplicon Surveyor assays and Sanger sequencing, the applicant confirmed that the chimeric RNA design promoted cleavage of the human EMX1 locus with a modification rate of approximately 4.7% (Figure 4).
[0247] By designing chimeric RNAs targeting multiple sites in human EMX1 and PVALB, as well as the mouse Th locus, the prevalence of CRISPR-mediated cleavage in eukaryotic cells was tested by targeting additional genomic loci in both human and mouse cells. Figure 15 Human PVALB ( Figure 15 A) and mouse Th( Figure 15 B) Screening for additional target prototypical spacers at the loci. Schematic diagrams of the gene loci within their respective last exons and the locations of the three prototypical spacers are provided. The underlined sequence includes the 30 bp prototypical spacer sequence and the 3 bp corresponding to the 3' end of the PAM sequence. Prototypical spacers on the sense and antisense strands are indicated above and below the DNA sequence, respectively. Modification rates of 6.3% and 0.75% were achieved at the human PVALB and mouse Th loci, respectively, demonstrating the broad applicability of the CRISPR system across multiple organisms for modifying different loci. Figure 3 B and 6). Although only one-third of the spacer cleavage was detected for each locus using the chimeric construct, all target sequences were cleaved when using the co-expressed pre-crRNA arrangement, with up to 27% effective indel generation. Figure 6 ).
[0248] Figure 13 provides an additional illustration of SpCas9, which can be reprogrammed to target multiple genomic loci in mammalian cells. Figure 13A A schematic diagram of the human EMX1 seat is provided, showing the positions of the five prototype spacers, indicated by an underlined sequence. Figure 13B A schematic diagram of the pre-crRNA / trcrRNA complex is provided, showing hybridization between the homologous repeat regions of pre-crRNA and tracrRNA (top), and a schematic diagram of the chimeric RNA design, including a 20bp guide sequence and tracr pairs and tracr sequences composed of partially homologous repeats, as well as the tracrRNA sequence hybridized in a hairpin structure (bottom). Surveyor assay results comparing the efficiency of Cas9-mediated cleavage at the five prototypical spacers in the human EMX1 locus are shown. Figure 13C In this study, processed pre-crRNA / tracrRNA complexes (crRNA) or chimeric RNA (chiRNA) were used to target each prototype spacer.
[0249] Since the secondary structure of RNA can be decisive for intermolecular interactions, a structure prediction algorithm based on minimum free energy and Boltzmann weighted structure sets was used to compare genome-targeting experiments. Figure 3The assumed secondary structures of all guide sequences used in B) are shown (see, for example, Gruber et al., 2008, Nucleic Acids Research, 36:W70). Analysis revealed that, in most cases, effective guide sequences in the context of chimeric crRNA are essentially devoid of secondary structure motifs, while ineffective guide sequences are more likely to form internal secondary structures that can prevent base pairing with the target prototypical spacer DNA. Therefore, it is possible that variability in spacer secondary structures may affect the efficiency of CRISPR-mediated interference when using chimeric crRNA.
[0250] Figure 3 An example of an instance representation carrier is provided. Figure 3 A schematic diagram is provided for a bicistronic vector used to drive the expression of a synthetic crRNA-tracrRNA chimera (chimeric RNA) along with SpCas9. The chimeric guide RNA contains a 20-bp guide sequence corresponding to a prototype spacer at a target site in the genome. Figure 3 B provides guide sequences targeting human EMX1, PVALB, and mouse Th loci, along with schematic diagrams of their predicted secondary structures. This indicates that modification efficiency at each target site is lower than that shown in the RNA secondary structure diagram (EMX1, n = 216 amplicon sequencing reads; PVALB, n = 224 reads; Th, n = 265 reads). The folding algorithm produces an output where each base is colored according to its assumed probability of the predicted secondary structure, as shown by... Figure 3 The rainbow scale, reproduced in gray in Figure B, is shown. Another vector design for SpCas9 is shown in Figure 44, illustrating a single expression vector incorporating a U6 promoter linked to the insertion site of the guiding oligonucleotide and a Cbh promoter linked to the SpCas9 coding sequence. The vector shown in Figure 44b includes a tracrRNA coding sequence linked to the H1 promoter.
[0251] To test whether spacers containing secondary structures could function in CRISPR-native prokaryotic cells, transformation interference with plasmids containing prototype spacers was tested in Escherichia coli strains heterologously expressing the Streptococcus pyogenes SF370 CRISPR locus 1. Figure 10 The CRISPR locus was cloned into a low-copy E. coli expression vector, and the crRNA array was replaced by a single spacer flanked by a pair of DR (pCRISPR) vectors. E. coli strains containing different pCRISPR plasmids were transformed with activation plasmids containing the corresponding prototype spacers and PAM sequences. Figure 10C). In bacterial assays, all spacers promoted effective CRISPR interference (C). Figure 4C These results suggest that there may be other factors affecting the efficiency of CRISPR activity in mammalian cells.
[0252] To investigate the specificity of CRISPR-mediated cleavage, a series of EMX1-targeted chimeric crRNAs with single-point mutations were used to analyze the impact of single nucleotide mutations in the guide sequence on prototypical spacer cleavage in the mammalian genome. Figure 4A ). Figure 4B The results of a Surveyor nuclease assay comparing the cleavage efficiency of Cas9 when paired with chimeric RNAs of different mutations are shown. A single base mismatch of up to 12 bp at the 5' of PAM essentially eliminated genome cleavage by SpCas9, while spacers with mutations further upstream retained activity against the original prototype spacer target. Figure 4B In addition to PAM, SpCas9 exhibits single-base specificity within the last 12 bp of the spacer. Furthermore, CRISPR can mediate genome cleavage as efficiently as a pair of TALE nucleases (TALEN) targeting the same EMX1 prototype spacer. Figure 4C A schematic diagram is provided showing the design of TALEN targeting EMX1 and Figure 4D The Surveyor gel (n=3) shows the efficiency of TALEN and Cas9.
[0253] A set of components for CRISPR-mediated gene editing in mammalian cells has been established using the error-prone NHEJ mechanism. The ability of CRISPR to stimulate homologous recombination (HR) was tested; homologous recombination is a high-fidelity gene repair pathway that enables precise editing in the genome. Wild-type SpCas9 can mediate site-specific lesions (DSBs), which can be repaired by both NHEJ and HR. Furthermore, the aspartic acid-to-alanine substitution (D10A) in the RuvC I catalytic domain of SpCas9 was engineered to convert the nuclease into a nicking enzyme (SpCas9n; shown in Figure 10). Figure 5A(See, for example, Sapranauskas et al., 2011, Nucleic Acids Research, 39:9275; Gasiunas et al., 2012, Proceedings of the National Academy of Sciences of the United States of America, 109:E2579), thus enabling the nicked genomic DNA to undergo high-fidelity homology-directed repair (HDR). Surveyor assays confirmed that SpCas9n does not generate indels at the EMX1 prototype spacer target. Figure 5B As shown, co-expression of the EMX1-targeting chimeric crRNA with SpCas9 produced indels at the target site, while co-expression with SpCas9n did not produce indels (n=3). Furthermore, sequencing of 327 amplicones did not detect any indels induced by SpCas9n. The same loci were selected to test CRISPR-mediated HR by co-transfection of HEK 293FT cells with chimeric RNA targeting EMX1, hSpCas9, or hSpCas9n, along with an HR template for introducing a pair of restriction enzyme sites (HindIII and NheI) near the prototype spacer. Figure 5C A schematic diagram of the HR strategy is provided, including the relative positions of the recombination sites and primer annealing sequences (arrows). SpCas9 and SpCas9n indeed catalyze the integration of the HR template into the EMX1 locus. PCR amplification of the target region, followed by restriction digestion with HindIII, revealed the cleavage products corresponding to the desired fragment size (shown in...). Figure 5D (Arrows in the restriction fragment length polymorphism analysis) where SpCas9 and SpCas9n mediate similar levels of HR efficiency. The applicant further confirmed HR using Sanger sequencing with genome amplicon sequencing. Figure 5E These results demonstrate the utility of CRISPR in facilitating targeted gene insertion in mammalian genomes. Given that the 14-bp nickase (12-bp from the spacer and 2-bp from PAM) specifically targets wild-type SpCas9, the availability of the nickase can significantly reduce the likelihood of off-target modifications, since single-strand breaks are not substrates for the error-prone NHEJ pathway.
[0254] An expression construct that simulates the natural structure of CRISPR loci with arrayed spacers. Figure 2A To test the potential of multivariate sequence targeting, a single CRISPR array encoding a pair of EMX1- and PVALB-targeting spacers was used, and effective cleavage was detected at both loci. Figure 5FThe schematic design of the crRNA array is shown, along with Surveyor blots demonstrating efficient modifications to the cuts. Targeted deletions of larger genomic regions were also detected via coexisting DSB using spacers targeting two targets within EMX1 separated by 119 bp, with a deletion efficiency of 1.6% (3 out of 182 amplicon). Figure 5G This demonstrates that the CRISPR system can mediate multi-stage editing within a single genome.
[0255] Example 2: CRISPR System Modifications and Alternatives
[0256] The ability to program RNA sequences specifically for DNA cleavage defines a new class of genome engineering tools for a variety of research and industrial applications. Several aspects of the CRISPR system can be further improved to increase the efficiency and versatility of CRISPR targeting. Optimized Cas9 activity can depend on free Mg2+ located in the nucleus of mammalian cells compared to that found in mammalian cell nuclei. 2+ High levels of free Mg 2+ The availability of CRISPR loci (see, for example, Jinek et al., 2012, Science, 337:816) and the preference for NGG motifs located precisely downstream of the prototype spacer limit the ability to target an average of every 12-bp in the human genome (Figure 11, evaluating both the positive and negative strands of human chromosome sequences). Some of these constraints can be overcome by exploring the diversity of CRISPR loci across microbial metagenomics (see, for example, Makarova et al., 2011, Nat Rev Microbiol, 9:467). Other CRISPR loci can be transplanted into the mammalian cellular environment using methods similar to those described in Example 1. For example, Figure 12 shows the adaptation of CRISPR 1 from Streptococcus thermophilus LMD-9 into a type II CRISPR system for heterologous expression in mammalian cells to achieve CRISPR-mediated genome editing. Figure 12A A schematic diagram of CRISPR 1 from Streptococcus thermophilus LMD-9 is provided. Figure 12B The design of the expression system for the Streptococcus thermophilus CRISPR system is shown. The human codon-optimized hStCas9 is expressed using the constitutive EF1α promoter. Mature versions of tracrRNA and crRNA are expressed using the U6 promoter, which promotes precise transcription initiation. Sequences from the mature crRNA and tracrRNA are shown. The single base indicated by the lowercase letter "a" in the crRNA sequence is used to remove the polyU sequence, which serves as the RNA polIII transcription terminator. Figure 12CProvides a guide sequence targeting the human EMX1 locus, along with a schematic diagram of its predicted secondary structures. This shows that modification efficiency at each target site is lower than the RNA secondary structures. The algorithm generating these structures colors each base according to its presumed probability of the predicted secondary structure, such as by... Figure 12C The rainbow scale, reproduced in gray, is shown. Results of hStCas9-mediated cleavage at target loci, determined using Surveyor, are presented. RNA-guided spacers 1 and 2 induced 14% and 6.4%, respectively. Statistical analysis of cleavage activity across biological replicas at these two prototypical spacer sites is also provided. Figure 6 In. Figure 16 Schematic diagrams of additional prototype spacers for the Streptococcus thermophilus CRISPR system and their corresponding PAM sequence targets at the human EMX1 locus are provided. The two prototype spacer sequences are highlighted and the corresponding PAM sequences satisfying the NNAGAAW motif are indicated by an underscore at 3' relative to the corresponding highlighted sequence. Both prototype spacers target the antisense strand.
[0257] Example 3: Sample Target Sequence Selection Algorithm
[0258] Design a software program to identify candidate CRISPR target sequences on both strands of an input DNA sequence based on a desired guide sequence length for a specified CRISPR enzyme and a CRISPR motif (PAM). For example, this could be achieved by searching for 5'-N on both the input sequence and its inverse complement. x The target site of Cas9 from Streptococcus pyogenes, containing the PAM sequence NGG, can be identified by searching for 5'-NGG-3'. Similarly, the target site of Cas9 from Streptococcus pyogenes CRISPR1, containing the PAM sequence NNAGAAW, can be identified by searching for 5'-Nx-NNAGAAW-3' on both the input sequence and its inverse complement. x -NGGNG-3' identifies the target site of NGGNG, a PAM sequence in Cas9 of CRISPR3 for identifying Streptococcus pyogenes. N can be fixed programmatically or specified by the user. x The value "x" in the value, such as 20.
[0259] Because multiple occurrences of a DNA target site in the genome can lead to non-specific genome editing, the program filters out potential sites based on the number of times the sequence appears in the relevant reference genome. For CRISPR enzymes whose sequence specificity is determined by a 'seed' sequence (such as the 11-12 bp 5' from the PAM sequence, including the PAM sequence itself), the filtering step can be based on the seed sequence. Therefore, to avoid editing at additional genomic loci, the filtering results are based on the number of times the seed:PAM sequence appears in the relevant genome. Users can be allowed to select the length of the seed sequence. For the purpose of filtering, users can also be allowed to specify the number of times the seed:PAM sequence appears in the genome. Unique sequences are filtered by default. The filtering level can be changed by altering both the length of the seed sequence and the number of times that sequence appears in the genome. The program can additionally or alternatively provide a guide sequence complementary to one or more reported target sequences by providing the reverse complement of one or more identified target sequences.
[0260] Further details of the methods and algorithms used to optimize sequence selection can be found in U.S. Application Serial No. 61 / 836,080 (Agent No. 44790.11.2022), which is incorporated herein by reference.
[0261] Example 4: Evaluation of multiple chimeric crRNA-tracrRNA hybrids
[0262] This example describes the results obtained for chimeric RNAs (chiRNAs; consisting of a guide sequence, a tracr pairing sequence, and a tracr sequence in a single transcript) with tracr sequences incorporating wild-type tracr RNA sequences of varying lengths. Figure 18a shows a schematic diagram of the bicistronic expression vector for the chimeric RNA and Cas9. Cas9 is driven by the CBh promoter and the chimeric RNA by the U6 promoter. The chimeric guide RNA consists of a 20 bp guide sequence (N) that binds to the tracr sequence (from the first “U” on the lower strand to the end of the transcript), which is truncated at the points indicated. The guide and tracr sequences are separated by the tracr-pairing sequence GUUUUAGAGCUA, followed by the loop sequence GAAA. The results of the Cas9-mediated indel SURVEYOR assay at the human EMX1 and PVALB loci are shown in Figures 18b and 18c, respectively. Arrows indicate the expected SURVEYOR fragments. ChiRNAs are indicated by their "+n" markers, and crRNAs refer to hybrid RNAs in which the guide sequence and tracr sequence are expressed as separate transcripts. Quantification of these results, performed in triplicate, is shown in histograms in Figures 19a and 19b, corresponding to Figures 18b and 18c, respectively ("ND" indicates no indel detected). Prototype spacer IDs and their corresponding genomic targets, prototype spacers, PAM sequences, and strand positions are provided in Table D. The guide sequence is designed to be complementary to the entire prototype spacer sequence (in the case of separate transcripts in hybrid systems), or only partially underlined (in the case of chimeric RNAs).
[0263] Table D:
[0264]
[0265] Cell culture and transfection
[0266] Human embryonic kidney (HEK) cell line 293FT (Lifetech) was maintained at 37°C with 5% CO2 in Durbecco's modified Eagle's Medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Lifetech), 100 U / mL penicillin, and 100 μg / mL streptomycin. 24 hours prior to transfection, 293FT cells were seeded at a density of 150,000 cells / well in 24-well plates (Corning). Cells were transfected using Lipofectamine 2000 (Lifetech) following the manufacturer's recommended protocol. A total of 500 ng of plasmid was used for each well of the 24-well plate.
[0267] SURVEYOR assay for genome modification
[0268] As described above, 293FT cells were transfected with plasmid DNA. Following transfection, cells were incubated at 37°C for 72 hours before genomic DNA extraction. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) following the manufacturer's protocol. In short, pelleted cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and then at 98°C for 10 minutes. For each gene, PCR amplification (primers listed in Table E) was performed on the genomic region flanking the CRISPR target site, and the products were purified using QiaQuick Spin columns (Qiagen) following the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10X Taq DNA polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 20 μl and subjected to a re-annealing process to allow heteroduplex formation: 95 °C for 10 min, cooled from 95 °C to 85 °C at -2 °C / s, cooled from 85 °C to 25 °C at -0.25 °C / s, and held at 25 °C for 1 min. After re-annealing, the product was treated with SURVEYOR nuclease and SURVEYOR enhancer S (Transgenomics) according to the manufacturer's recommended protocol and analyzed on 4%–20% Novex TBE polyacrylamide gels (Lifetech). The gels were stained with SYBR Gold DNA stain (Lifetech) for 30 min and imaged using a GelDoc gel imaging system (Bio-rad). Quantification was based on relative band intensity.
[0269] Table E:
[0270] Sp-EMX1-F EMX1 AAAACCACCCTTCTCTCTGGC Sp-EMX1-R EMX1 GGAGATTGGAGACACGGAGAG Sp-PVALB-F PVALB CTGGAAAGCCAATGCCTGAC Sp-PVALB-R PVALB GGCAGCAAACTCCTTGTCCT
[0271] Computer-aided identification of unique CRISPR target sites
[0272] To identify unique target sites for the Streptococcus pyogenes SF370 Cas9 (SpCas9) enzyme in the genomes of humans, mice, rats, zebrafish, fruit trees, and Caenorhabditis elegans, we developed a software package to scan both strands of a DNA sequence and identify all possible SpCas9 target sites. In this example, each SpCas9 target site was operationally defined as a 20 bp sequence, followed by an NGG prototype spacer adjacent motif (PAM) sequence, and we identified all chromosomes that satisfy this 5'-N... 20-NGG-3' defined sequences. To prevent non-specific genome editing, all target sites were filtered based on the number of times the sequence appeared in the relevant reference genome after all potential sites were identified. To utilize the sequence specificity of Cas9 activity conferred by a 'seed' sequence (which could be, for example, a sequence approximately 11-12 bp starting from the 5' of the PAM sequence), the 5'-NNNNNNNNNN-NGG-3' sequence was selected as unique in the relevant genome. All genome sequences were downloaded from the UCSC Genome Browser (human genome hg19, mouse genome mm9, rat genome rn5, zebrafish genome danRer7, fruitwort genome dm4, and Caenorhabditis elegans genome ce10). Comprehensive search results were available using the UCSC Genome Browser information. Exemplary visualizations of some target sites in the human genome are provided. Figure 21 middle.
[0273] Initially, three sites within the EMX1 locus were targeted in human HEK 293FT cells. The genomic modification efficiency of each chiRNA was assessed using the SURVEYOR nuclease assay, which detects mutations resulting from DNA double-strand breaks (DSBs) and their subsequent repair via the non-homologous end joining (NHEJ) DNA damage repair pathway. Co was designated as the construct indicator for chiRNA(+n), with the wild-type tracrRNA up to the +n nucleotide included in the chimeric RNA construct, where values of 48, 54, 67, and 85 were used for n. The chimeric RNAs contained long fragments of wild-type tracrRNA (chiRNA(+67) and chiRNA(+85)) mediated DNA cleavage at all three EMX1 target sites, with chiRNA(+85) exhibiting significantly higher levels of DNA cleavage than the corresponding crRNA / tracrRNA hybrids expressing the guide sequence and tracr sequence in separate transcriptomic form (Figs. 18b and 19a). Two sites in the PVALB locus were also targeted using chiRNAs that did not produce detectable cleavage using a hybrid system (the guide sequence and the tracr sequence expressed as separate transcripts). chiRNAs (+67) and (+85) were able to mediate significant cleavage at the two PVALB prototypical spacers (Figs. 18c and 19b).
[0274] For all five targets at the EMX1 and PVALB loci, a consistent increase in genomic modification efficiency was observed with the increase in tracr sequence length. Without being bound by any theory, the secondary structures formed by the 3' end of the tracrRNA may play a role in enhancing the CRISPR complex formation rate. Figure 21The diagram provides a schematic of the predicted secondary structure for each of the chimeric RNAs used in this example. This secondary structure was predicted using RNA folding (http: / / rna.tbi.univie.ac.at / cgi-bin / RNAfold.cgi) with a minimum free energy and partition function algorithm. The pseudocolor of each base (reproduced in grayscale) indicates the probability of pairing. Because chiRNAs with longer tracr sequences can cleave targets not cleaved by the native CRISPR crRNA / tracrRNA hybrid, it is possible that the chimeric RNA can be loaded onto Cas9 more efficiently than its native hybrid pair. To facilitate the application of Cas9 in site-specific genome editing in eukaryotic cells and organisms, all predicted unique target sites for *Streptococcus pyogenes* Cas9 were computer-identified in the genomes of humans, mice, rats, zebrafish, *C. elegans*, and *C. pyrenoidosa*. Chimeric RNAs can be designed to target Cas9 enzymes from other microorganisms to expand the target space of CRISPR RNA-programmable nucleases.
[0275] Figure 22 A schematic diagram of an exemplary bicistronic expression vector for expressing chimeric RNA comprising up to +85 nucleotides of wild-type tracr RNA sequence and SpCas9 with a nuclear localization sequence is shown. SpCas9 is expressed from a CBh promoter and terminated by a bGH polyA signal (bGH pA). The extended sequence shown in the schematic corresponds to the region surrounding the guide sequence insertion site and includes, from 5' to 3', the 3' portion of the U6 promoter (first masking region), the BbsI cleavage site (arrow), a partial homologous repeat (tracr pair sequence GTTTTAGAGCTA, underlined), the loop sequence GAAA, and the +85 tracr sequence (the underlined sequence following the loop sequence). An exemplary guide sequence insertion fragment is shown after the guide sequence insertion site, with the nucleotides of the guide sequence indicated by "N" for the selected target.
[0276] The sequences described in the above examples are as follows (the polynucleotide sequences are 5' to 3'):
[0277] U6-short tracrRNA (Streptococcus pyogenes SF370):
[0278] (Bold text = tracrRNA sequence; Underline = Termination sequence)
[0279] U6-long tracrRNA (Streptococcus pyogenes SF370):
[0280] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGGTAGTATTAAGTATTGTTTTATGGCTGATAAATTTCTTTGAATTTCTCCTTGATTATTTGTTATAAAAGTTATAAAATAATCTTGTTGGAACCATTCAAAACAGCATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT
[0281] U6-DR-BbsI backbone-DR (Streptococcus pyogenes SF370):
[0282] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGGGTTTTAGAGCTATGCTGTTTTGAATGGTCCCAAAACGGGTCTTCGAGAAGACGTTTTAGAGCTATGCTGTTTTGAATGGTCCCAAAAC
[0283] U6-chimeric RNA-BbsI backbone (Streptococcus pyogenes SF370)
[0284] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGGGTCTTCGAGAAGACCTGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCG
[0285] NLS-SpCas9-EGFP:
[0286]
[0287] SpCas9-EGFP-NLS:
[0288]
[0289] NLS-SpCas9-EGFP-NLS:
[0290]
[0291] NLS-SpCas9-NLS:
[0292]
[0293] NLS-mCherry-SpRNase 3:
[0294] MFLFLSLTSFLSSSRTLVSKGEEDNMAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGVVTVTQDSSLQDGEFIYKVKLRGTNFPSDGPVMQKKTMGWEASSERMYPEDGALKGEIKQRLKLKDGGHYDAEVKTTYKAKKPVQLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGMDELYKGSKQLEELLSTSFDIQFNDLTLLETAFTHTSYANEHRLLNVSHNERLEFLGDAVLQLIISEYLFAKYPKKTEGDMSKLRSMIVREESLAGFSRFCSFDAYIKLGKGEEKSGGRRRDTILGDLFEAFLGALLLDKGIDAVRRFLKQVMIPQVEKGNFERVKDYKTCLQEFLQTKGDVAIDYQVISEKGPAHAKQFEVSIVVNGAVLSKGLGKSKKLAEQDAAKNALAQLSEV
[0295] SpRNase 3-mCherry-NLS:
[0296] MKQLEELLSTSFDIQFNDLTLLETAFTHTSYANEHRLLNVSHNERLEFLGDAVLQLIISEYLFAKYPKKTEGDMSKLRSMIVREESLAGFSRFCSFDAYIKLGKGEEKSGGRRRDTILGDLFEAFLGALLLDKGIDAVRRFLKQVMIPQVEKGNFERVKDYKTCLQEFLQTKGDVAIDYQVISEKGPAHAKQFEVSIVVNGAVLSKGLGKSKKLAEQDAAKNALAQLSEVGSVSKGEEDNMAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGVVTVTQDSSLQDGEFIYKVKLRGTNFPSDGPVMQKKTMGWEASSERMYPEDGALKGEIKQRLKLKDGGHYDAEVKTTYKAKKPVQLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGMDELYKKRPAATKKAGQAKKKK
[0297] NLS-SpCas9n-NLS (the D10A nickase mutation is in lowercase):
[0298]
[0299] hEMX1-HR Template-HindII-NheI:
[0300]
[0301] NLS-StCsn1-NLS:
[0302]
[0303] U6-St_tracrRNA(7-97):
[0304] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGGAATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTGCAGTTTTAAAATTATGTTTTAAAAT GGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGTTACTTAAATCTTGCAGAAGCTACAAAGATAAGGCTTCATGCCGAAATCAACACCCTGTCATTTTATGGCAGGGTGTTTTCGTTATTTAA
[0305] U6-DR-spacer-DR (Streptococcus pyogenes SF370)
[0306] (Lowercase underscore = repetition in the same direction; N = guide sequence; bold = terminator)
[0307] Chimeric RNA, containing +48 tracr RNA (Streptococcus pyogenes SF370).
[0308] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0309] Chimeric RNA, containing +54tracr RNA (Streptococcus pyogenes SF370).
[0310] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0311] Chimeric RNA, containing +67tracr RNA (Streptococcus pyogenes SF370).
[0312] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0313] Chimeric RNA, containing +85 tracr RNA (Streptococcus pyogenes SF370).
[0314] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0315] CBh-NLS-SpCas9-NLS
[0316] CGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGTCGAGGTGAGCCCCACGTTCTGCTTCACTCTCCCCATCTCCCCCCCCTCCCCACCCCCAATTTTGTATTTATTTATTTTTTAATTATTTTGTGCAGCGATGGGGGCGGGGGGGGGGGGGGGGCGCGCGCCAGGCGGGGCGGGGCGGGGCGAGGGGCGGGGCGGGGCGAGGCGGAGAGGTGCGGCGGCAGCCAATCAGAGCGGCGCGCTCCGAAAGTTTCCTTTTATGGCGAGGCGGCGGCGGCGGCGGCCCTATAAAAAGCGAAGCGCGCGGCGGGCGGGAGTCGCTGCGACGCTGCCTTCGCCCCGTGCCCCGCTCCGCCGCCGCCTCGCGCCGCCCGCCCCGGCTCTGACTGACCGCGTTACTCCCACAGGTGAGCGGGCGGGACGGCCCTTCTCCTCCGGGCTGTAATTAGCTGAGCAAGAGGTAAGGGTTTAAGGGATGGTTGGTTGGTGGGGTATTAATGTTTAATTACCTGGAGCACCTGCCTGAAATCACTTTTTTTCAGGTTGGaccggtgccacc ATGGACTATAAGGACCACGACGGAGACTACAAGGATCATGATATTGATTACAAAGACGATGACGATAAGATGGCC CCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCCGACAAGAAGTACAGCATCGGCCTGGACATCG GCACCAACTCTGTGGGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAATTCAAGGTGCTGGGCAA CACCGACCGGCACAGCATCAAGAAGAACCTGATCGGAGCCCTGCTGTTCGACAGCGGCGAAACAGCCGAGGCCACC CGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAGAACCGGATCTGCTATCTGCAAGAGATCTTCAGCA ACGAGATGGCCAAGGTGGACGACAGCTTCTTCCACAGACTGGAAGAGTCCTTCCTGGTGGAAGAGGATAAGAAGCA CGAGCGGCACCCCATCTTCGGCAACATCGTGGACGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCACCTG AGAAAGAAACTGGTGGACAGCACCGACAAGGCCGACCTGCGGCTGATCTATCTGGCCCTGGCCCACATGATCAAGT TCCGGGGCCACTTCCTGATCGAGGGCGACCTGAACCCCGACAACAGCGACGTGGACAAGCTGTTCATCCAGCTGGT GCAGACCTACAACCAGCTGTTCGAGGAAAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCTGCC AGACTGAGCAAGAGCAGACGGCTGGAAAATCTGATCGCCCAGCTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGCA ACCTGATTGCCCTGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGCCGAGGATGCCAAACTGCA GCTGAGCAAGGACACCTACGACGACGACCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTGTTT CTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACATCCTGAGAGTGAACACCGAGATCACCAAGGCCC CCCTGAGCGCCTCTATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCTGAAAGCTCTCGTGCGGCA GCAGCTGCCTGAGAAGTACAAAGAGATTTTCTTCGACCAGAGCAAGAACGGCTACGCCGGCTACATTGACGGCGGA GCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGGAAAAGATGGACGGCACCGAGGAACTGCTCGTGA AGCTGAACAGAGAGGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATCCCCCACCAGATCCACCTGGG AGAGCTGCACGCCATTCTGCGGCGGCAGGAAGATTTTTACCCATTCCTGAAGGACAACCGGGAAAAGATCGAGAAG ATCCTGACCTTCCGCATCCCCTACTACGTGGGCCCTCTGGCCAGGGGAAACAGCAGATTCGCCTGGATGACCAGAA AGAGCGAGGAAACCATCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCTTCCGCCCAGAGCTTCATCGA GCGGATGACCAACTTCGATAAGAACCTGCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTACTTC ACCGTGTATAACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGC AGAAAAAGGCCATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAGCAGCTGAAAGAGGACTACTT CAAGAAAATCGAGTGCTTCGACTCCGTGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCACATAC CACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGACAATGAGGAAAACGAGGACATTCTGGAAGATATCG TGCTGACCCTGACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCCCACCTGTTCGACGA CAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGCAGGCTGAGCCGGAAGCTGATCAACGGCATC CGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGC TGATCCACGACGACAGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAGGGCGATAGCCTGCA CGAGCACATTGCCAATCTGGCCGGCAGCCCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGTGGACGAG CTCGTGAAAGTGATGGGCCGGCACAAGCCCGAGAACATCGTGATCGAAATGGCCAGAGAGAACCAGACCACCCAGA AGGGACAGAAGAACAGCCGCGAGAGAATGAAGCGGATCGAAGAGGGCATCAAAGAGCTGGGCAGCCAGATCCTGAA AGAACACCCCGTGGAAAACACCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGGGCGGGATATG TACGTGGACCAGGAACTGGACATCAACCGGCTGTCCGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGA AGGACGACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGGGGCAAGAGCGACAACGTGCCCTCCGA AGAGGTCGTGAAGAAGATGAAGAACTACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAGAGAAAGTTCGAC AATCTGACCAAGGCCGAGAGAGGCGGCCTGAGCGAACTGGATAAGGCCGGCTTCATCAAGAGACAGCTGGTGGAAA CCCGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATGAACACTAAGTACGACGAGAATGACAAGCT GATCCGGGAAGTGAAAGTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCGGAAGGATTTCCAGTTTTACAAA GTGCGCGAGATCAACAACTACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACCGCCCTGATCAAAA AGTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGTGTACGACGTGCGGAAGATGATCGCCAAGAG CGAGCAGGAAATCGGCAAGGCTACCGCCAAGTACTTCTTCTACAGCAACATCATGAACTTTTTCAAGACCGAGATT ACCCTGGCCAACGGCGAGATCCGGAAGCGGCCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGATA AGGGCCGGGATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCCAAGTGAATATCGTGAAAAAGACCGAGGTGCA GACAGGCGGCTTCAGCAAAGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCGCCAGAAAGAAGGACTGG GACCCTAAGAAGTACGGCGGCTTCGACAGCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAAGG GCAAGTCCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAAGAA TCCCATCGACTTTCTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATCAAGCTGCCTAAGTACTCC CTGTTCGAGCTGGAAAACGGCCGGAAGAGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAAACGAACTGGCCC TGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACTATGAGAAGCTGAAGGGCTCCCCCGAGGATAATGA GCAGAAACAGCTGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATCGAGCAGATCAGCGAGTTCTCCAAG AGAGTGATCCTGGCCGACGCTAATCTGGACAAAGTGCTGTCCGCCTACAACAAGCACCGGGATAAGCCCATCAGAG AGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAATCTGGGAGCCCCTGCCGCCTTCAAGTACTTTGACAC CACCATCGACCGGAAGAGGTACACCAGCACCAAAGAGGTGCTGGACGCCACCCTGATCCACCAGAGCATCACCGGC CTGTACGAGACACGGATCGACCTGTCTCAGCTGGGAGGCGACTTTCTTTTTCTTAGCTTGACCAGCTTTCTTAGTA GCAGCAGGACGCTTTAA (Underline = NLS-hSpCas9-NLS)
[0317] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0318] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0319] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0320] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0321] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0322] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0323] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0324] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0325] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0326] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0327] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0328] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0329] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0330] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0331] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0332] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0333] Example chimeric RNA targeting Streptococcus thermophilus LMD-9 CRISPR1 Cas9 (PAM with NNAGAAW)
[0334] (N = guide sequence; first underscore = TRACR pairing sequence; second underscore = TRACR sequence; bold = terminator)
[0335] Example chimeric RNA (PAM with NGNG) targeting Streptococcus thermophilus LMD-9 CRISPR3 Cas9
[0336] (N = guide sequence; first underlined = TRACR pairing sequence; second underlined = TRACR sequence; bold = terminator)
[0337] Cas9, a codon-optimized version of the LMD-9 CRISPR3 locus from Streptococcus thermophilus (with an NLS at both the 5' and 3' ends).
[0338]
[0339] Example 5: RNA-guided editing of bacterial genomes using the CRISPR-Cas system.
[0340] The applicant uses the CRISPR-associated endonuclease Cas9 to introduce precise mutations in the genomes of *Streptococcus pneumoniae* and *Escherichia coli*. This approach relies on Cas9-guided cleavage at the target site to kill unmutated cells and circumvents the need for selective markers or anti-selection systems. Cas9 specificity is achieved through reprogramming short CRISPR RNA (crRNA) sequences to allow for single and multinucleotide alterations on the editing template. The simultaneous use of two crRNAs enables multivariate mutagenesis. In *Streptococcus pneumoniae*, nearly 100% of cells surviving Cas9 cleavage contain the desired mutation, and this figure rises to 65% when combined with recombinant engineering in *E. coli*. The applicant comprehensively analyzed Cas9 target requirements to define the range of targetable sequences and demonstrated strategies for editing sites that do not meet these requirements, showcasing the diversity of this technique in bacterial genome engineering.
[0341] Understanding gene function depends on the ability to alter the DNA sequence within a cell in a controlled manner. Site-specific mutagenesis in eukaryotes is achieved using sequence-specific nucleases that promote homologous recombination of a template DNA containing the mutation of interest. Zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and homing meganucleases can be programmed to cut the genome at specific locations, but these pathways require the engineering of novel enzymes for each target sequence. In prokaryotes, mutagenesis methods either introduce a selection marker into the edited locus or require a two-step approach including an anti-selection system. Recently, phage recombinant proteins have been used for recombination engineering, a technique that promotes homologous recombination of linear DNA or oligonucleotides. However, because there is no screening for mutations, the efficiency of recombination engineering can be relatively low (from 0.1%–10% for point mutations to as low as 10% for larger modifications). -5 -10 -6 In many cases, screening large numbers of colonies is required. Therefore, there remains a need for new, affordable, easy-to-use, and effective technologies for genetic engineering in both eukaryotes and prokaryotes.
[0342] Recent work on the CRISPR (regularly spaced clustered short palindromic repeats) adaptive immune system in prokaryotes has led to the identification of sequence-specific nucleases programmed with small RNAs. CRISPR loci consist of a series of repetitive sequences separated by 'spacer' sequences that match the genomes of bacterial bacteriophages and other mobile genetic agents. This repeat-spacer array is transcribed into a long precursor and processed within the repeat sequences to produce a small crRNA that specifies the target sequence (also called the proto-spacer) to be cleaved by the CRISPR system. Necessary for cleavage is the presence of a sequence motif just downstream of the target region, called the proto-spacer adjacent motif (PAM). CRISPR-associated (cas) genes typically flank this repeat-spacer array and encode enzyme mechanisms responsible for crRNA biologicity and targeting. Cas9 is a dsDNA endonuclease that uses a crRNA to direct the cleavage site. Loading of this crRNA onto Cas9 occurs during the processing of the crRNA precursor and requires a small antisense RNA, tracrRNA, and RNase III. Compared to genome editing with ZFN or TALEN, altering Cas9 target specificity does not require protein engineering but only short crRNA-guided design.
[0343] The applicant recently disclosed that introducing a CRISPR system targeting a chromosomal locus in *Streptococcus pneumoniae* resulted in the killing of these transformed cells. Incidental survivors were observed to contain mutations in the target region, indicating that Cas9 dsDNA endonuclease activity targeting endogenous targets can be used for genome editing. The applicant disclosed that marker-free mutations can be introduced by transformation that recombines within the genome and eliminates the recognition of a template DNA fragment by the Cas9 target. Specific guidance of Cas9 with several different crRNAs allows for the introduction of multiple mutations simultaneously. The applicant also characterized the sequence requirements for Cas9 targeting in detail and disclosed that this pathway can be combined with recombination engineering for genome editing in *E. coli*.
[0344] Results: Genome editing via Cas9 cleavage of a chromosome target.
[0345] The Streptococcus pneumoniae strain crR6 contains a Cas9-based CRISPR system that cleaves a target sequence present in bacterial phage φ8232.5. This target was integrated into a second strain, R6. 8232.5 The target sequence of a mutation, located at the srtA chromosomal locus, was integrated into a third strain, R6, within the PAM region. 370.1 The srtA locus in this strain causes it to be 'immune' to CRISPR cleavage. Figure 28 a) The applicant transformed R6 cells with genomic DNA from crR6 cells. 8232.5 and R6 370.1 Cells, hoping to successfully transform R6 8232.5 Cells will cause cleavage of this target locus and cell death. Contrary to expectations, the applicant isolated R6. 8232.5 The transformant, although it is larger than R6 370.1 The converter has approximately 10 times less efficiency. Figure 28 b). Eight R6s 8232.5 Genetic analysis of transformants ( Figure 28 The results reveal that most of these are products of double recombination events. By replacing the φ8232.5 target with the wild-type srtA locus of the crR6 genome, this double recombination event eliminates the toxicity of Cas9 targeting, and this elimination does not include the prototype spacer required for Cas9 recognition. These results are evidence that the parallel introduction of a CRISPR system (target construct) targeting a genomic locus along with a template (edit template) used for recombination into the targeted locus leads to targeted genome editing. Figure 23 a).
[0346] To establish a simplified system for genome editing, the applicant modified the CRISPR loci in strain crR6 by deleting cas1, cas2, and csn2, genes that have been shown to be dispensable for CRISPR targeting, resulting in strain crR6M. Figure 28 a). This strain retains the same characteristics as crR6 ( Figure 28 b). To increase the efficiency of Cas9-based editing and demonstrate that a selected template DNA can be used to control the introduced mutation, the applicant used the wild-type srtA gene or the mutant R6. 370.1 Co-transformation of the target's PCR products into R6 8232.5 Cells, either of these two targets, should be resistant to Cas9 cleavage. This results in a 5 to 10-fold increase in transformation frequency compared to genomic crR6 DNA alone. Figure 23 b). Editing efficiency has also increased substantially; the transformants tested in 8 / 8 contained a wild-type srtA copy, and those in 7 / 8 contained copies present in R6. 370.1 PAM mutations in the target ( Figure 23 b and Figure 29 a). Taken together, these results suggest the potential of Cas9-assisted genome editing.
[0347] Analysis of Cas9 target requirements: To introduce specific alterations into the genome, an editing template carrying a mutation that eliminates Cas9-mediated cuts must be used, thereby preventing cell death. This is readily achievable when seeking a target absence or its replacement by another sequence (gene insertion). When the goal is to generate gene fusions or single nucleotide mutations, the elimination of Cas9 nuclease activity is only possible by introducing mutations into the editing template that alters the PAM or the prototype spacer. To determine the constraints of CRISPR-mediated editing, the applicant comprehensively analyzed the PAM and prototype spacer mutations that negate CRISPR targeting.
[0348] The aforementioned study proposed that *Streptococcus pyogenes* Cas9 requires an NGG PAM exactly downstream of the prototype spacer. However, because only a very limited number of PAM inactivation mutations have been described to date, the applicant conducted a systematic analysis to identify all 5-nucleotide sequences following the elimination of the CRISPR-cleaved prototype spacer. The applicant used randomized oligonucleotides to generate all 1,024 possible PAM sequences in a heterogeneous PCR product transformed into crR6 or R6 cells. Constructs carrying functional PAMs were expected to be recognized and disrupted by Cas9 in crR6 (but not R6 cells). Figure 24 a) More than 2×10 5 Colonies were pooled together to extract DNA, which was used as a template for co-amplification of all targets. Deep sequencing of the PCR products revealed all 1,024 sequences, covering a range of 5 to 42,472 reads (see the "Deep Sequencing Data Analysis" section). The function of each PAM was estimated by its relative proportion of reads in crR6 samples to those in R6 samples. Analysis of the first three bases of the PAMs (averaged by the last two bases) clearly showed that the NGG type was unrepresentative in the crR6 transformant. Figure 24 b). Furthermore, the last two bases had no detectable effect on NGG PAM (see the "Deep Sequencing Data Analysis" section), indicating that the NGGNN sequence is sufficient to authorize Cas9 activity. Partial targeting was observed against the NAG PAM sequence ( Figure 24b). Furthermore, the NNGGN motif partially deactivates CRISPR targeting (Table G), indicating that the NGG motif can still be recognized by Cas9 with reduced efficiency even with a 1 bp shift. These data elucidate the molecular mechanism of Cas9 target recognition and reveal that NGG (or CCN on the complementary strand) sequences are sufficient to target Cas9, and that NGG-to-NAG or NNGGN mutations should be avoided in this editing template. Due to the high frequency of these trinucleotide sequences (every 8 bp), this means that almost any location in the genome can be edited. Indeed, the applicant tested ten randomly selected targets carrying different PAMs and found that all were functional ( Figure 30 ).
[0349] Another way to disrupt Cas9-mediated cleavage is to introduce mutations in the prototype spacer region of the edited template. Point mutations within the 'seed sequence' (8 to 10 prototype spacer nucleotides adjacent to the PAM) are known to eliminate cleavage performed by the CRISPR nuclease. However, the precise length of this region is unknown, and it is unclear whether any mutation in the seed could disrupt Cas9 target recognition. The applicant followed the same deep sequencing approach described above to randomize the entire prototype spacer involved in the base pairs contacting the crRNA and to identify all sequences that disrupt the target. (The last sentence appears to be incomplete and possibly refers to a specific sequence or location.) 8232.5 Each of the 20 matching nucleotides (14) in the spc1 target in the cell ( Figure 23 a) were randomized and transformed into crR6 and R6 cells ( Figure 24 a). Consistent with the presence of a seed sequence, a mutation exactly 12 nucleotides upstream of this PAM negates the Cas9 cut ( Figure 24 c). However, different mutations exhibit significantly different effects. The terminal positions of the seed (12 to 7, starting from the PAM) are resistant to most mutations, and only one specific base substitution negates targeting. In contrast, mutations at the proximal positions (6 to 1, except 3) eliminate Cas9 activity for any nucleotide, although each specific substitution is at a different level. At position 3, only two substitutions affect CRISPR activity and at different intensities. The applicant infers that while mutations in the seed sequence prevent CRISPR targeting, there are limitations regarding the nucleotide changes that can occur at each position in the seed. Furthermore, these limitations are most likely to vary for different spacer sequences. Therefore, the applicant believes that mutations in the PAM sequence, if possible, should be the preferred editing strategy. Alternatively, multiple mutations in the seed sequence could be introduced to prevent Cas9 nuclease activity.
[0350] Cas9-mediated genome editing in Streptococcus pneumoniae: To develop a rapid and efficient method for targeted genome editing, the applicant engineered the strain crR6Rk, a strain in which spacers can be easily introduced via PCR. Figure 33 The applicant decided to edit the β-galactosidase (bgaA) gene of *Streptococcus pneumoniae*, whose activity can be easily measured. The applicant introduced alanine substitutions at the active site of this enzyme: R481A (R→A) and N563A, E564A (NE→AA) mutations. To illustrate different editing strategies, the applicant designed mutations in both the PAM sequence and the prototype spacer seed. In both cases, the same targeting construct (CCA in the complementary strand) was used, with a crRNA having a region complementary to a region of the β-galactosidase gene adjacent to the TGG PAM sequence. Figure 26 The R→A editing template produces a trinucleotide mismatch (CGT to GCA, also introducing a BtgZI restriction site) on the prototype spacer seed sequence. In the NE→AA editing template, the applicant simultaneously introduces a synonymous mutation that produces an inactive PAM (TGG to TTG) along with a mutation 218 nt downstream of the prototype spacer region (AAT GAA to GCT GCA, also producing a TseI restriction site). The latter editing strategy suggests the possibility of using a distal PAM to create a mutation at a location where a suitable target may be difficult to select. For example, although the Streptococcus pneumoniae R6 genome (which has a GC content of 39.7%) contains a PAM motif every 12 bp on average, some PAM motifs are isolated up to 194 bp ( Figure 33 In addition, the applicant designed a 6,664 bp ΔbgaA in-frame deletion. In all three cases, co-transformation with both the targeted and edited templates produced 10 times more kanamycin-resistant cells than co-transformation with a control edited template containing a wild-type bgaA sequence. Figure 25 b). The applicant genotyped 24 transformants (8 for each editing experiment) and found that all but one of them bound the desired alteration. Figure 25 c). DNA sequencing also confirmed the presence of both introduced mutations and secondary mutations in the target region. Figure 29 (b, c). Finally, the applicant measured β-galactosidase activity to confirm that all edited cells exhibited the desired phenotype. Figure 25 d).
[0351] Cas9-mediated editing can also be used to generate multiple mutations for studying biological pathways. The applicant decided to demonstrate this regarding the sorting enzyme-dependent pathway that anchors surface proteins to the envelope of Gram-positive bacteria. The applicant introduced a sorting enzyme deletion through co-transformation of a chloramphenicol resistance-targeting construct and a ΔsrtA editing template. Figure 33 a, b), and then a ΔbgaA deletion was introduced using a kanamycin resistance-targeting construct that replaced the aforementioned one. In Streptococcus pneumoniae, β-galactosidase is covalently linked to the cell wall via a sorting enzyme. Therefore, the deletion of srtA results in the release of surface proteins into the supernatant, and this double deletion does not produce detectable β-galactosidase activity. Figure 34 c). This sequential screening can be repeated many times as needed until multiple mutations are produced.
[0352] These two mutations can also be introduced at the same time. The applicant designed a targeting construct containing two spacers, one matching srtA and the other matching bgaA, and performed co-transformation with two edit templates at the same time. Figure 25 e). Genetic analysis of the transformants showed that editing occurred in 6 / 8 of the cases. Figure 25 f). Notably, the remaining two clones each contained either a ΔsrtA or a ΔbgaA deletion, indicating the possibility of combinatorial mutagenesis using Cas9. Ultimately, to eliminate the CRISPR sequence, the applicant introduced a plasmid containing the bgaA target and a spectinomycin resistance gene along with genomic DNA from the wild-type strain R6. The spectinomycin resistance transformants that retained this plasmid eliminated the CRISPR sequence. Figure 34 a, d).
[0353] Mechanism and efficiency of editing: To understand the underlying mechanisms of genome editing using Cas9, the applicant designed an experiment in which editing efficiency was measured independently of Cas9 cleavage. The applicant integrated the ermAM erythromycin resistance gene into the srtA locus and introduced a premature termination codon using Cas9-mediated editing. Figure 33 The resulting strain (JEN53) contains one ermAM (terminator) allele and is sensitive to erythromycin. This strain can be used to evaluate the efficiency by which the ermAM gene is restored by measuring the fraction of cells that recover antibiotic resistance with or without Cas9 cleavage. JEN53 was transformed with an edited template that restores the wild-type allele, along with a kanamycin-resistant CRISPR construct (CRISPR::ermAM(terminator)) targeting the ermAM (terminator) allele or a control construct without a spacer. ( Figure 26 (a, b) In the absence of kanamycin selection, the edited colony fraction is approximately 10-1. -2 (Erythromycin resistance CFU / Total CFU) Figure 26 c) represents the baseline frequency of recombinants without Cas9-mediated selection of unedited cells. However, if kanamycin selection was applied and the control CRISPR construct was co-transformed, the fraction of edited colonies increased to approximately 10-1. -1 (Kanamycin-and erythromycin-resistant CFU / Kanamycin-resistant CFU) Figure 26 c). This result shows that screening for recombination at this CRISPR locus, targeting Cas9 cuts independent of the genome, for recombination at the ermAM locus, indicates that a subset of cells is more readily transformed and / or recombined. Transformation of the CRISPR::ermAM (terminator) construct followed by kanamycin screening resulted in an increase in the fraction of erythromycin-resistant, edited cells to 99%. Figure 26 c). To determine whether this increase was caused by the killing of unedited cells, the applicant compared the results using the CRISPR::ermAM (terminator) or... These kanamycin-resistant colony-forming units (CFUs) were obtained after co-transformation of JEN53 cells with the construct.
[0354] The applicant counted 5.3 times fewer kanamycin-resistant colonies (2.5 × 10⁻⁶) after transformation of the ermAM (terminator) construct. 4 / 4.7×10 3 , Figure 35 a) This result indicates that Cas9 targeting a chromosomal locus does indeed lead to the killing of unedited cells. Finally, because it is known that introducing dsDNA breaks into the bacterial chromosome triggers a repair mechanism that increases the recombination rate of damaged DNA, the applicant investigated whether the Cas9-induced cleavage induced recombination of the edited template. The applicant co-transformed the CRISPR::erm (termination) construct with [the specific formula is missing from the original text]. The construct resulted in a 2.2-fold increase in colony count. Figure 26 d) indicates the presence of moderate induction of recombination. Taken together, these results show that co-selection of transformable cells, induction of recombination via Cas9-mediated cleavage, and selection of unedited cells each contribute to highly efficient genome editing in Streptococcus pneumoniae.
[0355] Because Cas9's cutting of the genome kills unedited cells, technicians would not expect to recover any kanamycin-resistant cells containing the Cas9 cassette but without the edit template. However, in the absence of the edit template, the applicant recovered numerous kanamycin-resistant colonies after transformation with the CRISPR::ermAM (terminator) construct. Figure 35 a) These 'escape' CRISPR-induced cell death generate a background that determines the limitations of this method. This background frequency can be calculated as the frequency of the CRISPR::ermAM (terminator) in this experiment. CFU, 2.6 × 10 -3 (7.1×10 1 / 2.7×10 4 The ratio of ) means that if the recombination frequency of the edited template is less than this value, CRISPR screening may not effectively restore the desired mutant in this context. To understand the origin of these cells, the applicant genotyped 8 background colonies and found 7 containing deletions of the target spacer ( Figure 35 b) and one of them contains a presumed passivation mutation in Cas9 ( Figure 35 c).
[0356] Genome editing using Cas9 in *E. coli*: Cas9 targeting can only be activated via chromosomal integration of the CRISPR-Cas system in organisms with high recombination rates. To develop a more general approach applicable to other microorganisms, the applicant decided to use a plasmid-based CRISPR-Cas system for genome editing in *E. coli*. Two plasmids were constructed: a pCas9 plasmid carrying the tracrRNA, Cas9, and a chloramphenicol resistance cassette. Figure 36 The applicant constructed a pCRISPR::rpsL plasmid containing an array of CRISPR spacers, and a pCRISPR kanamycin resistance plasmid. To measure editing efficiency independent of CRISPR selection, the applicant sought to introduce an A-to-C transversion into the rpsL gene, which confers streptomycin resistance. The applicant constructed a pCRISPR::rpsL plasmid containing a spacer that will guide Cas9 cleavage of the wild-type, rather than mutant, rpsL allele. Figure 27 b). First, the pCas9 plasmid was introduced into *E. coli* MG1655, and the resulting strain was co-transformed with the pCRISPR::rpsL plasmid and W542 (containing an edited oligonucleotide with the A-C mutation). Only streptomycin-resistant colonies were restored after transformation with the pCRISPR::rpsL plasmid, indicating that Cas9 cleavage induces recombination of the oligonucleotide. Figure 37However, the number of streptomycin-resistant colonies was two orders of magnitude lower than that of kanamycin-resistant colonies, presumably cells that escaped Cas9 cleavage. Therefore, under these conditions, Cas9 cleavage promotes mutation introduction, but is not efficient enough to select mutant cells from the 'escapee' background.
[0357] To improve the efficiency of genome editing in *E. coli*, the applicant used its CRISPR system for recombinant engineering, employing Cas9-induced cell death to select desired mutations. The pCas9 plasmid was introduced into the engineered strain HME63(31), which contains Gam, Exo, and β functions of a □-red phage. The resulting strain was then processed using the pCRISPR::rpsL plasmid (or a similar plasmid). (Control) and co-transformation with the W542 oligonucleotide ( Figure 27 a) The engineering efficiency is 5.3 × 10⁻⁶. -5 The fraction of total cells that became streptomycin resistant when the control plasmid was used was calculated. Figure 27 c). In contrast, transformation with this pCRISPR::rpsL plasmid increased the percentage of mutant cells to 65% ± 14%. Figure 27 c and 29f). The applicant observed that, after transformation with the pCRISPR::rpsL plasmid, the number of CFUs was reduced by approximately three orders of magnitude (4.8 × 10⁻⁶) compared to the control plasmid. 5 / 5.3×10 2 , Figure 38 a) This indicates screening for CRISPR-induced cell death caused by unedited cells. To measure the rate at which Cas9 cleavage was inactivated (an important parameter of the applicant's method), the applicant transformed cells with pCRISPR::rpsL or a control plasmid without the W542-edited oligonucleotide. Figure 38 a). Background of CRISPR's 'The Escapists' (measured as...) The ratio of CFU is 2.5 × 10⁻⁶. -4 (1.2×10 2 / 4.8×10 5 Genotyping of eight of these escapees revealed that, in all cases, the target spacer was absent. Figure 38 b). This background is higher than the recombination engineering efficiency of this rpsL mutation, 5.3 × 10⁻⁶. -5 This indicates that Cas9 cleavage must induce oligonucleotide recombination to obtain 65% edited cells. To confirm this, the applicant compared the results with those obtained using pCRISPR::rpsL or... The number of kanamycin and streptomycin resistant CFUs after conversion ( Figure 27d). Similar to the case with Streptococcus pneumoniae, the applicant observed moderate induction of recombinant activity, approximately 6.7-fold (2.0 × 10⁻⁶). -4 / 3.0×10 -5 Taken together, these results suggest that the CRISPR system provides a method for introducing selective mutations through recombination engineering.
[0358] The applicant discloses that the CRISPR-Cas system can be used to target edits in bacteria by co-introducing a target construct that kills wild-type cells and an editing template that eliminates both CRISPR cuts and introduces the desired mutation. Different types of mutations (insertions, deletions, or scarless single nucleotide substitutions) can be generated. Multiple mutations can be introduced simultaneously. The specificity and diversity of edits performed using this CRISPR system depend on several unique properties of the Cas9 endonuclease: (i) its target specificity can be programmed with a small RNA without enzyme engineering; (ii) the target specificity is very high, determined by interactions with a 20 bp RNA-DNA sequence that interacts with a low probability of non-target recognition; (iii) it can target almost any sequence, requiring only the presence of a neighboring NGG sequence; and (iv) almost any mutation in that NGG sequence, along with mutations in the seed sequence of the prototype spacer, eliminates the target.
[0359] The applicant demonstrates that genome engineering using this CRISPR system works not only in highly recombination-inducing bacteria such as Streptococcus pneumoniae, but also in Escherichia coli. Results in E. coli suggest that the method is applicable to other microorganisms where plasmids can be introduced. In E. coli, this pathway is complemented by recombination engineering of mutagenic oligonucleotides. To use this method in microorganisms where recombination engineering is not feasible, the host homologous recombination mechanism can be utilized by providing the editing template on a plasmid. Furthermore, given the growing evidence that CRISPR-mediated chromosome cutting leads to cell death in many bacteria and archaea, the use of an endogenous CRISPR-Cas system for editing purposes is foreseeable.
[0360] In both *Streptococcus pneumoniae* and *Escherichia coli*, the applicant observed that, despite co-selection of transformable cells and minor induction of recombination at the target site via Cas9 cleavage, the mechanism contributing most to editing was selection against unedited cells. Therefore, a major limitation of this method is the presence of a background of cells escaping CRISPR-induced cell death and lacking the desired mutant. The applicant showed that these 'escapees' primarily occur through the deletion of the target spacer, presumably after recombination of repetitive sequences flanking the target spacer. Further improvements could be made to the engineering of biogenic crRNAs that still support functional crRNAs but are sufficiently different from each other to eliminate flanking sequences of recombination. Alternatively, direct conversion of chimeric crRNAs could be explored. In the specific case of *E. coli*, the construction of this CRISPR-Cas system is not possible if this organism is also used as a cloning host. The applicant solved this problem by placing Cas9 and the tracrRNA in a plasmid different from the CRISPR array. The engineering of the inducible system also avoids this limitation.
[0361] While new DNA synthesis technologies offer the cost-effective ability to create any sequence with high throughput, integrating synthetic DNA into living cells to build a functional genome remains a challenge. Recently, the co-screening MAGE strategy has been shown to improve the mutagenic efficiency of recombination engineering by screening a subpopulation of cells with an increased likelihood of recombination at or around a given locus. In this approach, selectable mutations are introduced to increase the chance of generating nearby unselectable mutations. In contrast to the indirect screening provided by this strategy, the use of a CRISPR system enables the direct screening of desired mutations and their efficient recovery. These technologies, combined with the genetic engineer's toolbox and DNA synthesis, could substantially advance the ability to both decipher gene function and manipulate organisms for biotechnological purposes. Two other studies have also addressed CRISPR-related engineering of the mammalian genome. These crRNA-guided genome editing technologies are expected to have broad applications in basic and medical sciences.
[0362] Strains and Culture Conditions. Streptococcus pneumoniae strain R6 was provided by Dr. Alexander Tomasz. Strain crR6 appeared in the aforementioned study. Liquid cultures of Streptococcus pneumoniae were grown on THYE medium (30 g / L Todd-Hewitt agar, 5 g / L yeast extract). Cells were plated on tryptone-soybean agar (TSA) supplemented with 5% defibrinated sheep blood. Antibiotics were added as appropriate: kanamycin (400 μg / ml), chloramphenicol (5 μg / ml), erythromycin (1 μg / ml), streptomycin (100 μg / ml), or spectinomycin (100 μg / ml). β-galactosidase activity was measured using the Miller assay as previously described.
[0363] Escherichia coli strains MG1655 and HME63 (derived from MG1655, Δ(argF-lac)U169λcI857Δcro-bioA galKtyr 145UAG mutS<>amp) were obtained from Jeff Roberts and Donald Court, respectively (31). Liquid cultures of E. coli were grown in LB broth (Difco). Antibiotics were added as appropriate: chloramphenicol (25 μg / ml), kanamycin (25 μg / ml), and streptomycin (50 μg / ml).
[0364] Streptococcus pneumoniae transformation. Competent cells were prepared as described above (23). For all genome editing transformations, cells were slowly thawed on ice and resuspended in 10 volumes of M2 medium (supplemented with 100 ng / ml of the competent stimulating peptide CSP1 (40)) and then the editing construct was added (the editing construct was added to the cells at a final concentration between 0.7 ng / μl and 2.5 μg / μl). Cells were incubated at 37°C for 20 min before adding 2 μl of the targeting construct, and then at 37°C for 40 min. Serially diluted cells were plated on suitable medium to determine colony-forming unit (CFU) counts.
[0365] Recombinant engineering of *E. coli* λ-red. Strain HME63 was used for all recombinant engineering experiments. Recombinant cells were prepared and treated according to the previously disclosed protocol (6). Briefly, 2 ml of overnight culture (LB medium) inoculated from a single colony obtained from the plate was incubated at 30°C. The overnight culture was diluted 100-fold and incubated at 30°C with shaking (200 rpm) until the OD value was reached. 600The induction time is from 0.4-0.5 (approximately 3 hours). For λ-red induction, transfer the culture to a 42°C water bath and shake at 200 rpm for 15 min. Immediately after induction, vortex the culture in an ice-water slurry and freeze on ice for 5-10 min. Then wash the cells and aliquot them according to the protocol. For electroporation, mix 50 μl of cells with 1 mM of salt-free oligonucleotides (IDT) or 100-150 ng of plasmid DNA (prepared using the QIAprep Spin Miniprep kit, Qiagen). Electroporate the cells using a 1 mm Gene Pulser cuvette (Bio-rad) at 1.8 kV and immediately resuspend them in 1 ml of room temperature LB medium. Before plating onto LB agar with suitable antibiotic resistance, recover the cells at 30°C for 1-2 hours and incubate overnight at 32°C.
[0366] Preparation of Streptococcus pneumoniae genomic DNA. For transformation purposes, Streptococcus pneumoniae genomic DNA was extracted using the Wizard Genomic DNA Purification Kit, following the instructions provided by the manufacturer (Promega). For genotyping purposes, 700 μL of overnight Streptococcus pneumoniae culture was precipitated, resuspended in 60 μL of lysozyme solution (2 mg / mL), and incubated at 37°C for 30 min. Genomic DNA was extracted using the QIAprep Spin Miniprep Kit (Qiagen).
[0367] Strain Construction. All primers used in this study are provided in Table G. To generate *Streptococcus pneumoniae* crR6M, an intermediate strain, LAM226, was prepared. In this strain, the aphA-3 gene (providing kanamycin resistance) adjacent to the CRISPR array of the *Streptococcus pneumoniae* crR6 strain was replaced by a cat gene (providing chloramphenicol resistance). Briefly, the crR6 genomic DNA was amplified using primers L448 / L444 and L447 / L481, respectively. The cat gene was amplified from plasmid pC194 using primers L445 / L446. Each PCR product was gel purified and all three were fused using primers L448 / L481 via splicing PCR. The resulting PCR products were transformed into competent *Streptococcus pneumoniae* crR6 cells, and chloramphenicol-resistant transformants were screened. To generate *Streptococcus pneumoniae* crR6M, *Streptococcus pneumoniae* crR6 genomic DNA was amplified by PCR using primers L409 / L488 and L448 / L481, respectively. Each PCR product was gel-purified and fused using primer L409 / L481 via splicing PCR. The resulting PCR products were transformed into competent *Streptococcus pneumoniae* LAM226 cells, and kanamycin-resistant transformants were screened.
[0368] To generate *Streptococcus pneumoniae* crR6Rc, *Streptococcus pneumoniae* crR6M genomic DNA was amplified by PCR using primers L430 / W286, and *Streptococcus pneumoniae* LAM226 genomic DNA was amplified by PCR using primers W288 / L481. Each PCR product was gel purified and fused using primers L430 / L481 via splicing PCR. The resulting PCR products were transformed into competent *Streptococcus pneumoniae* crR6M cells, and chloramphenicol-resistant transformants were screened.
[0369] To generate *Streptococcus pneumoniae* crR6Rk, *Streptococcus pneumoniae* crR6M genomic DNA was amplified by PCR using primers L430 / W286 and W287 / L481, respectively. Each PCR product was gel-purified and fused using primer L430 / L481 via splicing PCR. The resulting PCR products were transformed into competent *Streptococcus pneumoniae* crR6Rc cells, and kanamycin-resistant transformants were screened.
[0370] To generate JEN37, *Streptococcus pneumoniae* crR6Rk genomic DNA was amplified by PCR using primers L430 / W356 and W357 / L481, respectively. Each PCR product was gel-purified and fused using primers L430 / L481 via splicing PCR. The resulting PCR products were transformed into competent *Streptococcus pneumoniae* crR6Rc cells, and kanamycin-resistant transformants were screened.
[0371] To generate JEN38, R6 genomic DNA was amplified using primers L422 / L461 and L459 / L426, respectively. 43 The ermAM gene (specifying erythromycin resistance) was amplified from plasmid pFW15 using primers L457 / L458. Each PCR product was gel-purified and all three were fused using primers L422 / L426 via splicing PCR. The resulting PCR products were transformed into competent Streptococcus pneumoniae crR6Rc cells, and erythromycin-resistant transformants were screened.
[0372] Streptococcus pneumoniae JEN53 is produced in two steps. First, as in... Figure 33 As shown, JEN43 was constructed. JEN53 was produced by transforming the genomic DNA of JEN25 into competent JEN43 cells and screening for both chloramphenicol and erythromycin.
[0373] To generate *Streptococcus pneumoniae* JEN62, *Streptococcus pneumoniae* crR6Rk genomic DNA was amplified by PCR using primers W256 / W365 and W366 / L403, respectively. Each PCR product was purified and ligated using Gibson splicing. The spliced products were transformed into competent *Streptococcus pneumoniae* crR6Rc cells, and kanamycin-resistant transformants were screened.
[0374] Plasmid construction. pDB97 was constructed by phosphorylation and annealing of oligonucleotides B296 / B297, followed by ligation into pLZ12spec digested with EcoRI / BamHI. The applicant fully sequenced pLZ12spec and deposited its sequence in a gene bank (accession number: KC112384).
[0375] pDB98 was obtained after cloning this CRISPR, with the leader sequence, along with a repeat-spacer-repeat unit, cloned into pLZ12spec. This was achieved by amplifying crR6Rc DNA using primers B298 / B320 and B299 / B321, followed by gene splicing PCR of the two products and cloning them into pLZ12spec with the restriction enzyme sites BamHI / EcoRI. In this way, the spacer sequence in pDB98 was engineered to contain two BsaI restriction enzyme sites in opposite directions, which allows for scarless cloning of new spacers.
[0376] pDB99 to pDB108 were constructed by annealing oligonucleotides B300 / B301 (pDB99), B302 / B303 (pDB100), B304 / B305 (pDB101), B306 / B307 (pDB102), B308 / B309 (pDB103), B310 / B311 (pDB104), B312 / B313 (pDB105), B314 / B315 (pDB106), B315 / B317 (pDB107), and B318 / B319 (pDB108), followed by BsaI linkage at the pDB98 cleavage.
[0377] The pCas9 plasmid was constructed as follows: The necessary CRISPR elements were amplified from the genomic DNA of *Streptococcus pyogenes* SF370, which possesses flanking homologous arms for Gibson assembly. The tracrRNA and Cas9 were amplified using oligonucleotides HC008 and HC010. The leader and CRISPR sequence were amplified using HC011 / HC014 and HC015 / HC009 to introduce two BsaI-type IIS sites between two homologous repeats to facilitate the insertion of the spacer.
[0378] pCRISPR was constructed by amplifying a subcloned pCas9CRISPR array in pZE21-MCS1 using oligomers B298+B299 and confining it with EcoRI and BamHI. The rpsL targeting spacer was cloned by annealing oligomers B352+B353 and cloning pCRISPR at the BsaI cut, resulting in pCRISPR::rpsL.
[0379] Generation of targeting and editing constructs. Targeting constructs for genome editing were prepared by Gibson splicing of left and right PCRs (Table G). Editing constructs, when applicable, were prepared by gene splicing PCR fusion of PCR product A (PCRA), PCR product B (PCR B), and PCR product C (PCR C) (Table G). The CRISPR::ermAM (terminator) targeting construct was generated by PCR amplification of JEN62 and crR6 genomic DNA (containing oligomers L409 and L481), respectively.
[0380] Generation of targets with randomized PAM or prototype spacers. The five nucleotides following the spacer target are amplified using primers W377 / L426 to amplify R6. 8232.5 Genomic DNA was randomized. This PCR product was then spliced with the cat gene and the upstream region of srtA amplified from the same template using primers L422 / W376. 80 ng of the spliced DNA was used to transform strains R6 and crR6. Samples used for this randomization target were prepared using the following primers: B280-B290 / L426 to randomize bases 1-10 of the target and B269-B278 / L426 to randomize bases 10-20. Primers L422 / B268 and L422 / B279 were used to amplify the cat gene and the upstream region of srtA, respectively, to be spliced with the initial and final 10 PCR products. These splice constructs were pooled and transformed into R6 and crR6 cells with 30 ng each. After transformation, cells were plated on chloramphenicol selection plates. For each sample, more than 2 × 10⁶ cells were selected. 5 Cells were pooled together in 1 ml of THYE, and genomic DNA was extracted using a Promega Wizard kit. Primers B250 / B251 were used to amplify the target region. PCR products were labeled and run on an Illumina MiSeq dual-end lane gel with 300 cycles.
[0381] Deep sequencing data analysis
[0382] Randomized PAM: For the randomized PAM assay, 3,429,406 reads were obtained for crR6 and 3,253,998 reads were obtained for R6. It was expected that only half of these would correspond to the PAM target, while the other half would be sequenced at the other end of the PCR product. 1,623,008 crR6 reads and 1,537,131 R6 reads contained an error-free target sequence. The occurrence of each possible PAM in these reads is shown in the supplementary document. To assess the functionality of the PAM, the relative proportions in the crR6 sample and R6 sample were calculated and expressed as r. ijklm Where I, j, k, l, and m are one of these four possible bases. Construct the following statistical model:
[0383] log(r ijklm )=μ+b2 i +b3 j +b4k +b2b3 i,j +b3b4 j,k +ε ijklm ,
[0384] Where ε is the residual, b2 is the effect of the second base of PAM, b3 is the effect of the third base, b4 is the effect of the fourth base, b2b3 is the interaction between the second and third bases, and b3b4 is the interaction between the third and fourth bases. Perform an analysis of variance:
[0385] Anova table
[0386]
[0387] When added to this model, b1 or b5 appear to be insignificant, and other interactions besides those included can be discarded. Model selection was performed more or less using the ANOVA method in R through successive comparisons of the complete model. Tukey's honest significance test was used to determine whether the differences between pairs of effects were significant.
[0388] The NGGNN type is significantly different from all other types and has the strongest influence (see table below).
[0389] To demonstrate that positions 1, 4, or 5 do not affect the NGGNN type, the applicant focused only on these sequences. Their effects appear to be normally distributed (see...). Figure 71 The model comparison using QQ plotting in R and the anova method in R shows that the null model is the best, i.e., there is no significant effect of b1, b4 and b5.
[0390] Model comparison using the ANOVA method in R for this NGGNN sequence
[0391]
[0392] NAGNN and NNGGN type partial interference
[0393] The NAGNN type is significantly different from all other types but has a much smaller effect than the NGGNN type (see Tukey significance test below).
[0394] Ultimately, the NTGGN and NCGGN types are similar and show significantly more CRISPR interference than the NTGHN and NCGHN types (where H is A, T, or C), as shown by the paired student test regulated by bonferroni.
[0395] The effect of b4 on the NYGNN sequence was compared using a t-test in pooled SD pairwise.
[0396]
[0397] Taken together, these results allow for the inference that NNGGN typically produces complete interference in the case of NGGGN, or partial interference in the case of NAGGN, NTGGN, or NCGGN.
[0398] Tukey multiple comparisons of the mean: 95% family-wise confidence level
[0399]
[0400] Randomized target
[0401] For randomized target assays, 540,726 reads were obtained for crR6 and 753,570 reads were obtained for R6. As previously mentioned, only half of these reads were expected to be sequenced at the ends of interest for the PCR product. After filtering, 217,656 and 353,141 reads carrying a single target with no error or a single point mutation were retained for crR6 and R6, respectively. The relative proportion of each mutant in the crR6 sample to that in the R6 sample was calculated. Figure 24 c). All mutations outside the seed sequence (13-20 bases from the PAM) showed complete interference. Those sequences were used as a reference to determine whether other mutations within the seed sequence could be considered significantly disruptive interference. These sequences were fitted with a normal distribution using the `fitdistr` function of the MASS R package. Figure 24 In c, the 0.99 quantile of the fitted distribution is shown as a dashed line. Figure 72 A histogram showing the data density with a fitted normal distribution (black line) and the .99 quantile (dashed line).
[0402] Table F. Relative abundance of PAM sequences in the crR6 / R6 sample (averaged by bases 1 and 5).
[0403]
[0404] Table G. Primers used in this study.
[0405]
[0406]
[0407]
[0408] Table H. Design of the targets and edit constructs used in this study
[0409]
[0410]
[0411] Example 6: Optimization of guide RNA for Streptococcus pyogenes Cas9 (referred to as SpCas9).
[0412] The applicant mutates the tracrRNA and the same-direction repeat sequence, or mutates the chimeric guide RNA to enhance RNA in the cell.
[0413] This optimization is based on the observation that thymine (T) elongation exists in both the tracrRNA and the guide RNA, which can lead to early transcriptional termination via the pol 3 promoter. Therefore, the applicant generated the following optimized sequence. The optimized tracrRNA and its corresponding optimized unidirectional repeat are shown in pairs.
[0414] Optimized tracrRNA 1 (mutation underlined):
[0415] GGAACCATTCA t AACAGCATAGCAAGTTA t AATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTT
[0416] Optimized same-direction repeat 1 (mutation underlined):
[0417] GTT a TAGAGCTATGCTGTT a TGAATGGTCCCAAAAC
[0418] Optimized tracrRNA 2 (mutant underlined):
[0419] GGAACCATTCAA t ACAGCATAGCAAGTTAA t ATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTT
[0420] Optimized same-direction repeat 2 (mutation underlined):
[0421] GT a TTAGAGCTATGCTGT a TTGAATGGTCCCAAAAC
[0422] To achieve optimal activity in eukaryotic cells, the applicant also optimized the chimeric guide RNA.
[0423] Original guide RNA:
[0424] NNNNNNNNNNNNNNNNNNGTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT
[0425] Optimized chimeric guide RNA sequence 1:
[0426] NNNNNNNNNNNNNNNNNNNNGT A TTAGAGCTAGAAATAGCAAGTTAA T ATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT
[0427] Optimized chimeric guide RNA sequence 2:
[0428] NNNNNNNNNNNNNNNNNNGTTTTAGAGCTATGCTGTTTTGGAAACAAAACAGCATGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT
[0429] Optimized chimeric guide RNA sequence 3:
[0430] NNNNNNNNNNNNNNNNNNNNGT A TTAGAGCTATGCTGT A TTGGAAACAA T ACAGCATAGCAAGTTAA T ATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT
[0431] The applicant revealed that optimized chimeric guide RNAs function better, such as in Figure 3 As indicated in the study, this experiment was performed by co-transfecting 293FT cells with Cas9 and a U6-guide RNA DNA cassette to express one of the four RNA forms shown above. The target of this guide RNA is the same site in the human EMX1 locus: “GTCACCTCCAATGACTAGGG”.
[0432] Example 7: Optimization of Streptococcus thermophilus LMD-9CRISPR1 Cas9 (referred to as St1Cas9).
[0433] The applicant designed the guided chimeric RNA shown in Figure 4.
[0434] By extending through the break of polythymine (T), the St1Cas9 guide RNA can undergo the same type of optimization as that performed on the SpCas9 guide RNA.
[0435] Example 8: Cas9 Diversity and Mutations
[0436] The CRISPR-Cas system is an acquired immune mechanism against invasive foreign DNA, utilized by diverse species across bacteria and archaea. The type II CRISPR-Cas9 system consists of a set of genes encoding proteins responsible for "acquiring" foreign DNA into the CRISPR locus and a set of proteins encoding the DNA cleavage mechanism; these include a DNA nuclease (Cas9), a non-coding trans-activating cr-RNA (tracrRNA), and an array of foreign DNA-derived spacers (crRNA) flanked by identically repeating sequences. Upon maturation by Cas9, the tracrRNA and crRNA duplexes guide the Cas9 nuclease to the target DNA sequence defined by the spacer guide sequence, and mediate a double-strand break in the DNA near a short sequence motif required for cleavage and specific to each CRISPR-Cas system. Type II CRISPR-Cas systems have been found throughout the bacterial kingdom and vary greatly in terms of the Cas9 protein sequence and size used for target cleavage, the tracrRNA and crRNA repetitive sequences, the genomic organization of these elements, and motif requirements. A single species can possess multiple distinct CRISPR-Cas systems.
[0437] The applicant evaluated 207 putative Cas9s from bacterial species identified based on sequence homology with known Cas9s and orthogonal homology with known subdomains, including the HNH endonuclease domain and the RuvC endonuclease domain [information from Eugene Koonin and Kira Makarova]. Phylogenetic analysis based on this set of protein sequence conserved sequences revealed five Cas9 families, including three groups of large Cas9s (approximately 1400 amino acids) and two groups of small Cas9s (approximately 1100 amino acids) (see Figures 39 and 40A-F).
[0438] In this example, the applicant demonstrated that the following mutations can convert SpCas9 into a nickase: D10A, E762A, H840A, N854A, N863A, D986A.
[0439] The applicant provided sequences showing where the sequence mutation sites are located within the SpCas9 gene (Figure 41). The applicant also revealed that these nickases are still able to mediate homologous recombination (as indicated in Figure 2). Furthermore, the applicant revealed that SpCas9 with these mutations (alone) does not induce double-strand breaks. Figure 47 ).
[0440] Example 9: Supplementing DNA targeting specificity of RNA-guided Cas9 nucleases
[0441] Cell culture and transfection
[0442] Human embryonic kidney (HEK) cell line 293FT (Lifetech Corporation) was maintained at 37°C with 5% CO2 in Dulbecco's modified Eagle's Medium (DMEM) supplemented with 10% fetal bovine serum (HyClone), 2 mM GlutaMAX (Lifetech Corporation), 100 U / mL penicillin, and 100 μg / mL streptomycin.
[0443] Seed 293FT cells into 6-well, 24-well, or 96-well plates (Corning) 24 hours prior to transfection. Transfect cells at 80%–90% confluence using Lipofectamine 2000 (Life Technologies) following the manufacturer’s recommended protocol. For each well of a 6-well plate, use a total of 1 µg of Cas9+sgRNA plasmid. For each well of a 24-well plate, use a total of 500 ng of Cas9+sgRNA plasmid unless otherwise specified. For each well of a 96-well plate, use 65 ng of Cas9 plasmid at a 1:1 molar ratio to the U6-sgRNA PCR product.
[0444] Human embryonic stem cell line HUES9 (Harvard Stem Cell Institutecore) was maintained in feeder-free conditions on GelTrex (Life Technologies) medium supplemented with 100 μg / ml Normocin (InvivoGen). HUES9 cells were transfected using the Amaxa P3 Primary Cell 4-D Nuclear Transfection Kit (Lonza) according to the manufacturer's protocol.
[0445] SURVEYOR nuclease assay for genome modification
[0446] As described above, 293FT cells were transfected with plasmid DNA. Following transfection, the cells were incubated at 37°C for 72 hours before genomic DNA extraction. Genomic DNA was extracted using QuickExtract DNA extraction solution (Epicentre) following the manufacturer's protocol. In short, the pelleted cells were resuspended in QuickExtract solution and incubated at 65°C for 15 minutes and then at 98°C for 10 minutes.
[0447] For each gene, PCR amplification (primers listed in Tables J and K) was performed on genomic regions flanking the CRISPR target site, and the products were purified using a QiaQuick Spin column (Qiagen) following the manufacturer's protocol. A total of 400 ng of purified PCR product was mixed with 2 μl of 10X Taq DNA polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 20 μl and subjected to a re-annealing process to allow for heteroduplex formation: 95 °C for 10 min, decreasing from 95 °C to 85 °C at -2 °C / s, decreasing from 85 °C to 25 °C at -0.25 °C / s, and holding at 25 °C for 1 min. After re-annealing, the products were treated with SURVEYOR nuclease and SURVEYOR enhancer S (Transgenomics) following the manufacturer's recommended protocol and analyzed on 4%–20% Novex TBE polyacrylamide gels (Life Technologies). The gels were stained with SYBR Gold DNA staining agent (Lifetech Corporation) for 30 minutes and imaged using a Gel Doc gel imaging system (Bio-rad Corporation). Quantification was performed based on relative band intensity.
[0448] Northern blot analysis of tracrRNA expression in human cells
[0449] Northern blotting was performed as described above. In short, RNA was heated to 95°C for 5 min before loading onto an 8% denaturing polyacrylamide gel (SequaGel, National Diagnostics). Afterward, the RNA was transferred to a pre-hybridized Hybond N+ membrane (GE Healthcare) and cross-linked with Stratagene UV cross-linking agent (Stratagene). The probe was labeled with [γ-32P]ATP (PerkinElmer) containing T4 polynucleotide kinase (New England Biolabs). After washing, the membrane was exposed to a fluorescent screen for one hour and scanned using a Typhoon imager.
[0450] Bisulfite sequencing to assess DNA methylation status
[0451] As described above, HEK 293FT cells were transfected with Cas9. Genomic DNA was isolated using the DNeasy Blood & Tissue Kit (Qiagen) and bisulfite-transformed using the EZ DNA Methylation-Lightning Kit (Zymo Research). Bisulfite PCR was performed using KAPA2G Robust HotStart DNA polymerase (KAPA Biosystems), with primers designed using Bisulfite PrimerSeeker (Zymo Research, Tables J and K). The resulting PCR amplicon was gel purified, digested with EcoRI and HindIII, and ligated into a pUC19 backbone prior to transformation. Individual clones were then subjected to Sanger sequencing to assess DNA methylation status.
[0452] In vitro transcription and cleavage assay
[0453] As described above, HEK 293FT cells were transfected with Cas9. Whole-cell lysates were then prepared using a lysis buffer (20 mM HEPES, 100 mM KCl, 5 mM MgCl2, 1 mM DTT, 5% glycerol, 0.1% Triton X-100) supplemented with a protease inhibitor mixture (Roche). T7-driven sgRNA was transcribed in vitro using a custom oligomer (Example 10) and the HiScribe T7 In Vitro Transcription Kit (NEB), following the manufacturer's recommended protocol. To prepare methylation target sites, the pUC19 plasmid was methylated with M.SssI and then linearized with NheI. The in vitro cleavage assay was performed as follows: For a 20 μL cleavage reaction, 10 μL of cell lysate was incubated with 2 μL of cleavage buffer (100 mM HEPES, 500 mM KCl, 25 mM MgCl2, 5 mM DTT, 25% glycerol), in vitro transcribed RNA, and 300 ng pUC19 plasmid DNA.
[0454] Deep sequencing to assess target specificity
[0455] Before genomic DNA extraction, HEK 293FT cells plated in 96-well plates were transfected with Cas9 plasmid DNA and a single guide RNA (sgRNA) PCR kit for 72 hours. Figure 72 A fusion PCR method was used to attach the Yimingda P5 aptamer along with a unique sample-specific barcode to the target amplicon, amplifying genomic regions flanking CRISPR target sites for each gene. Figure 74 (Example 10) Figure 73 (Illustrative description). Following the manufacturer's recommended protocol, PCR products were purified using an EconoSpin 96-well plate (Epoch Life Sciences).
[0456] Barcoded and purified DNA samples were quantified and pooled at equimolar ratios using a Quant-iT PicoGreen dsDNA assay kit or a Qubit 2.0 spectrophotometer (Lifetech Corporation). The sequencing libraries were then deep sequenced using a MiSeq personal sequencer (Lifetech Corporation).
[0457] Sequencing data analysis and indel detection
[0458] MiSeq reads were filtered by an average Phred quality (Q score) of at least 23, along with perfect sequence matching to the barcode and the amplicon forward primer. Reads from on-target and off-target loci were analyzed by first performing Smith-Waterman alignments on the amplicon sequences, including the target site (120 bp in total), up to and down 50 nucleotides. Simultaneously, alignments were performed on the indel from the target site (30 bp in total), from 5 nucleotides upstream to 5 nucleotides downstream. Target regions were discarded if the aligned portion fell outside the MiSeq read itself, or if the matched base pairs comprised less than 85% of its total length.
[0459] A metric is provided for the negative control for each sample, treating the inclusion or exclusion of indels as hypothetical cut events. For each sample, an indel is counted only if its quality score exceeds μ-σ, where μ is the mean quality score of the negative control corresponding to that sample, and σ is its standard deviation. This yields the overall target region indel rate for both the negative control and its corresponding sample. The per-target-region-per-read error rate, q, for the negative control, the observed indel count n for that sample, and its read count R (i.e., the maximum likelihood estimate of the read score of the target region with true indels), p, are derived by applying a binomial error model, as follows.
[0460] Let E be the (unknown) number of readings in a sample with a target region that is incorrectly counted as having at least one indel, which we can write as (without making any assumptions about the true number of indels).
[0461]
[0462] Because R(1-p) is the number of readings in the target region with no true indel. Meanwhile, since the observed number of readings with indel is n, n = E + Rp, in other words, the number of readings in the target region with errors but no true indel plus the number of readings in the target region that correctly have indels. We can then rewrite the above equation as follows:
[0463]
[0464] All values of the frequency of the target region with true indelp are made equally likely as the prior, Prob(n|p)∝Prob(p|n). Therefore, the maximum likelihood estimate (MLE) of the frequency of the target region with true indelp is set to maximize the value of p of Prob(n|p). This is evaluated numerically.
[0465] To set error limits for the true indel read frequencies within the sequencing library itself, the Wilson score interval (2) is calculated for each sample, taking into account the MLE estimate for the true indel target region Rp and the number of reads R. Specifically, the lower limit l and the upper limit u are calculated as follows:
[0466]
[0467]
[0468] The standard score required for confidence in a normal distribution with variance 1 is set to 1.96, which means a 95% confidence level.
[0469] qRT-PCR analysis of Cas9 and sgRNA expression
[0470] 293FT cells plated in 24-well plates were transfected as described above. 72 hours post-transfection, total RNA was harvested using the miRNeasy microreactor kit (Qiagen). Reverse strand synthesis of sgRNA was performed using the qScript Flex cDNA kit (VWR) and custom first-strand synthesis primers (Tables J and K). qPCR analysis was performed using Fast SYBR Green Master Mix (Lifetechnologies) and custom primers (Tables J and K), with GAPDH as an endogenous control. Relative quantification was calculated using the ΔΔCT method.
[0471] Table I | Target Site Sequences. Target sites were tested using the Streptococcus pyogenes type II CRISPR system with the essential PAM. For each target, cells were transfected with Cas9 and either crRNA-tracrRNA or chimeric sgRNA.
[0472]
[0473]
[0474] Table J | Primer Sequences
[0475] SURVEYOR measurement
[0476] Sp-EMX1-F1 EMX1 AAAACCACCCTTCTCTCTGGC Sp-EMX1-R1 EMX1 GGAGATTGGAGACACGGAGAG Sp-EMX1-F2 EMX1 CCATCCCCTTCTGTGAATGT Sp-EMX1-R2 EMX1 GGAGATTGGAGACACGGAGA Sp-PVALB-F PVALB CTGGAAAGCCAATGCCTGAC Sp-PVALB-R PVALB GGCAGCAAACTCCTTGTCCT
[0477] qRT-PCR for Cas9 and sgRNA expression
[0478]
[0479] Bisulfite PCR and sequencing
[0480]
[0481] Table K | Sequences of primers used to test sgRNA constructs. Primers hybridize to the reverse strand of the U6 promoter unless otherwise indicated. The U6 initiation site is italicized, the guiding sequence is represented as an N, the directing repeat sequence is bolded, and the tracrRNA sequence is underlined. The secondary structure of each sgRNA construct is shown in Figure 43.
[0482]
[0483] Table L | Target sites with alternating PAMs used to test PAM specificity of Cas9. All target sites used for PAM specificity testing were found to be located in the human EMX1 locus.
[0484]
[0485]
[0486] Example 10: Complementary Sequence
[0487] All sequences are oriented between 5' and 3'. For U6 transcription, the underlined string of Ts acts as a transcription terminator.
[0488] U6-short tracrRNA (Streptococcus pyogenes SF370)
[0489]
[0490] (tracrRNA sequence is bolded)
[0491] U6-DR-Guide Sequence-DR (Streptococcus pyogenes SF370)
[0492]
[0493] (The repeating sequence in the same direction is highlighted in gray, and the guiding sequence is a bold N.)
[0494] >sgRNA, containing +48tracrRNA (Streptococcus pyogenes SF370)
[0495]
[0496] (The guide sequence is a bolded N, and the tracrRNA fragment is also bolded.)
[0497] >sgRNA, containing +54tracrRNA (Streptococcus pyogenes SF370)
[0498]
[0499] (The guide sequence is a bolded N, and the tracrRNA fragment is also bolded.)
[0500] >sgRNA, containing +67tracrRNA (Streptococcus pyogenes SF370)
[0501]
[0502] (The guide sequence is a bolded N, and the tracrRNA fragment is also bolded.)
[0503] >sgRNA, containing +85tracrRNA (Streptococcus pyogenes SF370)
[0504]
[0505] (The guide sequence is a bolded N, and the tracrRNA fragment is also bolded.)
[0506] >CBh-NLS-SpCas9-NLS
[0507]
[0508]
[0509]
[0510] (NLS-hSpCas9-NLS is highlighted in bold)
[0511] Sequencing amplicon for EMX1 guidance in versions 1.1, 1.14, and 1.17.
[0512] CCAATGGGGAGGACATCGATGTCACCTCCAATGACTAGGGTGGGCAACCACAAACCCACGAGGGCAGAGTGCTGCTTGCTGCTGGCCAGGCCCCTGCGTGGGCCCAAGCTGGACTCTGGCCAC
[0513] Sequencing amplicon for EMX1 guidance in steps 1.2 and 1.16
[0514] CGAGCAGAAGAAGAAGGGCTCCCATCACATCAACCGGTGGCGCATTGCCACGAAGCAGGCCAATGGGGAGGACATCGATGTCACCTCCAATGACTAGGGTGGGCAACCACAAAACCCACGAG
[0515] Sequencing amplicon for EMX1 guidance in versions 1.3, 1.13, and 1.15.
[0516] GGAGGACAAAGTACAAACGGCAGAAGCTGGAGGAGGAAGGGCCTGAGTCCGAGCAGAAGAAGAAGGGCTCCCATCACATCAACCGGTGGCGCATTGCCACGAAGCAGGCCAATGGGGAGGACATCGAT
[0517] Sequencing amplicon for EMX1 guidance 1.6
[0518] AGAAGCTGGAGGAGGAAGGGCCTGAGTCCGAGCAGAAGAAGAAGGGCTCCCATCACATCAACCGGTGGCGCATTGCCACGAAGCAGGCCAATGGGGAGGACATCGATGTCACCTCCAATGACTAGGGTGG
[0519] Sequencing amplicon for EMX1 guidance 1.10
[0520] CCTCAGTCTTCCCATCAGGCTCTCAGCTCAGCCTGAGTGTTGAGGCCCCAGTGGCTGCTCTGGGGGCCTCCTGAGTTTCTCATCTGTGCCCCTCCCTCCCTGGCCCAGGTGAAGGTGTGGTTCCA
[0521] Sequencing amplicon for EMX1 guidance in steps 1.11 and 1.12
[0522] TCATCTGTGCCCCTCCCTCCCTGGCCCAGGTGAAGGTGTGGTTTCCAGAACCGGAGGACAAAGTACAAACGGCAGAAGCTGGAGGAGGAAGGGCCTGAGTCCGAGCAGAAGAAGAAGGGCTCCCATCACA
[0523] Sequencing amplicon for EMX1 guidance in steps 1.18 and 1.19
[0524] CTCCAATGACTAGGGTGGGCAACCACAAACCCACGAGGGCAGAGTGCTGCTTGCTGCTGGCCAGGCCCCTGCGTGGGCCCAAGCTGGACTCTGGCCACTCCCTGGCCAGGCTTTGGGGGAGGCCTGGAGT
[0525] Sequencing amplicon for EMX1 guidance 1.20
[0526] CTGCTTGCTGCTGGCCAGGCCCCTGCGTGGGCCCAAGCTGGACTCTGGCCACTCCCTGGCCAGGCTTTGGGGGAGGCCTGGAGTCATGGCCCCACAGGGCTTGAAGCCCGGGGCCGCCATTGACAGAG
[0527] T7 promoter F primer, used for target strand annealing
[0528] GAAATTAATACGACTCACTATAGGG
[0529] Oligomers containing pUC19 target site 1 for methylation (T7 reverse)
[0530] AAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAACAACGACGAGCGTGACACCACCCTATAGTGAGTCGTATTAATTTC
[0531] Oligomers containing pUC19 target site 2 for methylation (T7 reverse).
[0532] AAAAAAGCACCGACTCGGTGCCACTTTTTCAAGTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAACGCAACAATTAATAGACTGGACCTATAGTGAGTCGTATTAATTTC
[0533] Example 11: Oligomer-mediated Cas9-induced homologous recombination
[0534] Oligomeric homologous recombination assays compare the efficiency of different Cas9 variants and different HR templates (oligomers against plasmids).
[0535] 293FT cells were used. SpCas9 = wild-type Cas9 and SpCas9n = nickase Cas9 (D10A). The chimeric RNA target was the same EMX1 prototype spacer target 1 as in Examples 5, 9, and 10, and the oligomer was purified by IDT synthesis using PAGE.
[0536] Figure 44 illustrates one design of the oligoDNA used as a template for homologous recombination (HR) in this experiment. The long oligomer contained 100 bp homology to the EMX1 locus and the HindIII restriction site. 293FT cells were co-transfected with: first, a plasmid containing a chimeric RNA targeting the human EMX1 locus and wild-type cas9 protein; and second, the oligoDNA serving as the HR template. Samples were obtained from 293FT cells collected 96 hours after transfection with Lipofectamine 2000. All products were amplified using an EMX1 HR primer, purified by gel electrophoresis, and subsequently digested with HindIII to assess the efficiency of HR template integration into the human genome.
[0537] Figure 45 Figure 46 depicts a comparison of HR efficiencies induced by different combinations of Cas9 protein and HR templates. The Cas9 constructs used were wild-type Cas9 or the nicking enzyme version of Cas9 (Cas9n). The HR templates used were: antisense oligoDNA (antisense oligomers in the above figure), sense oligoDNA (sense oligomers in the above figure), or plasmid HR templates (HR templates in the above figure). The sense / antisense definition is that the strand of efficient transcription with a sequence corresponding to the transcribed mRNA is defined as the sense strand of the genome. HR efficiencies are shown as the percentage (base) of HindIII digestion bands against all genomic PCR amplification products.
[0538] Example 12: Autistic Mice
[0539] Recent large-scale sequencing projects have yielded a wealth of disease-associated genes. The discovery of a gene is only the beginning of understanding what it does and how it leads to pathological phenotypes. Current technologies and approaches for studying candidate genes are slow and laborious. The gold standard for gene targeting and genetic knockout requires significant investment in both financial resources and research personnel, as well as time and resources. The applicant has set out to utilize the hSpCas9 nuclease to target numerous genes and to do so with greater efficiency and lower turnaround time compared to any other technology. Because of the high efficiency of hSpCas9, the applicant can inject RNA into mouse zygotes and immediately obtain genome-modified animals without any initial gene targeting in mESC.
[0540] Cromote domain helicase DNA-binding protein 8 (CHD8) is a key gene involved in early spinal development and morphogenesis. Mice lacking CHD8 die during embryonic development. Mutations in this CHD8 gene have been associated with autism spectrum disorder in humans. This association was simultaneously published in three different articles in *Nature*. The same three studies identified a wide variety of genes associated with autism spectrum disorder. The applicant's goal is to create gene knockout mice targeting four genes identified in all the articles: Chd8, Katnal2, Kctd13, and Scn2a. In addition, the applicant selected two other genes associated with autism spectrum disorder, schizophrenia, and ADHD: GIT1, CACNA1C, and CACNB2. Finally, as a positive control, the applicant decided to target MeCP2.
[0541] For each gene, the applicant designed three gRNAs that are highly likely to knock out the gene. Knockout following the hSpCas9 nuclease induces a double-strand break, and the error-prone DNA repair pathway, through non-homologous end joining, corrects the break, establishing a mutation. The most likely outcome is a frameshift mutation that knocks out the gene. This targeting strategy involves finding a prototypical spacer sequence in the exons of the gene, which possesses a PAM sequence, NGG, and is unique in the genome. The prototypical spacer sequence is preferred in the first exon, as it is most detrimental to the gene.
[0542] Each gRNA was validated in the mouse cell line Neuro-N2a via transient co-transfection with hSpCas9 using liposomes. Genomic DNA was purified 72 hours post-transfection using QuickExtract DNA from Epicentre. PCR was performed to amplify loci of interest. Subsequently, the SURVEYOR mutation detection kit from Transgenomics was used. SURVEYOR results for each gRNA and its respective control are shown in Figure A1. A positive SURVEYOR result corresponds to a large band and two smaller bands in the genomic PCR, which are products of the SURVEYOR nuclease that causes double-strand breakage at a mutation site. The average cleavage efficiency of each gRNA was also determined. The selected gRNA for injection was the most efficient and unique within the genome.
[0543] RNA (hSpCas9+gRNA) was injected into the anterior nucleus of a zygote and then transplanted into a surrogate mother. The mother was allowed to reach full gestation, and the offspring were sampled via tail clipping 10 days after birth. DNA was extracted and used as a template for PCR, which was then processed using a SURVEYOR. The PCR products were then sequenced. The genomic PCR products from animals that tested positive in either the SURVEYOR assay or PCR sequencing were cloned into a pUC19 vector and sequenced to determine the putative mutations from each allele.
[0544] To date, mouse pups from the Chd8-targeting assay have been fully processed to the point of allele sequencing. Figure A2 shows Surveyor results for 38 live pups (lanes 1-38), 1 dead pup (lane 39), and 1 wild-type pup for comparison (lane 40). Pups 1-19 were injected with gRNA Chd8.2, and pups 20-38 were injected with gRNA Chd8.3. Thirteen of the 38 live pups were positive for one mutation. The dead pup also had one mutation. No mutation was detected in the wild-type sample. Genomic PCR sequencing findings were consistent with those of the Surveyor assay.
[0545] Example 13: CRISPR / Cas-mediated transcriptional modulation
[0546] Figure 67 This paper describes a design of a CRISPR-TF (transcription factor) with transcriptional activation activity. The chimeric RNA is expressed by the U6 promoter, while a human-codon-optimized double-mutant version of the Cas9 protein (hSpCas9m), operably linked to three NLS and VP64 domains, is expressed by an EF1a promoter. This double mutation, D10A and H840A, results in the Cas9 protein being unable to introduce any cleavage when directed by the chimeric RNA, but maintaining its ability to bind to the target DNA.
[0547] Figure 68Transcriptional activation of the human SOX2 gene was depicted using the CRISPR-TF system (chimeric RNA and Cas9-NLS-VP64 fusion protein). 293FT cells were transfected with a plasmid containing two components: (1) a U6-driven, different chimeric RNA targeting a 20-bp sequence within or around the human SOX2 genomic locus, and (2) an EF1a-driven hSpCas9m (double mutant)-NLS-VP64 fusion protein. 293FT cells were harvested 96 hours post-transfection, and activation levels were measured by introducing mRNA expression using a qRT-PCR assay. All expression levels were normalized against a control group (grey bars), representing results from cells transfected with the CRISPR-TF backbone plasmid (without the chimeric RNA). The qRT-PCR probe used to detect SOX2 mRNA was the Taqman Human Gene Expression Assay Probe (Life Technologies). All experiments represent data from three biological replicates, n=3, and error bars are shown as sem.
[0548] Example 14: NLS: Cas9 NLS
[0549] 293FT cells were transfected with plasmids containing two components: (1) the EF1a promoter, which drives the expression of Cas9 (wild-type human-codon-optimized Sp Cas9) with different NLS designs, and (2) the U6 promoter, which drives the expression of the same chimeric RNA targeting the human EMX1 locus.
[0550] Following the manufacturer's protocol, cells were collected 72 hours post-transfection and extracted with 50 μl of QuickExtract genomic DNA extraction solution. Target EMX1 genomic DNA was amplified by PCR and purified using a 1% agarose gel. The genomic PCR products were re-annealed and subjected to Surveyor assays, following the manufacturer's protocol. Genomic cleavage efficiencies of different constructs were measured using SDS-PAGE on a 4%–12% TBE-PAGE gel (LifeScience, Inc.), analyzed, and quantified using ImageLab (Bio-Rad) software, all following the manufacturer's protocol.
[0551] Figure 69A design for different Cas9 NLS constructs was described. All Cas9s were human-codon-optimized versions of Sp Cas9. The NLS sequence was ligated to the Cas9 gene at either the N-terminus or C-terminus. All Cas9 variants with different NLS designs were cloned into a backbone vector containing the EF1a promoter, thus being driven by the EF1a promoter. Within the same vector, a chimeric RNA targeting the human EMX1 locus and driven by the U6 promoter was present, forming a two-component system.
[0552] Table M. Cas9 NLS design test results. Genomic cleavage of different Cas9-NLS constructs was quantified using surveyor assays.
[0553]
[0554] Figure 70 The efficiency of genome cleavage induced by Cas9 variants with different NLS designs was depicted. The percentage indicates the portion of human EMX1 genomic DNA cleaved by each construct. All experiments were performed from 3 biological replicates. n=3, and the error is expressed as SEM.
[0555] Example 15: Using Cas9 to engineer microalgae
[0556] Methods of delivering Cas9
[0557] Method 1: The applicant used a vector that expresses Cas9 under the control of a constitutive promoter (such as Hsp70A-Rbc S2 or β2-tubulin) to deliver Cas9 and guide RNA.
[0558] Method 2: The applicant delivers Cas9 and T7 polymerases using a vector that expresses Cas9 and T7 polymerases under the control of a constitutive promoter (such as Hsp70A-Rbc S2 or β2-tubulin). The guide RNA will be delivered using a vector containing a T7 promoter that drives the guide RNA.
[0559] Method 3: The applicant delivers Cas9 mRNA and in vitro transcribed guide RNA into algal cells. The RNA can be transcribed in vitro. The Cas9 mRNA will consist of the coding region of Cas9 and the 3'UTR from Cop1 to ensure the stability of the Cas9 mRNA.
[0560] For homologous recombination, the applicant provided an additional homologous-directed repair template.
[0561] The sequence of the cassette that drives Cas9 expression is located after the β-2 tubulin promoter following the 3'UTR of Cop1.
[0562]
[0563] The sequence of the cassette that drives T7 polymerase expression under the control of the β-2 tubulin promoter located after the 3'UTR of Cop1:
[0564]
[0565] The sequence of the guide RNA driven by the T7 promoter (T7 promoter, N indicates the target sequence):
[0566] gaaatTAATACGACTCACTATA NNNNNNNNNNNNNNNNNNNN gttttag agctaGAAAtagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcttttttt
[0567] Gene delivery:
[0568] Chlamydomonas reinhardtii strains CC-124 and CC-125 from the Chlamydomonas Resource Center will be used for electroporation. The electroporation protocol follows the standard recommended protocol from the GeneArt Chlamydomonas Engineering kit.
[0569] Furthermore, the applicant generated Chlamydomonas reinhardtii strains constitutively expressing Cas9. This was accomplished by using pChlamy1 (linearized with PvuI) and selecting hygromycin-resistant colonies. The following is the sequence of pChlamy1 containing Cas9. Achieving gene knockout in this manner simply requires the delivery of guide RNA. For homologous recombination, the applicant delivered guide RNA along with a linearized homologous recombination template.
[0570] pChlamy1-Cas9:
[0571]
[0572] For all modified Chlamydomonas reinhardtii cells, the applicant used PCR, SURVEYOR nuclease assays, and DNA sequencing to verify successful modification.
[0573] Example 16: Using Cas9 as a transcriptional repressor in bacteria
[0574] The ability to artificially control transcription is crucial for both studying gene function and constructing synthetic gene networks with desired properties. The applicant here describes the use of the RNA-guided Cas9 protein as a programmable transcriptional repressor.
[0575] The applicant has previously demonstrated how to guide genome editing in Streptococcus pneumoniae using the Cas9 protein from Streptococcus pyogenes SF370. In this study, the applicant engineered a crR6Rk strain containing a minimal CRISPR system consisting of Cas9, tracrRNA, and a repeat. The D10A-H840 mutation was introduced into the Cas9 in this strain, resulting in strain crR6Rk**. Four spacers targeting different locations on the promoter of the bgaA β-galactosidase gene were cloned into a CRISPR array carried by the previously described pDB98 plasmid. The applicant observed an X to Y fold decrease in β-galactosidase activity dependent on this targeting location, demonstrating the potential of Cas9 as a programmable repressor. Figure 73 ).
[0576] To achieve Cas9** repression in *E. coli*, a green luminescence protein (GFP) reporter plasmid (pDB127) was constructed to express the gfpmut2 gene from a constitutive promoter. The promoter was designed to carry several NPPPAMs on both strands to measure the effect of Cas9** binding at different locations. The applicant introduced the D10A-H840 mutation into pCas9, and a simple clone targeting the novel spacer was designed to carry the described plasmid containing the tracrRNA, Cas9, and a minimal CRISPR array. Twenty-two different spacers were designed to target different regions of the gfpmut2 promoter and open reading frames. Approximately 20-fold reduction in luminescence was observed after targeting regions overlapping or adjacent to the -35 and -10 promoter elements and adjacent to the Shine-Dalgarno sequence. Targets on both strands showed similar levels of repression. These results suggest that Cas9** binding at any location within this promoter region prevents transcription initiation, presumably through steric repression of RNAP binding.
[0577] To determine whether Cas9** prevents transcriptional elongation, the applicant directed it to the reading frame of gpfmut2. Reduced expression was observed on both the coding and non-coding strands targeted, indicating that Cas9 binding is sufficiently strong to represent a barrier to RNAP. However, a 40% reduction in expression was observed when the coding strand was targeted, and a 20-fold reduction was observed for the non-coding strand. Figure 74 (b-74c, comparing T9, T10, and T11 with B9, B10, and B11). To directly determine the effect of Cas9** binding on transcription, the applicant extracted RNA from strains carrying T5, T10, B10, or a control construct not targeting pDB127, and subjected it to Northern blotting analysis using a probe that binds before (B477) or after (B510) the B10 and T10 target sites. Consistent with the applicant's blotting method, no gfpmut2 transcription was detected when Cas9** was directed to the promoter region (T5 target), and transcription was observed after targeting the T10 region. Interestingly, smaller transcripts were observed with the B477 probe. This band corresponds to a transcript of the expected size that would be interrupted by Cas9**, and is a direct indication of transcriptional termination caused by dgRNA::Cas9** binding to the coding strand. Surprisingly, the applicant did not detect transcripts when targeting the non-coding strand (B10). Since Cas9** bound to the B10 region cannot interfere with transcription initiation, this result indicates mRNA degradation. DgRNA::Cas9 showed in vitro binding to ssRNA. The applicant hypothesizes that binding can trigger mRNA degradation via host nucleases. Indeed, ribosome arrest can induce cleavage of translated mRNA in E. coli.
[0578] Some applications require precise coordination of gene expression rather than complete repression. The applicant seeks to achieve intermediate levels of repression by introducing mismatches that weaken crRNA / target interactions. The applicant has constructed a series of spacers based on B1, T5, and B10 constructs with an increasing number of mutations at the 5' end of the crRNA. Up to eight mutations in B1 and T5 do not affect this level of repression, and a gradual increase in repression is observed for additional mutations.
[0579] The repression observed with only 8 nt matches between crRNA and its target increases the off-target problem of using Cas9** as a transcriptional regulator. Since good PAM (NGG) is also required for Cas9 binding, the number of nucleotides required for a match to achieve a certain level of respiration is 10. 10 nt matches occur randomly every approximately 1 Mbp, and therefore such sites are even likely to be found in small bacterial genomes. However, for effective transcriptional repression, such sites need to be in the promoter region of a gene, making off-target effects less likely. The applicant also showed that gene expression can be affected if a non-coding strand of a gene is targeted. For this to happen, the random target must be correctly oriented, but such events are relatively more likely. In fact, during the course of this study, the applicant was unable to construct one of the spacers designed on pCas9**. The applicant subsequently found that this spacer showed a 12 bp match with a well-adjacent PAM in the required murC gene. Such off-target effects can be easily avoided by phylogenetic explosion of the designed spacers.
[0580] Several aspects of the invention are further described in the following numbered paragraphs:
[0581] 1. A carrier system comprising one or more carriers, wherein the system includes
[0582] a. A first regulatory element operatively linked to a traer pairing sequence and one or more insertion sites for inserting a guide sequence upstream of the traer pairing sequence, wherein, upon expression, the guide sequence directs a CRISPR complex in eukaryotic cells to bind specifically to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme complexed with: (1) a guide sequence hybridizing to the target sequence, and (2) a traer pairing sequence hybridizing to the traer sequence; and
[0583] b. A second regulatory element operatively linked to an enzyme-coding sequence encoding the CRISPR enzyme, the CRISPR enzyme comprising a nuclear localization sequence;
[0584] Components (a) and (b) are located on the same or different carriers in the system.
[0585] 2. The carrier system as described in paragraph 1, wherein component (a) further includes a traer sequence downstream of the traer pairing sequence under the control of the first regulating element.
[0586] 3. The vector system as described in paragraph 1, wherein component (a) further comprises two or more guide sequences operatively linked to the first regulatory element, wherein, when expressed, each of the two or more guide sequences guides the CRISPR complex to bind sequence-specifically to a different target sequence in eukaryotic cells.
[0587] 4. The carrier system as described in paragraph 1, wherein the system includes a traer sequence under the control of a third regulating element.
[0588] 5. The vector system as described in paragraph 1, wherein, when optimal alignment is performed, the traer sequence exhibits at least 50% sequence complementarity along the length of the traer pairing sequence.
[0589] 6. The vector system as described in paragraph 1, wherein the CRISPR enzyme comprises one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of a eukaryotic cell.
[0590] 7. The vector system as described in paragraph 1, wherein the CRISPR enzyme is a type II CRISPR system enzyme.
[0591] 8. The vector system as described in paragraph 1, wherein the CRISPR enzyme is a Cas9 enzyme.
[0592] 9. The vector system as described in paragraph 1, wherein the CRISPR enzyme is codon-optimized for expression in eukaryotic cells.
[0593] 10. The vector system as described in paragraph 1, wherein the CRISPR enzyme directs the cleavage of one or both strands at the target sequence location.
[0594] 11. The vector system as described in paragraph 1, wherein the CRISPR enzyme lacks DNA strand cutting activity.
[0595] 12. The vector system as described in paragraph 1, wherein the first regulatory element is a polymerase III promoter.
[0596] 13. The vector system as described in paragraph 1, wherein the second regulatory element is a polymerase II promoter.
[0597] 14. The vector system as described in paragraph 4, wherein the third regulatory element is a polymerase III promoter.
[0598] 15. The vector system as described in paragraph 1, wherein the guide sequence is at least 15 nucleotides in length.
[0599] 16. The vector system as described in paragraph 1, wherein, when optimally folded, less than 50% of the nucleotides of the guide sequence are involved in self-complementary base pairing.
[0600] 17. A vector comprising a regulatory element operatively linked to an enzyme-coding sequence encoding a CRISPR enzyme containing one or more nuclear localization sequences, wherein the regulatory element drives transcription of the CRISPR enzyme in a eukaryotic cell such that the CRISPR enzyme accumulates in the nucleus of the eukaryotic cell in a detectable amount.
[0601] 18. The vector as described in paragraph 17, wherein the regulatory element is a polymerase II promoter.
[0602] 19. The vector as described in paragraph 17, wherein the CRISPR enzyme is a type II CRISPR system enzyme.
[0603] 20. The vector as described in paragraph 17, wherein the CRISPR enzyme is a Cas9 enzyme.
[0604] 21. The vector as described in paragraph 17, wherein the CRISPR enzyme lacks the ability to cleave one or more strands of the target sequence to which it is bound.
[0605] 22. A CRISPR enzyme comprising one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of a eukaryotic cell.
[0606] 23. The CRISPR enzyme as described in paragraph 22, wherein the CRISPR enzyme is a type II CRISPR system enzyme.
[0607] 24. The CRISPR enzyme as described in paragraph 22, wherein the CRISPR enzyme is a Cas9 enzyme.
[0608] 25. A CRISPR enzyme as described in paragraph 22, wherein the CRISPR enzyme lacks the ability to cleave one or more strands of the target sequence to which it is bound.
[0609] 26. A eukaryotic host cell, comprising:
[0610] a. A first regulatory element operatively linked to a traer pairing sequence and one or more insertion sites for inserting a guide sequence upstream of the traer pairing sequence, wherein, when expressed, the guide sequence directs a CRISPR complex in eukaryotic cells to bind specifically to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme complexed with: (1) a guide sequence hybridizing to the target sequence, and (2) a traer pairing sequence hybridizing to the traer sequence; and / or
[0611] b. A second regulatory element operatively linked to an enzyme-coding sequence encoding the CRISPR enzyme, the CRISPR enzyme comprising a nuclear localization sequence.
[0612] 27. The eukaryotic host cell as described in paragraph 26, wherein the host cell comprises components (a) and (b).
[0613] 28. A eukaryotic host cell as described in paragraph 26, wherein component (a), component (b), or components (a) and (b) are stably integrated into a genome of the host eukaryotic cell.
[0614] 29. The eukaryotic host cell as described in paragraph 26, wherein component (a) further includes a traer sequence downstream of the traer pairing sequence under the control of the first regulatory element.
[0615] 30. The eukaryotic host cell as described in paragraph 26, wherein component (a) further comprises two or more guide sequences operatively linked to the first regulatory element, wherein, when expressed, each of the two or more guide sequences guides the CRISPR complex to bind sequence-specifically to a different target sequence in the eukaryotic cell.
[0616] 31. The eukaryotic host cell as described in paragraph 26, further comprising a third regulatory element operatively connected to the traer sequence.
[0617] 32. The eukaryotic host cell as described in paragraph 26, wherein, when optimal alignment is performed, the traer sequence exhibits at least 50% sequence complementarity along the length of the traer pairing sequence.
[0618] 33. The eukaryotic host cell as described in paragraph 26, wherein the CRISPR enzyme comprises one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of the eukaryotic cell.
[0619] 34. The eukaryotic host cell as described in paragraph 26, wherein the CRISPR enzyme is a type II CRISPR system enzyme.
[0620] 35. A eukaryotic host cell as described in paragraph 26, wherein the CRISPR enzyme is a Cas9 enzyme.
[0621] 36. A eukaryotic host cell as described in paragraph 26, wherein the CRISPR enzyme has been codon-optimized for expression in eukaryotic cells.
[0622] 37. A eukaryotic host cell as described in paragraph 26, wherein the CRISPR enzyme directs the cleavage of one or both strands at the target sequence location.
[0623] 38. A eukaryotic host cell as described in paragraph 26, wherein the CRISPR enzyme lacks DNA strand cutting activity.
[0624] 39. The eukaryotic host cell as described in paragraph 26, wherein the first regulatory element is a polymerase III promoter.
[0625] 40. The eukaryotic host cell as described in paragraph 26, wherein the second regulatory element is a polymerase II promoter.
[0626] 41. The eukaryotic host cell as described in paragraph 31, wherein the third regulatory element is a polymerase III promoter.
[0627] 42. A eukaryotic host cell as described in paragraph 26, wherein the guide sequence is at least 15 nucleotides in length.
[0628] 43. A eukaryotic host cell as described in paragraph 26, wherein, when folded optimally, less than 50% of the nucleotides of the guide sequence are involved in self-complementary base pairing.
[0629] 44. A non-human animal comprising any eukaryotic host cell described in any one of paragraphs 26-43.
[0630] 45. A kit comprising a vector system and instructions for using the kit, the vector system comprising:
[0631] a. A first regulatory element operatively linked to a traer pairing sequence and one or more insertion sites for inserting a guide sequence upstream of the traer pairing sequence, wherein, when expressed, the guide sequence directs a CRISPR complex in eukaryotic cells to bind specifically to a target sequence, wherein the CRISPR complex comprises a CRISPR enzyme complexed with: (1) a guide sequence hybridizing to the target sequence, and (2) a traer pairing sequence hybridizing to the traer sequence; and / or
[0632] b. A second regulatory element operatively linked to an enzyme-coding sequence encoding the CRISPR enzyme, the CRISPR enzyme comprising a nuclear localization sequence.
[0633] 46. The kit as described in paragraph 45, wherein the kit comprises components (a) and (b) located on the same or different carriers of the system.
[0634] 47. The kit as described in paragraph 45, wherein component (a) further comprises a traer sequence downstream of the traer pairing sequence under the control of the first regulatory element.
[0635] 48. The kit as described in paragraph 45, wherein component (a) further comprises two or more guide sequences operatively linked to the first regulatory element, wherein, when expressed, each of the two or more guide sequences guides the CRISPR complex to bind sequence-specifically to a different target sequence in eukaryotic cells.
[0636] 49. The kit as described in paragraph 45, wherein the system includes a traer sequence under the control of a third regulatory element.
[0637] 50. The kit as described in paragraph 45, wherein, when optimal alignment is performed, the traer sequence exhibits at least 50% sequence complementarity along the length of the traer paired sequence.
[0638] 51. The kit as described in paragraph 45, wherein the CRISPR enzyme comprises one or more nuclear localization sequences having sufficient strength to drive the CRISPR enzyme to accumulate in a detectable amount in the nucleus of a eukaryotic cell.
[0639] 52. The kit as described in paragraph 45, wherein the CRISPR enzyme is a type II CRISPR system enzyme.
[0640] 53. The kit as described in paragraph 45, wherein the CRISPR enzyme is a Cas9 enzyme.
[0641] 54. The kit as described in paragraph 45, wherein the CRISPR enzyme has been codon-optimized for expression in eukaryotic cells.
[0642] 55. The kit as described in paragraph 45, wherein the CRISPR enzyme directs the cleavage of one or both strands at the target sequence location.
[0643] 56. The kit as described in paragraph 45, wherein the CRISPR enzyme lacks DNA strand cutting activity.
[0644] 57. The kit as described in paragraph 45, wherein the first regulatory element is a polymerase III promoter.
[0645] 58. The kit as described in paragraph 45, wherein the second regulatory element is a polymerase II promoter.
[0646] 59. The kit as described in paragraph 49, wherein the third regulatory element is a polymerase III promoter....
Claims
1. An engineered, non-naturally occurring CRISPR-Cas system for genome editing in eukaryotic cells, comprising: a) A fusion protein comprising a Cas9 of a type II CRISPR-Cas system linked to one or more nuclear localization signals (NLS) and one or more heterologous protein domains, or a polynucleotide encoding the fusion protein that has been codon-optimized for expression in eukaryotic cells, wherein the heterologous protein domain has one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity; b) A CRISPR-Cas system chimerism guide, comprising from 5' to 3': (i) a guide sequence capable of hybridizing to a target sequence in the nucleus of a eukaryotic cell adjacent to a prototype spacer adjacent motif (PAM) and guiding the CRISPR-Cas complex to sequence-specific binding to the target sequence, wherein the guide sequence comprises 15-30 nucleotides in length, (ii) a tracr pairing sequence capable of hybridizing to a tracr sequence, and (iii) a tracr sequence comprising at least 30 nucleotides in length.
2. An engineered, non-naturally occurring CRISPR-Cas vector system for genome editing in eukaryotic cells, the vector system comprising one or more vectors, the one or more vectors containing a) A first regulatory element operatively linked to a nucleotide sequence encoding a fusion protein comprising a Cas9 system connected to one or more nuclear localization signals (NLS) of a type II CRISPR-Cas system and one or more heterologous protein domains, wherein the nucleotide sequence is codon-optimized for expression in eukaryotic cells, wherein the heterologous protein domain has one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity; b) A second regulatory element operatively linked to a nucleotide sequence encoding a CRISPR-Cas system chimerism guide, the CRISPR-Cas system chimerism guide comprising, from 5' to 3': (i) a guide sequence capable of hybridizing to a target sequence in the nucleus of a eukaryotic cell adjacent to a prototype spacer adjacent motif (PAM) and guiding the CRISPR-Cas complex to the target sequence in a sequence-specific manner, wherein the guide sequence comprises 15-30 nucleotides in length, (ii) a tracr pairing sequence capable of hybridizing to a tracr sequence, and (iii) a tracr sequence comprising at least 30 nucleotides in length; Components a) and b) are located on the same or different carriers in the system.
3. The system of claim 2, wherein the carrier is a viral carrier.
4. The system of claim 3, wherein the viral vector is an adeno-associated vector.
5. The system of claim 1 or 2, wherein the tracr sequence comprises a length of at least 40 nucleotides.
6. The system of claim 1 or 2, wherein the tracr sequence comprises a length of at least 50 nucleotides.
7. The system of claim 1 or 2, wherein the guide sequence comprises a length of at least 20 nucleotides.
8. The system of claim 1 or 2, wherein the fusion protein comprises at least one SV40 viral large T-antigen NLS (PKKKRKV) linked to Cas9.
9. The system of claim 1 or 2, wherein the fusion protein comprises at least two NLSs linked to Cas9.
10. The system of claim 1 or 2, wherein the fusion protein comprises at least one N-terminal NLS and at least one C-terminal NLS.
11. The system of claim 1 or 2, wherein the Cas9 is Streptococcus pyogenes Cas9 and the PAM is NGG; or the Cas9 is Streptococcus thermophilus Cas9 and the PAM is NNAGAAW.
12. The system of claim 1 or 2, wherein the Cas9 is Streptococcus pyogenes Cas9, which has a D10A mutation and is a nicking enzyme.
13. The system of claim 1 or 2, wherein the Cas9 is Streptococcus pyogenes Cas9, which has the H840A mutation and is a nicking enzyme.
14. An engineered, non-naturally occurring CRISPR-Cas system for transcriptional modulation in eukaryotic cells, comprising: a) A fusion protein comprising a Cas9 of a type II CRISPR-Cas system linked to one or more nuclear localization signals (NLS) and one or more heterologous protein domains, or a polynucleotide encoding the fusion protein that has been codon-optimized for expression in eukaryotic cells, wherein the Cas9 is Streptococcus pyogenes Cas9 with both D10A and H840A mutations, wherein the heterologous protein domain has one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity; b) A CRISPR-Cas system chimerism guide, comprising from 5' to 3': (i) a guide sequence capable of hybridizing to a target sequence in the nucleus of a eukaryotic cell adjacent to a prototype spacer adjacent motif (PAM) and guiding the CRISPR-Cas complex to sequence-specific binding to the target sequence, wherein the guide sequence comprises 15-30 nucleotides in length, (ii) a tracr pairing sequence capable of hybridizing to a tracr sequence, and (iii) a tracr sequence comprising at least 30 nucleotides in length.
15. The system of claim 1 or 2, wherein the system further comprises a template polynucleotide, wherein the template polynucleotide overlaps with at least 5 nucleotides of the target sequence.
16. The system of claim 1, wherein the polynucleotide encoding the fusion protein is codon-optimized for expression in mammalian or human cells, and wherein the polynucleotide encoding the fusion protein is mRNA.
17. Use of the system of claim 1 or 2 in the preparation of medicaments for gene therapy or gene editing.
18. Use of the system as described in claim 1 or 2 for in vitro or ex vivo modification of eukaryotic cells.
Citation Information
Patent Citations
Transgenic animals secreting desired proteins into milk
EP0264166A1
Regulation of endogenous gene expression in cells using zinc finger proteins
US20030087817A1
Transgenic plants with enhanced agronomic traits
US20090100536A1
Multiple domain proteins
US20110059502A1
Liposome carriers in chemotherapy of leishmaniasis
US4186183A