RNA-guided transcriptional regulation
By introducing specific RNA and DNA binding proteins to form a complex colocalized with the target DNA, the problem of difficult to regulate the expression of target nucleic acid in cells in the prior art is solved, and fine regulation of target nucleic acid expression and disease treatment are achieved.
Patent Information
- Application Number
- CN202110830978.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2013-06-04
- Filing Date
- 2014-06-04
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2034-06-04
AI Technical Summary
The prior art is difficult to effectively regulate the expression of target nucleic acids in cells, especially in the treatment of diseases or harmful conditions.
By introducing exogenous nucleic acids encoding RNA complementary to DNA, RNA-oriented nuclease-ineffective DNA binding proteins and transcriptional regulatory factor proteins or domains, a complex colocalized with the target DNA is formed to regulate the expression of the target nucleic acid.
The fine regulation of the expression of target nucleic acid is achieved, and the expression of target nucleic acid can be up-regulated or down-regulated to treat diseases or harmful conditions, improving the therapeutic effect and safety.
Smart Images

Figure CN113846096B_ABST
Abstract
Description
[0001] Related application data
[0002] This application claims priority to U.S. Provisional Patent Application No. 61 / 830,787, filed on June 4, 2013, which is hereby incorporated by reference in its entirety for all purposes.
[0003] Statement of Government Interests
[0004] This invention was made with government support under Grant No. P50 HG005550 from the National Institutes of Health and Grant No. DE-FG02-02ER63445 from the Department of Energy. The Government has certain rights in this invention. Background Art
[0005] Bacterial and archaeal CRISPR-Cas systems rely on short guide RNAs in complex with Cas proteins that direct the degradation of complementary sequences present within invasive foreign nucleic acids. See Deltcheva, E. et al., CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. Nature 471, 602-607 (2011); Gasiunas, G., Barrangou, R., Horvath, P. & Siksnys, V. Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria. Proceedings of the National Academy of Sciences of the United States of America 109, E2579-2586 (2012); Jinek, M. et al., A programmabledual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816-821 (2012); Sapranauskas, R. et al., The Streptococcus thermophilus CRISPR / Cassystem provides immunity in Escherichia coli.Nucleic acids research 39, 9275-9282 (2011); and Bhaya, D., Davison, M. & Barrangou, R. CRISPR-Cas systems in bacteria and archaea: versatile small RNAs for adaptive defense and regulation. Annual review of genetics 45, 273-297 (2011). Recent in vitro recombination of the type II CRISPR system of Streptococcus pyogenes (S. pyogenes) shows that crRNA ("CRISPR RNA") fused to the normally trans-encoded tracrRNA ("trans-activating CRISPR RNA") is sufficient to guide the Cas9 protein sequence to specifically cut the target DNA sequence matching the crRNA.Expression of gRNA homologous to the target site results in Cas9 recruitment and degradation of the target DNA. See H. Deveau et al., Phage response to CRISPR-encoded resistance in Streptococcus thermophilus. Journal of Bacteriology 190, 1390 (Feb, 2008). Summary of the invention
[0006] Aspects disclosed herein relate to a complex of a guide RNA, a DNA binding protein, and a double-stranded DNA target sequence. According to certain aspects, the DNA binding protein within the scope of the present disclosure includes a protein that forms a complex with a guide RNA and with a guide RNA that directs the complex to a double-stranded DNA sequence, wherein the complex is bound to the DNA sequence. This aspect disclosed herein can be referred to as the colocalization of RNA and DNA binding proteins to or with double-stranded DNA. In this way, the DNA binding protein-guide RNA complex can be used to locate a transcriptional regulatory factor protein or domain at a target DNA, thereby regulating the expression of the target DNA.
[0007] According to certain aspects, a method for regulating the expression of a target nucleic acid in a cell is provided, comprising introducing into the cell a first exogenous nucleic acid encoding one or more RNAs (ribonucleic acid) complementary to a DNA (deoxyribonucleic acid), wherein the DNA comprises the target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding an RNA guided nuclease-null DNA binding protein that binds to the DNA and is guided by the one or more RNAs, introducing into the cell a third exogenous nucleic acid encoding a transcriptional regulator protein or domain, wherein the one or more RNAs, the RNA guided nuclease-null DNA binding protein and the transcriptional regulator protein or domain are expressed, wherein the one or more RNAs, the RNA guided nuclease-null DNA binding protein and the transcriptional regulator protein or domain are co-localized to the DNA and wherein the transcriptional regulator protein or domain regulates the expression of the target nucleic acid.
[0008] According to one aspect, the exogenous nucleic acid encoding the RNA-guided nuclease-null DNA binding protein also encodes a transcriptional regulator protein or domain fused to the RNA-guided nuclease-null DNA binding protein. According to one aspect, the exogenous nucleic acid encoding one or more RNAs also encodes a target for the RNA binding domain, and the exogenous nucleic acid encoding the transcriptional regulator protein or domain also encodes an RNA binding domain fused to the transcriptional regulator protein or domain.
[0009] According to one aspect, the cell is a eukaryotic cell. According to one aspect, the cell is a yeast cell, a plant cell or an animal cell. According to one aspect, the cell is a mammalian cell.
[0010] According to one aspect, the RNA is about 10 to about 500 nucleotides. According to one aspect, the RNA is about 20 to about 100 nucleotides.
[0011] According to one aspect, the transcriptional regulatory factor protein or domain is a transcriptional activator. According to one aspect, the transcriptional regulatory factor protein or domain upregulates the expression of a target nucleic acid. According to one aspect, the transcriptional regulatory factor protein or domain upregulates the expression of a target nucleic acid to treat a disease or detrimental condition. According to one aspect, the target nucleic acid is associated with a disease or detrimental condition.
[0012] According to one aspect, one or more RNAs are guide RNAs. According to one aspect, one or more RNAs are tracrRNA-crRNA fusions. According to one aspect, the guide RNA includes a spacer sequence and a tracer mate sequence. The guide RNA may also include a tracr sequence, a portion of which hybridizes to the tracr mate sequence. The guide RNA may also include a linker nucleic acid sequence that connects the tracer mate sequence and the tracr sequence to produce a tracrRNA-crRNA fusion. The spacer sequence is bound to the target DNA, such as by hybridization.
[0013] According to one aspect, the guide RNA includes a truncated spacer sequence. According to one aspect, the guide RNA includes a truncated spacer sequence having a 1 base truncation at the 5' end of the spacer sequence. According to one aspect, the guide RNA includes a truncated spacer sequence having a 2 base truncation at the 5' end of the spacer sequence. According to one aspect, the guide RNA includes a truncated spacer sequence having a 3 base truncation at the 5' end of the spacer sequence. According to one aspect, the guide RNA includes a truncated spacer sequence having a 4 base truncation at the 5' end of the spacer sequence. Thus, the spacer sequence can have 1 to 4 base truncations at the 5' end of the spacer sequence.
[0014] According to certain embodiments, the spacer sequence may comprise about 16 to about 20 nucleotides that hybridize to the target nucleic acid sequence. According to certain embodiments, the spacer sequence may comprise about 20 nucleotides that hybridize to the target nucleic acid sequence.
[0015] According to certain aspects, the linker nucleic acid sequence can include about 4 to about 6 nucleic acids.
[0016] According to certain aspects, the tracr sequence may include about 60 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 64 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 65 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 66 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 67 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 68 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 69 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 70 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 80 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 90 to about 500 nucleic acids. According to certain aspects, the tracr sequence may include about 100 to about 500 nucleic acids.
[0017] According to certain aspects, the tracr sequence may include about 60 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 64 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 65 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 66 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 67 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 68 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 69 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 70 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 80 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 90 to about 200 nucleic acids. According to certain aspects, the tracr sequence may include about 100 to about 200 nucleic acids.
[0018] Exemplary guide RNAs are shown in Figure 5B.
[0019] According to one aspect, the DNA is genomic DNA, mitochondrial DNA, viral DNA or exogenous DNA.
[0020] According to certain aspects, a method for regulating the expression of a target nucleic acid in a cell is provided, comprising introducing into the cell a first exogenous nucleic acid encoding one or more RNAs (ribonucleic acid) complementary to a DNA (deoxyribonucleic acid), wherein the DNA comprises the target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding a DNA-binding protein of a type II CRISPR system that binds to DNA and is guided by the one or more RNAs, introducing into the cell a third exogenous nucleic acid encoding a transcriptional regulator protein or domain, wherein the one or more RNAs, the DNA-binding protein of the type II CRISPR system that is null for the RNA-guided nuclease, and the transcriptional regulator protein or domain are expressed, wherein the one or more RNAs, the DNA-binding protein of the type II CRISPR system that is null for the RNA-guided nuclease, and the transcriptional regulator protein or domain are co-localized to the DNA and wherein the transcriptional regulator protein or domain regulates the expression of the target nucleic acid.
[0021] According to one aspect, the exogenous nucleic acid encoding the RNA-guided nuclease-null DNA binding protein of the type II CRISPR system also encodes a transcriptional regulator protein or domain fused to the RNA-guided nuclease-null DNA binding protein of the type II CRISPR system. According to one aspect, the exogenous nucleic acid encoding one or more RNAs also encodes a target for the RNA binding domain, and the exogenous nucleic acid encoding the transcriptional regulator protein or domain also encodes an RNA binding domain fused to the transcriptional regulator protein or domain.
[0022] According to one aspect, the cell is a eukaryotic cell. According to one aspect, the cell is a yeast cell, a plant cell or an animal cell. According to one aspect, the cell is a mammalian cell.
[0023] According to one aspect, the RNA is about 10 to about 500 nucleotides. According to one aspect, the RNA is about 20 to about 100 nucleotides.
[0024] According to one aspect, the transcriptional regulator protein or domain is a transcriptional activator. According to one aspect, the transcriptional regulator protein or domain upregulates the expression of a target nucleic acid. According to one aspect, the transcriptional regulator protein or domain upregulates the expression of a target nucleic acid to treat a disease or adverse condition. According to one aspect, the target nucleic acid is associated with a disease or adverse condition.
[0025] According to one aspect, the one or more RNAs are guide RNAs. According to one aspect, the one or more RNAs are tracrRNA-crRNA fusions.
[0026] According to one aspect, the DNA is genomic DNA, mitochondrial DNA, viral DNA or exogenous DNA.
[0027] According to certain aspects, a method for regulating the expression of a target nucleic acid in a cell is provided, comprising introducing into the cell a first exogenous nucleic acid encoding one or more RNAs (ribonucleic acid) complementary to a DNA (deoxyribonucleic acid), wherein the DNA comprises a target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding a nuclease-ineffective Cas9 protein that binds to the DNA and is guided by the one or more RNAs, introducing into the cell a third exogenous nucleic acid encoding a transcriptional regulator protein or domain, wherein the one or more RNAs, the nuclease-ineffective Cas9 protein and the transcriptional regulator protein or domain are expressed, wherein the one or more RNAs, the nuclease-ineffective Cas9 protein and the transcriptional regulator protein or domain are co-localized to the DNA and wherein the transcriptional regulator protein or domain regulates the expression of the target nucleic acid.
[0028] According to one aspect, the exogenous nucleic acid encoding the nuclease-ineffective Cas9 protein also encodes a transcriptional regulator protein or domain fused to the nuclease-ineffective Cas9 protein. According to one aspect, the exogenous nucleic acid encoding one or more RNAs also encodes a target of an RNA binding domain, and the exogenous nucleic acid encoding a transcriptional regulator protein or domain also encodes an RNA binding domain fused to a transcriptional regulator protein or domain.
[0029] According to one aspect, the cell is a eukaryotic cell. According to one aspect, the cell is a yeast cell, a plant cell or an animal cell. According to one aspect, the cell is a mammalian cell.
[0030] According to one aspect, the RNA is about 10 to about 500 nucleotides. According to one aspect, the RNA is about 20 to about 100 nucleotides.
[0031] According to one aspect, the transcriptional regulator protein or domain is a transcriptional activator. According to one aspect, the transcriptional regulator protein or domain upregulates the expression of a target nucleic acid. According to one aspect, the transcriptional regulator protein or domain upregulates the expression of a target nucleic acid to treat a disease or adverse condition. According to one aspect, the target nucleic acid is associated with a disease or adverse condition.
[0032] According to one aspect, the one or more RNAs are guide RNAs. According to one aspect, the one or more RNAs are tracrRNA-crRNA fusions.
[0033] According to one aspect, the DNA is genomic DNA, mitochondrial DNA, viral DNA or exogenous DNA.
[0034] According to one aspect, a cell is provided, comprising a first exogenous nucleic acid encoding one or more RNAs complementary to a DNA, wherein the DNA comprises a target nucleic acid, a second exogenous nucleic acid encoding an RNA-guided nuclease-ineffective DNA-binding protein, and a third exogenous nucleic acid encoding a transcriptional regulator protein or domain, wherein the one or more RNAs, the RNA-guided nuclease-ineffective DNA-binding protein, and the transcriptional regulator protein or domain are elements (members) of a co-localized complex of the target nucleic acid.
[0035] According to one aspect, the exogenous nucleic acid encoding the RNA-guided nuclease-null DNA binding protein also encodes a transcriptional regulator protein or domain fused to the RNA-guided nuclease-null DNA binding protein. According to one aspect, the exogenous nucleic acid encoding one or more RNAs also encodes a target for the RNA binding domain, and the exogenous nucleic acid encoding the transcriptional regulator protein or domain also encodes an RNA binding domain fused to the transcriptional regulator protein or domain.
[0036] According to one aspect, the cell is a eukaryotic cell. According to one aspect, the cell is a yeast cell, a plant cell or an animal cell. According to one aspect, the cell is a mammalian cell.
[0037] According to one aspect, the RNA is about 10 to about 500 nucleotides. According to one aspect, the RNA is about 20 to about 100 nucleotides.
[0038] According to one aspect, the transcriptional regulator protein or domain is a transcriptional activator. According to one aspect, the transcriptional regulator protein or domain upregulates the expression of a target nucleic acid. According to one aspect, the transcriptional regulator protein or domain upregulates the expression of a target nucleic acid to treat a disease or adverse condition. According to one aspect, the target nucleic acid is associated with a disease or adverse condition.
[0039] According to one aspect, the one or more RNAs are guide RNAs. According to one aspect, the one or more RNAs are tracrRNA-crRNA fusions.
[0040] According to one aspect, the DNA is genomic DNA, mitochondrial DNA, viral DNA or exogenous DNA.
[0041] According to certain aspects, the RNA-guided nuclease-null DNA binding protein is an RNA-guided nuclease-null DNA binding protein of a type II CRISPR system. According to certain aspects, the RNA-guided nuclease-null DNA binding protein is a null Cas9 protein.
[0042] According to one aspect, a method for altering a DNA target nucleic acid in a cell is provided, the method comprising introducing into the cell a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in the DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one RNA-guided DNA-binding protein nickase and guided by the two or more RNAs, wherein the two or more RNAs and the at least one RNA-guided DNA-binding protein nickase are expressed and wherein the at least one RNA-guided DNA-binding protein nickase is co-localized with the two or more RNAs to the DNA target nucleic acid and nicks the DNA target nucleic acid, thereby resulting in the generation of two or more adjacent nicks.
[0043] According to one aspect, a method for altering a DNA target nucleic acid in a cell is provided, the method comprising introducing into the cell a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in the DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one RNA-guided DNA-binding protein nickase of a type II CRISPR system and guided by the two or more RNAs, wherein the two or more RNAs and the at least one RNA-guided DNA-binding protein nickase of the type II CRISPR system are expressed and wherein the at least one RNA-guided DNA-binding protein nickase of the type II CRISPR system co-localizes with the two or more RNAs to the DNA target nucleic acid and cleaves the DNA target nucleic acid, thereby resulting in the generation of two or more adjacent nicks.
[0044] According to one aspect, a method for changing a DNA target nucleic acid in a cell is provided, the method comprising introducing into the cell a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in the DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one Cas9 protein nickase, the nickase having an inactive nuclease domain and being guided by two or more RNAs, wherein the two or more RNAs and the at least one Cas9 protein nickase are expressed and wherein the at least one Cas9 protein nickase is co-localized with the two or more RNAs to the DNA target nucleic acid and cuts the DNA target nucleic acid, thereby resulting in the generation of two or more adjacent nicks.
[0045] According to the method for changing the DNA target nucleic acid, the two or more adjacent nicks are on the same strand of double-stranded DNA. According to one aspect, the two or more adjacent nicks are on the same strand of double-stranded DNA and result in homologous recombination. According to one aspect, the two or more adjacent nicks are on different strands of double-stranded DNA. According to one aspect, the two or more adjacent nicks are on different strands of double-stranded DNA and double-strand breaks are generated. According to one aspect, the two or more adjacent nicks are on different strands of double-stranded DNA and double-strand breaks are generated, resulting in non-homologous end joining. According to one aspect, the two or more adjacent nicks are on different strands of double-stranded DNA and are relatively offset from each other. According to one aspect, the two or more adjacent nicks are on different strands of double-stranded DNA and are relatively offset from each other, and double-strand breaks are generated. According to one aspect, the two or more adjacent nicks are on different strands of double-stranded DNA and are relatively offset from each other, and double-strand breaks are generated, resulting in non-homologous end joining. According to one aspect, the method further comprises introducing into the cell a third exogenous nucleic acid encoding a donor nucleic acid sequence, wherein the two or more nicks result in homologous recombination of the target nucleic acid with the donor nucleic acid sequence.
[0046] According to one aspect, a method for altering a DNA target nucleic acid in a cell is provided, the method comprising introducing into the cell a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in the DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one RNA-guided DNA-binding protein nickase and guided by the two or more RNAs, wherein the two or more RNAs and the at least one RNA-guided DNA-binding protein nickase are expressed, and wherein the at least one RNA-guided DNA-binding protein nickase co-localizes with the two or more RNAs to the DNA target nucleic acid and cleaves the DNA target nucleic acid, thereby resulting in the generation of two or more adjacent nicks, and wherein the two or more adjacent nicks are on different strands of the double-stranded DNA and generate double-strand breaks, thereby resulting in fragmentation of the target nucleic acid, thereby preventing expression of the target nucleic acid.
[0047] According to one aspect, a method for altering a DNA target nucleic acid in a cell is provided, the method comprising introducing into the cell a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in the DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one RNA-guided DNA-binding protein nickase of a type II CRISPR system and guided by the two or more RNAs, wherein the two or more RNAs and the at least one RNA-guided DNA-binding protein nickase of the type II CRISPR system are expressed, and wherein the at least one RNA-guided DNA-binding protein nickase of the type II CRISPR system co-localizes with the two or more RNAs to the DNA target nucleic acid and cleaves the DNA target nucleic acid, thereby resulting in two or more adjacent nicks, and wherein the two or more adjacent nicks are on different strands of the double-stranded DNA and produce double-strand breaks, thereby resulting in fragmentation of the target nucleic acid, thereby preventing expression of the target nucleic acid.
[0048] According to one aspect, a method for altering a DNA target nucleic acid in a cell is provided, the method comprising introducing into the cell a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in the DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one Cas9 protein nickase having an inactive nuclease domain and guided by the two or more RNAs, wherein the two or more RNAs and the at least one Cas9 protein nickase are expressed, and wherein the at least one Cas9 protein nickase is co-localized with the two or more RNAs to the DNA target nucleic acid and cuts the DNA target nucleic acid, thereby resulting in the generation of two or more adjacent nicks, and wherein the two or more adjacent nicks are on different strands of the double-stranded DNA and generate double-strand breaks, thereby resulting in fragmentation of the target nucleic acid, thereby preventing the expression of the target nucleic acid.
[0049] According to one aspect, a cell is provided, comprising a first exogenous nucleic acid encoding two or more RNAs, wherein each RNA is complementary to adjacent sites in a DNA target nucleic acid, and a second exogenous nucleic acid encoding at least one RNA-guided DNA-binding protein nickase, and wherein the two or more RNAs and the at least one RNA-guided DNA-binding protein nickase are elements of a co-localized complex of the DNA target nucleic acid.
[0050] According to one aspect, the RNA-guided DNA-binding protein nickase is an RNA-guided DNA-binding protein nickase of a type II CRISPR system. According to one aspect, the RNA-guided DNA-binding protein nickase is a Cas9 protein nickase with an inactive nuclease domain.
[0051] According to one aspect, the cell is a eukaryotic cell. According to one aspect, the cell is a yeast cell, a plant cell or an animal cell. According to one aspect, the cell is a mammalian cell.
[0052] According to one aspect, the RNA comprises from about 10 to about 500 nucleotides. According to one aspect, the RNA comprises from about 20 to about 100 nucleotides.
[0053] According to one aspect, the target nucleic acid is associated with a disease or deleterious condition.
[0054] According to one aspect, the two or more RNAs are guide RNAs. According to one aspect, the two or more RNAs are tracrRNA-crRNA fusions.
[0055] According to one aspect, the DNA target nucleic acid is genomic DNA, mitochondrial DNA, viral DNA or exogenous DNA.
[0056] Other features and advantages of certain embodiments of the invention will become more fully apparent in the following description of embodiments of the invention and the accompanying drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] This patent or patent application document contains color drawings. Copies of this patent or patent application publication with color drawings will be provided by the Office upon request and payment of the necessary fee. The above and other features and advantages of embodiments of the present invention will be more fully understood from the following detailed description of illustrative embodiments in conjunction with the accompanying drawings, in which:
[0058] Figure 1A and Figure 1B Schematic diagram of RNA-directed transcriptional activation. Figure 1C It is the design of the reporter construct. Figure 1D shows data indicating that the Cas9N-VP64 fusion shows RNA-guided transcriptional activation, as determined by both fluorescence activated cell sorting (fluorescence activated cell sorting, FACS) and immunofluorescence assay (IF). Figure 1E shows the assay data of FACS, and 1F shows gRNA sequence-specific transcriptional activation from the reporter construct in the presence of Cas9N, MS2-VP64 and gRNA with appropriate MS2 aptamer binding sites. Figure 1F Data showing transcriptional induction by single gRNA and multiple gRNAs are shown.
[0059] Figure 2A Methods for evaluating Cas9-gRNA complexes and TALE targeting landscapes are shown. Figure 2BData are shown indicating that, on average, Cas9-gRNA complexes tolerate 1-3 mutations in their target sequence. Figure 2C Data are shown indicating that the Cas9-gRNA complex is largely insensitive to point mutations, except those targeting the PAM sequence. Figure 2D Heat plot data are shown, indicating that the introduction of 2 base mismatches significantly impaired Cas9-gRNA complex activity. Figure 2E Data are shown indicating that, on average, 18-mer TALEs show tolerance to 1-2 mutations in their target sequences. Figure 2F Data are shown demonstrating that, similar to the Cas9-gRNA complex, the 18-mer TALE is largely insensitive to single base mismatches in its target. Figure 2G Heat map data are shown, indicating that the introduction of 2 base mismatches significantly impaired the activity of the 18-mer TALE.
[0060] Figure 3A A schematic diagram of guide RNA design is shown. Figure 3B Data showing the percentage of non-homologous end joining for off-set nicks resulting in 5' overhangs and off-set nicks resulting in 3' overhangs are shown. Figure 3C Data showing the percentage of targeting for offset nicks resulting in 5' overhangs and offset nicks resulting in 3' overhangs are shown.
[0061] Figure 4A Schematic representation of the metal coordinating residues in position D7 of RuvC PDB ID: 4EP4 (blue) (left panel), schematic representation of the HNH endonuclease domain from PDB IDs: 3M7K (orange) and 4H9D (cyan) including the coordinated Mg-ion (grey spheres) and DNA from 3M7K (purple) (middle panel), and the list of analyzed mutants (right panel). Figure 4B Data showing undetectable nuclease activity for Cas9 mutants m3 and m4, and their respective fusions to VP64 are shown. Figure 4C yes Figure 4B High-resolution inspection of data in .
[0062] Figure 5A is a schematic diagram of a homologous recombination assay to determine Cas9-gRNA activity. Figure 5B shows guide RNAs with random sequence insertions and the percentage of homologous recombination.
[0063] Fig. 6AThis is a schematic diagram of the guide RNA of the OCT4 gene. Figure 6B Transcriptional activation is shown for a promoter-luciferase reporter construct. Figure 6C Transcriptional activation by qPCR of endogenous genes is shown.
[0064] Fig. 7A This is a schematic diagram of the guide RNA of the REX1 gene. Figure 7B Transcriptional activation is shown for a promoter-luciferase reporter construct. Figure 7C Transcriptional activation by qPCR of endogenous genes is shown.
[0065] Fig. 8A A schematic diagram of the high-level specificity analysis processing pipeline used to calculate normalized expression levels is shown. Figure 8B Distribution data of the percentage of binding sites for the number of mismatches generated within the biased construct library are shown. Left panel: theoretical distribution. Right panel: observed distribution from the actual TALE construct library. Figure 8C Data are shown for the percentage distribution of tagcounts aggregated to binding sites for the number of mismatches. Left: Distribution observed from positive control samples. Right: Distribution observed from samples in which non-control TALEs were induced.
[0066] Fig. 9A Shown are data from analysis of the target profile of Cas9-gRNA complexes, which show tolerance to 1-3 mutations in their target sequences. Fig. 9B Data from analysis of the target profile of the Cas9-gRNA complex are shown, showing insensitivity to point mutations except those mapped to the PAM sequence. Fig. 9C Heat map data of target profiling analysis for Cas9-gRNA complexes are shown, showing that introduction of 2 base mismatches significantly impairs activity. Fig.9D Data from a nuclease-mediated HR assay are shown, confirming that the predicted PAM for S. pyogenes Cas9 is NGG, and also NAG.
[0067] Figure 10A shows data from a nuclease-mediated HR assay confirming that the 18-mer TALE tolerates multiple mutations in its target sequence. Fig. 10B Analytical data of the targeting profiles of TALEs of three different sizes (18-mer, 14-mer, and 10-mer) are shown. Fig. 10C Data are shown for a 10-mer TALE, demonstrating resolution close to single-base mismatches. Fig. 10D Heat map data showing a 10-mer TALE, demonstrating resolution close to single-base mismatches.
[0068] Fig.11A The designed guide RNAs are shown. Fig. 11B The percentage of non-homologous end joining for multiple guide RNAs is shown.
[0069] Fig. 12A The Sox2 gene is shown. Fig. 12B The Nanog gene is shown.
[0070] Figures 13A-13F Targeting maps of two other Cas9-gRNA complexes are shown.
[0071] Fig.14A Specificity profiles of two gRNAs, wild type (SEQ ID NO: 88) and mutant (SEQ ID NO: 89-90), are shown. Sequence differences are highlighted in red. Fig. 14B and 14C This shows that the assay is specific for the gRNA being evaluated (according to Fig.13D replotted data).
[0072] Fig.15A -15D shows gRNA2 with single or double base mismatches (highlighted in red) in the spacer sequence relative to the target ( Fig.15A -B) and gRNA3( Fig. 15C -D). The sequences are shown in SEQ ID NOs: 91-131.
[0073] Fig.16A -16D shows 2 independent gRNAs tested using the nuclease assay: gRNA1 ( Fig.16A -B) and gRNA3( Fig. 16C -D). The sequences are shown in SEQ ID NOs: 66, 185-186 and 133-140.
[0074] Figures 17A-17B It is shown that the PAM of S. pyogenes Cas9 was confirmed to be NGG and NAG using a nuclease-mediated HR assay. The sequences are shown in SEQ ID NOs: 67-69 and 141.
[0075] Figures 18A-18B The use of a nuclease-mediated HR assay to confirm that the 18-mer TALE tolerates multiple mutations in its target sequence is shown. The sequences are shown in SEQ ID NOs: 70-73.
[0076] Fig.19A -19C shows a comparison of TALE monomer specificity relative to TALE protein specificity. The sequences are shown in SEQ ID NOs: 142-150.
[0077] Figures 20A-20B Data related to off-set nicking are shown. The sequences are shown as SEQ ID NOs: 151-158.
[0078] Figures 21A-21C The offset nicking and NHEJ spectra are shown. The sequences are shown in SEQ ID NOs: 159-184 and 187. DETAILED DESCRIPTION
[0079] Embodiments disclosed in the present invention are based on the use of DNA binding proteins that colocalize transcriptional regulatory factor proteins or domains to DNA in a manner that regulates target nucleic acids. It is readily known to those skilled in the art that these DNA binding proteins bind to DNA for a variety of purposes. These DNA binding proteins may be naturally occurring. DNA binding proteins included within the scope of the present invention include those that can be guided by RNA referred to herein as guide RNA (guide RNA). According to this aspect, guide RNA and RNA-guided DNA binding proteins form a colocalization complex on DNA. According to certain aspects, the DNA binding protein may be a nuclease-null DNA binding protein. According to this aspect, the nuclease-null DNA binding protein may be produced by altering or modifying a DNA binding protein with nuclease activity. These DNA binding proteins with nuclease activity are known to those skilled in the art, and include naturally occurring DNA binding proteins with nuclease activity, such as the Cas9 protein present in (for example) a type II CRISPR system. These Cas9 proteins and type II CRISPR systems are well documented in the art. See Makarova et al., Nature Reviews, Microbiology, Vol. 9, June 2011, pp. 467-477 (including all supplementary information), which is incorporated herein by reference in its entirety.
[0080] Exemplary DNA binding proteins with nuclease activity act to cause nicks or cut double-stranded DNA. These nuclease activities can be produced by DNA binding proteins with one or more polypeptide sequences showing nuclease activity. These exemplary DNA binding proteins can have two different nuclease domains, each of which is responsible for cutting a specific strand of double-stranded DNA or causing nicks therefor. Exemplary polypeptide sequences with nuclease activity known to those skilled in the art include McrA-HNH nuclease-associated domains and RuvC-like nuclease domains. Therefore, exemplary DNA binding proteins are those that essentially contain one or more of McrA-HNH nuclease-associated domains and RuvC-like nuclease domains. According to some aspects, the DNA binding protein is changed or otherwise modified to inactivate nuclease activity. These changes or modifications include changing one or more amino acids to inactivate nuclease activity or nuclease domain. These modifications include removing one or more polypeptide sequences (i.e., nuclease domains) showing nuclease activity, so that one or more polypeptide sequences (i.e., nuclease domains) showing nuclease activity are not present in the DNA binding protein. Based on the disclosure of the present invention, other modifications that inactivate nuclease activity will be apparent to those skilled in the art. Therefore, a nuclease-ineffective DNA binding protein comprises a modified polypeptide sequence that inactivates the nuclease activity or the removal of one or more polypeptide sequences that inactivate the nuclease activity. Even if the nuclease activity is inactivated, the nuclease-ineffective DNA binding protein retains the ability to bind to DNA. Therefore, a DNA binding protein comprises one or more polypeptide sequences required for DNA binding, but may lack one or more or all nuclease sequences that exhibit nuclease activity. Therefore, a DNA binding protein comprises one or more polypeptide sequences required for DNA binding, but may have one or more or all nuclease sequences that exhibit inactivated nuclease activity.
[0081] According to one aspect, the DNA binding protein with two or more nuclease domains can be modified or changed to inactivate the nuclease domains except one of the nuclease domains. The DNA binding protein modified or changed is called a DNA binding protein nickase, so that the DNA binding protein only cuts or causes nicks to a chain in the double-stranded DNA. When guided to DNA by RNA, the DNA binding protein nickase is called the RNA-guided DNA binding protein nickase.
[0082] An exemplary DNA binding protein is an RNA-guided DNA binding protein of a type II CRISPR system lacking nuclease activity. An exemplary DNA binding protein is a nuclease-inactive Cas9 protein. An exemplary DNA binding protein is a Cas9 protein nickase.
[0083] In Streptococcus pyogenes (S. pyogenes), Cas9 produces a blunt-ended double-strand break 3 bp upstream of the protospacer-adjacent motif (PAM) through a process mediated by two catalytic domains in the protein: the HNH domain that cuts the complementary DNA strand and the RuvC-like domain that cuts the non-complementary strand. See Jinke et al., Science 337, 816-821 (2012), which is incorporated herein by reference in its entirety. Cas9 proteins are known to be present in a variety of type II CRISPR systems, including the following: as identified in the supplementary information of Makarova et al., Nature Reviews, Microbiology, Vol. 9, June 2011, pp. 467-477: Methanococcus maripaludis C7; Corynebacterium diphtheriae; Corynebacterium efficiens YS-314; Corynebacterium glutamicum ATCC 13032 Kitasato; Corynebacterium glutamicum ATCC 13032 Bielefeld; Corynebacterium glutamicum R; Corynebacterium kroppenstedtii DSM 44385; Mycobacterium abscessus ATCC 19977; Nocardia farcinica IFM10152; Rhodococcus erythropolis PR4; Rhodococcus jostii RHA1; Rhodococcus opacus B4 uid36573; Acidothermus cellulolyticus 11B; Arthrobacter chlorophenolicus A6; Kribbella flavida DSM17836 uid43465; Thermomonospora curvata DSM 43183; Bifidobacterium dentium Bdl;Bifidobacterium longum DJO10A; Slackiaheliotrinireducens DSM 20476; Persephonella marina EX HI; Bacteroides fragilis NCTC 9434; Capnocytophaga ochracea DSM 7271; Flavobacterium psychrophilum JIP02 86; Akkermansia muciniphila ATCC BAA 835; Roseiflexus castenholzii DSM 13941; Roseiflexus RSI; Synechocystis PCC6803; Elusimicrobium minutum Peil91; uncultured termite group 1 bacterial phylotypes Termite group 1bacterium phylotype)Rs D17; Fibrobacter succinogenes S85; Bacillus cereus ATCC 10987; Listeria innocua; Lactobacillus casei; Lactobacillus rhamnosus GG; Lactobacillus salivarius UCC118; Streptococcus agalactiae A909; Streptococcus agalactiae NEM316; Streptococcus agalactiae 2603; Streptococcus dysgalactiae equisimilis GGS 124; Streptococcus equi zooepidemicus MGCS10565; Streptococcus gallolyticus UCN34 uid46061; Streptococcus gordonii Challis subspecies CHI;Streptococcus mutans NN2025 uid46353; Streptococcus mutans; Streptococcus pyogenes Ml GAS; Streptococcus pyogenes MGAS5005; Streptococcus pyogenes MGAS2096; Streptococcus pyogenes MGAS9429; Streptococcus pyogenes MGAS10270; Streptococcus pyogenes MGAS6180; Streptococcus pyogenes MGAS315; Streptococcus pyogenes pyogenes SSI-1; Streptococcus pyogenes MGAS10750; Streptococcus pyogenes NZ131; Streptococcus thermophiles CNRZ1066; Streptococcus thermophiles LMD-9; Streptococcus thermophiles LMG 18311; Clostridium botulinum A3 Loch Maree; Clostridium botulinum BEklund 17B; Clostridium botulinum Ba4 657; Clostridium botulinum F Langeland; Clostridium cellulolyticum cellulolyticum)H10; Finegoldia magna ATCC 29328; Eubacterium rectale ATCC33656; Mycoplasmamagallisepticum; Mycoplasma mobile 163K; Mycoplasma penetrans; Mycoplasma synoviae 53;Streptobacillus moniliformis DSM 12112; Bradyrhizobium BTAil; Nitrobacter hamburgensis X14; Rhodopseudomonas palustris BisB18; Rhodopseudomonas palustris BisB5; Parvibaculum lavamentivorans DS-1; Dinoroseobacter shibae DFL 12; Gluconacetobacter diazotrophicus Pal 5FAPERJ; Gluconacetobacter diazotrophicus Pal 5JGI; Azospirillum B510uid46085; Rhodospirillum rubrum ATCC 11170; Diaphorobacter TPSYuid29975; Verminephrobacter eiseniae EF01-2; Neisseria meningitides 053442; Neisseria meningitides αl4; Neisseria meningitides Z2491; Desulfovibriosalexigens DSM 2638; Campylobacter jejuni doylei 269 97; Campylobacter jejuni 81116; Campylobacter jejuni; Campylobacter lari RM2100; Helicobacter hepaticus; Wolinella succinogenes; Tolumonas auensis DSM9187; Pseudoalteromonas atlantica T6c; Shewanella pealeana ATCC700345; Legionella pneumophila Paris;Actinobacillus succinogenes 130Z; Pasteurella multocida; Francisella tularensis novicida U112; Francisella tularensis holarctica; Francisella tularensis FSC 198; Francisella tularensis tularensis; Francisella tularensis WY96-3418; and Treponema denticola ATCC35405. Thus, aspects of the present disclosure relate to Cas9 proteins present in a type II CRISPR system that render nucleases ineffective or render them nickases as described herein. ;
[0084] In the literature, those skilled in the art may refer to the Cas9 protein as Csn1. The sequence of the S. pyogenes Cas9 protein, which is the subject of the experiments described herein, is shown below. See Deltcheva et al., Nature 471, 602-607 (2011), which is incorporated herein by reference in its entirety.
[0085]
[0086] According to certain aspects of the RNA-guided genome regulation method described herein, Cas9 is changed to reduce, significantly reduce or eliminate nuclease activity. According to one aspect, the Cas9 nuclease activity is reduced, significantly reduced or eliminated by changing the RuvC nuclease domain or the HNH nuclease domain. According to one aspect, the RuvC nuclease domain is inactivated. According to one aspect, the HNH nuclease domain is inactivated. According to one aspect, the RuvC nuclease domain and the HNH nuclease domain are inactivated. According to other aspects, a Cas9 protein is provided, wherein the RuvC nuclease domain and the HNH nuclease domain are inactivated. According to other aspects, a Cas9 protein with an inactive nuclease of the degree of inactivation of the RuvC nuclease domain and the HNH nuclease domain is provided. According to other aspects, a Cas9 nickase is provided, wherein any one of the RuvC nuclease domain or the HNH nuclease domain is inactivated, whereby the remaining nuclease domain is kept active for nuclease activity. In this way, only one chain of double-stranded DNA is cut or nicked.
[0087] According to other aspects, a nuclease-ineffective Cas9 protein is provided, wherein one or more amino acids in Cas9 are changed or otherwise removed to provide a nuclease-ineffective Cas9 protein. According to one aspect, the amino acids include D10 and H840. See Jinke et al., Science 337, 816-821 (2012). According to other aspects, the amino acids include D839 and N863. According to one aspect, one or more of D10, H840, D839 and H863 are replaced with amino acids that reduce, substantially eliminate or eliminate nuclease activity. According to one aspect, one or more or all of D10, H840, D839 and H863 are replaced with alanine. According to one aspect, a Cas9 protein having one or more or all of D10, H840, D839, and H863 substituted with an amino acid that reduces, substantially eliminates, or eliminates nuclease activity (e.g., alanine) is referred to as a nuclease-ineffective Cas9 or Cas9N and exhibits reduced or eliminated nuclease activity, or nuclease activity is absent or substantially absent within a detection level range. According to this aspect, the nuclease activity of Cas9N can be undetectable, i.e., below the detection level of a known assay, using a known assay.
[0088] According to one aspect, the nuclease-ineffective Cas9 protein includes its homologs and orthologs that retain the ability of the protein to bind DNA and be guided by RNA. According to one aspect, the nuclease-ineffective Cas9 protein includes a sequence as described for the naturally occurring Cas9 from Streptococcus pyogenes (S.pyogenes) and having one or more or all of D10, H840, D839 and H863 substituted with alanine, and a protein sequence having at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% homology thereto and being a DNA binding protein, such as an RNA-guided DNA binding protein.
[0089] According to one aspect, the nuclease-ineffective Cas9 protein includes a sequence as described for the naturally occurring Cas9 from Streptococcus pyogenes (S.pyogenes) except the protein sequence of the RuvC nuclease domain and the HNH nuclease domain, and a protein sequence having at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98% or 99% homology thereto and being a DNA binding protein, such as an RNA-guided DNA binding protein. In this way, aspects disclosed herein include protein sequences responsible for DNA binding, for example, responsible for colocalization with the guide RNA and binding to DNA, and protein sequences homologous thereto, and do not need to include the protein sequence of the RuvC nuclease domain and the HNH nuclease domain (to the extent that DNA binding is not required), because these domains can be inactivated or removed from the protein sequence of the naturally occurring Cas9 protein to produce a nuclease-ineffective Cas9 protein.
[0090] For the purpose of this disclosure, Figure 4A Metal coordinating residues in known protein structures with homology to Cas9 are shown. Residues are labeled based on their position in the Cas9 sequence. Left: RuvC structure, PDB ID: 4EP4 (blue) position D7, which corresponds to D10 in the Cas9 sequence, and is highlighted in the Mg-ion coordination position. Middle: Structure of the HNH endonuclease domain from PDB ID: 3M7K (orange) and 4H9D (cyan), including the coordinated Mg-ion (grey sphere) and DNA from 3M7K (purple). Residues D92 and N113 in positions D53 and N77 of 3M7K and 4H9D, which have sequence homology to Cas9 amino acids D839 and N863, are shown as sticks. Right: List of mutants prepared and analyzed for nuclease activity: Cas9 wild type; Cas9m1 with D10 substituted with alanine; Cas9m2 with D10 substituted with alanine and H840 substituted with alanine; Cas9m3 with D10 substituted with alanine, H840 substituted with alanine, and D839 substituted with alanine; and Cas9m4 with D10 substituted with alanine, H840 substituted with alanine, D839 substituted with alanine, and N863 substituted with alanine.
[0091] like Figure 4B As shown, Cas9 mutants m3 and m4 and their respective fusions with VP64 showed undetectable nuclease activity by deep sequencing of the target site. The graph shows mutation frequency relative to genomic position, with the red line demarcating the gRNA target. Figure 4C yes Figure 4BThis figure confirms that the mutation landscape shows a comparable pattern to unmodified sites.
[0092] According to one aspect, an engineered Cas9-gRNA system is provided, which enables RNA-guided genomic regulation in human cells by connecting a transcriptional activation domain to a nuclease-invalid Cas9 or guide RNA. According to one aspect disclosed in the present invention, one or more transcriptional regulatory proteins or domains (these terms are used interchangeably) are combined or otherwise connected to a nuclease-deficient Cas9 or one or more guide RNAs (gRNAs). The transcriptional regulatory domain corresponds to a target site. Therefore, aspects disclosed in the present invention include methods and materials for locating a transcriptional regulatory domain to a target site by fusing, connecting or binding these domains to any one of Cas9N or gRNA.
[0093] According to one aspect, a Cas9N-fusion protein capable of transcriptional activation is provided. According to one aspect, the VP64 activation domain (see Zhang et al., Nature Biotechnology 29, 149-153 (2011), which is incorporated herein by reference in its entirety) is bound, fused, connected or otherwise linked (tethered) to the C-terminus of Cas9N. According to one method, a transcriptional regulatory domain is provided to the site of the target genomic DNA by the Cas9N protein. According to one method, a Cas9N fused to a transcriptional regulatory domain is provided in a cell together with one or more guide RNAs. The Cas9N with a transcriptional regulatory domain fused thereon is bound to or near the target genomic DNA. One or more guide RNAs are bound to or near the target genomic DNA. The transcriptional regulatory domain regulates the expression of the target gene. According to a specific aspect, when bound to a gRNA target sequence near a promoter, the Cas9N-VP64 fusion activates the transcription of a reporter construct, thereby showing RNA-guided transcriptional activation.
[0094] According to one aspect, a gRNA-fusion protein capable of transcriptional activation is provided. According to one aspect, the VP64 activation domain is bound, fused, connected or otherwise linked to the gRNA. According to one method, a transcriptional regulatory domain is provided to a site of a target genomic DNA by a gRNA. According to one method, a gRNA fused to a transcriptional regulatory domain is provided in a cell together with a Cas9N protein. Cas9N is bound to or near the target genomic DNA. One or more guide RNAs having a transcriptional regulatory protein or domain fused thereto are bound to or near the target genomic DNA. The transcriptional regulatory domain regulates the expression of the target gene. According to a specific aspect, the Cas9N protein and the gRNA fused to the transcriptional regulatory domain activate the transcription of the reporter construct, thereby showing RNA-guided transcriptional activation.
[0095] By inserting random sequences into the gRNA and assaying the function of Cas9, regions of the gRNA that will tolerate modification are identified, thereby constructing a gRNA tether capable of transcriptional regulation. gRNAs with random sequence insertions at the 5' end of the crRNA portion of the chimeric gRNA or the 3' end of the tracrRNA portion retain functionality, while insertions in the tracrRNA scaffold portion of the chimeric gRNA result in loss of function. See Figure 5A -B, which summarizes the flexibility of gRNA for random base insertion. Figure 5A It is a schematic diagram of a homologous recombination (HR) assay to determine the activity of Cas9-gRNA. As shown in Figure 5B, the gRNA with a random sequence inserted at the 5' end of the crRNA portion of the chimeric gRNA or the 3' end of the tracrRNA portion retains functionality, while the insertion in the tracrRNA scaffold portion of the chimeric gRNA causes loss of function. The insertion point in the gRNA sequence is shown by red nucleotides. Without wishing to be bound by scientific theory, the increased activity of random base insertion at the 5' end may be caused by the increase in the half-life of a longer gRNA.
[0096] To connect VP64 to the gRNA, two copies of the MS2 phage coat protein binding RNA stem loop are added (attached, appended) to the 3' end of the gRNA. See Fusco et al., Current Biology. CB13, 161-167 (2003), which is incorporated herein by reference in its entirety. These chimeric gRNAs were expressed together with Cas9N and MS2-VP64 fusion proteins. Sequence-specific transcriptional activation from the reporter construct was observed in the presence of all three components.
[0097] Figure 1A Figure 1 is a schematic diagram of RNA-directed transcriptional activation. Figure 1AAs shown, to generate a Cas9N-fusion protein capable of transcriptional activation, the VP64 activation domain was directly linked to the C-terminus of Cas9N. Figure 1B As shown, to generate a gRNA strand capable of transcriptional activation, two copies of the MS2 bacteriophage coat protein-binding RNA stem-loop were added to the 3' end of the gRNA. These chimeric gRNAs were expressed together with Cas9N and MS2-VP64 fusion proteins. Figure 1C The design of the reporter construct for determining transcriptional activation is shown. The two reporters have different gRNA target sites and share a control TALE-TF target site. As shown in Figure 1D, the Cas9N-VP64 fusion shows RNA-guided transcriptional activation, as measured by both fluorescence-activated cell sorting (FACS) and immunofluorescence assay (IF). Specifically, although the control TALE-TF activates two reporters, the Cas9N-VP64 fusion activates the reporter in a gRNA sequence-specific manner. As shown in Figure 1E, only in the presence of all 3 components: Cas9N, MS2-VP64 and gRNA with an appropriate MS2 aptamer binding site, gRNA sequence-specific transcriptional activation from the reporter construct was observed by both FACS and IF.
[0098] According to certain aspects, methods for regulating endogenous genes using Cas9N, one or more gRNAs, and transcriptional regulatory proteins or domains are provided. According to one aspect, the endogenous gene can be any desired gene, referred to herein as a target gene. According to an exemplary aspect, the gene targets for regulation include ZFP42 (REX1) and POU5F1 (OCT4), both of which tightly regulate genes involved in the maintenance of pluripotency. Figure 1F As shown, 10 gRNAs targeting the 5 kb extension of DNA upstream of the transcription start site were designed for the REX1 gene (DNA enzyme hypersensitive sites were highlighted in green). Transcriptional activation was determined using a promoter-luciferase reporter construct (see Takahashi et al., Cell 131 861-872 (2007), which is incorporated herein by reference in its entirety) or directly by qPCR of endogenous genes.
[0099] Fig. 6A -C involves RNA-guided regulation of OCT4 using Cas9N-VP64. Fig. 6A As shown, 21 gRNAs were designed for the OCT4 gene targeting a ~5 kb stretch of DNA upstream of the transcription start site. DNAse hypersensitive sites are highlighted in green. Figure 6B Transcriptional activation using a promoter-luciferase reporter construct is shown. Figure 6C Direct transcriptional activation by qPCR of endogenous genes is shown. Although introduction of a single gRNA modestly stimulated transcription, multiple gRNAs acted synergistically to stimulate robust multi-fold transcriptional activation.
[0100] Fig. 7A -C involves RNA-guided regulation of REX1 using Cas9N, MS2-VP64, and gRNA+2X-MS2 aptamer. Fig. 7A As shown, 10 gRNAs targeting a ~5 kb stretch of DNA upstream of the transcription start site were designed for the REX1 gene. DNAse hypersensitive sites are highlighted in green. Figure 7B Transcriptional activation using a promoter-luciferase reporter construct is shown. Figure 7C Transcriptional activation directly by qPCR of endogenous genes is shown. Although the introduction of a single gRNA moderately stimulates transcription, multiple gRNAs work synergistically to stimulate robust multiple transcriptional activation. In one aspect, the lack of 2X-MS2 aptamers on gRNA does not lead to transcriptional activation. See Maeder et al., Nature Methods 10, 243-245 (2013) and Perez-Pinera et al., Nature Methods 10, 239-242 (2013), each of which is incorporated herein by reference in its entirety.
[0101] Thus, the method involves the use of multiple guide RNAs with Cas9N protein and transcriptional regulatory proteins or domains to regulate the expression of target genes.
[0102] Both Cas9 and gRNA linking methods are effective, with the former showing ~1.5-2 times higher efficacy. In contrast to the 3-component complex assembly (complex assembly, complex assembly), this difference may be due to the need for 2-components. However, in principle, the gRNA linking method enables different gRNAs to recruit different effector domains, as long as each gRNA uses a different RNA-protein interaction pair. See Karyer-Bibens et al., Biology of the Cell / Under the Auspices of the European Cell Biology Organization 100, 125-138 (2008), which is incorporated herein by reference in its entirety. According to one aspect disclosed in the present invention, specific guide RNAs and general Cas9N proteins (i.e., for different target genes, the same or similar Cas9N proteins) can be used to regulate different target genes. According to one aspect, a multiple gene regulation (multiplex gene regulation) method using the same or similar Cas9N is provided.
[0103] The method disclosed in the present invention also involves editing target genes using Cas9N proteins and guide RNAs as described herein to provide multiple genetic engineering and epigenetic engineering (epigenetic engineering) of human cells. Taking Cas9-gRNA targeting as a problem (see Jiang et al., Nature Biotechnology 31, 233-239 (2013), which is incorporated herein by reference in its entirety), a method for deeply interrogating the affinity of Cas9 for target sequence changes in a large space is provided. Therefore, aspects disclosed in the present invention provide direct high-throughput data reading of Cas9 targeting in human cells, while avoiding the complexity of problems caused by dsDNA cutting toxicity and mutagenic repair caused by specific testing of Cas9 using natural nuclease-activity.
[0104] Other aspects of the present disclosure relate to the use of DNA binding proteins or systems generally used for transcriptional regulation of target genes. Based on the present disclosure, those skilled in the art will easily identify exemplary DNA binding systems. These DNA binding systems do not need to have any nuclease activity, such as that of naturally occurring Cas9 proteins. Therefore, these DNA binding systems do not need to inactivate nuclease activity. An exemplary DNA binding system is TALE. As a genome editing tool, TALE-FokI dimers are commonly used, and for genome regulation, TAEL-VP64 fusions have been shown to be highly effective. According to one aspect, the use of Figure 2A The method shown evaluates TALE specificity. A construct library was designed in which each element of the library contained a minimal promoter driving a dTomato fluorescent protein. A 24bp (A / C / G) random transcript tag was inserted downstream of the transcription start site m, and two TF binding sites were arranged upstream of the promoter: one was a constant DNA sequence common to all library elements, and the second was a variable feature of a "biased" binding site library, which was engineered to span a large sequence set of multiple mutation combinations away from the target sequence, and a programmable DNA targeting complex was designed for binding. This was achieved by using engineered degenerate oligonucleotides, which had a nucleotide frequency at each position such that the target sequence nucleotides appeared at a frequency of 79%, and each other nucleotides appeared at a frequency of 7%. See Patwardhan et al., Nature Biotechnology 30, 265-270 (2012), which is incorporated herein by reference in its entirety. Then, the reporter library was sequenced to show the binding between the 24bp dTomato transcript tags and their corresponding "biased" target sites in the library elements. The greater diversity of transcript tags ensures that the commonality of tags between different targets will be extremely rare, and the biased construction of target sequences means that sites with fewer mutations will bind more tags than sites with more mutations. Then, the transcription of the dTomato reporter gene is stimulated by engineering the control-TF that binds to the common DNA site or engineering the target-TF that binds to the target site. By performing RNAseq on the stimulated cells, the abundance of each expressed transcript tag in each sample is measured, and then the previously established binding table is used to map it back to their corresponding binding sites. It is expected that control-TF will excite all library members to the same extent, because all library elements have their binding sites in common, and it is expected that the distribution of members expressed by target-TF is biased towards those that it preferentially targets. By dividing the label counts obtained for target-TF by those label counts obtained for control-TF (control-TF, control-TF), this assumption is used in step 5 to calculate the normalized expression level of each binding site.
[0105] like Figure 2B As shown in Figure 2, the targeting profile of the Cas9-gRNA complex shows that, on average, it tolerates 1-3 mutations in its target sequence. Figure 2C As shown, the Cas9-gRNA complex is also largely insensitive to point mutations, except for those that map to the PAM sequence. Notably, the data show that the predicted PAM for S. pyogenes Cas9 is not only NGG, but also NAG. Figure 2DAs shown, the introduction of 2 base mismatches significantly impaired the activity of the Cas9-gRNA complex, however this was only when these were positioned 8-10 bases near the 3' end of the gRNA target sequence (in the heat map, the target sequence positions are marked from 1 to 23, starting from the 5' end).
[0106] The mutation tolerance of another widely used genome editing tool, the TALE domain, was determined using the transcription-specific assay described here. Figure 2E As shown in , the TALE off-targeting data for the 18-mer TALE showed that on average it could tolerate 1-2 mutations in its target sequence, but was unable to activate most of the 3-base mismatch variants in its target. Figure 2F As shown in Figure 2, similar to the Cas9-gRNA complex, the 18-mer TALE is largely insensitive to single base mismatches in its target. Figure 2G As shown, the introduction of a 2-base mismatch significantly impairs the activity of the 18-mer TALE. TALE activity is more sensitive to mismatches closer to the 5' end of its target sequence (in the heat map, the target sequence positions are marked from 1 to 18 starting from the 5' end).
[0107] The results were confirmed using targeting experiments in nuclease assays, which are the subject of Figures 10A-C to evaluate the targeting profiles of TALEs of different sizes. As shown in Figure 10A, the 18-mer TALE was confirmed to tolerate multiple mutations in its target sequence using a nuclease-mediated HR assay. Fig. 10B As shown, the target profiles of TALEs of three different sizes (18-mer, 14-mer, and 10-mer) were analyzed using the method described in Figure 2. Shorter TALEs (14-mer and 10-mer) are progressively more specific in their targeting, but are also almost an order of magnitude lower in activity. Fig. 10C and 10DAs shown, the 10-mer TALE shows a mismatch resolution close to a single base, which loses almost all activity against targets with 2 mismatches (in the heat map, starting from the 5' end, the target sequence positions are marked from 1 to 10). In general, these data indicate that shorter TALEs designed for engineering can achieve higher specificity in genome engineering applications, and the requirement for FokI dimerization in TALE nuclease applications is necessary to avoid off-target effects. See Kim et al., Proceedings of the National Academy of Sciences of the United States of America 93, 1156-1160 (1996) and Pattanayak et al., Nature Methods 8, 765-770 (2011), each of which is incorporated herein by reference in its entirety.
[0108] Fig. 8A -C relates to a high level specific analysis process for calculating the normalized expression levels illustrated by the examples from the experimental data. Fig. 8AAs shown, a construct library is generated by the biased distribution of binding site sequences and the random sequence 24bp tag (top) that will be introduced into the reporter gene transcript. The transcribed tags are highly degenerate, so that they should draw many-to-one diagrams for Cas9 or TALE binding sequences. The construct library is sequenced (the third level, left figure) to establish a tag that co-occurs with the binding site, thereby resulting in the generation of a binding site relative to the transcription tag (the 4th level, left figure) association table. The library barcode (represented in this article with light blue and light yellow; Levels 1-4, left figure) can be used immediately for sequencing multiple construct libraries constructed by different binding sites. Then, the construct library is transfected into a cell population, and a group of different Cas9 / gRNA or TALE transcription factors (the 2nd level, right figure) are induced in the population sample. A sample (highest level, green frame) is always induced with a fixed TALE activator that fixes the binding site sequence in the targeting construct; the sample is used as a positive control (green sample, also represented by a + symbol). The cDNA generated from the reporter mRNA molecules in the induced samples is then sequenced and analyzed to obtain the tag counts for each tag in the sample (levels 3 and 4, right). As with construct library sequencing, multiple samples (including positive controls) are sequenced and analyzed by appending sample barcodes. Here, light red represents a non-control sample that has been sequenced and analyzed with a positive control (green). Since only transcription tags appear in each reading and construct binding sites do not appear, the total tag counts expressed for each binding site in each sample are calculated relative to the tag binding table using the binding sites obtained from construct library sequencing (level 5). The counts (scores, tallies) for each non-positive control sample are then converted to normalized expression levels for each binding site by dividing them by the tags (scores, tallies) obtained in the positive control sample. In Figure 2B and 2E and in Fig. 9A and Fig. 10B An example of a plot of normalized expression levels by number of mismatches is provided in . In this overall processing pipeline, some level of filtering for incorrect tags, tags that are not bindable to the construct library, and tags that are clearly shared with multiple binding sites is not covered. Figure 8B Example distributions of the percentage of binding sites for the number of mismatches generated within a biased construct library are shown. Left: theoretical distribution. Right: observed distribution from a real TALE construct library. Figure 8CExample distributions of the percentage of tag counts aggregated to the binding site for the number of mismatches are shown. Left: Distribution observed from the positive control sample. Right: Distribution observed from the sample in which a non-control TALE (non-control TALE) was induced. Since the positive control TALE binds to a fixed site in the construct, the distribution of aggregated tag counts closely reflects Figure 8B : Distribution of binding sites in the target-TF, while for non-control TALE samples the distribution is skewed to the left due to higher expression levels resulting from sites with fewer mismatches. Bottom: The relative enrichment between these was calculated by dividing the tag counts obtained for the target-TF by those obtained for the control-TF, which shows the average expression level relative to the number of mutations in the target site.
[0109] Specificity data generated using different Cas9-gRNA complexes reconfirmed these results. Fig. 9A As shown, different Cas9-gRNA complexes tolerated 1-3 mutations in their target sequences. Fig. 9B As shown, the Cas9-gRNA complex is also largely insensitive to point mutations, except those localized to the PAM sequence. Fig. 9C As shown in Figure 2, however, the introduction of a 2-base mismatch significantly impaired activity (in the heat map, target sequence positions are marked from 1 to 23 starting from the 5' end). Fig.9D As shown, the predicted PAM for S. pyogenes Cas9 was confirmed to be NGG, as well as NAG, using a nuclease-mediated HR assay.
[0110] According to certain aspects, binding specificity is improved according to the method described herein. Since the synergy between multiple complexes is a factor in the activation of Cas9N-VP64 target genes, the transcriptional regulation application of Cas9N naturally has considerable specificity, because a single off-target binding event should have a minimal impact. According to one aspect, an offset nick is used in the genome-editing method. Most nicks rarely lead to NHEJ events (see Certo et al., Nature Methods 8, 671-676 (2011), which is incorporated herein by reference in its entirety), so the impact of off-target nicks is minimized. On the contrary, for causing gene disruption, it is very effective to cause an offset nick to produce a double-strand break (DSB). According to certain aspects, compared with 3' protrusions, 5' protrusions produce more significant NHEJ events. Similarly, compared to NHEJ events, 3' protrusions are conducive to HR, although when 5' protrusions are produced, the total number of HR events is significantly lower. Thus, methods are provided for using nicking for homologous recombination and using offset nicking to generate double-strand breaks to minimize the effects of off-target Cas9-gRNA activity.
[0111] Figure 3A -C involves multiple off-set nicking and the use of guide RNA to reduce off-target binding. Figure 3A As shown, once a targeted nick or break is introduced, a traffic light reporter is used to simultaneously measure HR and NHEJ events. DNA cleavage events resolved by the HDR pathway restore the GFP sequence, while mutagenic NHEJ results in a frameshift, placing GFP out of frame and the downstream mCherry sequence in frame. For this assay, 14 gRNAs covering a 200bp DNA stretch were designed: 7 targeting the sense strand (U1-7) and 7 targeting the antisense strand (D1-7). Using the Cas9D10A mutant that creates a nick on the complementary strand, different two-way combinations of gRNAs were used to cause a range of programmed 5' or 3' overhangs (the sites where the 14 gRNAs create nicks are indicated). As Figure 3B As shown, causing an offset nick to generate a double-strand break (DSB) is very effective for causing gene disruption. Notably, an offset nick that results in a 5' overhang, as opposed to a 3' overhang, results in more NHEJ events. Figure 3C As shown, creating a 3' overhang also favors the ratio of HR to NHEJ events, but the total number of HR events is significantly lower when a 5' overhang is created.
[0112] Fig.11A -B involves Cas9D10A nickase-mediated NHEJ. Fig.11A As shown, once a targeted nick or double-strand break is introduced, a traffic light reporter is used to measure NHEJ events. Briefly, once a DNA cleavage event is introduced, if the break is caused by mutagenic NHEJ, GFP is translated out of frame, while the downstream mCherry sequence is in frame, resulting in red fluorescence. 14 gRNAs covering a 200bp DNA stretch were designed: 7 targeting the sense strand (U1-7) and 7 targeting the antisense strand (D1-7). As Fig. 11B As shown, it was observed that unlike wild-type Cas9, which resulted in the generation of DSBs and robust NHEJ in all targets, the majority of nicks (using the Cas9D 10A mutant) resulted in few NHEJ events. All 14 sites were located within a contiguous 200 bp DNA stretch, and more than 10-fold differences in targeting efficacy were observed.
[0113] According to certain aspects, described herein is a method for regulating the expression of a target nucleic acid in a cell, comprising introducing one or more, two or more, or multiple exogenous nucleic acids into the cell. The exogenous nucleic acid introduced into the cell encodes one or more guide RNAs, one or more nuclease-invalid Cas9 proteins, and transcriptional regulatory factor proteins or domains. Generally speaking, guide RNAs, nuclease-invalid Cas9 proteins, and transcriptional regulatory factor proteins or domains are referred to as colocalization complexes, so that those skilled in the art understand the term as the extent to which guide RNAs, nuclease-invalid Cas9 proteins, and transcriptional regulatory factor proteins or domains bind to DNA and regulate the expression of target nucleic acids. According to certain other aspects, the exogenous nucleic acid introduced into the cell encodes one or more guide RNAs and Cas9 protein nickases. Generally speaking, guide RNA Cas9 protein nickases are referred to as colocalization complexes, so that those skilled in the art understand the term as the extent to which guide RNAs and Cas9 protein nickases bind to DNA and cause the target nucleic acid to produce nicks.
[0114] Cells disclosed in accordance with the present invention include any cells in which exogenous nucleic acids can be introduced and expressed as described herein. It should be understood that the basic concepts disclosed in accordance with the present invention as described herein are not limited by cell types. Cells disclosed in accordance with the present invention include eukaryotic cells, prokaryotic cells, animal cells, plant cells, fungal cells, archaeal cells, true bacterial cells, etc. Cells include eukaryotic cells, such as yeast cells, plant cells and animal cells. Specific cells include mammalian cells. In addition, cells include any cells in which it is beneficial or desirable to regulate target nucleic acids. These cells may include those in which specific proteins are not expressed enough, thereby causing diseases or harmful conditions. Those skilled in the art are readily aware of these diseases or harmful conditions. According to the present invention, nucleic acids responsible for expressing specific proteins can be targeted by the methods described herein and transcriptional activators, thereby causing the target nucleic acids to be upregulated and the corresponding expression of specific proteins. In this way, the methods described herein provide therapeutic treatments.
[0115] The target nucleic acid includes any nucleic acid sequence that the colocalization complex as described herein can be used to regulate or produce a nick. The target nucleic acid includes a gene. For the purposes disclosed in the present invention, a DNA (such as double-stranded DNA) can include a target nucleic acid, and the colocalization complex can be bound to or colocalized with the DNA to the target nucleic acid or adjacent to or near it, and in which the colocalization complex can have a desired effect on the target nucleic acid. These target nucleic acids can include endogenous (or naturally occurring) nucleic acids and exogenous (or foreign) nucleic acids. Based on the disclosure of the present invention, a technician will be able to easily identify or design a guide RNA and Cas9 protein colocalized to DNA (including target nucleic acids). A technician will be able to further identify transcriptional regulatory factor proteins or domains that are also colocalized to DNA (including target nucleic acids). DNA includes genomic DNA, mitochondrial DNA, viral DNA or exogenous DNA.
[0116] Foreign nucleic acids (i.e., those that are not part of the cell's native nucleic acid composition) can be introduced into cells using any method known to those skilled in the art for such introduction. These methods include transfection, transduction, viral transduction, microinjection, lipofection, nucleofection, nanoparticle bombardment, transformation, conjugation, and the like. Using readily identifiable literature sources, those skilled in the art will readily understand and adapt these methods.
[0117] Transcriptional regulatory factor proteins or domains that serve as transcriptional activators include VP16 and VP64, etc., which can be easily identified by those skilled in the art based on the disclosure of the present invention.
[0118] Diseases and harmful conditions are those characterized by abnormal loss of expression of specific proteins. These diseases or harmful conditions can be treated by upregulation of specific proteins. Therefore, a method for treating a disease or harmful condition is provided, wherein the colocalization complex as described herein associates or otherwise binds to DNA (including target nucleic acids), and the transcriptional activator of the colocalization complex upregulates the expression of the target nucleic acid. For example, upregulation of PRDM16 and other genes that promote brown fat differentiation and increased metabolic absorption can be used to treat metabolic syndrome or obesity. Activation of anti-inflammatory genes is useful for autoimmune and cardiovascular diseases. Activation of tumor suppressor genes is useful for treating cancer. Based on the disclosure of the present invention, those skilled in the art will easily identify these diseases and harmful conditions.
[0119] The following examples are described as representative of the present disclosure. These examples should not be viewed as limiting the scope of the present disclosure, as these and other equivalent embodiments will be apparent with reference to the present disclosure, the drawings, and the appended claims.
[0120] Embodiment 1
[0121] Cas9 mutants
[0122] Sequences with known structures homologous to Cas9 were investigated to identify candidate mutations in Cas9 that can remove the natural activity of its RuvC and HNH domains. Using HHpred (world wide web toolkit.tuebingen.mpg.de / hhpred), the full length sequence of Cas9 was queried for the full protein database (full Protein Data Bank) (January 2013). The query returned two different HNH endonucleases with significant sequence homology to the HNH domain of Cas9: PacI and hypothetical endonucleases (PDB IDs are: 3M7K and 4H9D, respectively). These proteins were checked to find residues involved in magnesium ion coordination. Then, the corresponding residues were identified in the sequence comparison with Cas9. In each structure, two Mg-coordinated side chains were identified, which were aligned with the same amino acid type in Cas9. They are 3M7K D92 and N113, and 4H9D D53 and N77. These residues correspond to Cas9 D839 and N863. It was also reported that mutations of PacI residues D92 and N113 to alanine render the nuclease catalytically inadequate. Based on this analysis, Cas9 mutations D839A and N863A were made. In addition, HHpred also predicted homology between the N-terminus of Cas9 and Thermus thermophilus RuvC (PDB ID: 4EP4). The sequence alignment covers the previously reported mutation D10A that eliminates the function of the RuvC domain in Cas9. To confirm it as a suitable mutation, the metal binding residues were determined as described above. In 4EP4, D7 helps coordinate magnesium ions. This position has sequence homology corresponding to Cas9 D10, confirming that the mutation helps eliminate metal binding and, therefore, catalytic activity from the Cas9 RuvC domain.
[0123] Example II
[0124] Plasmid construction
[0125] Cas9 mutants were generated using the Quikchange kit (Agilent technologies). Target gRNA expression constructs (1) were ordered directly from IDT as single gBlocks and cloned into pCR-Bluntll-TOPO vectors (Invitrogen); or (2) were synthesized by Genewiz routine (custom); or (3) were spliced into gRNA cloning vectors (plasmid #41824) using Gibson assembly of oligonucleotides. Vectors for HR reporter assays involving GFP with breakage were constructed by fusion PCR assembly of GFP sequences with stop codons and appropriate fragments spliced into EGIP lentiviral vectors from Addgene. These lentiviral vectors were then used to establish a GFP reporter stable system. TALEN used in this study was constructed using standard protocols. See Sanjana et al., Nature Protocols 7, 171-192 (2012), which is incorporated herein by reference in its entirety. Cas9N and MS2 VP64 were fused using standard PCR fusion protocol procedures. Promoter luciferase constructs for OCT4 and REX1 were obtained from Addgene (plasmid #17221 and plasmid #17222).
[0126] Example III
[0127] Cell culture and transfection
[0128] HEK 293T cells were cultured in high glucose Dulbecco's modified Eagle's medium (DMEM, Invitrogen) supplemented with 10% fetal bovine serum (FBS, Invitrogen), penicillin / streptomycin (pen / strep, Invitrogen) and non-essential amino acids (NEAA, Invitrogen). The cells were maintained at 37°C and 5% CO2 in a humidified incubator.
[0129] Transfections for nuclease assays were performed as follows: 0.4 × 10 cells were transfected with 2 μg of Cas9 plasmid, 2 μg of gRNA, and / or 2 μg of DNA donor plasmid using Lipofectamine 2000 according to the manufacturer's protocol. 6 Cells were harvested 3 days after transfection and analyzed by FACS, or for direct determination of genomic cleavage, ~1 × 10 6The genomic DNA of each cell was collected. For these assays, PCR was performed to amplify the target region by genomic DNA derived from the cells, and the amplicons were deeply sequenced by a MiSeq Personal Sequencer (Illumina) with a coverage of >200,000 reads. The sequencing data was analyzed to estimate the NHEJ efficiency.
[0130] For transfections involving transcriptional activation assays: 0.4 × 10 cells were transfected with (1) 2 μg Cas9N-VP64 plasmid, 2 μg gRNA, and / or 0.25 μg reporter construct; or (2) 2 μg Cas9N plasmid, 2 μg MS2-VP64, 2 μg gRNA-2XMS2 aptamer, and / or 0.25 μg reporter construct. 6 cells. Cells were harvested 24-48 hours after transfection and assayed using FACS or immunofluorescence, or their total RNA was extracted and the extracts were subsequently analyzed by RT-PCR. In this paper, standard taqman probes for OCT4 and REX1 from Invitrogen were used, and for each sample, GAPDH was normalized.
[0131] For transfection of transcriptional activation assays involving Cas9-gRNA complexes and TALE-specific profiles: 0.4×10 cells were transfected with (1) 2 μg Cas9N-VP64 plasmid, 2 μg gRNA, and 0.25 μg reporter library; or (2) 2 μg TALE-TF plasmid and 0.25 μg reporter library; or (3) 2 μg control-TF plasmid and 0.25 μg reporter library. 6 Cells were harvested 24 hours after transfection (to avoid stimulation of the reporter in saturated form). Total RNA was extracted using RNAeasy-plus kit (Qiagen), and standard RT-PCR was performed using Superscript-III (Invitrogen). Libraries for next generation sequencing were generated by targeted PCR amplification of transcript-tags.
[0132] Example IV
[0133] Computational and sequence analysis for Cas9-TF and TALE-TF reporter expression level calculations
[0134] Fig. 8A A high-level logical flow of this process is shown in , and additional details are provided herein. For details on construct library composition, see Fig. 8A (Level 1) and 8B.
[0135] Sequencing:For Cas9 experiments, construct libraries were obtained as 150 bp overlapping paired-end reads on an Illumina MiSeq ( Fig. 8A , level 3, left) and reporter gene cDNA sequence ( Fig. 8A , level 3, right panel), while for TALE experiments the corresponding sequences were obtained as 51 bp non-overlapping paired-end reads on Illumina MiSeq.
[0136] Construct library sequence processing: Alignment: For Cas9 experiments, paired reads were aligned to a set of 250 bp reference sequences corresponding to 234 bp constructs flanked by paired 8 bp library barcodes using novoalign V2.07.17 (novocraft.com / main / index / php). Fig. 8A , level 3, left). In the reference sequence provided by novoalign, the 23 bp degenerate Cas9 binding site region and the 24 bp degenerate transcript tag region (see Fig. 8A , level 1) is called Ns, and the construct library barcodes are explicitly provided. For TALE experiments, the same procedure was used except that the reference sequence length was 203bp and the length of the degenerate binding site region was 18bp (relative to 23bp). Validity check: Novoalign output for the included files in which the left and right reads of each read pair are aligned to the reference sequence respectively. Additional validity conditions are applied to the only read pairs that are uniquely aligned to the reference sequence, and the only read pairs that pass all these conditions are retained. Validity conditions include: (i) Each of the two construct library barcodes must be aligned to the reference sequence barcode at least 4 positions, and for the same construct library, the two barcodes must be aligned to the barcode. (ii) All bases aligned to the N region of the reference sequence must be called As, Cs, Gs, Gs or Ts by novoalign. Note: In the reference N region, neither Cas9 nor TALE experiments have left and right read overlaps, so there is no possibility of ambiguous novoalign calls for these N bases. (iii) Likewise, no insertions or deletions called by novoalign must be present in these regions. (iv) No Ts must be present in the transcript tag regions (since these random sequences are generated only from As, Cs, and Gs). Read pairs that violate any of these conditions are collected in a read pair rejection file. These validity checks are implemented using a conventional (custom) perl script.
[0137] Inducible sample reporter gene cDNA sequence processing: Alignment: Overlapping read pairs were first merged into a 79 bp common segment using SeqPrep (downloaded from the World Wide Web at github.com / jstjohn / SeqPrep), and these 79 bp common segments were then aligned as unpaired single reads to a set of reference sequences using novoalign (see Fig. 8A , level 3, right), where (for construct library sequences) the 24 bp degenerate transcript tag is referred to as Ns, while the sample barcode is explicitly provided. Both TALE and Cas9 cDNA sequence regions correspond to the same 63 bp cDNA region flanked by paired 8 bp sample barcode sequences. Validity check: The same conditions used for construct library sequencing were applied (see above) except that: (a) in the present invention, due to the prior SeqPrep merging of read pairs, the validity processing does not have to filter unique alignments of both reads in a read pair, but only unique alignments of the merged reads; (b) only transcript tags appear in cDNA sequence reads, so the validity processing only applies to these tag regions of the reference sequence and not to individual binding site regions.
[0138] Compilation of a table of binding sites relative to transcript tag binding: These tables were generated from validated construct library sequences using regular perl ( Fig. 8A , level 4, left panel). Although the 24 bp tag sequence consisting of A, C, and G bases should be essentially unique across the entire construct library (shared probability = ~2.8e-11), early analysis of binding sites versus tag binding revealed that in fact multiple binding sequences share a non-negligible portion of the tag sequence, which may be primarily caused by a combination of sequence errors in the binding sequence or oligo synthesis errors in the oligos used to generate the construct library. In addition to tag sharing, tags found to be bound to binding sites in validated read pairs can also be present in the construct library read pair rejection file if they are unclear due to barcode mismatches in the construct library from which they may have come. Finally, the tag sequence itself may contain sequence errors. To address these sources of error, tags are classified using three attributes: (i) safe versus unsafe, where unsafe indicates that a tag can be present in the construct library rejected read pair file; shared versus non-shared, where shared indicates that a tag is found to be associated with multiple binding site sequences, and 2+ versus only 1-, where 2+ indicates tags that appear at least twice in the validated construct library sequence and are therefore considered unlikely to contain sequence errors. Combining these three criteria yielded eight classes of tags associated with each binding site, with the safest (but least abundant) class containing only safe, non-shared, 2+ tags; and the least safe (but most abundant) class containing all tags regardless of safety, share, or number of occurrences.
[0139] Calculation of normalized expression levels: Implemented using regular Perl code Fig. 8A , steps indicated in levels 5-6. First, using the binding site relative to transcript tag table previously calculated for the construct library, the sum of tag counts obtained for each induced sample is calculated (aggregated) for each binding site (see Figure 8C ). For each sample, the sum of the tag counts for each binding site was then divided by the sum of the tag counts for the positive control sample to produce a normalized expression level. Other considerations for these calculations include:
[0140] 1. For each sample, a subset of "novel" tags were found in the validity-checked cDNA gene sequences that could not possibly exist in the binding site-to-transcript tag binding table. These tags were ignored in subsequent calculations.
[0141] 2. In the binding site relative to transcript tag binding table, the tag counts are aggregated as described above for each of the eight categories of tags as described above. Because binding site bias in the construct library frequently produces sequences similar to the central sequence, but the number of sequences with an increased number of mismatches gradually decreases, binding sites with fewer mismatches are generally clustered into a large number of tags, while binding sites with more mismatches are clustered into a smaller number. Therefore, although it is generally desirable to use the safest tag species, binding sites with two or more mismatches can be evaluated based on the small number of tags per binding site, making safety counts and ratios less statistically reliable, even if the tags themselves are more reliable. In these cases, all tags were used. For this reason, some compensation was obtained from the fact that for n mismatch positions, the number of tag counts for a separate set grows with the number of combinations of mismatch positions (equal to ). ), and the compensation increases significantly with n; therefore, for different numbers of mismatches n, the average value of the tag counts of the set (as shown in Figures 2b, 2e and Fig. 9A and 10B (shown) based on label counts for a collection of statistically very large groups when n ≥ 2.
[0142] 3. Finally, the binding sites built into the TALE construct library were 18 bp and tag binding was assigned based on these 18 bp sequences, but some experiments were performed using TALEs that were programmed to bind to the central 14 bp or 10 bp region within the 18 bp construct binding site region. When calculating expression levels for these TALEs, tags were grouped into the binding site based on the corresponding region of the 18 bp binding site in the binding table, thereby ignoring binding site mismatches outside of this region.
[0143] Example V
[0144] Using Cas9 N- RNA-guided SOX2 and NANOG regulation of VP64
[0145] The sgRNA (aptamer-modified single guide RNA) tethering method described herein enables different sgRNAs to recruit different effector domains, as long as each sgRNA uses a different RNA-protein interaction pair, thus enabling multiplexed gene regulation using the same Cas9N-protein. Fig. 12A SOX2 and Fig. 12B NANOG gene, 10 gRNAs were designed targeting a DNA stretch of ~1 kb upstream of the transcription start site. Hypersensitive sites of DNA enzymes are highlighted in green. Transcriptional activation by qPCR of endogenous genes was measured. In both cases, although the introduction of a single gRNA moderately stimulated transcription, multiple gRNAs acted synergistically to stimulate robust multifold transcriptional activation. Data are mean + / - SEM (N = 3). Fig. 12A -B, two other genes, SOX2 and NANOG, were targeted and regulated by sgRNAs within a ~1 kb stretch upstream of the promoter DNA. sgRNAs placed adjacent to the transcription start site resulted in robust gene activation.
[0146] Example VI
[0147] Evaluating the targeting profile of Cas9-gRNA complexes
[0148] Using the method described in Figure 2, two other Cas9-gRNA complexes were analyzed ( Fig.13A -C) and ( Fig.13D -F) targeting profile. The two gRNAs have very different specificity profiles, with gRNA2 tolerating up to 2-3 mismatches and gRNA3 tolerating only up to 1. Fig. 13B , 13E ) and two base mismatch maps ( Fig. 13C , 13F These aspects are reflected in Fig. 13C and 13F In the figure, gray boxes containing "x" indicate base mismatches for which the available data are insufficient to calculate the normalized expression level. To improve data display, yellow boxes containing asterisks "*" indicate mismatches for which the normalized expression level beyond the top of the color scale is an outlier. Statistical significance symbols are: *** for p < .0005 / n; ** for P < .005 / n; * for p < .05 / n; NS (non-significant) for P > = .05 / n, where n is the number of comparisons (see Table 2).
[0149] Example VII
[0150] Validation, specificity of reporter assay
[0151] like Fig.14A Specificity data were generated using two different sgRNA:Cas9 complexes as shown in Figure 3-C. This confirms that the assay is specific for the sgRNA being evaluated, as the corresponding mutant sgRNA is unable to stimulate the reporter library. Fig.14A : Specificity profiles of two gRNAs (wild-type and mutant; sequence differences highlighted in red) were evaluated using a reporter library designed against the wild-type gRNA target sequence. Fig. 14B : Confirm that the assay is specific for the gRNA being evaluated (based on Fig.13D Replotted data), because the corresponding mutant gRNA cannot stimulate the reporter library. Statistical significance symbols are: for p < .0005 / n, ***; for P < .005 / n, **; for p < .05 / n, *; for P > = .05 / n, NS (non-significant), where n is the number of comparisons (see Table 2). Different sgRNAs can have different specificity profiles ( Fig.13A , 13D ), specifically, sgRNA2 tolerated up to 3 mismatches, while sgRNA3 tolerated only up to 1. The greatest sensitivity to mismatches was localized to the 3' end of the spacer, although mismatches were also observed to affect activity at other positions.
[0152] Example VIII
[0153] Verification, single-base and double-base gRNA mismatches
[0154] like Fig.15A -D, targeting was confirmed by assaying that single base mismatches within 12 bp of the 3' end of the spacer in the sgRNA resulted in detectable targeting. However, a 2 bp mismatch within this region resulted in a significant loss of activity. Using the nuclease assay, 2 independent gRNAs were tested: gRNA 2 ( Fig.15A -B) and gRNA3( Fig. 15C -D). It was confirmed that a single base mismatch within 12 bp of the 3' end of the spacer in the gRNAs tested resulted in detectable targeting, whereas a 2 bp mismatch within this region resulted in a rapid loss of activity. Consistent with the results of Figure 13, these results also highlight the differences in specificity profiles between different gRNAs. Data are mean + / - SEM (N = 3).
[0155] Example IX
[0156] Verification, 5' gRNA truncation
[0157] like Fig.16A -D, truncations in the 5' portion of the spacer resulted in retained sgRNA activity. Using the nuclease assay, two independent gRNAs were tested: gRNA1 ( Fig.16A -B) and gRNA3( Fig. 16C -D). It was observed that 1-3 bp 5' truncations were well tolerated, but larger deletions resulted in a loss of activity. Data are mean + / - SEM (N=3).
[0158] Example X
[0159] Validation, Streptococcus pyogenes (S.pyogenes) PAM
[0160] like Fig.17A -B shows that the PAM of Streptococcus pyogenes (S.pyogenes) Cas9 was confirmed to be NGG and NAG using a nuclease-mediated HR assay. Data are mean + / - SEM (N = 3). According to additional studies, a set of approximately 190K Cas9 targets generated in human exons that do not contain alternating NGG targets in the last 13nt of the common target sequence was scanned to determine whether there are alternating NAG sites or NGG sites with mismatches in the previous 13nt. Only 0.4% were found to have no such alternating targets.
[0161] Example XI
[0162] Validation, TALE mutation
[0163] Using nuclease-mediated HR assay ( Fig.18A -B) Confirmation that the 18-mer TALE tolerates multiple mutations in its target sequence. Fig.18A -B, certain mutations in the middle of the target resulted in higher TALE activity, as determined by targeted experiments in a nuclease assay.
[0164] Example XII
[0165] TALE monomer specificity versus TALE protein specificity
[0166] To remove the role of variable diresidues (RVDs) in individual repeat sequences, we confirmed that the choice of RVDs does contribute to base specificity, but overall TALE specificity is still a function of protein binding energy. Fig.19A -C shows the comparison of TALE monomer specificity relative to TALE protein specificity. Fig.19A : Using the modification method described in Figure 2, the target maps of two 14-mer TALE-TFs with adjacent groups of 6 NI or 6 NH repeats were analyzed. In this method, a reduced reporter library with a degenerate 6-mer sequence in the middle was generated and used to determine the TALE-TF specificity. Figure 19B-C: In both cases, it was noted that the expected target sequence was enriched (i.e., for NI repeats, sequences with 6 As, and for NH repeats, sequences with 6 Gs). Each of these TALEs still tolerates 1-2 mismatches in the central 6-mer target sequence. Although the choice of monomers does contribute to base specificity, TALE specificity is generally a function of protein binding energy. According to one aspect, shorter engineered TALEs or TALEs with high and low affinity monomer compositions in genome engineering applications result in higher specificity, and when shorter TALEs are used, FokI dimerization in nuclease applications enables further reduction of off-target effects.
[0167] Example XIII
[0168] Offset incision, natural site
[0169] Fig. 20A -B shows data related to offset nicking. In the context of genome editing, offset nicking is produced to generate DSBs. Most nicking does not result in non-homologous end joining (NHEJ)-mediated indels, and therefore when offset nicking is caused, off-target single nicking events will likely result in extremely low indel rates. It is effective to cause offset nicking to produce DSBs for causing gene disruption at two merged reporter sites and at the native AAVS1 genomic site. Fig. 20A : Targeting the native AAVS1 locus with 8 gRNAs covering a 200 bp DNA stretch: 4 targeting the sense strand (s1-4), 4 targeting the antisense strand (as1-4). Using the Cas9D10A mutant that nicks on the complementary strand, different two-way combinations of gRNAs were used to induce a range of programmed 5' or 3' overhangs. Fig. 20B :Using the mensuration based on Sanger sequencing, it is observed that although a single gRNA does not cause detectable NHEJ events, it is very effective to cause a skewed cut to produce DSB for causing gene disruption. It is worth noting that, contrary to 3' protrusions, the skewed cut that causes 5' protrusions leads to more NHEJ events. The number of Sanger sequencing clones is highlighted above the bar, and the predicted protruding length is shown below the corresponding X-axis legend.
[0170] Example XIV
[0171] Offset nicking, NHEJ spectrum
[0172] Fig.21A -C involves offset nicking and NHEJ spectra. Representative Sanger sequencing results for three different offset nicking combinations are shown, with the location of the targeting gRNA highlighted by a box. In addition, consistent with the standard model of homologous recombination (HR)-mediated repair, engineering of 5' overhangs by offset nicking resulted in more robust NHEJ events than 3' overhangs ( Figure 3B In addition to NHEJ stimulation, robust HR induction was observed when 5' overhangs were generated. Generation of 3' overhangs did not result in an improvement in HR rates ( Figure 3C ).
[0173] Example XV
[0174] Table 1
[0175] gRNA targets for endogenous gene regulation
[0176] Targets in the REX1, OCT4, SOX2, and NANOG promoters used in the Cas9-gRNA-mediated activation experiments are listed.
[0177]
[0178] Example XVI
[0179] Table 2
[0180] Summary of statistical analysis of Cas9-gRNA and TALE specificity data
[0181] Table 2 (a) is used to compare the P-values of the normalized expression levels of TALE or Cas9-VP64 activators bound to target sequences with a specific number of target site mutations. The normalized expression levels are shown by the box plots in the figures specified in the "Figure" column, where the boxes represent the distribution of these levels by the number of mismatches in the target site. The P-value is calculated for each consecutive pair of mismatch numbers in each box plot using a t-test, where the t-test is a one-sample or two-sample t-test (see Methods). Statistical significance is evaluated using a Bonferroni-corrected P-value threshold, where the correction is based on the number of comparisons within each box plot. Statistical significance symbols are: *** for p<.0005 / n; ** for P<.005 / n; * for p<.05 / n; NS (non-significant) for P>=.05 / n, where n is the number of comparisons. Table 2 (b) Figure 2DStatistical characteristics of the seed region in the 20bp target site: log10 (P-value) shows the separation between the expression values of Cas9N VP64+gRNA bound to the target sequence with two mutations for those positions that are mutated in the candidate seed region at the 3' end of the 20bp target site relative to all other positions. The greatest separation was found in the last 8-9bp of the target site, which is represented by the maximum -log10 (P-value) (highlighted above). These positions can be understood as representing the start of the "seed" region of the target site. For information on how to calculate the P-value, see the section "Statistical characteristics of the seed region" in the method.
[0182]
[0183] Example XVII
[0184] Protein and RNA sequences in the examples
[0185] A. The following shows the Cas9 based on the m4 mutant N- Sequences of VP64 activator constructs. Cas9 with the highest activity m4 VP64 and Cas9 m4 VP64 Three forms of N fusion protein were constructed. The corresponding vectors of m3 and m2 mutants were also constructed ( Figure 4A ) (NLS and VF64 domains are highlighted).
[0186] >Cas9 m4 VP64
[0187]
[0188]
[0189]
[0190] >Cas9 m4 VP64 N sequence
[0191]
[0192]
[0193] >Cas9 m4 VP64 C
[0194]
[0195]
[0196]
[0197] B. Provided below are the sequences of MS2-activator constructs with 2X MS2 aptamer domains and the corresponding gRNA backbone vectors (highlighting the NLS, VP64, gRNA spacer, and MS2-binding RNA stem-loop domains). Two versions of the former were constructed in the form of MS2vp64N fusion proteins that showed the highest activity.
[0198] >MS2 VP64 N
[0199]
[0200] >MS2 VP64 C
[0201]
[0202] >gRNA 2XMS2
[0203]
[0204] C. The dTomato fluorescence-based transcriptional activation reporter sequences are listed below (IScel control-TF target, gRNA target, minCMV promoter and FLAG tag+dTomato sequences are highlighted).
[0205] >TF reporter 1
[0206]
[0207] >TF reporter 2
[0208]
[0209] D. The general format of the reporter library for TALE and Cas9-gRNA specificity assays is provided below (IScel control-TF target, gRNA / TALE target site (23 bp for gRNA, 18 bp for TALE), minCMV promoter, RNA barcode, and dTomato sequence are highlighted).
[0210] >Specific reporter library
[0211] Sequence Listing <110> President and Fellows of Harvard College <120> RNA-guided transcriptional regulation <130> 010498.00503 <140> PCT / US14 / 040868 <141> 2014-06-04 <150> US 61 / 830787 <151> 2013-06-04 <160> 184 <170> PatentIn version 3.5 <210> 1 <211> 1368 <212> PRT <213> Streptococcus pyogenes <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Asp Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp His Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 2 <211> 4332 <212> DNA <213> Artificial Sequence <220> <223> VP64-Activator Construct <400> 2 gccaccatgg acaagaagta ctccattggg ctcgctatcg gcacaaacag cgtcggctgg 60 gccgtcatta cggacgagta caaggtgccg agcaaaaaat tcaaagttct gggcaatacc 120 gatcgccaca gcataaagaa gaacctcatt ggcgccctcc tgttcgactc cggggagacg 180 gccgaagcca cgcggctcaa aagaacagca cggcgcagat atacccgcag aaagaatcgg 240 atctgctacc tgcaggagat ctttagtaat gagatggcta aggtggatga ctctttcttc 300 cataggctgg aggagtcctt tttggtggag gaggataaaa agcacgagcg ccacccaatc 360 tttggcaata tcgtggacga ggtggcgtac catgaaaagt acccaaccat atatcatctg 420 aggaagaagc ttgtagacag tactgataag gctgacttgc ggttgatcta tctcgcgctg 480 gcgcatatga tcaaatttcg gggacacttc ctcatcgagg gggacctgaa cccagacaac 540 agcgatgtcg acaaactctt tatccaactg gttcagactt acaatcagct tttcgaagag 600 aacccgatca acgcatccgg agttgacgcc aaagcaatcc tgagcgctag gctgtccaaa 660 tcccggcggc tcgaaaacct catcgcacag ctccctgggg agaagaagaa cggcctgttt 720 ggtaatctta tcgccctgtc actcgggctg acccccaact ttaaatctaa cttcgacctg 780 gccgaagatg ccaagcttca actgagcaaa gacacctacg atgatgatct cgacaatctg 840 ctggcccaga tcggcgacca gtacgcagac ctttttttgg cggcaaagaa cctgtcagac 900 gccattctgc tgagtgatat tctgcgagtg aacacggaga tcaccaaagc tccgctgagc 960 gctagtatga tcaagcgcta tgatgagcac caccaagact tgactttgct gaaggccctt 1020 gtcagacagc aactgcctga gaagtacaag gaaattttct tcgatcagtc taaaaatggc 1080 tacgccggat acatgacgg cggagcaagc caggagat tttacaatt tattaagccc 1140 atcttggaaa aaatggacgg caccgaggag ctgctggtaa agcttacag agagatctg 1200 ttgcgcaaac agcgcacttt cgacaatgga agcatccccc accagatca cctgggcgaa 1260 ctgcacgcta tcctcaggcg gcagaggat tttaccct tttgaaga taacagggaa 1320 aagattgaga aaatcctcac atttcggata ccctactatg taggccccct cgcccgggga 1380 aattccagat tcgcgtggat gactcgcaa tcagagaga ccatcactcc ctggaacttc 1440 gaggaagtcg tggataaggg ggcctctgcc cagtccttca tcgaaggat gactacttt 1500 gataaaaatc tgcctaacga aaggtgctt cctaacact ctctgctgta cgagtacttc 1560 acagtttata acgagctcac caggtcaa tacgtcacag aagggag aaagccagca 1620 ttcctgtctg gagagcagaa gaaagctatc gtggacctcc tctcagac gaaccggaaa 1680 gttaccgtga aacagctcaa agagactat ttcaaaga ttgaatgttt cgactctgtt 1740 gaatcagcg gagtggagga tcgctcaac gcatccctgg gaacgtatca cgatctcctg 1800 aaaatcatta agachagga cttcctggac agaggaga agaggacat tctgaggac 1860 attgtcctca cccttacgtt gtttgaagat agggagatga tgagaacg cttgaaaact 1920 tacgctcatc tctcgacga caagtcatg aacagctca agaggcgccg atacagga 1980 tggggcggc tgtcaagaaa actgatcaat gggatccgag acagcag tggaagaca 2040 atcctggatt ttcttaagtc cgatggattt gccaaccgga acttcatgca gttgatccat 2100 gatgactctc tcaccttatta ggaggacatc cagaagcac aagtttctgg ccagggggac 2160 agtcttcacg agcacatcgc taatcttgca ggtagcccag ctcaaaaa gggaatactg 2220 cagaccgtta agtcgtgga tgaactcgtc aaagtaatgg gaaggcataa gccgagaat 2280 atcgttatcg agatggcccg aggaaccaa actacccaga agggacagaa gaacagtagg 2340 gaaaggatga agaggattga agagggtata aagaactgg ggtcccaat ccttaaggaa 2400 cacccagttg aaaacaccca gctcagaat gagaagctct acctgtacta cctgcagaac 2460 ggcagggaca tgtacgtgga tcaggactg gatacaatc ggctctccga ctacgacgtg 2520 gctgctatcg tgccccagtc ttttctcaaa gatgattcta ttgataata agtgttgaca agatccgata aagctagagg gaagagtgat aacgtcccct cagaagagt tgtcaagaaa atgaaaatt attggcggca gctgctgac gccaaactga tcacacaacg gaagttcgat aatctgacta aggctgaacg aggtggcctg tctgagttgg ataaagccgg cttcatcaaa aggcagcttg ttgagacacg ccagatcacc aagcacgtgg cccaaattct cgattcacgc atgaacacca agtacgatga aaatgacaaa ctgattcgag aggtgaaagt tattactctg aagtctaagc tggtctcaga tttcagaaag gactttcagt tttataaggt gagagagatc aacaattacc accatgcgca tgatgcctac ctgaatgcag tggtaggcac tgcacttatc aaaaaatatc ccaagcttga atctgaatt gtttacggag actataaagt gtacgatgtt aggaaatga tcgcaaagtc tgagcagga ataggcaagg ccaccgctaa gtacttcttt 3180. tttgaattt tttcaagacc gagattacac tggccaatgg agagattcgg aagcgaccac ttatcgaaac aaacggagaa acaggagaaa tcgtgtggga caagggtagg gatttcgcga cagtccggaa ggtcctgtcc atgccgcagg tgaacatcgt taaaaagacc 3300 gaagtacaga ccggaggtt ctccaaggaa agtatcctcc cgaaaggaa cagcgacaag 3360 ctgatcgcac gcaaaaaga ttgggacccc aagaatacg gcggattcga ttctcctaca 3420 gtcgcttaca gtgtactggt tgtggccaaa gtggagaaag ggaagtctaa aaaactcaa 3480 agcgtcagg aactgctggg catcacaatc atggagcgat caagcttcga aaaaaacccc 3540 atcgactttc tcaggcgaa aggatataaa gaggtcaaa aagacctcat cattaagctt 3600 cccaagtact ctctctttga gcttgaaac ggccggaaac gatgctcgc tagtgcgggc 3660 gagctgcaga aaggtaacga gctggcactg cccctaaat acgttaattt cttgtatctg 3720 gccagccact atgaaagct caaagggtct cccgaagata atgagcagaa gcagctgttc 3780 gtggaacaac acaacacta ccttgatgag atcatcgagc aaataagcga attctccaaa 3840 agagtgatcc tcgccgacgc taacctcgat aaggtgctt ctgcttacaa taagcacagg 3900 gataagccca tcagggagca ggcagaaaac attatccact tgtttactct gaccaacttg 3960 ggcgcgcctg cagccttcaa gtacttcgac accaccatag acagaaagcg gtacacctct 4020 acaaaggagg tcctggacgc cacactgatt catcagtcaa ttacggggct ctatgaaaca 4080 agaatcgacc tctctcagct cggtggagac agcagggctg accccaagaa gaagaggaag 4140 gtggaggcca gcggttccgg acgggctgac gcattggacg attttgatct ggatatgctg 4200 ggaagtgacg ccctcgatga ttttgacctt gacatgcttg gttcggatgc ccttgatgac 4260 tttgacctcg acatgctcgg cagtgacgcc cttgatgatt tcgacctgga catgctgatt 4320 aactctagat ga 4332 <210> 3 <211> 4365 <212> DNA <213> Artificial sequence <220> <223> VP64 - activator construct <400> 3 gccaccatgc ccaagaagaa gaggaaggtg ggaaggggga tggacaagaa gtactccatt 60 gggctcgcta tcggcacaaa cagcgtcggc tgggccgtca ttacggacga gtacaaggtg 120 ccgagcaaaa aattcaaagt tctgggcaat accgatcgcc acagcataaa gaagaacctc 180 attggcgccc tcctgttcga ctccggggag acggccgaag ccacgcggct caaaagaaca 240 gcacggcgca gatacccg cagaaagaat cggatctgct acctgcagga gatctttagt 300 aatgagatgg ctaggtgga tgactttc ttccataggc tggaggtc ctttttgtg 360 gaggaggata aaaagcacga gcgccaccca atctttggca atatcgtgga cgaggtggcg 420 taccatgaaa agtacccaac catatatcat ctgaggaaga agcttgtaga cagtactgat 480 aaggctgact tgcggttgat ctatctcgcg ctggcgcata tgatcaatt tcggggacac 540 ttcctcatcg agggggacct gaacccagac aacagcgatg tcgacaacct ctttatccaa 600 ctggttcaga cttacaatca gctttcgaa gagaacccga tcaacccatc cggagttgac 660 gccaaagcaa tcctgagcgc taggctgtcc aaatcccggc ggctcgaaa cctcatcgca 720 cagctccctg gggagaagaa gaacggcctg ttggtaatc ttacgccct gtcactcggg 780 ctgaccccca actttaatc taacttcgac ctggccgaag atgccaagct tcaacgagc 840 aaagacacct acgatgatga tctcgacaat ctgctggccc agatcggcga ccagtacgca 900 gaccttttt tggcggcaa gaacctgtca gacgccattc tgctgagtga tattctgcga 960 gtgaacacgg agatcaccaa agctccgctg agcgctagta tgatcaagcg ctatgatgag 1020 caccaccaag acttgacttt gctgaaggcc cttgtcagac agcaactgcc tgagaagtac 1080 aaaggaaattt tcttcgatca gtctaaaaat ggctacgccg gatacattga cggcggagca 1140 agccaggagg aattttacaa atttattaag cccatcttgg aaaaaatgga cggcaccgag 1200 gagctgctgg taagcttaa cagaagaat ctgttgcgca aacagcgcac tttcgacaat 1260 ggaagcatcc cccaccagat tcacctgggc gaactgcacg ctatcctcag gcggcaagag 1320 gatttctacc cctttttgaa agataacagg gaaaagattg agaaaatcct cacattttcgg 1380 ataccctact atgtaggccc cctcgccccgg ggaaattcca gattcgcgtg gatgactcgc 1440 aaatcagaag agaccatcac tccctggaac ttcgaggaag tcgtggataa gggggcctct 1500 gcccagtcct tcatcgaaag gatgactaac tttgataaaa atctgcctaa cgaaaaggtg 1560 cttcctaaac actctctgct gtacgagtac ttcacagttt ataacgagct caccaaggttc 1620 aaatacgtca cagaagggat gagaaagcca gcattcctgt ctggagagca gaaagct 1680 atcgtggacc tcctcttcaa gacgaaccgg aaagttaccg tgaaacagct caaagaac 1740 tatttcaaaa agattgaatg tttcgactct gttgaaatca gcggagtgga ggatcgcttc 1800 aacgcatccc tgggaacgta tcacgatctc ctgaaaatca ttaaagacaa ggacttcctg 1860 1920 1980 atgaaacagc tcaagaggcg ccgatataca ggatggggc ggctgtcaag aaaactgatc 2040 aatgggatcc gagacaagca gagtggaaag acaatcctgg attttcttaa gtccgatgga 2100 tttgccaacc ggaacttcat gcagttgatc catgatgact ctctcacctt tagggaggac 2160 atccagaaag cacaagtttc tggccagggg gacagtcttc acgagcacat cgctaatctt 2220 gcaggtagcc cagctatcaa aaagggataa ctgcagaccg ttaaggtcgt ggatgaactc 2280 gtcaaagtaa tgggaaggca taagcccgag aatatcgtta tcgagatggc ccgagagaac 2340 caaactaccc agaagggaca gaacagt agggaaagga tgaagaggat tgaagagggt 2400 ataaaagaac tggggtccca aatccttaag gaacacccag ttgaaaacac ccagcttcag 2460 aatgagaagc tctacctgta ctacctgcag aacggcaggg acatgtacgt ggatcaggaa 2520 ctggacatca atcggctctc cgactacgac gtggctgcta tcgtgcccca gtcttttctc 2580 aaagatgatt ctattgataa taaagtgttg acaagatccg ataaagctag agggaagagt 2640 gataacgtcc cctcagaaga agttgtcaag aaaatgaaaa attattggcg gcagctgctg 2700 aacgccaaac tgatcacaca acggaagttc gataatctga ctaaggctga acgaggtggc 2760 ctgtctgagt tggataaagc cggcttcatc aaaaggcagc ttgttgagac acgccagatc 2820 accaagcacg tggcccaaat tctcgattca cgcatgaaca ccaagtacga tgaaaatgac 2880 aaactgattc gagaggtgaa agttattact ctgaagtcta agctggtctc agatttcaga 2940 aaggactttc agttttataa ggtgagagag atcaacaatt accaccatgc gcatgatgcc 3000 tacctgaatg cagtggtagg cactgcactt atcaaaaaat atcccaagct tgaatctgaa 3060 tttgtttacg gagactataa agtgtacgat gttaggaaaa tgatcgcaaa gtctgagcag 3120 gaataggca aggccaccgc taagtacttc ttttacagca atattatgaa tttttcaag 3180 3240 gaaagggag aaatcgtgtg ggacaagggt agggatttcg cgacagtccg gaaggtcctg 3300 tccatgccgc aggtgaacat cgttaaaaag accgaagtac agaccggagg cttctccaag 3360 gaagtatcc tcccgaaaag gaacagcgac aagctgatcg cacgcaaaaa agattgggac 3420 cccaagaaat acggcggatt cgattctcct acagtcgctt acagtgtact ggttgtggcc 3480 aaagtggaga agggaagtc taaaaaactc aaaagcgtca aggaactgct gggcatcaca 3540 atcatggagc gatcaagctt cgaaaaaaaac cccatcgact ttctcgaggc gaaaggatat 3600 aaagaggtca aaaaagacct catcattaag cttcccaagt actctctctt tgagcttgaa 3660 aacggggga aacgaatgct cgctagtgcg ggcgagctgc agaaaggtaa cgagctggca 3720 ctgccctcta aatacgttaa tttcttgtat ctggccagcc actatgaaaa gctcaaaggg 3780 tctcccgaag atatgagca gaagcagctg ttcgtggaac aacaaaca ctaccttgat 3840 gagatcatcg agcaaataag cgaattctcc aaaagagtga tcctcgccga cgctaacctc 3900 gataaggtgc tttctgctta caataagcac agggataagc ccatcaggga gcaggcagaa 3960 aacattatcc acttgtttac tctgaccaac ttgggcgcgc ctgcagcctt caagtacttc 4020 gacaccacca tagacagaaa gcggtacacc tctacaaagg aggtcctgga cgccacactg 4080 attcatcagt caattacggg gctctatgaa acaagaatcg acctctctca gctcggtgga 4140 gacagcaggg ctgaccccaa gaagaagagg aaggtggagg ccagcggttc cggacgggct 4200 gacgcattgg acgattttga tctggatatg ctgggaagtg acgccctcga tgattttgac 4260 cttgacatgc ttggttcgga tgcccttgat gactttgacc tcgacatgct cggcagtgac 4320 gcccttgatg atttcgacct ggacatgctg attaactcta gatga 4365 <210> 4 <211> 4425 <212> DNA <213> Artificial sequence <220> <223> VP64-Activator construct <400> 4 gccaccatgg acaagaagta ctccattggg ctcgctatcg gcacaaacag cgtcggctgg 60 gccgtcatta cggacgagta caaggtgccg agcaaaaaat tcaaagttct gggcaatacc 120 gatcgccaca gcataaagaa gaacctcatt ggcgccctcc tgttcgactc cggggagacg 180 gccgaagcca cgcggctcaa aagaacagca cggcgcagat atacccgcag aaagaatcgg 240 atctgctacc tgcaggagat ctttagtaat gagatggcta aggtggatga ctctttcttc 300 cataggctgg aggagtcctt tttggtggag gaggataaaa agcacgagcg ccacccaatc 360 tttggcaata tcgtggacga ggtggcgtac catgaaaagt acccaaccat atatcatctg 420 aggaagaagc ttgtagacag tactgataag gctgacttgc ggttgatcta tctcgcgctg 480 gcgcatatga tcaaatttcg gggacacttc ctcatcgagg gggacctgaa cccagacaac 540 agcgatgtcg acaaactctt tatccaactg gttcagactt acaatcagct tttcgaagag 600 aacccgatca acgcatccgg agttgacgcc aaagcaatcc tgagcgctag gctgtccaaa 660 tcccggcggc tcgaaaacct catcgcacag ctccctgggg agaagaagaa cggcctgttt 720 ggtaatctta tcgccctgtc actcgggctg acccccaact ttaaatctaa cttcgacctg 780 gccgaagatg ccaagcttca actgagcaaa ccacctacg atgatgatct cgacaatctg 840 ctggcccaga tcggcgacca gtacgcagac ctttttgg cggcaagaa cctgtcagac 900 gccattctgc tgagtgatat tctgcgagtg aacacggaga tcaccaagc tccgctgagc 960 gctagtatga tcaagcgcta tgatgagcac caccagact tgacttgct gaaggccctt 1020 gtcagacagc aactgcctga gaagtacaag gaaatttttct tcgatcagtc taaaaatggc 1080 tacgccggat acatgacgg cggagcaagc caggagat tttacaatt tattaagccc 1140 atcttggaaa aaatggacgg caccgaggag ctgctggtaa agcttacag agagatctg 1200 ttgcgcaaac agcgcacttt cgacaatgga agcatccccc accagatca cctgggcgaa 1260 ctgcacgcta tcctcaggcg gcagaggat tttaccct tttgaaga taacagggaa 1320 aagattgaga aaatcctcac atttcggata ccctactatg taggccccct cgcccgggga 1380 aattccagat tcgcgtggat gactcgcaa tcagagaga ccatcactcc ctggaacttc 1440 gaggaagtcg tggataaggg ggcctctgcc cagtccttca tcgaaggat gactacttt 1500 gataaaaatc tgcctaacga aaggtgctt cctaacact ctctgctgta cgagtacttc 1560 acagtttata acgagctcac caggtcaa tacgtcacag aagggag aaagccagca 1620 ttcctgtctg gagagcagaa gaaagctatc gtggacctcc tctcagac gaaccggaaa 1680 gttaccgtga aacagctcaa agagactat ttcaaaga ttgaatgttt cgactctgtt 1740 gaatcagcg gagtggagga tcgctcaac gcatccctgg gaacgtatca cgatctcctg 1800 aaaatcatta agachagga cttcctggac agaggaga agaggacat tctgaggac 1860 attgtcctca cccttacgtt gtttgaagat agggagatga tgagaacg cttgaaaact 1920 tacgctcatc tctcgacga caagtcatg aacagctca agaggcgccg atacagga 1980 tggggcggc tgtcaagaaa actgatcaat gggatccgag acagcag tggaagaca 2040 atcctggatt ttcttaagtc cgatggattt gccaaccgga acttcatgca gttgatccat 2100 gatgactctc tcaccttatta ggaggacatc cagaagcac aagtttctgg ccagggggac 2160 agtcttcacg agcacatcgc taatcttgca ggtagcccag ctcaaaaa gggaatactg 2220 cagaccgtta agtcgtgga tgaactcgtc aaagtaatgg gaaggcataa gccgagaat 2280 atcgttatcg agatggcccg aggaaccaa actacccaga agggacagaa gaacagtagg 2340 gaaaggatga agaggattga agagggtata aagaactgg ggtcccaat ccttaaggaa 2400 cacccagttg aaaacaccca gctcagaat gagaagctct acctgtacta cctgcagaac 2460 ggcagggaca tgtacgtgga tcaggactg gatacaatc ggctctccga ctacgacgtg 2520 gctgctatcg tgccccagtc ttttctcaa gatgattcta ttgataataa agtgttgaca 2580 agatccgata aagctagagg gagagtgat aacgtcccct cagagaagt tgtcaagaaa 2640 atgaaaaatt attggcggca gctgctgaac gccaactga tcacaacg gaagttcgat 2700 aatctgacta aggctgaacg aggtggctg tctgagttgg ataaagccgg cttcatcaa 2760 agcagcttg ttgagacacg ccagatcacc aagcacgtgg cccaattct cgatcacgc 2820 atgaacacca agtacgatga aaatgacaaa ctgattcgag agtgaaagt tattactctg 2880 aagtctaagc tggtctcaga ttcagaag gactttcagt tttaaggt gagagagatc 2940 aacaattacc accatgcgca tgatgcctac ctgaatgcag tggtaggcac tgcacttatc 3000 aaaaaatatc ccaagcttga atctgaattt gtttacggag actataaagt gtacgatgtt 3060 aggaaaatga tcgcaagtc tgagcaggaa atggcaag ccaccgcta gtactcttt 3120 tacagcaata ttatgaattt ttcagacc gagattacac tggccaatgg agagattcgg 3180 aagcgaccac ttatcgaac aaacggagaa acggagaaa tcgtgtggga caagggtagg 3240 gatttcgcga cagtccggaa ggtcctgtcc atgccgcagg tgaacatcgt taaaaagacc 3300 gaagtacaga ccggaggtt ctccaaggaa agtatcctcc cgaaaggaa cagcgacaag 3360 ctgatcgcac gcaaaaaga ttgggacccc aagaatacg gcggattcga ttctcctaca 3420 gtcgcttaca gtgtactggt tgtggccaaa gtggagaaag ggaagtctaa aaaactcaa 3480 agcgtcagg aactgctggg catcacaatc atggagcgat caagcttcga aaaaaacccc 3540 atcgactttc tcaggcgaa aggatataaa gaggtcaaa aagacctcat cattaagctt 3600 cccaagtact ctctctttga gcttgaaac ggccggaaac gatgctcgc tagtgcgggc 3660 gagctgcaga aaggtaacga gctggcactg ccctctaaat acgttaattt cttgtatctg 3720 gccagccact atgaaaagct caaagggtct cccgaagata atgagcagaa gcagctgttc 3780 gtggaacac acaaacacta ccttgatgag atcatcgagc aaataagcga attctccaaa 3840 agagtgatcc tcgccgacgc taacctcgat aaggtgcttt ctgcttacaa taagcacagg 3900 gataagccca tcagggagca ggcagaaaac attatccact tgtttactct gaccaacttg 3960 ggcgcgcctg cagccttcaa gtacttcgac accaccatag acagaaagcg gtacacctct 4020 acaaaggagg tcctggacgc cacactgatt catcagtcaa ttacggggct ctatgaaaca 4080 agaatcgacc tctctcagct cggtggagac agcagggctg accccaagaa gaagaggaag 4140 gtggaggcca gcggttccgg acgggctgac gcattggacg atttgatct ggatatgctg 4200 ggaagtgacg ccctcgatga ttttgacctt gacatgcttg gttcggatgc ccttgatgac 4260 tttgacctcg acatgctcgg cagtgacgcc cttgatgatt tcgacctgga catgctgatt 4320 aactctagag cggccgcaga tccaaaaaag aagagaaagg tagatccaaa aaagaagaga 4380 aaggtagatc caaaaaagaa gagaaaggta gatacggccg catag 4425 <210> 5 <211> 587 <212> DNA <213> Artificial Sequence <220> <223> MS2-Activator Construct <400> 5 ccaccatggg acctaagaaa aagaggaagg tggcggccgc ttctagaatg gcttctaact 60 ttactcagtt cgttctcgtc gacaatggcg gaactggcga cgtgactgtc gccccaagca 120 acttcgctaa cgggatcgct gaatggatca gctctaactc gcgttcacag gcttacaaag 180 taacctgtag cgttcgtcag agctctgcgc agaatcgcaa atacaccatc aaagtcgagg 240 tgcctaaagg cgcctggcgt tcgtacttaa atatggaact aaccattcca attttcgcca 300 cgaattccga ctgcgagctt attgttaagg caatgcaagg tctcctaaaa gatggaaacc 360 cgattccctc agcaatcgca gcaaactccg gcatctacga ggccagcggt tccggacggg 420 ctgacgcatt ggacgatttt gatctggata tgctgggaag tgacgccctc gatgattttg 480 accttgacat gcttggttcg gatgcccttg atgactttga cctcgacatg ctcggcagtg 540 acgcccttga tgatttcgac ctggacatgc tgattaactc tagatga 587 <210> 6 <211> 681 <212> DNA <213> Artificial Sequence <220> <223> MS2 - Activator Construct <400> 6 gccaccatgg gacctaagaa aaagaggaag gtggcggccg cttctagaat ggcttctaac 60 tttactcagt tcgttctcgt cgacaatggc ggaactggcg acgtgactgt cgccccaagc 120 aacttcgcta acgggatcgc tgaatggatc agctctaact cgcgttcaca ggcttacaaa 180 gtaacctgta gcgttcgtca gagctctgcg cagaatcgca aatacaccat caaagtcgag 240 gtgcctaaag gcgcctggcg ttcgtactta aatatggaac taaccattcc aattttcgcc 300 acgaattccg actgcgagct tattgttaag gcaatgcaag gtctcctaaa agatggaaac 360 ccgattccct cagcaatcgc agcaaactcc ggcatctacg aggccagcgg ttccggacgg 420 gctgacgcat tggacgattt tgatctggat atgctgggaa gtgacgccct cgatgatttt 480 gaccttgaca tgcttggttc ggatgccctt gatgactttg acctcgacat gctcggcagt 540 gacgcccttg atgatttcga cctggacatg ctgattaact ctagagcggc cgcagatcca 600 aaaaagaaga gaaaggtaga tccaaaaaag aagagaaagg tagatccaaa aaagaagaga 660 aaggtagata cggccgcata g 681 <210> 7 <211> 557 <212> DNA <213> Artificial Sequence <220> <223> MS2 - Activator Construct <220> <221> misc_feature <222> (320)..(339) <223> Where N is G, A, T or C <400> 7 tgtacaaaaa agcaggcttt aaaggaacca attcagtcga ctggatccgg taccaaggtc 60 gggcaggaag agggcctatt tcccatgatt ccttcatatt tgcatatacg atacaaggct 120 gttagagaga taattagaat taatttgact gtaaacacaa agatattagt acaaaatacg 180 tgacgtagaa agtaataatt tcttgggtag tttgcagttt taaaattatg ttttaaaatg 240 gactatcata tgcttaccgt aacttgaaag tatttcgatt tcttggcttt atatatcttg 300 tggaaaggac gaaacaccgn nnnnnnnnnn nnnnnnnnng ttttagagct agaaatagca 360 agttaaaata aggctagtcc gttatcaact tgaaaaagtg gcaccgagtc ggtgctctgc 420 aggtcgactc tagaaaacat gaggatcacc catgtctgca gtattcccgg gttcattaga 480 tcctaaggta cctaattgcc tagaaaacat gaggatcacc catgtctgca ggtcgactct 540 agaaattttt tctagac 557 <210> 8 <211> 882 <212> DNA <213> Artificial Sequence <220> <223> Activation Reporter Construct <400> 8 tagggataac agggtaatag tgtcccctcc accccacagt ggggcgaggt aggcgtgtac 60 ggtgggaggc ctatataagc agagctcgtt tagtgaaccg tcagatcgcc tggagaattc 120 gccaccatgg actacaagga tgacgacgat aaaacttccg gtggcggact gggttccacc 180 gtgagcaagg gcgaggaggt catcaaagag ttcatgcgct tcaaggtgcg catggagggc 240 tccatgaacg gccacgagtt cgagatcgag ggcgagggcg agggccgccc ctacgagggc 300 acccagaccg ccaagctgaa ggtgaccaag ggcggccccc tgcccttcgc ctgggacatc 360 ctgtcccccc agttcatgta cggctccaag gcgtacgtga agcaccccgc cgacatcccc 420 gattacaaga agctgtcctt ccccgagggc ttcaagtggg agcgcgtgat gaacttcgag 480 gacggcggtc tggtgaccgt gacccaggac tcctccctgc aggacggcac gctgatctac 540 aaggtgaaga tgcgcggcac caacttcccc cccgacggcc ccgtaatgca gaagaagacc 600 atgggctggg aggcctccac cgagcgcctg tacccccgcg acggcgtgct gaagggcgag 660 atccaccagg ccctgaagct gaaggacggc ggccactacc tggtggagtt caagaccatc 720 tacatggcca agaagcccgt gcaactgccc ggctactact acgtggacac caagctggac 780 atcacctccc acaacgagga ctacaccatc gtggaacagt acgagcgctc cgagggccgc 840 caccacctgt tcctgtacgg catggacgag ctgtacaagt aa 882 <210> 9 <211> 882 <212> DNA <213> Artificial Sequence <220> <223> Activated Reporter Construct <400> 9 tagggataac agggtaatag tggggccact agggacagga ttggcgaggt aggcgtgtac 60 ggtgggaggc ctatataagc agagctcgtt tagtgaaccg tcagatcgcc tggagaattc 120 gccaccatgg actacaagga tgacgacgat aaaacttccg gtggcggact gggttccacc 180 gtgagcaagg gcgaggaggt catcaaagag ttcatgcgct tcaaggtgcg catggagggc 240 tccatgaacg gccacgagtt cgagatcgag ggcgagggcg agggccgccc ctacgagggc 300 acccagaccg ccaagctgaa ggtgaccaag ggcggccccc tgcccttcgc ctgggacatc 360 ctgtcccccc agttcatgta cggctccaag gcgtacgtga agcaccccgc cgacatcccc 420 gattacaaga agctgtcctt ccccgagggc ttcaagtggg agcgcgtgat gaacttcgag 480 gacggcggtc tggtgaccgt gacccaggac tcctccctgc aggacggcac gctgatctac 540 aaggtgaaga tgcgcggcac caacttcccc cccgacggcc ccgtaatgca gaagaagacc 600 atgggctggg aggcctccac cgagcgcctg tacccccgcg acggcgtgct gaagggcgag 660 atccaccagg ccctgaagct gaaggacggc ggccactacc tggtggagtt caagaccatc 720 tacatggcca agaagcccgt gcaactgccc ggctactact acgtggacac caagctggac 780 atcacctccc acaacgagga ctacaccatc gtggaacagt acgagcgctc cgagggccgc 840 caccacctgt tcctgtacgg catggacgag ctgtacaagt aa <210> 10 <211> 912 <212> DNA <213> The snowstorm <220> <223> Thanks for reading <220> <221> misc_feature <222> (22)..(44) <223> Door NWG, A, TWC <220> <221> misc_feature <222> (154)..(177) <223> Door NWG, A, TWC <400> 10 tagggataac agggtaatag tnnnnnnnnn nnnnnnnnnn nnnncgaggt aggcgtgtac ggtgggaggc ctataagc agagctcgtt tagtgaaccg tcagatcgcc tggagaattc gccaccatgg actacaagga tgacgacgat aaannnnnnn nnnnnnnnnn nnnnnnnact tccggtggcg gactgggttc caccgtgagc aagggcgagg aggtcatcaa agagttcatg 240 cgcttcaagg tgcgcatgga gggctccatg aacggccacg agttcgagat cgagggcgag 300 ggcgagggcc gcccctacga gggcacccag accgccaagc tgaaggtgac caagggcggc 360 cccctgccct tcgcctggga catcctgtcc ccccagttca tgtacggctc caaggcgtac 420 gtgaagcacc ccgccgacat ccccgattac aagaagctgt ccttccccga gggcttcaag 480 tgggagcgcg tgatgaactt cgaggacggc ggtctggtga ccgtgaccca ggactcctcc 540 ctgcaggacg gcacgctgat ctacaaggtg aagatgcgcg gcaccaactt cccccccgac 600 ggccccgtaa tgcagaagaa gaccatgggc tgggaggcct ccaccgagcg cctgtacccc 660 cgcgacggcg tgctgaaggg cgagatccac caggccctga agctgaagga cggcggccac 720 tacctggtgg agttcaagac catctacatg gccaagaagc ccgtgcaact gcccggctac 780 tactacgtgg acaccaagct ggacatcacc tcccacaacg aggactacac catcgtggaa 840 cagtacgagc gctccgaggg ccgccaccac ctgttcctgt acggcatgga cgagctgtac 900 aagtaagaat tc 912 <210> 11 <211> 23 <212> DNA <213> Artificial sequence <220> <223> Targeting probe <400> 11 ctggcggatc actcgcggtt agg 23 <210> 12 <211> 23 <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 12 cctcggcctc caaaagtgct agg 23 <210> 13 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 13 acgctgattc ctgcagatca ggg 23 <210> 14 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 14 ccaggaatac gtatccacca ggg 23 <210> 15 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 15 gccacaccca agcgatcaaa tgg 23 <210> 16 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 16 aaataataca ttctaaggta agg 23 <210> 17 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 17 gctactgggg aggctgaggc agg 23 <210> 18 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 18 tagcaataca gtcacattaa tgg 23 <210> 19 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 19 ctcatgtgat ccccccgtct cgg 23 <210> 20 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 20 ccgggcagag agtgaacgcg cgg 23 <210> twenty one <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> twenty one ttccttccct ctcccgtgct tgg 23 <210> twenty two <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> twenty two tctctgcaaa gcccctggag agg 23 <210> twenty three <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> twenty three aatgcagttg ccgagtgcag tgg 23 <210> twenty four <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> twenty four cctcagcctc ctaaagtgct ggg 23 <210> 25 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 25 gagtccaaat cctctttact agg 23 <210> 26 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 26 gagtgtctgg atttgggata agg 23 <210> 27 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 27 cagcacctca tctcccagtg agg 23 <210> 28 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 28 tctaaaaccc agggaatcat ggg 23 <210> 29 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 29 cacaaggcag ccagggatcc agg 23 <210> 30 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 30 gatggcaagc tgagaaacac tgg 23 <210> 31 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 31 tgaaatgcac gcatacaatt agg 23 <210> 32 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 32 ccagtccaga cctggccttc tgg 23 <210> 33 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 33 cccagaaaaa cagaccctga agg 23 <210> 34 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 34 aagggttgag cacttgttta ggg 23 <210> 35 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 35 atgtctgagt tttggttgag agg 23 <210> 36 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 36 ggtcccttga aggggaagta ggg 23 <210> 37 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 37 tggcagtcta ctcttgaaga tgg 23 <210> 38 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 38 ggcacagtgc cagaggtctg tgg 23 <210> 39 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 39 taaaaataaa aaaactaaca ggg 23 <210> 40 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 40 tctgtggggg acctgcactg agg 23 <210> 41 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 41 ggccagaggt caaggctagt ggg 23 <210> 42 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 42 cacgaccgaa acccttctta cgg 23 <210> 43 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 43 gttgaatgaa gacagtctag tgg 23 <210> 44 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 44 taagaacaga gcaagttacg tgg 23 <210> 45 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 45 tgtaaggtaa gagaggagag cgg 23 <210> 46 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 46 tgacacacca actcctgcac tgg 23 <210> 47 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 47 tttacccact tccttcgaaa agg 23 <210> 48 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 48 gtggctggca ggctggctct ggg 23 <210> 49 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 49 ctccccccggc ctcccccgcg cgg 23 <210> 50 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 50 caaaacccgg cagcgaggct ggg 23 <210> 51 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 51 aggagccgcc gcgcgctgat tgg 23 <210> 52 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 52 cacacacacc cacacgagat ggg 23 <210> 53 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 53 gaagaagcta aagagccaga ggg 23 <210> 54 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 54 atgagaattt caataacctc agg 23 <210> 55 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 55 tcccgctctg ttgcccaggc tgg 23 <210> 56 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 56 cagacaccca ccaccatgcg tgg 23 <210> 57 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 57 tcccaattta ctgggattac agg 23 <210> 58 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 58 tgatttaaaa gttggaaacg tgg 23 <210> 59 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 59 tctagttccc cacctagtct ggg 23 <210> 60 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 60 gattaactga gaattcacaa ggg 23 <210> 61 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Targeted Probes <400> 61 cgccaggagg ggtgggtcta agg 23 <210> 62 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Reporter construct <400> 62 gtcccctcca ccccacagtg ggg 23 <210> 63 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Reporter construct <400> 63 ggggccacta gggacaggat tgg 23 <210> 64 <211> 71 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 64 taatactttt atctgtcccc tccaccccac agtggggcca ctagggacag gattggtgac 60 agaaaagccc c 71 <210> 65 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 65 ggggccacta gggacaggat 20 <210> 66 <211> 80 <212> RNA <213> Artificial sequence <220> <223> Guide RNA <400> 66 guuuuagagc uagaaauagc aaguuaaaau aaggcuagcu uguuaucaac uugaaaaagu 60 ggcaccgagu cggugcuuuu 80 <210> 67 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 67 gtcccctcca ccccacagtg cag 23 <210> 68 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 68 gtcccctcca ccccacagtg caa 23 <210> 69 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 69 gtcccctcca ccccacagtg cgg 23 <210> 70 <211> 52 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 70 tgtcccctcc accccacagt ggggccacta gggacaggat tggtgacaga aa 52 <210> 71 <211> 52 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 71 tgtccccccc accccacagt ggggccacta gggacaggat tggtgacaga aa 52 <210> 72 <211> 52 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 72 aaaaccctcc accccacagt ggggccacta gggacaggat tggtgacaga aa 52 <210> 73 <211> 52 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 73 tgtcccctcc ttttttcagt ggggccacta gggacaggat tggtgacaga aa 52 <210> 74 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 74 caccggggtg gtgcccatcc tgg 23 <210> 75 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 75 ggtgcccatc ctggtcgagc tgg 23 <210> 76 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 76 cccatcctgg tcgagctgga cgg 23 <210> 77 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 77 ggccacaagt tcagcgtgtc cgg 23 <210> 78 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 78 cgcaaataag agctcaccta cgg 23 <210> 79 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 79 ctgaagttca tctgcaccac cgg 23 <210> 80 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 80 ccggcaagct gcccgtgccc tgg 23 <210> 81 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 81 gaccaggatg ggcaccaccc cgg 23 <210> 82 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 82 gccgtccagc tcgaccagga tgg 23 <210> 83 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 83 ggccggacac gctgaacttg tgg 23 <210> 84 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 84 taacagggta atgtcgaggc cgg 23 <210> 85 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 85 aggtgagctc ttatttgcgt agg 23 <210> 86 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 86 cttcagggtc agcttgccgt agg 23 <210> 87 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 87 gggcacgggc agcttgccgg tgg 23 <210> 88 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 88 gagatgatcg ccccttcttc tgg 23 <210> 89 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 89 gagatgatcg ccccttcttc 20 <210> 90 <211> 20 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 90 gtgatgaccg gccgttcttc 20 <210> 91 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 91 gtcccctcca ccccacagtg ggg 23 <210> 92 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 92 gagatgatcg cccgttcttc tgg 23 <210> 93 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 93 guccccucca ccccacagug 20 <210> 94 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 94 guccccucca ccccacaguc 20 <210> 95 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 95 guccccucca ccccacagag 20 <210> 96 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 96 guccccucca ccccacacug 20 <210> 97 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 97 guccccucca ccccacugug 20 <210> 98 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 98 guccccucca ccccagagug 20 <210> 99 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 99 guccccucca ccccucagug 20 <210> 100 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 100 guccccucca cccgacagug 20 <210> 101 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 101 guccccucca ccgcacagug 20 <210> 102 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 102 guccccucca cgccacagug 20 <210> 103 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 103 guccccucca gcccacagug 20 <210> 104 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 104 guccccuccu ccccacagug 20 <210> 105 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 105 guccccucga ccccacagug 20 <210> 106 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 106 guccccucca ccccacagac 20 <210> 107 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 107 guccccucca ccccacucug 20 <210> 108 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 108 guccccucca ccccugagug 20 <210> 109 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 109 guccccucca ccggacagug 20 <210> 110 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 110 guccccucca ggccacagug 20 <210> 111 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 111 guccccucgu ccccacagug 20 <210> 112 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 112 ggggccacta gggacaggat ggg 23 <210> 113 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 113 gagaugaucg ccccuucuuc 20 <210> 114 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 114 gagaugaucg ccccuucuug 20 <210> 115 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 115 gagaugaucg ccccuucuac 20 <210> 116 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 116 gagaugaucg ccccuucauc 20 <210> 117 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 117 gagaugaucg ccccuuguuc 20 <210> 118 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 118 gagaugaucg ccccuacuuc 20 <210> 119 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 119 gagaugaucg ccccaucuuc 20 <210> 120 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 120 gagaugaucg cccguucuuc 20 <210> 121 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 121 gagaugaucg ccgcuucuuc 20 <210> 122 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 122 gagaugaucg cgccuucuuc 20 <210> 123 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 123 gagaugaucg gcccuucuuc 20 <210> 124 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 124 gagaugaucc ccccuucuuc 20 <210> 125 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 125 gagaugaugg ccccuucuuc 20 <210> 126 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 126 gagaugaucg ccccuucuag 20 <210> 127 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 127 gagaugaucg ccccuugauc 20 <210> 128 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 128 gagaugaucg ccccaacuuc 20 <210> 129 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 129 gagaugaucg ccgguucuuc 20 <210> 130 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 130 gagaugaucg ggccuucuuc 20 <210> 131 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 131 gagaugaugc ccccuucuuc 20 <210> 132 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 132 gagatgatcg ccccttcttc tgg 23 <210> 133 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 133 ggggccacua gggacaggau 20 <210> 134 <211> 19 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 134 gggccacuag ggacaggau 19 <210> 135 <211> 18 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 135 ggccacuagg gacaggau 18 <210> 136 <211> 17 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 136 gccacuaggg acaggau 17 <210> 137 <211> 20 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 137 gagaugaucg ccccuucuuc 20 <210> 138 <211> 18 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 138 gaugaucgcc ccuucuuc 18 <210> 139 <211> 15 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 139 gaucgccccu ucuuc 15 <210> 140 <211> 11 <212> RNA <213> Artificial sequence <220> <223> RNA target sequence <400> 140 gcccccuucuu c 11 <210> 141 <211> twenty one <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 141 gtcccctcca ccccacagtg c 21 <210> 142 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <220> <221> misc_feature <222> (5)..(10) <223> Where N is G, A, T or C <400> 142 tgtcnnnnnn accc 14 <210> 143 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 143 tgtcaaaaaa accc 14 <210> 144 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 144 tgtcgggggg accc 14 <210> 145 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 145 tgtcaaaaaa accc 14 <210> 146 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 146 tgtcgggggg accc 14 <210> 147 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 147 tgtcccccccc accc 14 <210> 148 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 148 tgtctttttt accc 14 <210> 149 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 149 tgtcccccccc accc 14 <210> 150 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 150 tgtctttttt accc 14 <210> 151 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 151 ggatcctgtg tccccgagct ggg 23 <210> 152 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 152 gttaatgtgg ctctggttct ggg 23 <210> 153 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 153 ggggccacta gggacaggat tgg 23 <210> 154 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 154 cttcctagtc tcctgatatt ggg 23 <210> 155 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 155 tggtcccagc tcggggacac agg 23 <210> 156 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 156 agaaccagag ccacattaac cgg 23 <210> 157 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 157 gtcaccaatc ctgtccctag tgg 23 <210> 158 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 158 agacccaata tcaggagact agg 23 <210> 159 <211> 75 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 159 gggatcctgt gtccccgagc tgggaccacc ttatattccc agggccggtt aatgtggctc 60 tggttctggg tactt 75 <210> 160 <211> 69 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 160 gggatcctgt gtccccgagc tgggaccacc ttatattccc agggccggtt aatgtggttc 60 tgggtactt 69 <210> 161 <211> 113 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 161 gggatcctgt gtccccgagc tgggaccacc ttatattccc agggcagggc cggttggacc 60 accttatatt cccagggcag ggccggttaa tgtggctctg gttctgggta ctt 113 <210> 162 <211> 34 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 162 gggatcctgt gtccccgtct ggttctgggt actt 34 <210> 163 <211> 47 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 163 gggatcctgt gtccccgagc tgggaccacc ttatattctg ggtactt 47 <210> 164 <211> 17 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 164 gggatcctgt ggtactt 17 <210> 165 <211> 93 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 165 agggccggtt aatgtggctc tggttctggg tacttttatc tgtcccctcc accccacagt 60 ggggccacta gggacaggat tggtgacaga aaa 93 <210> 166 <211> 83 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 166 agggccggtt aatgaatgtg gctctggttc tgggtacttt tatctgtccc ctccacccca 60 cagtggggcc actagacaga aaa 83 <210> 167 <211> 76 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 167 agggccggtt aatgtggctc tggttctggg tacttttatc tgtcccccag tggggccact 60 gattggtgac agaaaa 76 <210> 168 <211> 29 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 168 agggccggtt caggattggt gacagaaaa 29 <210> 169 <211> 34 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 169 agggccggtt aatgtggcga ttggtgacag aaaa 34 <210> 170 <211> 63 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 170 agggccggtt aatgtggctc tggttctggg tacttttatc tgtccccgat tggtgacaga 60 aaa 63 <210> 171 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 171 agggccggtt aatgtggctc tggttctggg tacttttatc tgtcccctcc accccacagt 60 ggggacagga ttggtgacag aaaa 84 <210> 172 <211> 27 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 172 agggccggtt aatgtggtga cagaaaa 27 <210> 173 <211> 105 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 173 agggccggtt aatgtggctc tggttctggg tacttttatc tgtcccctcc accccagggg 60 acagtctgtc ccctccaccc cagggacagg attggtgaca gaaaa 105 <210> 174 <211> 80 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 174 agggccggtt aatgtggctc tggttctggg tacttttatc tgtcccctcc accactaggg 60 acaggattgg tgacagaaaa 80 <210> 175 <211> 53 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 175 cccacagtgg ggccactagg gacaggattg gtgacagaaa agccccatac ccc 53 <210> 176 <211> twenty two <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 176 cccacagtgg ggccactacc cc 22 <210> 177 <211> 96 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 177 cccacagtgg ggccactagt agaaaagccc catccttagg cctcccccat ccttaggcct 60 cctccttcct agtctcctga tattgggtct aacccc 96 <210> 178 <211> 94 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 178 cccacagtgg ggccactagg gacaggattg gtgacagaaa agccccatcc ttaggcctcc 60 tccttcctag tctcctgata ttgggtctaa cccc 94 <210> 179 <211> 62 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 179 cccacagtgg ggccaccctt aggcctcctc cttcctagtc tcctgatatt gggtctaacc 60 cc 62 <210> 180 <211> 38 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 180 cccacagtgg ggccactagt gatattgggt ctaacccc 38 <210> 181 <211> 94 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 181 cccacagtgg ggccactagg gacaggattg gtgacaaaaa agccccatcc ttacgcctcc 60 tccttcctag tctcctgata ttgggtctaa cccc 94 <210> 182 <211> 65 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 182 cccacagtgg ggccactagg gacaggcctc ctccttccta gtctcctgat attgggtcta 60 acccc 65 <210> 183 <211> 102 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 183 cccacagtgg ggccactagg gacaggggga caggattggt gacagaaaag ccccatcctt 60 aggcctcctc cttcctagtc tcctgatatt gggtctaacc cc 102 <210> 184 <211> 76 <212> DNA <213> Artificial sequence <220> <223> Target oligonucleotide sequence <400> 184 cccacaggat tggtgacaga aaagccccat ccttaggcct cctccttcct agtctcctga 60 tattgggtct aacccc 76
Claims
1. A method for changing a double-stranded target nucleic acid of DNA in a cell, comprising: A first exogenous nucleic acid encoding two or more tracrRNA-crRNA fusion guide RNAs is introduced into the cell, wherein each tracrRNA-crRNA fusion guide RNA comprises a spacer sequence, a tracing partner sequence, and a tracr sequence, a portion of the tracr sequence hybridizes to the tracing partner sequence, the tracing partner sequence and the tracr sequence are connected by a linker nucleic acid sequence, and each spacer sequence is complementary to an adjacent site in the double-stranded DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one Cas9 protein nickase having an inactive nuclease domain and guided by the two or more tracrRNA-crRNA fusion guide RNAs, and wherein the two or more tracrRNA-crRNA fusion guide RNAs and the at least one Cas9 protein nickase are expressed, and wherein the at least one Cas9 protein nickase is co-localized with the two or more tracrRNA-crRNA fusion guide RNAs to the double-stranded DNA target nucleic acid, and cleaves the double-stranded DNA target nucleic acid, thereby resulting in two or more adjacent nicks, wherein the two or more adjacent nicks are on different strands of the double-stranded target nucleic acid, The method is used for non-disease treatment purposes.
2. The method of claim 1, wherein the two or more adjacent nicks produce a double-strand break.
3. The method of claim 1, wherein the two or more adjacent nicks produce double-strand breaks, thereby resulting in non-homologous end joining. The method of claim 1 , wherein the two or more adjacent cuts are offset relative to each other. The method of claim 1 , wherein the two or more adjacent nicks are offset relative to each other and produce double-strand breaks.
6. The method of claim 1, wherein the two or more adjacent nicks are offset relative to each other and produce double-strand breaks, thereby resulting in non-homologous end joining.
7. The method of claim 1, further comprising introducing into the cell a third exogenous nucleic acid encoding a donor nucleic acid sequence, wherein the two or more nicks result in homologous recombination of the double-stranded target nucleic acid with the donor nucleic acid sequence.
8. A cell comprising a first exogenous nucleic acid encoding two or more tracrRNA-crRNA fusion guide RNAs, wherein each tracrRNA-crRNA fusion guide RNA comprises a spacer sequence, a tracing partner sequence, and a tracr sequence, a portion of the tracr sequence hybridizes to the tracing partner sequence, the tracing partner sequence and the tracr sequence are connected by a linker nucleic acid sequence, and each spacer sequence is complementary to an adjacent site in a double-stranded DNA target nucleic acid, and A second exogenous nucleic acid encoding at least one Cas9 protein nickase having an inactive nuclease domain, and wherein the two or more tracrRNA-crRNA fusion guide RNAs and the at least one Cas9 protein nickase are elements of a co-localized complex of the DNA double-stranded target nucleic acid.
9. The cell of claim 8, wherein the cell is a eukaryotic cell.
10. The cell according to claim 8, wherein the cell is a yeast cell, a plant cell or an animal cell.
11. The cell of claim 8, wherein the tracrRNA-crRNA fusion guide RNA comprises 10 to 500 nucleotides.
12. The cell of claim 8, wherein the tracrRNA-crRNA fusion guide RNA comprises 20 to 100 nucleotides.
13. The cell of claim 8, wherein the target nucleic acid is associated with a disease or deleterious condition.
14. The cell according to claim 8, wherein the DNA double-stranded target nucleic acid is genomic DNA, mitochondrial DNA or exogenous DNA.
15. A method for changing a double-stranded target nucleic acid of DNA in a cell, comprising: A first exogenous nucleic acid encoding two or more tracrRNA-crRNA fusion guide RNAs is introduced into the cell, wherein each tracrRNA-crRNA fusion guide RNA comprises a spacer sequence, a tracing partner sequence, and a tracr sequence, a portion of the tracr sequence hybridizes to the tracing partner sequence, the tracing partner sequence and the tracr sequence are connected by a linker nucleic acid sequence, and each spacer sequence is complementary to an adjacent site in the double-stranded DNA target nucleic acid, introducing into the cell a second exogenous nucleic acid encoding at least one Cas9 protein nickase having an inactive nuclease domain and guided by two or more tracrRNA-crRNA fusion guide RNAs, and wherein the two or more tracrRNA-crRNA fusion guide RNAs and the at least one Cas9 protein nickase are expressed, and wherein the at least one Cas9 protein nickase co-localizes with the two or more tracrRNA-crRNA fusion guide RNAs to the double-stranded target nucleic acid and cleaves the double-stranded target nucleic acid, thereby resulting in two or more adjacent nicks, and wherein the two or more adjacent nicks are on different strands of the double-stranded target nucleic acid and produce double-strand breaks, thereby causing the double-stranded target nucleic acid to be fragmented, thereby preventing the expression of the double-stranded target nucleic acid, The method is used for non-disease treatment purposes.
16. Use of a first exogenous nucleic acid and a second exogenous nucleic acid in preparing a product for changing a double-stranded target nucleic acid of DNA in a cell, wherein The first exogenous nucleic acid encodes two or more tracrRNA-crRNA fusion guide RNAs, wherein each tracrRNA-crRNA fusion guide RNA comprises a spacer sequence, a tracing partner sequence, and a tracr sequence, a portion of the tracr sequence hybridizes to the tracing partner sequence, the tracing partner sequence and the tracr sequence are connected by a linker nucleic acid sequence, and each spacer sequence is complementary to an adjacent site in the double-stranded DNA target nucleic acid, The second exogenous nucleic acid encodes at least one Cas9 protein nickase having an inactive nuclease domain and guided by the two or more tracrRNA-crRNA fusion guide RNAs, and wherein the two or more tracrRNA-crRNA fusion guide RNAs and the at least one Cas9 protein nickase are expressed by the exogenous nucleic acid introduced into the cell, and wherein the at least one Cas9 protein nickase is co-localized with the two or more tracrRNA-crRNA fusion guide RNAs to the DNA target nucleic acid, and cuts the DNA double-stranded target nucleic acid, thereby resulting in two or more adjacent nicks, wherein the two or more adjacent nicks are on different strands of the double-stranded target nucleic acid.
17. The use according to claim 16, wherein the two or more adjacent nicks produce a double-strand break.
18. The use according to claim 16, wherein the two or more adjacent nicks generate double-strand breaks, thereby resulting in non-homologous end joining.
19. The use according to claim 16, wherein the two or more adjacent cutouts are offset relative to each other.
20. The use according to claim 16, wherein the two or more adjacent nicks are offset relative to each other and produce double-strand breaks.
21. The use according to claim 16, wherein the two or more adjacent nicks are offset relative to each other and produce double-strand breaks, thereby resulting in non-homologous end joining.
22. The use according to claim 16, further comprising introducing a third exogenous nucleic acid encoding a donor nucleic acid sequence into the cell, wherein the two or more nicks result in homologous recombination of the double-stranded target nucleic acid with the donor nucleic acid sequence.
23. Use of a first exogenous nucleic acid and a second exogenous nucleic acid in preparing a product for altering a double-stranded target nucleic acid of a DNA in a cell, The first exogenous nucleic acid encodes two or more tracrRNA-crRNA fusion guide RNAs, wherein each tracrRNA-crRNA fusion guide RNA comprises a spacer sequence, a tracing partner sequence, and a tracr sequence, a portion of the tracr sequence hybridizes to the tracing partner sequence, the tracing partner sequence and the tracr sequence are connected by a linker nucleic acid sequence, and each spacer sequence is complementary to an adjacent site in the double-stranded DNA target nucleic acid, The second exogenous nucleic acid encodes at least one Cas9 protein nickase having an inactive nuclease domain and guided by two or more tracrRNA-crRNA fusion guide RNAs, and wherein the two or more tracrRNA-crRNA fusion guide RNAs and the at least one Cas9 protein nickase are expressed by the heterologous nucleic acid introduced into the cell, and wherein the at least one Cas9 protein nickase co-localizes with the two or more tracrRNA-crRNA fusion guide RNAs to the double-stranded DNA target nucleic acid and cleaves the double-stranded DNA target nucleic acid, thereby generating two or more adjacent nicks, and The two or more adjacent nicks are on different strands of the double-stranded target nucleic acid and generate double-strand breaks, thereby causing the double-stranded target nucleic acid to be fragmented, thereby preventing the expression of the double-stranded target nucleic acid.