Increasing the specificity of RNA-guided genome editing using truncated guide RNAs (tru-gRNAs)

Truncated guide RNAs with 17-19 nucleotides enhance the specificity of CRISPR/Cas9 genome editing by reducing off-target effects, ensuring precise genomic modifications.

JP7812830B2Active Publication Date: 2026-02-10THE GENERAL HOSPITAL CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023186663
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-12-26
Filing Date
2023-10-31
Publication Date
2026-02-10
Estimated Expiration
2034-03-14

AI Technical Summary

Technical Problem

The CRISPR/Cas9 system for genome editing often suffers from off-target effects due to mismatches in the guide RNA/target site junction, making it difficult to predict and control the specificity of genome editing.

Method used

The use of truncated guide RNAs (tru-gRNAs) with 17-19 nucleotides of target complementarity, combined with dCas9-heterologous functional domain fusion proteins, to enhance the specificity of genome editing by reducing off-target effects.

Benefits of technology

Tru-gRNAs significantly improve the specificity of genome editing by minimizing off-target mutations, maintaining high on-target activity, and allowing precise modifications to genomic sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007812830000139
    Figure 0007812830000139
  • Figure 0007812830000140
    Figure 0007812830000140
  • Figure 0007812830000141
    Figure 0007812830000141
Patent Text Reader

Abstract

To provide methods for increasing specificity of RNA-guided genome editing.SOLUTION: The present invention provides a method of increasing specificity of RNA-guided genome editing in a cell, the method comprising contacting the cell with a guide RNA that includes a complementarity region consisting of 17-18 nucleotides that are complementary to 17-18 consecutive nucleotides of the complementary strand of a selected target genomic sequence.SELECTED DRAWING: Figure 2-6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority Claim) This application is a continuation of U.S. Patent Application No. 61 / 799,647, filed March 15, 2013; U.S. Patent Application No. 61 / 838,178, filed June 21, 2013; U.S. Patent Application No. 61 / 838,148 filed June 21 and December 2013 This application claims the benefit of U.S. Patent Application No. 61 / 921,007, filed on the 26th. The entire contents of the above application are incorporated herein by reference.

[0002] (Federally sponsored research or development) This invention was made possible through grant number DP1 GM105378 awarded by the National Institutes of Health. This invention was made with government support under the terms of the Federal Register. The government has certain rights in this invention.

[0003] Truncated guide RNAs (tru-gRNAs) are used to perform RNA-guided genome editing, e.g., C A method for increasing the specificity of editing using the RISPR / Cas9 system. [Background technology]

[0004] Recent studies have shown that clustered, regularly spaced, short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems (Wiedenheft et al., Nature re 482, 331-338(2012); Horvath et al., Science 32 7,167-170(2010); Terns et al., Curr Opin Microbi ol 14, 321-327 (2011)) has been shown to be effective in bacteria, yeast, and human cells, as well as in Schwann cells. In vivo studies in intact organisms such as fruit flies, zebrafish, and mice It has been revealed that this could serve as a platform for performing genome editing in ivo. (Wang et al., Cell 153, 910-918 (2013); Shen et al., C ell Res(2013); Dicarlo et al., Nucleic Acids Res (2013); Jiang et al., Nat Biotechnol 31, 233-239( 2013); Jinek et al., Elife 2, e00471 (2013); Hwang et al. ,Nat Biotechnol 31,227-229(2013);Cong et al.,S science 339,819-823(2013);Mali et al.Science 3 39,823-826(2013c);Cho et al., Nat Biotechnol 31 ,230-232(2013);Gratz et al.,Genetics 194(4):10 29-35(2013)). Cas9 nucleic acid derived from Streptococcus pyogenes (S. pyogenes) Cas9 is a gene encoding an artificially designed guide RNA (gRNP). A) the first 20 nucleotides and a protospacer adjacent motif (PAM), e.g. The desired target genomic DNA sequence flanked by PAMs matching the sequence NGG or NAG This can be induced through base pair complementarity between the complementary strands of the 2013);Dicarlo et al., Nucleic Acids Res(2013);J iang et al., Nat Biotechnol 31, 233-239(2013); Ji nek et al., Elife 2, e00471 (2013); Hwang et al., Nat Bio technol 31,227-229(2013);Cong et al.Science 3 39,819-823(2013);Mali et al., Science 339,823-8 26(2013c);Cho et al., Nat Biotechnol 31, 230-232 (2013); Jinek et al., Science 337, 816-821(2012)) . in vitro (Jinek et al., Science 337, 816-821(201 2)), bacteria (Jiang et al., Nat Biotechnol 31, 233-239 ( 2013)) and human cells (Cong et al., Science 339, 819-823 ( Previous studies conducted in 2013) showed that Cas9-mediated cleavage can sometimes The gRNA / target site junction, specifically the 3′ end of the 20-nucleotide (nt) gRNA complementary region, 'A single mismatch in the last 10-12 nt at the end can invalidate It is shown that: Summary of the Invention

[0005] In CRISPR-Cas genome editing, complementary regions (those that bind to the target DNA by base pairing) are Cas9 nuclease was introduced using a guide RNA containing both the Cas9 binding region and the Cas9-binding region. This nuclease targets the target DNA (see Figure 1). Some mismatches (up to 5 as shown here) are allowed in the Although it is possible to cleave any single mismatch or combination of mismatches, It is difficult to predict the effect on the genotype of these nucleases. Off-target effects may occur, the location of which may be difficult to predict. Genome editing using CRISPR / Cas systems, e.g., Cas9 or Cas9-based Methods for increasing the specificity of genome editing using fusion proteins of the present invention are described. For example, a truncated target-complementary region (i.e., less than 20 nt, e.g., 17-19 nt or is a 17-18 nt target complementarity, e.g., a 17 nt, 18 nt, or 19 nt target complementarity ) and methods of using them. As used herein, "17-18" or "17-19" refers to 17 nucleotides, 18 nucleotide or 19 nucleotides.

[0006] In one aspect, the present invention provides a target of 17-18 nucleotides or 17-19 nucleotides. The complementary region, for example, a target consisting of 17 to 18 nucleotides or 17 to 19 nucleotides A region of complementary sequence, e.g., a continuous region of 17-18 nucleotides or 17-19 nucleotides a guide RNA molecule (e.g., a single guide RNA molecule) having a target-complementary region consisting of a target complementarity In some embodiments, the guide RNA is a selected A sequence of 17-18 or 17-19 consecutive nucleotides on the complementary strand of the target genome sequence A complementary sequence consisting of 17-18 nucleotides or 17-19 nucleotides complementary to the nucleotide In some embodiments, the target-complementary region comprises a 17-18 nucleotide ( In some embodiments, the region of complementarity consists of a region of complementarity corresponding to a selected target sequence. In some embodiments, the complementary strand is complementary to 17 consecutive nucleotides of the complementary strand of The region is complementary to 18 consecutive nucleotides of the complementary strand of a selected target sequence.

[0007] In another aspect, the present invention provides a method for producing a polypeptide comprising the sequence: (X 17~18 or X 17~19 )GUUUUAGAGCUA(SEQ ID NO:2404); (X 17~18 or X17~19 )GUUUUAGAGCUAUGCUGUUUUG( SEQ ID NO: 2407); or (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(SEQ ID NO: 24 08); (X 17~18 or X 17~19 )GUUUUAGAGCUAGAAAUAGCAAG UUAAAAUAAGGCUAGUCCG(X N )(SEQ ID NO:1); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGAAAAGC AUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N )(Array number No. 2); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUGG AAACAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGU UAUC(X N )(SEQ ID NO:3); (X 17~18 )GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAA GGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCG GUGC(X N ) (SEQ ID NO: 4), (X 17~18 or X 17~19 )GUUUAAGAGCUAGAAAUAGCAAG UUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGC ACCGAGUCGGUGC (SEQ ID NO: 5); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 6); or (X 17~18 or X 17~19 )GUUUAAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7); wherein X 17~18 or X 17~19 is the standard to be selected. a target sequence, preferably a protospacer adjacent motif (PAM), such as NGG, NAG, or or NNGG, complementary to the complementary strand of the target sequence adjacent to the 5' side (17-18 nucleotides) or 17-19 nucleotides) (see, e.g., the structure in Figure 1 ), X N is any sequence that does not interfere with the binding of ribonucleic acid to Cas9, and N (in RNA) can be 0 to 200, for example, 0 to 100, 0 to 50, or 0 to 20. 17~18 or X 17~19 is identical to the naturally occurring sequence adjacent to the rest of the RNA In some embodiments, a terminator that stops RNA Pol III transcription is The RNA is cleaved with the optional presence of one or more T's used as termination signals. one or more Us at the 3' end of the oligonucleotide, e.g., 1 to 8 or more Us (e.g., U , UU, UUU, UUUU, UUUUU, UUUUUU, UUUUUUU, UUUUUU In some embodiments, the RNA comprises a target sequence at the 5' end of the RNA molecule. one or more, e.g., up to three, that are not complementary, e.g., one, two, or three additional In some embodiments, the target-complementary region comprises 17-18 nucleotides. In some embodiments, the complementary region consists of a region complementary to the target of choice. In some embodiments, the target sequence is complementary to 17 consecutive nucleotides of the complementary strand of the target sequence. The region of complementarity is complementary to 18 consecutive sequences.

[0008] In another aspect, the present invention provides DNA molecules encoding the ribonucleic acids described herein and A host cell harboring or expressing the nucleic acid or vector is provided.

[0009] In a further aspect, the present invention provides a method for increasing the specificity of RNA-guided genome editing in cells. The present invention provides a method for detecting genomic DNA fragments comprising the steps of: ~18 nucleotides or 17-18 nucleotides complementary to 17-19 nucleotides or a guide RNA described herein containing a complementary region consisting of 17 to 19 nucleotides. The method includes contacting the

[0010] In yet another aspect, the present invention provides a method for targeting double-stranded DNA molecules, e.g., within the genomic sequence of a cell. The present invention provides a method for inducing a single- or double-strand break in a region of a Cas9 nucleic acid. and a selected target sequence, preferably adjacent to the protospacer. The target sequence is flanked 5' by a PAM, e.g., NGG, NAG, or NNGG. 17, 18, or 19 nucleotides complementary to the complementary strand of the sequence a guide RNA comprising a sequence consisting of, for example, a ribonucleic acid described herein, This includes transferring or introducing the gene into a cell.

[0011] Also provided herein are methods for modifying a target region of a double-stranded DNA molecule in a cell. This method involves the use of dCas9-heterologous functional domain fusion proteins (dCas9-HFD and 17 to 18 consecutive nucleotides of the complementary strand of the selected target genomic sequence, or 17-18 nucleotides or 17-19 nucleotides complementary to 17-19 nucleotides A guide RNA described herein comprising a complementary region consisting of a This includes introducing the compound into the cell.

[0012] In some embodiments, the guide RNA is (i) the complement of a selected target genomic sequence. 17-18 nucleotides or 17-19 nucleotides complementary to the A single guide R containing a complementary region of 8 nucleotides or 17-19 nucleotides or (ii) 17-18 consecutive nucleotides of the complementary strand of the selected target genomic sequence. 17-18 nucleotides or 17-19 nucleotides complementary to The crRNA and tracrRNA contain a complementary region of 9 nucleotides.

[0013] In some embodiments, the target complementarity region is 17-18 nucleotides (of target complementarity) In some embodiments, the region of complementarity consists of a contiguous region of the complementary strand of a selected target sequence. In some embodiments, the region of complementarity is complementary to the 17 nucleotides of the sequence It is complementary to 18.

[0014] X of any molecule described herein 17~18 or X 17~19 The remaining RNA None of the adjacent sequences in the segment are identical to the naturally occurring sequence. In this study, one or more of the nucleotides used as termination signals to stop RNA Pol III transcription were identified. Since multiple Ts are optionally present, the RNA may contain one or more Us at the 3' end of the molecule, e.g. For example, 1 to 8 or more U's (e.g., U, UU, UUU, UUUU, UUUUU) , UUUUUU, UUUUUUU, UUUUUUUU). In some embodiments , the RNA may contain one or more, e.g., up to three, e.g., one, that are not complementary to the target sequence Two or three additional nucleotides are included at the 5' end of the RNA molecule.

[0015] In some embodiments, one or more nucleotides of the RNA, e.g., the target complement sexual area 17~18 or X 17~19 One or more nucleotides inside or outside of The bonds are modified, e.g., locked (2'-O-4'-C methylene bridge), 5'-methylcytidine, 2'-O-methyl-pseudouridine, or In some embodiments, the bose phosphate backbone is replaced by a polyamide chain. The crRNA or part or all of the crRNA may be, for example, X 17~18 or X 17~1 9. Contains deoxyribonucleotides (e.g., all nucleotides) within or outside the target-complementary region. or some are DNA, e.g., DNA / RNA hybrids).

[0016] In another aspect, the present invention provides a method for detecting a target region of a double-stranded DNA molecule, e.g., within the genomic sequence of a cell. The present invention provides a method for modifying a dCas9-heterologous functional domain fusion protein. (dCas9-HFD); and a selected target sequence, preferably adjacent to the protospacer The target sequence adjacent to the 5' side of the motif (PAM), e.g., NGG, NAG, or NNGG A sequence consisting of 17 to 18 nucleotides or 17 to 19 nucleotides complementary to the complementary strand of A guide RNA, e.g., a ribonucleic acid described herein, is expressed in a cell or X 17~18 or X 17~19 is next to the rest of the RNA In some embodiments, R NA is one or more, e.g., up to three, e.g., one, two, or more, that are not complementary to the target sequence. or three additional nucleotides at the 5' end of the RNA molecule.

[0017] In another aspect, the present invention provides a method for identifying a target region of a double-stranded DNA molecule, e.g., within the genomic sequence of a cell. Methods for modifying, for example, introducing sequence-specific cleavage within such regions are provided. The method is: Cas9 nuclease or nickase or dCas9-heterologous functional domain fusion protein Protein (dCas9-HFD) and tracrRNA, e.g., the sequence GGAACCAUUCAAAACAGCAUAGCAA GUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGG CACCGAGUCGGUGC (SEQ ID NO: 8) or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAA AAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2405) or an active portion thereof ; AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUU GAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2407) or its active part; CAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAU CAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2409) or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAA AAGUG (SEQ ID NO: 2410) or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA (SEQ ID NO: 241 1) or an active portion thereof; or UAGCAAGUUAAAAUAAGGCUAGUCCG (SEQ ID NO: 2412) or a tracrRNA comprising or consisting of an active portion thereof; A selected target sequence, preferably a protospacer adjacent motif (PAM), such as NG G, NAG, or NNGG, and a 17-18 nucleotide sequence complementary to the complementary strand of the target sequence adjacent to the 5' side. crRNA containing a sequence of nucleotides or 17-19 nucleotides In some embodiments, the method further comprises expressing or introducing into a cell a The RNA sequence is: (X 17~18 or X 17~19 )GUUUUAGAGCUA(SEQ ID NO:2404); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG( SEQ ID NO: 2407); or (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(SEQ ID NO: 24 08) It has.

[0018] In some embodiments, the crRNA is (X 17~18 or X 17~19 )GUUU UAGAGCUAUGCUGUUUUG (SEQ ID NO: 2407) and tracrR NA is GGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGG CUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGU GC (SEQ ID NO: 8); 17~18 or X 17~19 )GUUUU AGAGCUA (SEQ ID NO: 2404) and the tracrRNA is UAGCAAGU UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCA CCGAGUCGGUGC (SEQ ID NO: 2405); or the cRNA is (X 17~ 18 or X 17~19 )GUUUUAGAGCUAUGCU (SEQ ID NO: 2408) and tracrRNA is AGCAUAGCAAGUUAAAAUAAGGCUAGU CCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(sequence Number 2406).

[0019] X 17~18 or X 17~19 is a naturally occurring sequence adjacent to the rest of the RNA. In some embodiments, the sequence is the same as the sequence of the RNA Pol III transcript. There is optionally one or more T's present which serve as termination signals to stop The RNA (e.g., tracrRNA or crRNA) contains one or more or multiple Us, e.g., 2 to 8 or more Us (e.g., U, UU, UUU, UU UU, UUUUU, UUUUUU, UUUUUUU, UUUUUUUU). In some embodiments, the RNA (e.g., tracrRNA or crRNA) is one or more, e.g., up to three, e.g., one, fragment at the 5' end of the nucleic acid sequence that is not complementary to the target sequence In some embodiments, the crRNA or or one or more nucleotides of tracrRNA, e.g., sequence X 17~18 Also is X 17~19 one or more nucleotides within or outside of the For example, locked (2'-O-4'-C methylene bridge), 5'-methylcytidine 2'-O-methyl-pseudouridine, or the ribose phosphate backbone is poly( In some embodiments, the tracrRNA or crR If some or all of NA is 17~18 or X 17~19 Within the target complementarity region or externally, containing deoxyribonucleotides (e.g., entirely or partially DNA). (e.g., DNA / RNA hybrids).

[0020] In some embodiments, a dCas9-heterologous functional domain fusion protein (dCas9 -HFD) is an HFD that modifies gene expression, histones, or DNA, e.g., transcriptional activity. activation domain, transcriptional repressors (e.g., heterochromatin protein 1 (HP1), silencers, such as HP1α or HP1β, modify the methylation status of DNA. enzymes (e.g., DNA methyltransferases (DNMTs) or TET proteins) that enzymes that modify histone subunits (e.g., histone amino acids, e.g., TET1) or histone cleavage enzymes (e.g., histone amino acids, e ... cetyltransferase (HAT), histone deacetylase (HDAC) or histone In a preferred embodiment, the heterologous functional domain comprises a transcriptional activation domains, such as VP64 or NF-κB p65 transcriptional activation domains; DNA demerger Enzymes that catalyze thiolation, such as members of the TET protein family or the catalytic domain of one of the members of the family; or histone modifications (e.g., L SD1, histone methyltransferase, HDAC or HAT) or transcription factor a silencing domain, e.g., heterochromatin protein 1 (HP1), e.g., HP a transcriptional silencing domain of HP1α or HP1β; or a biological tether, e.g., MS2, CRISPR / Cas subtype Ypest protein 4 (Csy4) or The lambda N protein. dCas9-HFD was published on March 15, 2013. U.S. Provisional Patent Application No. 61 / 799, filed with attorney docket number 00786-0882P02, No. 647, U.S. Patent Application No. 61 / 838,148 filed June 21, 2013, and and International Application No. PCT / US14 / 27335, all of which are incorporated herein by reference. The entirety of which is incorporated herein by reference.

[0021] In some embodiments, the target genomic sequence selected by the methods described herein Indel mutations or sequence changes occur in the

[0022] In some embodiments, the cell is a eukaryotic cell, e.g., a mammalian cell, e.g., a human cell. do.

[0023] Unless otherwise specified, all technical and scientific terms used herein are It has the same meaning as commonly understood by a person skilled in the art to which the invention pertains. describes methods and materials for use in the present invention, but other methods known in the art may also be used. Any suitable methods and materials may be used. The materials, methods and examples are merely illustrative. The following publications are mentioned in this specification and are not intended to be limiting. All references, including patent applications, patents, sequences, and database entries, are The entire disclosure is incorporated by reference. In the event of a conflict, the present specification, including definitions, will control. .

[0024] Other features and advantages of the invention are set forth in the following detailed description, drawings, and claims. It will become clear from [Brief explanation of the drawings]

[0025] [Figure 1] Figure 1: Schematic diagram showing the gRNA / Cas9 nuclease complex bound to the target DNA site. The scissors indicate the approximate cleavage point of the Cas9 nuclease at the genomic DNA target site. Note that the numbering of the nucleotides in the guide RNA proceeds backwards from 5' to 3'. [Figure 2-1] Figure 2A: Schematic diagram showing the principle of shortening the 5' complementarity region of a gRNA. Thick gray line = target DNA site, thin dark gray structure = gRNA, and black lines indicate base pairing (or lack thereof) between the gRNA and target DNA site. [Figure 2-2] Figure 2B: Schematic overview of the EGFP disruption assay. Repair of a targeted Cas9-mediated double-strand break in a single integrated EGFP-PEST reporter gene by error-prone NHEJ-mediated repair causes cells to generate a frameshift mutation that disrupts the coding sequence and concomitant loss of fluorescence. [Figure 2-3]Figures 2C-2F: Activity of RNA-guided nuclease (RGN) carrying single guide RNAs (gRNAs) with (C) a single mismatch, (D) adjacent double mismatches, (E) double mismatches spaced at various intervals, and (F) increasing numbers of adjacent mismatches, assayed at three different target sites in an EGFP reporter gene sequence. Shown are the average values ​​of replicate activities normalized to the activity of a perfectly matched single gRNA. Error bars represent the standard error of the mean. The location of the mismatch for each single gRNA is highlighted in gray in the grid below. The sequences of the three EGFP target sites were as follows: EGFP Site 1 GGGCACGGGCAGCTTGCCGGTGG (SEQ ID NO: 9); EGFP Site 2 GATGCCGTTCTTCTGCTTGTCGG (SEQ ID NO: 10); EGFP Site 3 GGTGGTGCAGATGAACTTCAGGG (SEQ ID NO: 11). [Figure 2-4] Figures 2C-2F: Activity of RNA-guided nuclease (RGN) carrying single guide RNAs (gRNAs) with (C) a single mismatch, (D) adjacent double mismatches, (E) double mismatches spaced at various intervals, and (F) increasing numbers of adjacent mismatches, assayed at three different target sites in an EGFP reporter gene sequence. Shown are the average values ​​of replicate activities normalized to the activity of a perfectly matched single gRNA. Error bars represent the standard error of the mean. The location of the mismatch for each single gRNA is highlighted in gray in the grid below. The sequences of the three EGFP target sites were as follows: EGFP Site 1 GGGCACGGGCAGCTTGCCGGTGG (SEQ ID NO: 9); EGFP Site 2 GATGCCGTTCTTCTGCTTGTCGG (SEQ ID NO: 10); EGFP Site 3 GGTGGTGCAGATGAACTTCAGGG (SEQ ID NO: 11). [Figure 2-5]Figures 2C-2F: Activity of RNA-guided nuclease (RGN) carrying single guide RNAs (gRNAs) with (C) a single mismatch, (D) adjacent double mismatches, (E) double mismatches spaced at various intervals, and (F) increasing numbers of adjacent mismatches, assayed at three different target sites in an EGFP reporter gene sequence. Shown are the average values ​​of replicate activities normalized to the activity of a perfectly matched single gRNA. Error bars represent the standard error of the mean. The location of the mismatch for each single gRNA is highlighted in gray in the grid below. The sequences of the three EGFP target sites were as follows: EGFP Site 1 GGGCACGGGCAGCTTGCCGGTGG (SEQ ID NO: 9); EGFP Site 2 GATGCCGTTCTTCTGCTTGTCGG (SEQ ID NO: 10); EGFP Site 3 GGTGGTGCAGATGAACTTCAGGG (SEQ ID NO: 11). [Figure 2-6]Figure 2G: Mismatches at the 5' end of the gRNA increase CRISPR / Cas sensitivity compared to mismatches at the 3' end. The gRNA forms Watson-Crick base pairs between the RNA and DNA, except at positions designated "m," which are mismatched using Watson-Crick transversion (i.e., EGFP site #2 M18-19 is mismatched by converting the gRNA to its Watson-Crick partner at positions 18 and 19). Positions near the 5' end of the gRNA are generally highly tolerant, but matches at these positions are critical for nuclease activity when other residues are mismatched. When all four positions are mismatched, nuclease activity is no longer detectable. This further indicates that matches at this 5' end can compensate for mismatches at other 3' ends. Note that these experiments were performed with a non-codon-optimized version of Cas9, which may exhibit lower absolute levels of nuclease activity than the codon-optimized version. Figure 2H: Efficiency of Cas9 nuclease activity induced by gRNAs with complementary regions of various lengths ranging from 15 to 25 nt in a human cell-based U2OS EGFP decay assay. Because gRNA expression from the U6 promoter requires the presence of a 5' G, only gRNAs with specific lengths of complementary region (15 nt, 17 nt, 19 nt, 20 nt, 21 nt, 23 nt, and 25 nt) to the target DNA site could be evaluated. [Figure 3-1] Figure 3A: Efficiency of EGFP decay in human cells mediated by Cas9 and full-length or truncated gRNAs against four target sites in the EGFP reporter gene. The length of the complementary region and the corresponding target DNA site are indicated. Ctrl = control gRNA lacking the complementary region. [Figure 3-2]Figure 3B: Efficiency of targeted indel mutations introduced into seven different human endogenous gene targets by matched standard RGN (Cas9 and standard full-length gRNA) and tru-RGN (Cas9 and gRNA with a truncation in the 5' complementary region). The length of the gRNA complementary region and the corresponding target DNA site are indicated. Indel frequencies were measured by T7EI assay. Ctrl = control gRNA lacking the complementary region. [Figure 3-3] Figure 3C: DNA sequences of insertion-deletion mutations induced by RGN using tru-gRNA or matching full-length gRNAs targeting the EMX1 site. The portion of the target DNA site interacting with the gRNA complementarity region is highlighted in gray, and the first base of the PAM sequence is shown in lowercase. Deletions are represented by dashed lines highlighted in gray, and insertions are represented by italicized letters highlighted in gray. The total number of deleted or inserted bases and the number of times each sequence was isolated are shown on the right. [Figure 3-4] Figure 3D: Efficiency of precise HDR / ssODN-mediated changes introduced into two endogenous human genes by matched standard RGN and tru-RGN. %HDR was measured using a BamHI restriction digestion assay (see Experimental Procedures in Example 2). Control gRNA = empty U6 promoter vector. [Figure 3-5]Figure 3E: U2OS.EGFP cells were transfected with varying amounts of full-length gRNA expression plasmid (top row) or tru-gRNA expression plasmid (bottom row) along with a constant amount of Cas9 expression plasmid, and the percentage of cells with reduced EGFP expression was assayed. Mean values ​​from duplicate experiments are shown along with the standard error of the mean. Note that for these three EGFP target sites, the data obtained with tru-gRNA closely matched the data from experiments performed with full-length gRNA expression plasmids instead of tru-gRNA plasmids. Figure 3F: U2OS.EGFP cells were transfected with varying amounts of Cas9 expression plasmid along with a constant amount of full-length gRNA expression plasmids for each target (top row) or tru-gRNA expression plasmids (bottom row) (the amount of each tru-gRNA was determined from the experiment in Figure 3E). Mean values ​​from duplicate experiments are shown along with the standard error of the mean. Note that for these three EGFP target sites, the data obtained with tru-gRNA closely matched the data from experiments performed with full-length gRNA expression plasmids instead of tru-gRNA plasmids. The results of these titrations determined the concentrations of plasmids used in the EGFP decay assays performed in Examples 1 and 2. [Figure 4]Figure 4A: Schematic diagram showing the location of VEGFA sites 1 and 4 targeted by gRNA for paired double nicks. The target site of the full-length gRNA is underlined, and the first base of the PAM sequence is shown in lowercase. The location of the BamHI restriction site inserted by HDR using an ssODN donor is indicated. Figure 4B: tru-gRNA can be used with a paired nickase strategy to efficiently induce indel mutations. Replacing the full-length gRNA targeting VEGFA site 1 with tru-gRNA does not reduce the efficiency of indel mutations observed with a paired full-length gRNA targeting EGFA site 4 and Cas9-D10A nickase. The control gRNA used is a gRNA lacking the complementary region. Figure 4C: tru-gRNA can be used with a paired nickase strategy to efficiently induce precise HDR / ssODN-mediated sequence changes. Replacing the full-length gRNA against VEGFA site 1 with the tru-gRNA does not reduce the efficiency of indel mutations observed with a paired full-length gRNA against VEGFA site 4 using a ssODN donor template and Cas9-D10A nickase. The control gRNA used is a gRNA lacking the complementary region. [Figure 5-1] Figure 5A: Activity of RGN targeting each of the three EGFP sites using full-length gRNAs (top row) or tru-gRNAs (bottom row) with a single mismatch at each position (except the 5'-most base, which must remain G for efficient expression from the U6 promoter). The gray boxes in the bottom grid represent the position of the Watson-Crick transversion mismatch. The empty gRNA control used is a gRNA lacking the complementary region. RGN activity was measured using an EGFP decay assay, and the values ​​shown represent the percentage of EGFP negativity observed relative to RGN using perfectly matched gRNAs. Experiments were performed in duplicate, and the average values ​​are shown with error bars representing the standard error of the mean. [Figure 5-2]Figure 5B: Activity of RGN targeting each of the three sites in EGFP using full-length gRNAs (top row) or tru-gRNAs (bottom row) with double mismatches adjacent to each position (except the 5'-most base, which must be G for efficient expression from the U6 promoter). Data are presented as in Figure 5A. [Figure 6-1] Figure 6A: Absolute frequencies of on-target and off-target indel mutations induced by RGN targeting three different endogenous human gene sites, as measured by deep sequencing. Indel frequencies are shown for the three target sites in cells expressing targeted RGN with full-length gRNA, tru-gRNA, or a control gRNA lacking the complementary region. The absolute numbers of indel mutations used to generate these graphs can be found in Table 3B. [Figure 6-2]Figure 6B: Fold improvement in off-target site specificity for three tru-RGNs. The values ​​shown represent the ratio of on-target activity / off-target activity of tru-RGN to that of standard RGN, calculated using the data in (A) and Table 3B for the indicated off-target sites. For sites marked with an asterisk (*), no indels were observed with tru-RGN, so the values ​​shown represent conservative statistical estimates of specificity for these off-target sites (see Results and Experimental Procedures). Figure 6C, top: Comparison of on-target and off-target sites for tru-RGN targeting VEGFA site 1 identified by T7EI assay (more sites were identified by deep sequencing). Note that the full-length gRNA mismatches two nucleotides at the 5' end of the target site, two nucleotides that are not present in the tru-gRNA target site. Mismatches at the off-target site relative to the on-target site are highlighted in bold and underlined. Mismatches between the gRNA and the off-target site are indicated by an X. Figure 6C, bottom panel: Frequency of indel mutations induced at off-target sites by RGN with full-length or truncated gRNAs. Indel mutation frequencies were determined by T7EI assay. Note that the off-target site in this figure is the site designated OT1-30, which we previously investigated for indel mutations induced by standard RGN targeting VEGFA site 1 (see Example 1 and Fu et al., Nat Biotechnol. 31(9):822-6 (2013)). Because the frequency of indel mutations appears to be within the reliable detection limit of the T7EI assay (2-5%), it is possible that off-target mutations were not identified at this site in our previous experiments. [Figure 7-1] Figures 7A-7D: DNA sequences of indel mutations induced by RGN using tru-gRNA or matching full-length gRNAs targeting VEGFA sites 1 and 3. Sequences are depicted as in Figure 3C. [Figure 7-2]Figures 7A-7D: DNA sequences of indel mutations induced by RGN using tru-gRNA or matching full-length gRNAs targeting VEGFA sites 1 and 3. Sequences are depicted as in Figure 3C. [Figure 7-3] Figures 7A-7D: DNA sequences of indel mutations induced by RGN using tru-gRNA or matching full-length gRNAs targeting VEGFA sites 1 and 3. Sequences are depicted as in Figure 3C. [Figure 7-4] Figures 7A-7D: DNA sequences of indel mutations induced by RGN using tru-gRNA or matching full-length gRNAs targeting VEGFA sites 1 and 3. Sequences are depicted as in Figure 3C. [Figure 7-5] Figure 7E: Frequency of indel mutations induced by tru-gRNAs with mismatched 5' G nucleotides. Shown are the frequencies of indel mutations induced in human U2OS.EGFP cells by Cas9 directed by tru-gRNAs with 17-nt, 18-nt, or 20-nt regions of complementarity to VEGFA sites 1 and 3 and EMX1 site 1. The three gRNAs contain a mismatched 5' G (denoted by the position marked in bold). Bars represent the results of experiments with full-length gRNA (20 nt), tru-gRNA (17-nt or 18-nt), and tru-gRNA with a mismatched 5' G nucleotide (17-nt or 18-nt with a bold T at the 5' end). (Note that no activity was detected with the mismatched tru-gRNA against EMX1 site 1.) [Figure 8-1] Figures 8A-8C: Sequences of off-target indel mutations induced by RGN in human U2OS.EGFP cells. Wild-type genomic off-target sites recognized by RGN (including PAM sequences) are highlighted in gray and numbered as in Tables 1 and B. Note that the complementary strand is shown for some sites. Deletions are indicated by dashed lines on a gray background. Insertions are italicized and highlighted in gray. [Figure 8-2]Figures 8A-8C: Sequences of off-target indel mutations induced by RGN in human U2OS.EGFP cells. Wild-type genomic off-target sites recognized by RGN (including PAM sequences) are highlighted in gray and numbered as in Tables 1 and B. Note that the complementary strand is shown for some sites. Deletions are indicated by dashed lines on a gray background. Insertions are italicized and highlighted in gray. [Figure 8-3] Figures 8A-8C: Sequences of off-target indel mutations induced by RGN in human U2OS.EGFP cells. Wild-type genomic off-target sites recognized by RGN (including PAM sequences) are highlighted in gray and numbered as in Tables 1 and B. Note that the complementary strand is shown for some sites. Deletions are indicated by dashed lines on a gray background. Insertions are italicized and highlighted in gray. [Figure 9-1] Figures 9A-9C: Sequences of off-target indel mutations induced by RGN in human HEK293 cells. Wild-type genomic off-target sites recognized by RGN (including PAM sequences) are highlighted in gray and numbered as in Tables 1 and B. Note that the complementary strand is shown for some sites. Deletions are indicated by dashed lines on a gray background. Insertions are italicized and highlighted in gray. *Many single-bp indels were generated. [Figure 9-2] Figures 9A-9C: Sequences of off-target indel mutations induced by RGN in human HEK293 cells. Wild-type genomic off-target sites recognized by RGN (including PAM sequences) are highlighted in gray and numbered as in Tables 1 and B. Note that the complementary strand is shown for some sites. Deletions are indicated by dashed lines on a gray background. Insertions are italicized and highlighted in gray. *Many single-bp indels were generated. [Figure 9-3]Figures 9A-9C: Sequences of off-target indel mutations induced by RGN in human HEK293 cells. Wild-type genomic off-target sites recognized by RGN (including PAM sequences) are highlighted in gray and numbered as in Tables 1 and B. Note that the complementary strand is shown for some sites. Deletions are indicated by dashed lines on a gray background. Insertions are italicized and highlighted in gray. *Many single-bp indels were generated. DETAILED DESCRIPTION OF THE INVENTION

[0026] (Detailed explanation) CRISPR RNA-guided nuclease (RGN) is a simple and efficient method for genome editing. It has rapidly emerged as a platform. t Biotechnol 31,233-239(2013)) has recently reported that Ca Although the specificity of s9 RGNs has been systematically investigated, the specificity of RGNs in human cells has not been fully elucidated. These nucleases are widely used in research and therapeutic applications. If this is the case, the extent of off-target effects of RGN in eukaryotic cells, including humans, may be unknown. It is extremely important to understand the range of the phenotype. We used SEI to characterize off-target cleavage by Cas9-based RGN. Depending on the position along the guide RNA (gRNA)-DNA junction, single and multiple nucleotide sequences are expressed to varying degrees. Two or more mismatches were allowed. By examining the partial mismatched sites, Oocytes induced by four of six RGNs targeted to endogenous loci in mouse cells Off-target alterations were rapidly detected. Up to five off-target sites were identified. Many of these match with the frequency observed at the intended on-target site. Therefore, RGN is mutated at frequencies equal to or greater than It also showed high activity against imperfectly matched RNA-DNA, and this observation was This can complicate their use in therapeutic applications.

[0027] The results described herein make it possible to predict the specificity profile of any RGN. This shows that the EGFP reporter assay is neither easy nor simple. Experiments have shown that single and double mismatches can have different effects on RGN activity in human cells. The effect depends strictly on the position(s) within the target site of the mismatch. For example, as previously reported, Generally, changes found in the 3' half of the sgRNA / DNA junction are more frequent than those found in the 5' half. Although the impact is greater than the changes that can be observed (Jiang et al., Nat Biotechnol 31 ,233-239(2013);Cong et al.,Science 339,819-823 (2013); Jinek et al., Science 337, 816-821(2012)) However, single and double mutations at the 3' end appear to be well tolerated in some cases. In contrast, double mutations at the 5' end can significantly reduce activity. The magnitude of the effect of mismatches (one or more) appears to be site-dependent. Nucleotide substitutions (Watson-Kuhn-Schmidt used in our EGFP reporter experiments) We have tested a wide range of RGNs (beyond the RGN transformation) and Conducting profiling provides further insight into the extent of off-target interactions. In this regard, the recent detailed study by Marraffini et al. Bacterial cell-based method (Jiang et al., Nat Biotechnol 31, 233-2 39 (2013)) or the in vitro recombination previously applied to ZFNs by Liu et al. A natorial library-based cleavage site selection method (Pattanayak et al., Nat Methods 8, 765-770 (2011)) has demonstrated even greater RGN specificity. This is thought to be useful for creating files.

[0028] Although there are difficulties in predicting RGN specificity on a broad scale, By examining a subset of genomic regions that differ from the target site by 1 to 5 mismatches, This allowed us to identify the true off-target effects of RGN. Under the conditions of these experiments, RGN-induced mutations at many of these off-target sites were not observed. The frequency is similar to (or higher than) that observed at the intended on-target site and the T7EI assay (performed in our laboratory) yielded reliable mutation frequencies. It was possible to detect these sites using a method with a detection limit of approximately 2-5%. These mutation rates are so high that ZFN-induced off-target mutations are far less frequent than previously thought. Deep sequencing required for detection of alterations and TALEN-induced off-target alterations The use of the lacing method could be avoided (Pattanayak et al., Nat Method s 8,765-770(2011); Perez et al., Nat Biotechnol 26, 808-816(2008); Gabriel et al., Nat Biotechnol 29, 816-823 (2011); Hockemeyer et al., Nat Biotec hnol 29,731-734(2011)). In addition, RGN activity in human cells Analysis of targeted mutagenesis confirms the difficulty of predicting RGN specificity All single and double mismatched off-target sites showed evidence of mutations. Although not all of the mutations were present, some mutations were also observed at sites with up to five mismatches. True off-target sites identified include those that are involved in metastasis or transformation compared to the intended target sequence. There is no obvious bias towards differences due to the study (Table E; grey highlighted rows).

[0029] Although off-target sites were found in some RGNs, the identification of such sites was extensive. The six RGNs studied were Only a small fraction of the much larger total number of potential off-target sequences in the human genome (Sites that differ from the intended target site by 3 to 6 nucleotides; compare Table E with Table C. We investigated the inheritance of such a large number of off-target mutations using the T7EI assay. Although considering fertility is neither a practical nor cost-effective strategy, future research could address this issue. Using high-throughput sequencing, many potential off-target sites can be identified. This allows for the detection of true off-target mutations with higher sensitivity. For example, if such a method is used, the inventors can For the two types of RGN for which no off-target mutations were identified, further off-target mutations were identified. Furthermore, it is thought that it will be possible to clarify the RGN activity in cells. RGN specificity and epigenomic factors (e.g., DNA methylation and As our understanding of both the chromatin state and the chromatin domain improves, other potential sites will emerge that need to be explored. This reduces the number of RGN off-targets, making genome-wide assessment of RGN off-targets more practical. This means that the price is likely to be affordable.

[0030] As described herein, to minimize the frequency of genomic off-target mutations, Several strategies can be used, for example, to optimize the specific selection of RGN target sites. It is possible to target the intended target site and the off-target site at up to five different positions. If a position can be efficiently mutated by RGN, the mismatch counts will determine whether the position is mutated efficiently. It is unlikely that it would be effective to select a target site with the minimum number of off-target sites that would be Any RGN that targets a sequence within the human genome typically contains a 20 bp RNA:DNase. There are thousands of potential off-target sites that vary by four or five positions within the A-complementarity region. (See, for example, Table C.) In addition, the nucleic acid sequence of the gRNA complementarity region Nucleotide content may influence the extent of potential off-target effects. For example, it has been shown that a high GC content stabilizes RNA:DNA hybrids (S Ugimoto et al., Biochemistry 34, 11211-11216 (199 5)), thus, the stability of gRNA / genomic DNA hybridization and It is expected that the tolerance for mismatch will also increase. The number of mismatch sites within the genome and the stability of RNA:DNA hybrids across the genome To assess the potential and mechanisms by which gRNAs affect RGN specificity across multiple RGNs, Further experiments with larger numbers are needed, but determining these predictive parameters is likely to be successful. Even if this is possible, the impact of implementing such guidelines will likely be limited to a limited scope of RGN targeting. It is important to note that this may result in further limitations.

[0031] One promising general strategy to reduce RGN-induced off-target effects is intracellular One possible solution is to reduce the concentration of gRNA and Cas9 nuclease expressed in the u We tested this idea by using RGN at VEGFA target sites 2 and 3 in 2OS.EGFP cells. The amount of sgRNA and Cas9 expression plasmids transfected was reduced. The mutation rate at the on-target site was reduced, but the relative proportion of off-target mutations was The results showed that the expression of the α-glucan in the α-glucan in the β ... The absolute ratio of on-target mutagenesis was also significantly higher in both cytoplasmic types (HEK293 and K562 cells). Although the rate was lower than in U2OS.EGFP cells, a high level of off-target mutagenesis was observed. Therefore, even if the expression levels of gRNA and Cas9 in cells were reduced, It is unlikely to be a solution to reduce off-target effects. However, the high rate of off-target mutagenesis observed in human cells is due to the presence of gRNA and / or C This suggests that this is not due to overexpression of as9.

[0032] [Table 1]

[0033] [Table 2]

[0034] RGN induces significant off-target mutagenesis in three different human cell types The observation that this can be achieved has important implications for the use of this genome editing platform. When applied to research, the distribution of unwanted changes is a particular issue in cultured cells with long generation times. Cellular or organismal experiments address the potentially complex effects of frequent off-target mutations. Off-target effects are not random and must be taken into account. Since this effect is related to the position of the gene, one way to control this effect is to use different DNA sequences. It is conceivable that multiple RGNs could be targeted to induce the same genomic alterations. When applied to therapy, the observations presented here suggest that these nucleases may be useful in the treatment of human diseases. If safe long-term use is to be achieved, RGN specificity must be carefully defined and / or Or clearly indicates a need for improvement.

[0035] How to improve specificity As shown herein, the Streptococcus pyogenes (S. pyogenes) Cas9 protein Protein-based CRISPR-Cas RNA-guided nucleases are used to target target genes. May have significant off-target mutagenic effects equal to or greater than target activity (Example 1) Such off-target effects are a significant concern in research applications, especially in future therapeutic applications. Therefore, CRISPR-Cas RNA-guided nucleases ( There is a need for methods to improve the specificity of RGN.

[0036] As described in Example 1, Cas9 RGN is highly expressed at off-target sites in human cells. It can induce frequent indel mutations (see also Cradick et al., 2013; Fu et al., 2014). (See Hsu et al., 2013; Pattanayak et al., 2013). Such undesirable changes are caused by as many as five mismatches with the intended on-target site. gRNA complementarity can occur in different genomic sequences (see Example 1). Mismatches at the 5' end of the functional region are generally more tolerated than mismatches at the 3' end. However, such a relationship is not absolute and shows site dependence (see Example 1 and Fu et al., 2013; Hsu et al., 2013; Pattanayak et al., 2013). As a result, currently, computational methods that depend on the number and / or position of mismatches are available. The method is of reduced predictive value for identifying true off-target sites. Therefore, if RNA-guided nucleases are to be used in research and therapeutic applications, off-target Methods to reduce the frequency of target mutations remain a key priority.

[0037] Truncated guide RNAs (tru-gRNAs) increase specificity Generally speaking, there are two systems of guide RNAs that work together to induce transcription by Cas9. System 1 and two separate systems using separate crRNA and tracrRNA to induce cleavage Chimeric crRNA-tracrRNA hybrids combining guide RNAs into a single system System 2 uses a single guide RNA (called sgRNA). (See, e.g., J. Med., 2012;337:816-821). The acrRNA can be truncated to various lengths, and the various lengths are expressed in distinct systems ( It has been shown to function in both systems (system 1) and chimeric gRNA systems (system 2). For example, In some embodiments, the tracrRNA is at least 1 nt, 2 nt from its 3' end. nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 15nt , 20nt, 25nt, 30nt, 35nt or 40nt shortened. In some embodiments, the tracrRNA molecule comprises at least 1 nt, 2n t, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 15nt, It may be shortened by 20 nt, 25 nt, 30 nt, 35 nt or 40 nt. In this case, the tracrRNA molecule may be fragmented from both the 5' and 3' ends, e.g., At least 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10 nt, 15 nt or 20 nt, at least 1 nt, 2 nt, 3 nt, 4 nt at the 3' end nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 15nt, 20nt, 25 The sequence may be shortened by nt, 30 nt, 35 nt, or 40 nt. Science 2012;337:816-821;Mali et al.Science.2 013 Feb 15;339(6121):823-6;Cong et al., Science 2013 Feb 15;339(6121):819-23; and Hwang Yo and Fu et al., Nat Biotechnol. 2013 Mar;31(3):227- 9; see Jinek et al., Elife 2, e00471 (2013). Generally, it has been shown that the longer the chimeric gRNA, the higher the on-target activity. However, the relative specificities of gRNAs of various lengths have not yet been clarified. Therefore, in some cases it may be desirable to use a shorter gRNA. In this state, the gRNA is located within approximately 100 to 800 bp upstream of the transcription start site. For example, within about 500 bp upstream of the transcription start site, or within about 100 to 150 bp downstream of the transcription start site. It is complementary to a region that is within 800 bp, e.g., within about 500 bp. In some embodiments, vectors (e.g., plasmids) encoding two or more gRNAs, e.g., , two, three, four, five or more targeting different sites in the same region of the target gene Use a plasmid encoding a larger gRNA.

[0038] This application aims to shorten the complementary region of the gRNA rather than lengthen it, which may seem counterintuitive. This paper describes strategies to improve RGN specificity based on these ideas. A short gRNA was integrated into a single EGFP reporter gene and the endogenous human gene Similar (and in some cases even better) efficacy than full-length gRNAs at multiple sites in the Cas9-mediated on-target genome editing events can be induced in various types, exhibiting different efficiencies. Furthermore, RGNs using such short gRNAs have been shown to target a small number of gRNA-target DNA junctions. The sensitivity to mismatches is increased. Most importantly, the use of truncated gRNAs results in The incidence of genomic off-target effects in mouse cells was significantly reduced, and 5 This truncated gRNP fragment provides a 200-fold improvement in specificity. The strategy using A is to avoid the potential for mutation induction without compromising on-target activity. Highly efficient method that reduces off-target effects without the need to express a second gRNA This method alone can be used to reduce the off-target effects of RGN in human cells. This can be done either as a single antibody or in conjunction with other strategies such as paired nickases. It is also possible.

[0039] Therefore, one way to increase the specificity of CRISPR / Cas nucleases is to The length of the guide RNA (gRNA) species used to direct the specificity of the nuclease can be shortened. The 5' end of the genomic DNA target site has a complementary strand of 17-18 nt. Id RNA, e.g., a single gRNA or crRNA (paired with tracrRNA) to insert additional adjacent, for example, protospacer adjacent motifs (PAMs) of the sequence NGG. ) to guide the Cas9 nuclease to specific 17-18 nt genomic targets. This is possible (Figure 1).

[0040] Although it may be expected that increasing the length of the complementary region of the gRNA would improve specificity, The present inventors (Hwang et al., PLoS One. 2013 Jul 9;8(7):e6 8708) and others (Ran et al., Cell. 2013 Sep 12;154(6) :1380-9) previously reported that increasing the target region complementary to the 5' end of gRNA increases the onset of We have observed that the efficiency of function at the target site actually decreases.

[0041] On the other hand, in the experiment of Example 1, multiple mismatches were found within the standard length 5' complementary targeting region. The gRNA containing the nucleotide sequence still reliably induces Cas9-mediated cleavage of its target site. Therefore, it was shown that it is possible to synthesize short fragments lacking these 5'-terminal nucleotides. It was thought that truncated gRNAs might exhibit activity comparable to their full-length counterparts (Figure 2A). Furthermore, these 5' nucleotides usually occupy positions other than the gRNA-target DNA junction. Shorter gRNAs are more sensitive to mismatches because they are thought to cancel out the mismatches. , and thus was predicted to induce low levels of off-target mutations (Figure 2A). .

[0042] Alternatively, if the length of the target DNA sequence is shortened, gRNA:DNA hybrids can be used. This reduces the stability of the target and reduces mismatch tolerance, thereby increasing targeting specificity. In other words, shortening the gRNA sequence to recognize shorter DNA targets is expected to increase the This allows for low tolerance of even single nucleotide mismatches and therefore high specificity. This actually yields RNA-guided nucleases with reduced off-target effects. It is thought that...

[0043] This strategy of shortening the gRNA complementarity region has been demonstrated for other Cas proteins of bacterial or archaeal origin. Cas9 mutants that nick proteins and single strands of DNA, or either or both Nuclease activity of dCas9 and other proteins with catalytically inactivating mutations in the nuclease domain Streptococcus pyogenes (S. pyogenes) Cas9 and its related variants, including Cas9 mutants lacking the This strategy could potentially be used for other RNA-guided proteins. In addition to systems using gRNAs, dual gRNAs (e.g., crRNA and transRNA as seen in natural systems) have also been used. This can be applied to systems that use crRNA.

[0044] Therefore, herein, fusion with normally transcoded tracrRNA is described. Single guide RNA containing a crRNA, e.g., Mali et al., Science 201 3 Feb 15;339(6121):823-6 It is a proto-RNA but contains a protospacer adjacent motif (PAM), e.g., NGG, NAG, or or less than 20 nucleotides (nt) of the complementary strand of the target sequence immediately 5' to the NNGG; For example, complementary to 19nt, 18nt or 17nt, preferably 17nt or 18nt In some embodiments, guide RNAs are described that have a truncated C sequence at the 5' end. The as9 guide RNA has the sequence: (X 17~18 or X 17~19 )GUUUUAGAGCUA(SEQ ID NO:2404); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG( SEQ ID NO: 2407); or (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(SEQ ID NO: 24 08); (X 17~18 or X 17~19 )GUUUUAGAGCUAGAAAUAGCAAG UUAAAAUAAGGCUAGUCCG(X N )(SEQ ID NO:1); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGAAAAGC AUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N )(Array number No. 2); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUGG AAACAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGU UAUC(X N ) (SEQ ID NO: 3); (X 17~18 or X 17~19 )GUUUUAG AGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC AACUUGAAAAAGUGGCACCGAGUCGGUGC(X N ) (SEQ ID NO: 4) , (X 17~18 or X 17~19 )GUUUAAGAGCUAGAAAUAGCAAG UUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGC ACCGAGUCGGUGC (SEQ ID NO: 5); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 6); or (X 17~1 8 or X 17~19 )GUUUAAGAGCUAUGCUGGGAAACAGCAUAG CAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAG UGGCACCGAGUCGGUGC (SEQ ID NO: 7); In the sequence, X 17~18 or X 17~19 Each of the A nucleotide sequence complementary to 7-18 nucleotides or 17-19 nucleotides In addition, the present specification also includes the above-mentioned references (Jinek et al., Science. 337(6096) ):816-21(2012) and Jinek et al., Elife.2:e00471(2 DNA encoding the truncated Cas9 guide RNAs described in .

[0045] The guide RNA can be any sequence that does not interfere with the binding of the ribonucleic acid to Cas9. N wherein N (in the RNA) is 0 to 200, for example, 0 to 100, 0 to 50, or 0 It can be ~20.

[0046] In some embodiments, the guide RNA has one or more adenines (A ) or uracil (U) nucleotides. In some embodiments, the RNA Pol III. One or more T's are optionally used as termination signals to stop transcription. The RNA contains one or more Us, e.g., 1 to 8 or more, at the 3' end of the molecule. More than one U (e.g., U, UU, UUU, UUUU, UUUUU, UUUUUU, UUU UUUU, UUUUUUUU).

[0047] Modified RNA oligonucleotides, such as locked nucleic acids (LNA), are used to by locking the RNA-DNA fragment into a more favorable (stable) conformation. It has been shown that 2'-O- -methyl RNA is a modified base with an additional covalent bond between the 2' oxygen and the 4' carbon. Incorporation of this into oligonucleotides improves overall thermal stability and selectivity It is possible to (Formula I). [ka]

[0048] Thus, in some embodiments, the tru-gRNAs disclosed herein comprise: The modified RNA oligonucleotides may comprise one or more of the short RNA oligonucleotides described herein. A shortened guide RNA molecule is a 17-18 nt or 17-20 nt guide RNA molecule complementary to the target sequence. One, part or all of the 19n5' region may be modified, e.g., locked. (2'-O-4'-C methylene bridge), may be 5'-methylcytidine, 2' -O-methyl-pseudouridine, or the ribose phosphate backbone may be polyamino It may be a ribonucleic acid (peptide nucleic acid) or may be, for example, a synthetic ribonucleic acid.

[0049] In other embodiments, one, some, or all nucleotides of the tru-gRNA sequence are modified. It may be decorated, for example, locked (2'-O-4'-C methylene bridge). , 5'-methylcytidine, and 2'-O-methyl-pseudouridine. Alternatively, the ribose phosphate backbone may be replaced by a polyamide chain (peptide nucleic acid). ), for example, may be a synthetic ribonucleic acid.

[0050] In the context of cells, the complex of Cas9 and the synthetic gRNA listed above produces CRISP. It is possible to improve the genome-wide specificity of the R / Cas9 nuclease system.

[0051] An exemplary modified or synthetic tru-gRNA has the following sequence: (X 17~18 or X 17~19 )GUUUUAGAGCUA(X N ) (SEQ ID NO: 24 04); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG( X N )(SEQ ID NO:2407); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(X N)(array No. 2408); (X 17~18 or X 17~19 )GUUUUAGAGCUAGAAAUAGCAAG UUAAAAUAAGGCUAGUCCG(X N )(SEQ ID NO:1); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGAAAAGC AUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N )(Array number No. 2); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUGG AAACAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGU UAUC(X N )(SEQ ID NO:3); (X 17~18 )GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAA GGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCG GUGC(X N ) (SEQ ID NO: 4), (X 17~18 or X 17~19 )GUUUAAGAGCUAGAAAUAGCAAG UUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGC ACCGAGUCGGUGC (SEQ ID NO: 5); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 6); or (X 17~18 or X 17~19)GUUUAAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7); In the sequence, X 17~18 or X 17~19 is that Each of these is a target sequence, preferably a protospacer adjacent motif (PAM), such as NGG, 17-18nt or 17-19nt of the target sequence adjacent to the 5' side of NAG or NNGG t, and further comprising one or more nucleotides in the sequence, e.g., the sequence X 17~18 Or X 17~19 One or more nucleotides in the sequence X N Within One or more nucleotides or one or more sequences within any sequence of the tru-gRNA Multiple nucleotides are locked. X N interferes with the binding of ribonucleic acid to Cas9 N (in RNA) is any sequence not containing N, where N is 0 to 200, e.g., 0 to 100, 0 to 50 Alternatively, the number may be 0 to 20. In some embodiments, the number of transcription cycles is 1 to 20. There are optionally one or more T's used as termination signals to NA has one or more Us at the 3' end of the molecule, e.g., 1 to 8 or more Us (e.g., For example, U, UU, UUU, UUUU, UUUUU, UUUUUU, UUUUUUU, UU UUUUUU).

[0052] Although some of the examples described herein use a single gRNA, other applications of this method may also be used. Used with dual gRNAs (e.g., crRNA and tracrRNA found in natural systems) In this case, a single tracrRNA may be used to express multiple tracrRNAs using the system of the present invention. Used with different crRNAs, for example: (X 17~18 or X 17 ~19 )GUUUUAGAGCUA(SEQ ID NO:2404); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG( SEQ ID NO: 2407); or (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(SEQ ID NO: 24 08); and the tracrRNA sequence, where the crRNA is The method and molecule described herein can be used as a guide RNA to tracrRNA with the same or different D In some embodiments, the method comprises: providing a cell with a nucleic acid sequence encoding the sequence GGA ACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUC CGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence number No. 8) or an active portion thereof (an active portion is a molecule that forms a complex with Cas9 or dCas9) and tracrRNA containing or consisting of a portion of the tracrRNA that retains the ability to bind to the tracrRNA. In some embodiments, the tracrRNA molecule comprises a 3' end At least 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt from the end, 9nt, 10nt, 15nt, 20nt, 25nt, 30nt, 35nt or 40nt In another embodiment, the tracrRNA molecule may be truncated by adding a small amount of At least 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 15nt, 20nt, 25nt, 30nt, 35nt or 40nt shortened Alternatively, the tracrRNA molecule may be For example, at least 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt at the 5' end t, 8nt, 9nt, 10nt, 15nt or 20nt, at least 1nt at the 3' end , 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 15 nt, 20 nt, 25 nt, 30 nt, 35 nt, or 40 nt shortened. Exemplary tracrRNA sequences include SEQ ID NO:8 as well as the following: UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAA AAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2405) or an active portion thereof ; AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUU GAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2407) or its active part; CAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAU CAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2409) or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAA AAGUG (SEQ ID NO: 2410) or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA (SEQ ID NO: 241 1) or an active portion thereof; or UAGCAAGUUAAAAUAAGGCUAGUC CG (SEQ ID NO: 2412) or an active portion thereof.

[0053] (X17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG( In some embodiments, the following trac Using rRNA: GGAACCAUUCAAAACAGCAUAGCAAGUUAAA AUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGC (SEQ ID NO: 8) or an active portion thereof. (X 17~18 or X 17~1 9) Some clones using GUUUUAGAGCUA (SEQ ID NO: 2404) as crRNA In this embodiment, the following tracrRNA is used: UAGCAAGUUAAAAUAA GGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCG GUGC (SEQ ID NO: 2405) or an active portion thereof. (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU (SEQ ID NO: 2408) is used as crRNA. In some embodiments, the following tracrRNA is used: AGCAUAGCAAGUU AAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCAC CGAGUCGGUGC (SEQ ID NO: 2406) or an active portion thereof.

[0054] Furthermore, in systems using separate crRNA and tracrRNA, one or both of them It is synthetic and contains one or more modified (e.g., locked) nucleotides or deoxyribonucleotides. It may contain polynucleotides.

[0055] In some embodiments, a single guide RNA and / or crRNA and / or tracrRNA contains one or more adenines (A) or uracils (U) at the 3' end. It may contain nucleotides.

[0056] Existing Cas9-based RGNs are useful for directing targeting to desired genomic sites. However, the RNA-DNA heteroduplex formation is A duplex can form a more promiscuous range of structures than its DNA-DNA counterpart. The DNA-DNA duplex is more sensitive to mismatches, which is why DNA-induced RNA-guided nucleases do not readily bind to off-target sequences and are Therefore, the short guide RNAs described herein have high specificity compared to the short guide RNAs described herein. A is a hybrid, i.e., one or more deoxyribonucleotides, e.g., a short The DNA oligonucleotide is complementary to all or part of the gRNA, e.g., the complementary region of the gRNA. The DNA-based molecule may be a hybrid in which all or part of It may replace all or part of a single gRNA system or may be a dual crRNA / tr It may replace all or part of the crRNA of the acrRNA system. Because NA duplexes have a lower tolerance for mismatches than RNA-DNA duplexes, Such a system that integrates DNA into the complementary region allows for more reliable targeting of the genome. The target DNA sequence should be a DNA sequence. Methods for creating such duplexes are well known in the art. and is known in, for example, Barker et al., BMC Genomics. 2005 Apr. 22;6:57; and Sugimoto et al., Biochemistry. 2000 See Sep 19;39(37):11270-81.

[0057] An exemplary modified or synthetic tru-gRNA has the following sequence: (X 17~18 or X 17~19 )GUUUUAGAGCUAGAAAUAGCAAG UUAAAAUAAGGCUAGUCCG(X N )(SEQ ID NO:1); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGAAAAGC AUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N )(Array number No. 2); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUGG AAACAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGU UAUC(X N ) (SEQ ID NO: 3); (X 17~18 or X 17~19 )GUUUUAG AGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC AACUUGAAAAAGUGGCACCGAGUCGGUGC(X N ) (SEQ ID NO: 4) , (X 17~18 or X 17~19 )GUUUAAGAGCUAGAAAUAGCAAG UUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGC ACCGAGUCGGUGC (SEQ ID NO: 5); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 6); or (X 17~18 or X 17~19 )GUUUAAGAGCUAUGCUGGAAACA GCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7); In the sequence, X 17~18 or X 17~19 is that Each of these is a target sequence, preferably a protospacer adjacent motif (PAM), such as NGG, 17-18nt or 17-19nt of the target sequence adjacent to the 5' side of NAG or NNGG t, and further comprising one or more nucleotides in the sequence, e.g., the sequence X 17~18 Or X 17~19 One or more nucleotides in the sequence X N Within One or more nucleotides or one or more sequences within any sequence of the tru-gRNA Multiple nucleotides are deoxyribonucleotides. N ribonucleic acid and Cas9 N (in RNA) is 0 to 200, e.g., 0 to 1 In some embodiments, the RNA PolI II. Optionally, one or more T's are present, which serve as termination signals to stop transcription. Therefore, RNA has one or more Us, e.g., 1 to 8 or more, at the 3' end of the molecule. More than one U (e.g., U, UU, UUU, UUUU, UUUUU, UUUUUU, UUU UUUU, UUUUUUUU).

[0058] Furthermore, in systems using separate crRNA and tracrRNA, one or both of them It may be synthetic and contain one or more deoxyribonucleotides.

[0059] In some embodiments, a single guide RNA or crRNA or tracrRNA contains one or more adenine (A) or uracil (U) nucleotides at the 3' end. nothing.

[0060] In some embodiments, gRNAs are used in conjunction with genome sequences to minimize off-target effects. targeting a portion of the genome that differs from the sequence of the rest of the genome by at least three mismatches .

[0061] The described method involves transfecting cells with the truncated Cas9 gRNA (tru-gR) described herein. NA) (optionally a modified tru-gRNA or a DNA / RNA hybrid tru- gRNA) and nucleases that can be guided by this truncated Cas9 gRNA, such as For example, a Cas9 nuclease, such as the Cas9 nuclease described in Mali et al. Cas9 nickase as described in Jinek et al., 2012; or dCas9-heterologous The functional domain fusion (dCas9-HFD) is expressed or contacted with cells. This may include:

[0062] Cas9 Several bacteria express Cas9 protein variants. Currently, Streptococcus pyogenes (St Reptococcus pyogenes Cas9 is the most commonly used, but other Cas9 protein also expresses high levels of Streptococcus pyogenes Cas9. Some share sequence identity with the same guide RNA. The gRNAs used are diverse, and the PAM sequences they recognize (defined by the RNA) are different. The sequence of 2 to 5 nucleotides (defined by the protein adjacent to the sequence) also differs. Chylinski et al. classified Cas9 proteins from many bacterial groups (RNA Bi ology 10:5,1-12;2013), and Appendix 1 and Appendix 1 contain numerous Cas Nine proteins are listed, which are incorporated herein by reference. For the s9 protein, see Esvelt et al., Nat Methods. 2013 No. v;10(11):1116-21 and Fonfara et al., “Phylogeny o f Cas9 determines functional exchangeabi lity of dual-RNA and Cas9 among ortholog ous type II CRISPR-Cas systems.” Nucleic Acids Res.2013 Nov 22. [Pre-print electronic publication] doi:10.1 It is described in 093 / nar / gkt1074.

[0063] The methods and compositions described herein can be used with various species of Cas9 molecules. Streptococcus pyogenes (S. pyogenes) and Streptococcus thermophilus (S. thermophilus) While the Cas9 molecule of C. philus is the subject of most of this disclosure, other Cas9 molecules are also disclosed herein. Use Cas9 molecules derived from or based on the Cas9 proteins of other species, as listed In other words, most of the present description is based on the and those using the Cas9 molecule of S. thermophilus However, Cas9 molecules from other species can be used instead. The following table is based on the accompanying figure from Chylinski et al., 2013. Species include:

[0064] [Table 3] JPEG0007812830000005.jpg238166JPEG0007812830000006.jpg176166

[0065] The constructs and methods described herein can be used with any of the Cas9 proteins listed above and This may involve the use of the corresponding guide RNA or other equivalent guide RNAs. In addition, Streptococcus thermophilus Cong et al. showed that the Cas9 CRISPR1 system of LMD-9 functions in human cells. (Science 339, 819 (2013)). gitides) Cas9 orthologs are reported in Hou et al., Proc Natl Acad Sc i US A.2013 Sep 24;110(39):15644-9 and Es velt et al., Nat Methods.2013 Nov;10(11):1116-2 1. Furthermore, Jinek et al. hilus and L. innocua Cas9 orthologues (different Neisseria meningitidis and Neisseria meningitidis, which appear to utilize guide RNAs (not the C. jejuni Cas9 ortholog) but with slightly higher efficiency. Although the expression was reduced, it was induced by S. pyogenes double gRNA. It has been demonstrated in vitro that this enzyme can cleave target plasmid DNA.

[0066] In some embodiments, the systems of the invention are encoded in bacteria or expressed in mammalian cells. Currently codon-optimized, D10, E762, H983 or D986 and H840 or mutations at N863, e.g., D10A / D10N and H840A / H840N / H84 Protein expression using Streptococcus pyogenes (S. pyogenes) Cas9 protein containing 0Y Substitutions at these positions inactivate the catalytic activity of the nuclease portion of the protein; As described in Himasu et al., Cell 156, 935-949 (2014) 2) alanine, but may also be other residues, e.g., glutamine, asparagine, tyrosine , serine or aspartic acid, e.g., E762Q, H983N, H983Y, D986 N, N863D, N863S, or N863H (FIG. 1C). Catalytically inactivated Streptococcus pyogenes (S) that can be used in the methods and compositions of the present invention. pyogenes) Cas9 are as follows, exemplary D10A and H840A Mutations are shown in bold and underlined.

[0067] [Table 4] JPEG0007812830000008.jpg62159

[0068] In some embodiments, the Cas9 nuclease used herein is selected from the group consisting of pyogenic strains, lactic acid bacteria, and lactic acid bacteria. It is at least about 50% identical to the sequence of S. pyogenes Cas9, i.e., i.e., at least 50% identical to SEQ ID NO: 33. In some embodiments, the nucleotide The sequence is approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 90% identical to SEQ ID NO: 33. 85%, 90%, 95%, 99% or 100% identical. All of the differences from SEQ ID NO: 33 are in non-conserved regions, as described by Chylinski et al., RN A Biology 10:5,1-12;2013 (e.g., Appendix 1 and Appendix 1 );Esvelt et al., Nat Methods.2013 Nov;10(11):11 16-21 and Fonfara et al., Nucl. Acids Res. (2014) 42 (4):2577-2590.[Preprint electronic publication 2013 Nov 22]doi:1 By sequence alignment of the sequences listed in 0.1093 / nar / gkt1074 This is confirmed.

[0069] To determine the percent identity of two sequences, the sequences are aligned for optimal comparison purposes (optimally If necessary for proper alignment, the first and second amino acid or nucleic acid sequences may be modified to include a nucleotide sequence. (This allows you to introduce gaps and ignore non-homologous sequences.) References to align for comparison purposes The length of the sequence is at least 50% (in some embodiments, about 5% of the length of the reference sequence). 0%, 55%, 60%, 65%, 70%, 75%, 85%, 90%, 95% or 100 The nucleotides or residues at corresponding positions are then compared. A position in a sequence is occupied by the same nucleotide or residue as the corresponding position in a second sequence. If the two sequences are identical, then the two molecules are identical at that position. The uniformity is the number of gaps and the size of each gap that need to be introduced for optimal alignment of two sequences. It is a function of the number of identical positions common to the two sequences, taking into account the length of the loop.

[0070] The comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. For the purposes of this application, the GAP program of the GCG software package is used. Needleman and Wunsch (1970) J. Mol. Biol. 48:444-453) algorithm, and gap penalties were used. Bl with a penalty of 12, gap extension penalty of 4, and frameshift gap penalty of 5 The ossum62 scoring matrix is ​​used to determine the percent identity between two amino acid sequences. do.

[0071] Cas9-HFD Cas9-HFD is described in U.S. Provisional Patent Application No. 6, filed March 15, 2013. No. 1 / 799,647, U.S. Patent Application No. 61 / 838 filed June 21, 2013 ,148 and International Application No. PCT / US14 / 27335, all of which are incorporated herein by reference in their entireties.

[0072] Cas9-HFD can be used to express heterologous functional domains (e.g., transcription activation domains, e.g., VP6 4 or the transcriptional activation domain of NF-κB p65) and the catalytically inactivated Cas9 protein. It is produced by fusing the N-terminus or C-terminus of a protein (dCas9). In the case of Streptococcus pyogenes, as described above, the dCas9 may be of any species, but preferably is of Streptococcus pyogenes. (S. pyogenes) dCas9, and in some embodiments, the Cas9 is For example, as shown in SEQ ID NO: 33 above, the catalytic activity of the nuclease portion of the protein is Mutation of D10 and H840 residues to inactivate the protein, e.g., D10N / D10A and and H840A / H840N / H840Y.

[0073] The transcription activation domain may be fused to the N-terminus or C-terminus of Cas9. Although the present description uses transcription activation domains as an example, other domains known in the art may also be used. Other heterologous functional domains (e.g., transcriptional repressors (e.g., KRAB, ERD, SID) For example, the ets2 repressor factor (ERF) repressor domain (ERD) Amino acids 473–530, amino acids 1–97 of the KRAB domain of KOX1, or Mad Amino acids 1-36 of the mSIN3 interaction domain (SID); Beerli et al., PNA S USA 95:14628-14633 (1998)) or heterologous Chromatin protein 1 (HP1, also known as swi6), e.g., HP1α or silencers, such as HP1β; fixed RNA binding sequences, such as those found in the MS2 coat protein RNA-binding sequences bound by protein, endoribonuclease Csy4, or lambda N protein Proteins or peptides that can recruit long non-coding RNAs (lncRNAs) fused with nucleotide sequences, etc. peptides; enzymes that modify the methylation state of DNA (e.g., DNA methyltransferases) or enzymes that modify histone subunits. enzymes (e.g., histone acetyltransferases (HATs), histone deacetylases (HDAC), histone methyltransferases (e.g., lysine or arginine residues), groups) or histone demethylases (e.g., methylating lysine or arginine residues) Several sequences of such domains, e.g. For example, the sequence of a domain that catalyzes the hydroxylation of methylated cytosine in DNA is known in the art. Exemplary proteins are known in the art to recognize 5-methylcytosine (5-mC) in DNA. Ten-El, an enzyme that converts 5-hydroxymethylcytosine (5-hmC) to 5-hydroxymethylcytosine (5-hmC), Includes the even-translocation (TET) 1-3 family.

[0074] The sequences of human TET1-3 are known in the art and are shown in the table below.

[0075] [Table 5]

[0076] In some embodiments, all or part of the full-length sequence of the catalytic domain, e.g., 2OGFeD is encoded by a tein-rich extension and seven highly conserved exons a catalytic module comprising an O domain, e.g., a Tet domain comprising amino acids 1580 to 2052; 1 catalytic domain, Tet2 containing amino acids 1290–1905 and amino acids 966–1 For example, the Tet3 gene may contain the Tet3 gene, including 678. For alignment showing catalytic residues, see Iyer et al., Cell Cycle. 2009 Jun. 1;8(11):1698-710.Epub 2009 Jun 27, Figure 1, For full-length sequences, see the accompanying documents (ftp site ftp.ncbi.nih.gov / pub / aravind / DONS / supplementary_material_ DONS.html) (see, e.g., seq 2c) ); in some embodiments, the sequence is amino acids 1418 to 2136 of Tet1 or Te Includes the corresponding region of t2 / 3.

[0077] Other catalytic modules are derived from proteins identified in Iyer et al., 2009. It can be something that.

[0078] In some embodiments, the heterologous functional domain is a biological tether, all or part of the protein, endoribonuclease Csy4 or lambda N protein These proteins contain (e.g., their DNA-binding domains). RNA molecules containing suitable stem-loop structures are identified by the dCas9 gRNA targeting sequence. For example, MS2 coat protein, endoribonuclease Using dCas9 fused to Csy4 or lambda N, Csy4 binding sequences, MS Long non-binding sequences such as XIST or HOTAIR linked to Lambda 2 or Lambda N binding sequences It can recruit coding RNAs (lncRNAs) (e.g., Keryer-Bib See Ens et al., Biol. Cell 100:125-138 (2008) Alternatively, a Csy4 binding sequence, an MS2 binding sequence, or a lambda N protein binding sequence may be used. For example, as described in Keryer-Bibens et al. (supra), another protein and the proteins can be linked to dC using the methods and compositions described herein. In some embodiments, Csy4 is catalytically inactive. It's sexualized.

[0079] In some embodiments, the fusion protein comprises a linker between dCas9 and a heterologous functional domain. These fusion proteins (or fusion proteins with linked structures) contain The linker may comprise any sequence that does not interfere with the function of the fusion protein. In embodiments, the linker is short, e.g., 2-20 amino acids, and typically flexible (e.g., (i.e., containing highly flexible amino acids such as glycine, alanine, and serine). In some embodiments, the linker is GGGS (SEQ ID NO: 34) or GGGGS (SEQ ID NO: 3 5) one or more units, e.g., repeated two, three, four or more times GGGS (SEQ ID NO: 34) or GGGGS (SEQ ID NO: 35) units. Other linker sequences can be used.

[0080] Expression system To use the described guide RNAs, they must be expressed from a nucleic acid encoding them. This can be done in a variety of ways, for example, by using guide RNAs The nucleic acid encoding the vector is cloned into an intermediate vector and transformed into a prokaryotic or eukaryotic cell. , replication and / or expression. The intermediate vector typically contains a nucleic acid encoding a guide RNA. Prokaryotic vectors, e.g., plasmids, for storing or manipulating the vector and producing guide RNA. Alternatively, it may be a shuttle vector or an insect vector. The nucleic acid is cloned into an expression vector and then expressed in a plant cell, an animal cell, preferably a mammalian cell. Alternatively, it may be administered to a human cell, a fungal cell, a bacterial cell or a protozoan cell.

[0081] For expression, the sequence encoding the guide RNA is usually placed in a promoter that directs transcription. The vector is subcloned into an expression vector containing a suitable bacterial promoter and a eukaryotic promoter. Motors are well known in the art and are described, for example, in Sambrook et al., Molecula r Cloning, A Laboratory Manual (3rd edition, 2001); Kriegler,Gene Transfer and Expression:A Laboratory Manual (1990); and Current Protocol columns in Molecular Biology (eds. Ausubel et al., 2010 Bacterial expression systems for expressing engineered proteins are described in, for example, E. coli, Bacillus species and Salmonella nella) (Palva et al., 1983, Gene 22:22 9-235). Kits for such expression systems are commercially available. Eukaryotic expression systems for insect cells are well known in the art and are also commercially available.

[0082] The promoter used to direct expression of a nucleic acid will depend on the particular application. For example, the expression and purification of fusion proteins usually uses a strong constitutive promoter. In contrast, when guide RNAs are administered in vivo for gene regulation, Depending on the specific use of the guide RNA, it can be a constitutive or inducible promoter. Furthermore, preferred promoters for administering guide RNA include: Weak promoters such as HSV TK or promoters with similar activity The promoter may also contain elements responsive to transactivation, e.g., hypoxia response. response element, Gal4 response element, lac repressor response element and tetracycline regulation These may include small molecule regulatory systems such as the RU-486 system and the RU-486 system (see, e.g., Gossen and B ujard,1992,Proc.Natl.Acad.Sci.USA,89:554 7; Oligino et al., 1998, Gene Ther., 5:491-496; Wan g et al., 1997, Gene Ther., 4:432-441; Neering et al., 19 96, Blood, 88:1147-55; and Rendahl et al., 1998, Na (See, Biotechnol., 16:757-761).

[0083] In addition to a promoter, expression vectors usually contain vectors for expressing nucleic acids in prokaryotic or eukaryotic host cells. The expression cassette includes a transcription unit or expression cassette that contains all other elements necessary to produce the desired product. A typical expression cassette is operably linked to a nucleic acid sequence encoding, for example, a gRNA. The promoter and the transcription factors responsible for efficient polyadenylation of the transcript, transcription termination, and ribosome binding are The cassette also contains a transcription factor (TGF-γ) binding site and any signal required for translation termination. For example, enhancers and heterologous splice intron signals may be included.

[0084] The specific expression vector used to deliver the genetic information into the cell will be the intended gRNA The gene is selected taking into consideration its intended use, for example, expression in plants, animals, bacteria, fungi, protozoa, etc. Standard bacterial expression vectors include pBR322-based plasmids, pSKF, and pE Plasmids such as T23D and commercially available tag fusion expression systems such as GST and LacZ are available. Examples include:

[0085] Eukaryotic expression vectors include expression vectors containing regulatory elements from eukaryotic viruses, e.g., SV40 vectors, papillomavirus vectors and Epstein-Barr virus-derived vectors Other exemplary eukaryotic vectors include pMSG and pAV00. 9 / A+, pMTO10 / A+, pMAMneo-5, baculovirus pDSVE and and SV40 early promoter, SV40 late promoter, metallothionein promoter -, mouse mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrosis The promoters that are effective for expression in eukaryotic cells, including the phosphopromoter, Examples of vectors include any other vectors that can be used to express proteins.

[0086] The vector for expressing the guide RNA contains an RNA that drives the expression of the guide RNA. It may comprise a PolIII promoter, for example, an H1, U6 or 7SK promoter. These human promoters express the gR promoter in mammalian cells after plasmid transfection. Alternatively, for example, a T7 promoter may be used for in vitro transcription. RNA can be transcribed in vitro and then purified. For example, vectors suitable for expressing small RNAs such as siRNA and shRNA are used. It can be used.

[0087] Some expression systems incorporate a marker for selection of stably transfected cell lines, e.g. For example, thymidine kinase, hygromycin B phosphotransferase, and dihydrofolate In addition, it has been demonstrated that the baculovirus vector can be used to propagate the vector in insect cells. Under the direction of strong baculovirus promoters such as the rehedrin promoter High-yield expression systems, such as those that place the gRNA coding sequence in the .

[0088] Other elements typically included in expression vectors include a gene that functions in E. coli. a replicon carrying the recombinant plasmid, an antibiotic resistance coding sequence that allows for selection of bacteria carrying the recombinant plasmid The gene to be loaded and a unique region in a non-essential region of the plasmid that allows for the insertion of recombinant sequences. Restriction sites are included.

[0089] Using standard transfection methods, we have developed bacterial, mammalian, and microbial transfection systems that express large amounts of protein. The protein is then purified using standard techniques. (e.g., Colley et al., 1989, J. Biol. Chem., 264:17 619-22;Guide to Protein Purification,in Methods in Enzymology, vol. 182 (Deutscher et al. (See, e.g., 1990). Transformation of eukaryotic and prokaryotic cells is carried out by standard techniques. (e.g., Morrison, 1977, J. Bacteriol. 132:3 49-351; Clark-Curtiss and Curtiss, Methods i See Enzymology 101:347-362 (Wu et al., eds., 1983). sea ​​bream).

[0090] Any known method for introducing foreign nucleotide sequences into host cells may be used. Such methods include calcium phosphate transfection, polybrene, and protoplast transfection. fusion, electroporation, nucleofection, liposomes, microinjection vectors, naked DNA, plasmid vectors, viral vectors (episomal and including cloned genomic DNA, cDNA, and synthetic DNA. and the use of any other well-known method for introducing foreign genetic material into a host cell. (See, e.g., Sambrook et al., supra). The only requirement is the specific Genetically engineer at least one gene capable of expressing gRNA in a host cell. It is possible to introduce it into

[0091] The present invention includes vectors and cells containing the vectors. [Example]

[0092] The invention is further described in the following examples, which are set forth in the claims. It is not intended to limit the scope of the invention as claimed.

[0093] Example 1. Evaluation of the specificity of RNA-guided endonucleases CRISPR RNA-guided nuclease (RGN) is a simple and efficient method for genome editing. This example demonstrates the rapid emergence of a human cell-based platform. Characterizing Cas9-based off-target cleavage of RGN using reporter assays This article describes what to do.

[0094] material and method In Example 1, the following materials and methods were used.

[0095] Guide RNA construction DNA oligonucleotides carrying variable 20 nt sequences for Cas9 targeting (Table A ) was annealed to form a BsmBI-digested plasmid pMLM3 A short double-stranded DNA fragment compatible with ligation to 636 was generated. Cloning of the annealed oligonucleotides resulted in the expression of 20 Plasmid encoding a chimeric +103 single-stranded guide RNA with a variable 5' nucleotide (Hwang et al., Nat Biotechnol 31, 227-229) 2013); Mali et al., Science 339, 823-826 (2013)). pMLM3636 and expression plasmid pJDS246 (codon optimized) used for the study The plasmids encoding the Cas9 variants are both available from the non-profit plasmid distribution service Addgene (a The cascade of sequences from the cascade of genes encoding ...

[0096] [Table 6] JPEG0007812830000011.jpg246161JPEG0007812830000012.jpg246158JPEG0007812830000013.jpg245161JPEG0007812830000014.j pg245159JPEG0007812830000015.jpg246161JPEG0007812830000016.jpg246159JPEG0007812830000017.jpg246161JPEG0007812830 000018.jpg245158JPEG0007812830000019.jpg246161JPEG0007812830000020.jpg246159JPEG0007812830000021.jpg245160JPEG00 07812830000022.jpg245158JPEG0007812830000023.jpg246160JPEG0007812830000024.jpg245156JPEG0007812830000025.jpg24599

[0097] EGFP activity assay U2OS.EGFP cells, which contain a single copy of the EGFP-PEST fusion gene Culture was performed as previously described (Reyon et al., Nat Biotech 30, 4 60-465(2012)). For transfection, SE Cell Line The 4D-Nucleofector™ X kit (Lonza) was used according to the manufacturer's protocol. The indicated amounts of sgRNA expression plasmid and pJDS246 were used according to the protocol. Nucleofect 200,000 cells with 30 ng of the tomato-encoding plasmid. Two days after transfection, cells were analyzed using a BD LSRII flow cytometer. We analyzed the transfection to optimize the concentration of gRNA / Cas9 plasmid. All other transfections were performed in duplicate.

[0098] PCR amplification and sequence verification of endogenous human genomic loci Phusion Hot Start II high-fidelity DNA polymerase (NEB) and PCR reactions were carried out using the PCR primers and conditions listed in Table B. Most loci were analyzed by touchdown PCR (98°C, 10 seconds; 72–62°C, -1°C / second). 10 cycles of [98°C, 10 sec; 62°C, 15 sec; 72°C, 30 sec; Amplification was successful using 25 cycles of 72°C for 30 seconds. The remaining annealing was performed using a constant annealing temperature of 72°C and 3% DMSO or 1M betaine. PCR was performed for 35 cycles. The PCR products were analyzed using a QIAXCEL capillary electrophoresis system. The size and purity of the product were verified by electrophoresis. The DNA was processed using Sap-IT (Affymetrix) and analyzed by Sanger sequencing (MGH DNA Seq Each target site was verified by sequencing by the Quantifying Core.

[0099] [Table 7] JPEG0007812830000027.jpg254149JPEG0007812830000028.jpg254151JPEG0007812830000029.jpg254147JPEG0007812830000030.jpg254151JPEG0007812830000031.jpg254145JPEG0007812830000032.jpg254152JPEG0007812830000033.jpg254144JPEG0007812830000034.jpg254144JPEG0007812830000035.jpg254152JPEG0007812830000036.jpg254147JPEG0007812830000037.jpg254152JPEG0007812830000038.jpg254149JPEG0007812830000039.jpg254152JPEG0007812830000040.jpg254147JPEG0007812830000041.jpg254145JPEG0007812830000042.jpg254146JPEG0007812830000043.jpg254149JPEG0007812830000044.jpg254144JPEG0007812830000045.jpg254146JPEG0007812830000046.jpg254144JPEG0007812830000047.jpg254147JPEG0007812830000048.jpg254149JPEG0007812830000049.jpg254142JPEG0007812830000050.jpg254152JPEG0007812830000051.jpg254151JPEG0007812830000052.jpg254149JPEG0007812830000053.jpg254146JPEG0007812830000054.jpg254147JPEG0007812830000055.jpg254141JPEG0007812830000056.jpg253144JPEG0007812830000057.jpg254148JPEG0007812830000058.jpg254144JPEG0007812830000059.jpg254151JPEG0007812830000060.jpg254148.

[0100] Determining the frequency of RGN-induced on-target and off-target mutations in human cells For U2OS.EGFP and K562 cells, 4D Nucleofector System (Lonza) according to the manufacturer's instructions. 5 g in cells 250 ng of RNA expression plasmid or empty U6 promoter plasmid (negative control); Transfect 750 ng of Cas9 expression plasmid and 30 ng of td-Tomato expression plasmid. HEK293 cells were infected with Lipofectamine LTX reagent (L 1.65 × 10 5 Inject gRNA expression plasmid or empty U6 promoter plasmid (negative control) into cells. 125ng, Cas9 expression plasmid 375ng and td-Tomato expression plasmid 30 ng was transfected using the QIAamp DNA Blood Mini Kit ( Transfected U2OS.EG cells were transfected using QIAGEN according to the manufacturer's instructions. Genomic DNA was collected from FP cells, HEK293 cells, or K562 cells. Three rounds of nucleofection were performed to obtain enough genomic DNA to amplify the target candidate sites. nucleofection (U2OS.EGFP cells), two nucleofections (K562 cells) or DNA obtained by two Lipofectamine LTX transfections A was pooled and then subjected to T7EI. This procedure was performed twice for each condition tested. Two pools of the same genomic DNA were prepared by transfection, and each transfection was carried out four times. Then, PCR was performed using these genomic DNAs as templates as described above. Ampure XP beads (Agencourt) were used according to the manufacturer's instructions. The T7EI assay was performed as previously described (Reyon et al., 2012, supra).

[0101] DNA sequencing of NHEJ-mediated indel mutations The purified PCR products used in the T7EI assay were cloned into Zero Blunt TOPO vectors. (Life Technologies) and cloned into MGH DNA Auto Plasmid DNA was extracted using the alkaline lysis miniprep method by the Information Core. M13 forward primer (5'-GTAAAACGACGGCCAG-3') SEQ ID NO: 1059) using the Sanger method (MGH DNA Sequencing Co. The plasmid was sequenced by PCR.

[0102] Example 1a. Single nucleotide mismatch To begin to uncover RGN specificity determinants in human cells, several The impact of systematically introducing mismatches at various positions within the gRNA / target DNA junction To do this, we conducted a large-scale study to evaluate the impact of the A quantitative human cell-based, highly sensitive green fluorescent assay for rapid quantification of target nuclease activity Protein (EGFP) decay assay (see Methods above and Reyon et al., 2012, above) (See the description below) (Figure 2B). In this assay, nuclease-induced double-strand breaks Frames introduced by error-prone non-homologous end joining (NHEJ) repair of deep stub breaks (DSBs) Human U2OS.E caused by inactivating shift insertion / deletion (indel) mutations By assessing the fluorescent signal within GFP cells, we confirmed the presence of a single integrated EGFP receptor. The activity of the nuclease targeting the target gene can be quantified (Figure 2B). The study described here identified three approximately 100 different targeting sequences within EGFP: A single gRNA of 100 nt was used: EGFP site 1 GGGCACGGGCAGCTTGCCGGTGG (SEQ ID NO: 9) EGFP site 2 GATGCCGTTCTTCTGCTTGTCGG (SEQ ID NO: 10) EGFP site 3 GGTGGTGCAGATGAACTTCAGGG (SEQ ID NO: 11) . Each of the above gRNAs was shown to efficiently induce Cas9-mediated disruption of EGFP expression. (See Examples 1e and 2a and Figures 3E (top row) and 3F (top row) I want to be.

[0103] In the first experiment, 20 nucleotides of the complementary targeting regions of three EGFP-targeting gRNAs were transfected into the GFP-targeting vector. The effect of single nucleotide mismatches in 19 of the nucleotides was examined. To do this, positions 1 to 19 (3' to 5') were selected for each of the three target sites. The directions are numbered 1 to 20; see Figure 1) in the Watson-Crick transform. We generated mutant gRNAs carrying version mismatches and confirmed that these various gRNAs are human The ability of the nucleotide at position 20 to induce Cas9-mediated EGFP decay in cells was tested. The nucleotide is part of the U6 promoter sequence and must be a guanine to avoid affecting expression. (Because this is necessary, no mutant gRNAs with substitutions at this position were generated.)

[0104] EGFP target site #2 tolerates more mismatches at the 5' end of the gRNA than at the 3' end. Previous research suggests that it is well tolerated (Jiang et al., Nat Biotech nol 31,233-239(2013);Cong et al.,Science 339,8 19-823(2013); Jinek et al., Science 337, 816-821( (2012)), a single mismatch at positions 1–10 of the gRNA results in the concomitant Ca The activity of s9 was significantly affected (Fig. 2C, middle panel). However, the EGFP target site In #1 and #3, even if there is a single mismatch at any position except for a part of the gRNA, Even within the 3' end of the sequence, high tolerance was observed. There were differences in the specific locations where sensitivity to the match was high (Figure 2C, top and bottom panels). For example, target site #1 is particularly sensitive to mismatches at position 2. Target site #3 showed the highest sensitivity to mismatches at positions 1 and 8. The degree was shown.

[0105] Example 1b. Multiple Mismatches To test the effect of two or more mismatches at the gRNA / DNA junction, Contains two Watson-Crick transversion mismatches at positions 1 and 2 apart. We generated a series of mutant gRNAs and demonstrated the expression of these gRNAs in human cells using an EGFP decay assay. We tested the ability of RNA to induce Cas9 nuclease activity. Both types of mismatches occur in the 3' half of the gRNA targeting region. However, the magnitude of this effect differs depending on the region. Differences were observed, with target site #2 showing the highest sensitivity to the two mismatches, and target site #3 showing the highest sensitivity to the two mismatches. Position #1 showed the lowest sensitivity overall. To achieve this, positions 19–15 of the 5' end of the gRNA targeting region (single and two mismatches) were Mutant gRNs with increasing numbers of mismatched positions in the range of positions likely to be more tolerant of mismatches A was constructed.

[0106] In this way, gRNAs with increasing mismatches were tested, and the results showed that they were highly effective at three target sites. In both cases, the introduction of three or more adjacent smatches significantly abolished RGN activity. It was revealed that the mismatch occurs starting from position 19 at the 5' end and proceeds towards the 3' end. The activity of three different EGFP-targeting gRNAs suddenly decreased with increasing addition of thiol. Specifically, gRNAs containing mismatches at positions 19 and 19+18 were found to be substantially Showing full activity, whereas positions 19+18+17, 19+18+17+16 and 1 The gRNA with mismatches at positions 9, 18, 17, 16, and 15 significantly increased the cytotoxicity compared to the negative control. No difference was observed (Figure 2F). (Position 20 is the U6 promoter driving the expression of gRNA.) Since the gRNA is part of the nucleotide sequence, it must be G. Note that no mismatch was introduced at position 20.)

[0107] Rationale for obtaining RGNs with increased specificity by shortening gRNA complementarity were also obtained in the following experiments: with four different EGFP-targeting gRNAs (Figure 2H). The introduction of two mismatches at positions 18 and 19 did not significantly affect activity. However, by introducing two more mismatches at positions 10 and 11 of the gRNA, The introduction of only two mismatches (10 / 11) resulted in almost complete loss of activity. It is interesting to note that the activity of the α-glucan-containing β-glucan is not significantly affected by the α-glucan-containing β-glucan.

[0108] Taken together, these results obtained in human cells suggest that RGN activity is conserved within gRNA targets. This supports the idea that the 3′ half of the nucleotide sequence may be more sensitive to mismatches. However, these data also suggest that RGN specificity is complex and target site dependent. Single and double mismatches, one or more in the 3' half of the RNA targeting region This clearly shows that even if a mismatch occurs, there is often a high tolerance for it. These data also suggest that any mismatches in the 5' half of the gRNA / DNA junction This suggests that it is not necessarily well tolerated.

[0109] Furthermore, these results suggest that gR, which has a shorter region of complementarity (specifically, approximately 17 nt), This strongly suggests that NA has higher activity characteristics. By combining the heterozygosity with the 2-nt specificity conferred by the PAM sequence, sufficient to be unique within a large, complex genome such as that found in a human cell Note that a specification of one 19 bp sequence of length is obtained.

[0110] Example 1c. Off-target mutations Can off-target mutations of RGN targeting endogenous human genes be identified? To clarify whether this is the case, we performed a multivariate analysis of three different sites in the VEGFA gene, one in the EMX1 gene, and It targets a site in the RNF2 gene and a site in the FANCF gene. Six types of single gRNAs were used (Table 1 and Table A). Human U2OS.EG, as detected by nuclease I (T7EI) assay Efficiently induce Cas9-mediated indels at each endogenous locus in FP cells (Methods and Table 1 above). We then analyzed these six RGNs. For each, nuclease-induced NHEJ-mediated indels in U2OS.EGFP cells To obtain evidence of mutations, dozens of candidate off-target sites (ranging from 46 to 64) were identified. The evaluated loci contained all genome sequences that differed by one or two nucleotides. In addition to the genomic sites, a subset of genomic sites differing by 3–6 nucleotides were included. The RNA targeting sequence is overlaid with one or more such mismatches in the 5' half. Using the T7EI assay, VEGFA site 1 (of 53 sites examined) was significantly higher than that of the control site (Table B). Four off-target sites were identified among the candidate sites, and VEGFA site 2 was identified among the 46 sites examined. 12 (of 64 sites examined) in VEGFA sites3, 7 (of 64 sites examined) in EMX In one site (out of 46 sites examined), one was easily identified (Table 1 and Table B). 43 and 50 candidates were examined in the RNF2 or FANCF genes, respectively. No off-target mutations were detected in the complementary sites (Table B). The mutation rate at the site was extremely high, ranging from 5.6% to 12% of the mutation rate observed at the intended target site. The true off-targets ranged from 5% (average 40%) to 10% (Table 1). The target site contains a mismatch at the 3' end, with a total of five mismatches. Most off-target sites were found within protein-coding genes (Table 1 ) DNA sequencing of some off-target sites revealed predicted RGN cleavage sites. Further molecular support was obtained for the occurrence of indel mutations in the .

[0111] [Table 8] JPEG0007812830000062.jpg89170

[0112] Example 1d. Off-target mutations in other cell types It was confirmed that RGN can induce off-target mutations at high frequency in U2OS.EGFP cells. Having confirmed this, we next investigated whether these nucleases are also present in other types of human cells. The present inventors have previously investigated whether TALENs have such effects. 15 Because U2OS.EGFP cells were used to evaluate the activity of U2OS.EGFP cells were chosen for this experiment, but human nucleases were used to test the activity of the targeted nuclease. HEK293 cells and K562 cells are more widely used. also reported that VEGFA sites 1, 2, and 3 were expressed in HEK293 and K562 cells. We evaluated four types of RGNs targeted to the EMX1 site and the EMX2 site. RGN, but the mutation frequency is somewhat lower than that observed in U2OS.EGFP cells The two additional human cell lines also expressed NHE receptors at their intended on-target sites. We found that J-mediated indel mutations were efficiently induced (evaluated by T7EI assay). ) (Table 1). 24 RGNs of the four types mentioned above were initially identified in U2OS.EGFP cells. When off-target sites were evaluated, many sites were identified in HEK293 cells and K56 2 cells, the corresponding on-target site was mutated at a similar frequency. As expected, these off-target sites in HEK293 cells were DNA sequencing of a portion of the genome revealed changes at predicted genomic loci. Further molecular evidence was obtained (Figures 9A-9C). Of the detected off-target sites, 4 were detected in HEK293 cells and 11 in K562 cells. The exact reason why these off-target genes did not show detectable mutations is unknown. It is noteworthy that many of the mutation sites showed relatively low mutation frequencies even in U2OS.EGFP cells. Therefore, in our experiments, HEK293 cells were more sensitive than U2OS.EGFP cells. Because RGN activity appears to be generally lower in HEK29 and K562 cells, The mutation rates at these sites in K562 and K3 cells were significantly higher than those in our T7EI assay. It is estimated that the concentration of iodine will be below the reliable detection limit (approximately 2-5%). The results obtained by the present inventors in HEK293 and K562 cells are consistent with the results observed in RGN. Evidence that high frequency of off-target mutations is a general phenomenon observed in multiple human cell types This shows that.

[0113] Example 1e. gRNA Expression and Cas9 Expression Plus Used in EGFP Decay Assay Gradually increasing the amount of mido Induction of a frameshift mutation via non-homologous end joining reliably disrupts EGFP expression. Three different sequences (shown above) located upstream of EGFP nucleotide 502, the position where EGFP is expressed. A single gRNA was generated against the EGFP sites 1 to 3 (Maeder, ML et al. ,Mol Cell 31,294-301(2008);Reyon, D. et al., Nat Biotech 30,460-465(2012)).

[0114] For the three target sites, first prepare various amounts of gRNA expression plasmid (12.5–25 0 ng) was transfected with a single copy of the constitutively expressing EGFP-PEST reporter gene. We transfected our U2OS.EGFP reporter cells with codon-optimized Cas9 nucleases. The highest concentration of gRNA was transfected with 750 ng of the gRNA expression plasmid. In the case of the code plasmid (250 ng), RGN efficiently disrupted EGFP expression in all three forms. However, a smaller amount of gRNA expression plasmid was used in the transfection. When transfected with β-glucan, RGNs for target sites #1 and #3 showed similar levels of disruption. For RGN activity at target site #2, transfecting gRNA expression plasmid The decrease was immediate when the amount of α-glucan was reduced (Fig. 3E (top row)).

[0115] Cas9 code transfected into our U2OS.EGFP reporter cells. EGFP decay was assayed using increasing amounts of the plasmid (50 ng to 750 ng). As shown in Figure 3F (top), at target site #1, the transfected Cas9 A one-third reduction in the amount of code plasmid was tolerated without substantial loss of EGFP decay activity. However, the activity of RGN targeting target sites #2 and #3 was not observed in the transfected cells. The effect was immediately reduced when the amount of Cas9 plasmid used was reduced by one-third (Figure 3F (top)). Based on the above results, the experiments described in Examples 1a-1d included EGFP targeting site #1, For #2 and #3, 2 gRNA expression plasmids / Cas9 expression plasmids were used, respectively. 5ng / 250ng, 250ng / 750ng and 200ng / 750ng were used.

[0116] Some gRNA / Cas9 combinations are more effective at disrupting EGFP expression than others The reason for this is that some of these combinations depend on the amount of plasmid used for transfection. The reasons for the greater or lesser sensitivity are also not understood. The range of off-target sites present may affect the activity of each However, 1-6 bp of these specific target sites may account for the difference in behavior of the three gRNAs. No differences were observed in the number of genomic sites that differed by only one sequence (Table C).

[0117] [Table 9]

[0118] Example 2: Reducing the length of gRNA complementarity to improve RGN cleavage specificity Although this approach may seem counterintuitive at first, simply shortening the length of the gRNA-DNA junction This minimizes the off-target effects of RGN without compromising its on-target activity. We hypothesized that longer gRNAs would actually function better at the on-target site. The efficiency of the enzyme is low (see below and Hwang et al., 2013a; Ran et al., 2013). In contrast, as shown in Example 1 above, multiple mismatches at the 5' end Even gRNAs with a nucleotide sequence can still induce reliable cleavage at their target site. (Figures 2A and 2C-2F), and full on-target activity requires these nucleotides. Therefore, the truncated g gene lacking these 5' nucleotides may be unnecessary. We hypothesized that the RNA would exhibit activity equivalent to that of the full-length gRNA (Fig. 2A). If the 5' nucleotides of the full-length gRNA are not required for on-target activity, their presence Could the presence of gRNA also compensate for mismatches at other positions in the gRNA-target DNA junction? If this assumption is correct, the sensitivity to gRNA mismatches will be high. In addition to this, the level of Cas9-mediated off-target mutations may be substantially reduced. We hypothesized that this may be the case (Figure 2A).

[0119] Experimental procedure The following experimental procedure was used in Example 2.

[0120] Plasmid construction All gRNA expression plasmids contain complementary regions as described above (Example 1). Design, synthesize, and anneal a pair of oligonucleotides (IDT) to construct the plasmid pML The construct was constructed by cloning into M3636 (available from Addgene). The resulting gRNA expression vector contains approximately 100 kDa of the nucleotide sequence 110 kDa, with expression driven by the human U6 promoter. All oligonucleotides used in the construction of the gRNA expression vector encode a 100-nt gRNA. The nucleotides are listed in Table D. The QuikChange kit (Agilent Technology) (logies) with the following primers: Cas9 D10A sense primer 5'-tgg ataaaaagtattctattggtttagccatcggcactaattc cg-3' (SEQ ID NO: 1089); Cas9 D10A antisense primer 5'- cggaattagtgccgatggctaaaccaatagaatacttttt Plasmid pJDS246 was transformed with atcca-3' (SEQ ID NO: 1090). By using the Cas9 D gene with a mutation in the RuvC endonuclease domain, A 10A nickase expression plasmid (pJDS271) was constructed. Both the targeted gRNA plasmid and the Cas9 nickase plasmid are non-commercially available. Midi distribution service Addgene (addgene.org / crispr-cas) It is available from

[0121] [Table 10] JPEG0007812830000065.jpg253156JPEG0007812830000066.jpg253152JPEG0007812830000067.jpg253155JPEG0007812830000068.jpg25215 2JPEG0007812830000069.jpg252158JPEG0007812830000070.jpg252157JPEG0007812830000071.jpg253156JPEG0007812830000072.jpg25339

[0122] Human cell-based EGFP decay assay U2OS carrying a single copy of the integrated EGFP-PEST gene reporter. EGFP cells have been previously described (Reyon et al., 2012). 0% FBS, 2 mM GlutaMax (Life Technologies), penicillin-resistant Staphylococcus aureus (PEA) Advanc supplemented with cilin / streptomycin and 400 μg / ml G418 The cells were maintained in ed DMEM (Life Technologies). To assay for disruption, the LONZA 4D-Nucleofector™ was used to inject SE solution. Using the DN100 program and following the manufacturer's instructions, 2 x 10 5 U2O Transfection of S.EGFP cells with gRNA expression plasmids or negative control empty U6 promoter plasmids Smid, Cas9 expression plasmid (pJDS246) (see Example 1 and Fu et al., 2013 ) and 10 ng of td-Tomato expression plasmid (a control for transfection efficiency) The inventors transfected the EGFP sites #1, #2, #3 and # For each experiment, use 25ng of gRNA expression plasmid and 25ng of Cas9 expression plasmid. 0ng, 250ng / 750ng, 200ng / 750ng and 250ng / 750n Two days after transfection, cells were treated with trypsin and 10% (vol Dulbecco's modified Eagle's medium (DMEM) supplemented with 1000 mg / vol fetal bovine serum (FBS) The samples were resuspended in PBS (Invitrogen) and analyzed using a BD LSRII flow cytometer. Transfection and flow cytometry measurements were performed in duplicate for each sample. carried out.

[0123] Transfection of human cells and isolation of genomic DNA On-target and off-target responses induced by RGN targeting endogenous human genes To evaluate the insertion-deletion mutations, we cultured U2OS.EGFP cells or HEK293 cells were transfected with the plasmid: U2OS.EGFP cells were transfected with the plasmid as described above. HEK293 cells were transfected using the same conditions as in the EGFP decay assay. , in a CO2 incubator at 37°C, containing 10% FBS and 2 mM GlutaMax ( Advanced DMEM (Lif) supplemented with Life Technologies e Technologies) per well in a 24-well plate. x10 5 Cells were transfected by seeding at a density of 1000 x g for 22-24 hours. After incubation, Lipofectamine LTX reagent was added according to the manufacturer's instructions (Li The cells were transfected with gRNA expression plasmids or 125 ng of empty U6 promoter plasmid (negative control), 125 ng of Cas9 expression plasmid (p JDS246) (Example 1 and Fu et al., 2013) 375 ng and td-Tomato expression 10 ng of plasmid was transfected. 16 hours after transfection, the medium Both cell types were replaced with Agencourt DN 2 days after transfection. Advance genomic DNA isolation kit (Beckman) according to the manufacturer's instructions. Genomic DNA was collected using 12 separate 4D transfections for each RGN sample assayed. Replicate transfections were performed, and genomes were extracted from each of the 12 transfections. After isolating the genome DNA, these samples were combined and each was divided into six pooled genome DNA samples. Two "duplicate" pools of NA samples were then generated. assay, Sanger sequencing and / or deep sequencing Evaluate indel mutations at on-target and off-target sites in the duplicated samples. did.

[0124] Assessing the frequency of precise changes introduced by HDR using ssODN donor templates Therefore, 2 × 10 5Transfection of U2OS.EGFP cells with gRNA expression plasmid or empty U6 250 ng of promoter plasmid (negative control), Cas9 expression plasmid (pJDS2 46) 750 ng, 50 pmol of ssODN donor (or no ssODN for control) and 10 ng of td-Tomato expression plasmid (transfection control) was transfected. Three days after transfection, Agencourt DNA advance was applied. Purify genomic DNA using genomic DNA fragments and probe for Bam at the locus of interest as described below. The introduction of the HI site was assayed. All transfections were performed in duplicate. Ta.

[0125] In experiments involving Cas9 nickase, 2 × 10 5 U2OS.EGFP cells per g 125 ng of RNA expression plasmid (if paired gRNA is used) or gRNA expression plasmid 250 ng of the current plasmid (if using a single gRNA), Cas9-D10A nickase 750 ng of expression plasmid (pJDS271), 10 ng of td-Tomato plasmid, and (If performing HDR) 50p of ssODN donor template (encoding a BamHI site) mol were transfected. All transfections were performed in duplicate. Two days after transfection (if assaying indel mutations) or after transfection Three days after the reaction (when assaying HDR / ssODN-mediated changes), Genome DNA was isolated using the Ourt DNA Advance genomic DNA isolation kit (Beckman). Mu DNA was recovered.

[0126] T7EI assay for quantifying the frequency of indel mutations T7EI was performed as previously described (Example 1 and Fu et al., 2013). Briefly, Phusion high-fidelity DNA polymerase (New England Biolabs) with one of the following programs: (1) Touchdown PCR program Ram [(98℃, 10 seconds; 72-62℃, -1℃ / cycle, 15 seconds; 72℃, 30 seconds) × 10 cycles, (98℃, 10 seconds; 62℃, 15 seconds; 72℃, 30 seconds) × 25 cycles ] or (2) a constant Tm PCR program [(98°C, 10 s; 68°C or 72°C, 15 sec; 72°C, 30 sec) × 35 cycles] with 3% DMSO or 1M beta-glucan as needed. PCR used with ELISA to amplify specific on-target or off-target sites All primers used in the above amplifications are listed in Table E. The resulting PCR The products were 300-800 bp in size and were purified using Ampur according to the manufacturer's instructions. e Purified with XP beads (Agencourt). 200 ng of purified PCR product. were hybridized in a total volume of 19 μl of 1× NEB buffer 2 and denatured using the following conditions: Heteroduplex formation was achieved by: 95°C, 5 min; 95 to 85°C, -2°C / sec; 85 to 2 5°C, -0.1°C / sec; hold at 4°C. The hybridized PCR product was then incubated with the T7 endonuclease. Add 1 μl of ase I (New England Biolabs, 10 units / μl); The mixture was incubated at 37°C for 15 minutes. 2 μl of 0.25 M EDTA solution was added. The T7EI reaction was stopped by ATP and purified using AMPure XP beads (Agencourt). The reaction product was purified by eluting with 20 μl of 0.1× EB buffer (QIAgen). The reaction products were then analyzed using a QIAXCEL capillary electrophoresis system, and the results were consistent with those previously described. The frequency of indel mutations was calculated using the same formula as previously described (Reyon et al., 2012). I calculated it.

[0127] [Table 11] JPEG0007812830000074.jpg254155JPEG0007812830000075.jpg253144JPEG0007812830000076.jpg253154JPEG0007812830000077.jpg253146JPEG0007812830000078.jpg254138JPEG0007812830000079.jpg253142JPEG0007812830000080.jpg253144JPEG0007812830000081.jpg254152JPEG0007812830000082.jpg253147JPEG0007812830000083.jpg253152JPEG0007812830000084.jpg253147JPEG0007812830000085.jpg253138JPEG0007812830000086.jpg253154JPEG0007812830000087.jpg254144JPEG0007812830000088.jpg254154JPEG0007812830000089.jpg253138JPEG0007812830000090.jpg253138JPEG0007812830000091.jpg254153JPEG0007812830000092.jpg253133JPEG0007812830000093.jpg254138JPEG0007812830000094.jpg254141JPEG0007812830000095.jpg253138JPEG0007812830000096.jpg254148JPEG0007812830000097.jpg254134JPEG0007812830000098.jpg254138JPEG0007812830000099.jpg253139JPEG0007812830000100.jpg253139JPEG0007812830000101.jpg254156JPEG0007812830000102.jpg253149JPEG0007812830000103.jpg253147JPEG0007812830000104.jpg253153JPEG0007812830000105.jpg254134JPEG0007812830000106.jpg254142JPEG0007812830000107.jpg253139JPEG0007812830000108.jpg253147JPEG0007812830000109.jpg253145JPEG0007812830000110.jpg254154JPEG0007812830000111.jp g254138JPEG0007812830000112.jpg253142JPEG0007812830000113.jpg254147JPEG00078 12830000114.jpg254154JPEG0007812830000115.jpg253139JPEG0007812830000116.jpg25 4137JPEG0007812830000117.jpg254143JPEG0007812830000118.jpg254133JPEG00078128 30000119.jpg253139JPEG0007812830000120.jpg254142JPEG0007812830000121.jpg2541 42JPEG0007812830000122.jpg253152JPEG0007812830000123.jpg254138JPEG0007812830 000124.jpg253133JPEG0007812830000125.jpg254134JPEG0007812830000126.jpg253113.

[0128] Sanger sequencing to quantify the frequency of indel mutations The purified PCR products used in the T7EI assay were cloned into Zero Blunt TOPO vectors. -(Life Technologies) and is listed as one of the Top 10 chemically competent companies The plasmid DNA was isolated and transferred to the Massachusetts General Hospital (MGH). GH) DNA Automation Core provided the M13 forward primer (5'-GT AAAACGACGGCCAG-3' (SEQ ID NO: 1059).

[0129] Restriction digestion to quantify specific changes induced by HDR using ssODN Assay Phusion high-fidelity DNA polymerase (New England Biolab s) was used to perform PCR reactions for specific on-target sites. Program (98°C, 10 seconds; 72°C to 62°C, -1°C / cycle, 15 seconds; 72°C, 30 seconds) 98℃, 10 seconds; 62℃, 15 seconds; 72℃, 30 seconds) × 10 cycles, (98℃, 10 seconds; 62℃, 15 seconds; 72℃, 30 seconds) × 25 cycles The VEGF and EMX1 gene loci were amplified using 100% DMSO with 3% DMSO. The primers used in these PCR reactions are listed in Table E. Amp was used according to the manufacturer's instructions. The PCR product was purified using ure beads (Agencourt). For detection of the BamHI restriction site encoded by the type, 200 ng of purified PCR product was used in 3 The DNA was digested with BamHI at 7°C for 45 minutes using Ampure XP beads (Agenco Purify the digestion product by eluting with 20 μl 0.1×EB buffer using and analyzed and quantified using a QIAXCEL capillary electrophoresis system.

[0130] TruSeq library generation and sequencing data analysis Locus-specific primers were selected for on-target and candidate validated off-target sequences. The target site was designed to flank the target site, creating a PCR product approximately 300-400 bp in length. Genomic DNA from the above pooled duplicate samples was used as a template for PCR. PCR products were purified with Ampure XP beads (Agencourt) according to the manufacturer's instructions. The purified PCR products were quantified using a QIAXCEL capillary electrophoresis system. PCR products for each locus were amplified and purified from each of the pooled duplicate samples (above). After quantification, equal amounts were pooled for deep sequencing. The cones were subjected to dual-index I as previously described (Fisher et al., 2011). After adapter ligation, the library was The purified protein was then subjected to electrophoresis on a QIAXCEL capillary electrophoresis system to confirm the change in size. The adapter-ligated library was quantified by qPCR and then analyzed using Illumina Microarray. Seq-Sequencing was performed at Dana-Farber Cancer Institute using 250bp paired-end reads. The study was conducted by the Molecular Biology Core Facility. We performed 75,000-1,000 PCR for each sample. 270,000 reads (average approximately 422,000) were analyzed. As previously described (Sander et al., 2013) analyzed the indel mutagenesis rate of TruSeq reads. Specificity ratios observed at on-target loci determined by deep sequencing Calculated as the ratio of the total mutagenesis observed at a specific off-target locus to the total mutagenesis observed at a specific off-target locus. The fold improvement in specificity of tru-RGN for each off-target site was calculated using matched Ratio of specificity of the same target using tru-gRNA versus the corresponding full-length gRNA The specificity was calculated as a ratio of observed specificity. As mentioned in the text, some off-target No indel mutations were observed with tru-gRNA at these sites. Using a Poisson calculator, we calculated the upper limit of the actual number of variant sequences to be 3. The upper limit was calculated with a 95% confidence level. The inventors then used this upper limit to estimate the off-target The minimum fold improvement in specificity of the RT-site was estimated.

[0131] Example 2a. Truncated gRNAs efficiently direct Cas9-mediated genome editing in human cells can 5'-truncated gRNAs do not function as efficiently as their full-length counterparts To test this hypothesis, we first identified the following sequence: 5'- GG C G A G GGCGATG of the EGFP reporter gene with CCACCTAcGG-3' (SEQ ID NO: 2241) A series of increasingly shorter gRNAs were constructed against a single target site as described above. The specific EGFP sites were chosen because each has a G at the 5' end (used in these experiments). 15nt, 17nt containing the nucleotide sequence (necessary for efficient expression from the U6 promoter) Thus, 19nt and 20nt complementary gRNAs can be generated for this site. This is because there was a single enhanced green fluorescent protein (E) that was integrated and constitutively expressed. Quantifying the frequency of RGN-induced indels by assessing the disruption of the GFP gene Human cell-based reporter assays that can be used for eyon et al., 2012) (Figure 2B) to confirm that the gRNAs of different lengths were inserted into the target site. The ability to induce Cas9-induced indels was measured.

[0132] As mentioned above, gRNAs with longer complementarity (21 nt, 23 nt, and 25 nt) showed lower activity than the standard full-length gRNA containing a 20-nt complementary sequence (Figure 2H ), this result is in line with results recently reported by others (Ran et al., Cell 2013). However, gRNAs with 17-nt or 19-nt target complementarity were not fully The activity was equal to or greater than that of the long gRNA, whereas the shorter gRNA (only 15 nt) No significant activity was observed with the complementary gRNA (Figure 2H).

[0133] To verify the generality of these initial findings, we used full-length gRNAs and four additional EGFRs. For FP reporter gene sites (EGFP sites #1, #2, #3, and #4; Figure 3A) The corresponding gRNAs with 18nt, 17nt, and / or 16nt complementarity were then up-regulated. At each of the four target sites, 17nt and / or 18nt complementarity was obtained. gRNAs with the nucleotide sequence cleaved as efficiently as their corresponding full-length gRNAs (or in some cases , and more efficiently) to induce Cas9-mediated disruption of EGFP expression (Figure 3A However, with gRNAs with only 16 nt of complementarity, only two sites were created. The inventors observed significantly reduced or undetectable activity in the α-glucan-1-phosphate dehydrogenase (α-glucan-1-phosphate dehydrogenase) (Fig. 3A). For each of the various sites identified, equal amounts of full-length or truncated gRNA expression plasmids were used. The EGFP sites #1, #2 and #3 were transfected with the GFP and Cas9 expression plasmids. For #1 and #2, transfect the Cas9 expression plasmid and truncated gRNA expression plasmid. Control experiments with varying amounts of plasmid demonstrated that truncated gRNAs functioned as well as their full-length counterparts. (Figure 3E (bottom) and 3F (bottom)). It was suggested that the same amount of plasmid could be used when carrying out the above procedure. These results indicate that truncated gRNAs with 17nt or 18nt complementarity generally produce fully This shows that the complementarity of gRNA can function as efficiently as that of long gRNA. A truncated gRNA of the required length is called a "tru-gRNA." An RGN using this is called a "tru-RGN."

[0134] tru-RGN then efficiently induces indels at chromatinized endogenous gene targets. We have already verified whether it is possible to synthesize three endogenous human genes (VEGFA, EMX Four sites are targeted with standard full-length gRNAs within the VEG (VEG1 and CLTA): FA site 1, VEGFA site 3, EMX1 and CTLA (see Example 1 and Fu et al., 2014). 13; Hsu et al., 2013; Pattanayak et al., 2013). tru-gRNAs were constructed against seven sites within the tru gene (Figure 3B). For VEGFA site 2, this target sequence is involved in gRNA expression from the U6 promoter. Since there is no G at position 17 or 18 of the complementary region required for -It was not possible to validate the gRNA.) Using the tyrosine kinase I (T7EI) genotyping assay (Reyon et al., 2012), The frequency of Cas9-mediated indel mutations induced by various gRNAs at their target sites was calculated. Quantification was performed in human U2OS.EGFP cells. Five of the seven four sites were negative. This also indicates that tru-RGNs are similar to the indel mutations mediated by the corresponding canonical RGNs. The insertion-deletion mutations were reliably induced with an efficiency of approximately 100% (Fig. 3B). For the two sites that showed lower activity than the reactant, the absolute mutagenesis rate was still very low. It is noteworthy that the values ​​were high (average 13.3% and 16.6%), which may be useful for various applications. Three of the target sites (VEGFA sites 1 and 3 and EMX1) Sanger sequencing of the tru-RGN-induced indels predicted These mutations are derived from the cleavage sites induced by canonical RGN. It was confirmed that the mutations were virtually indistinguishable (Fig. 3C and Figs. 7A to 7D).

[0135] In addition, the present invention shows that a minimum of 17 nt of complementarity is required for efficient RGN activity. As previously reported, tru-gR has a mismatched 5'G and an 18-nt complementary region. NA can efficiently induce Cas9-induced indels, whereas mismatch 5 The tru-gRNA, which contains a 'G' and a 17-nt complementary region, is complementary to the corresponding full-length gRNA. It was found that the activity was low or undetectable compared to the control (Fig. 7E).

[0136] To further evaluate the genome editing ability of tru-RGN, we investigated whether tru-RGN was a ssODN driver. We tested the ability to induce precise sequence changes via HDR using a base template. Studies have shown that in human cells, Cas9-induced cleavage occurs from the homologous ssODN donor to the endogenous locus. It has been shown that the introduction of nucleotide sequences can stimulate the transfer of nucleotide sequences (Cong et al., 2013; Mali et al., 2014). 013c; Ran et al., 2013; Yang et al., 2013). Therefore, the homologous ssOD The ability to introduce the N-encoded BamHI restriction site into these endogenous genes was determined by VEGFA The corresponding full-length gRNAs and tru-gRNAs targeting site 1 and EMX1 sites were tru-RGN contains standard gRNAs carrying full-length gRNA counterparts at both sites. The BamHI site was introduced with a similar efficiency to that of RGN (Fig. 3D). The data show that tru-RGN activates both sites as efficiently as standard RGN in human cells. We demonstrate that it can function to induce indels and precise HDR-mediated genome editing events. are.

[0137] Example 2b. tru-RGN Shows Increased Sensitivity to Mismatches at the gRNA / DNA Junction Indicates large tru-RGN can function efficiently to induce on-target genome editing changes. These nucleases were confirmed to be effective against mismatches at the gRNA / DNA junction. To evaluate this, we tested whether EGFP site #1, A systematic series of mutants of the tru-gRNA (upper part of Figure 3A) tested in #2 and #3 This mutant gRNA was constructed at each position within the complementary region (expression from the U6 promoter). It carries a single Watson-Crick substitution at position 1 (excluding the required 5' G) (Figure 5A). Using the same three-part EGFP-based decay assay as described in Example 1 These mutant tru-gRNAs and the corresponding series of similar mutations generated for positions The relative abilities of all full-length gRNAs to induce Cas9-mediated indels were assessed. Therefore, at all three EGFP target sites, tru-RGN generally exhibited a higher activity than the corresponding Higher sensitivity to single mismatches than standard RGN carrying full-length gRNAs (Compare the upper and lower panels in Figure 5A.) The sensitivity differs depending on the region. The largest difference was observed between sites #2 and #3, where tru-gRNA shares 17 nt of complementarity. It was seen.

[0138] Boosted by increased sensitivity of tru-RGN to single nucleotide mismatches Therefore, we next systematically identified two adjacent positions of the gRNA-DNA junction. Therefore, the inventors tried to investigate the effect of matching the have Watson-Crick transversion substitutions at two adjacent nucleotide positions, Mutants of tru-gRNA targeting GFP target sites #1, #2, and #3 were generated. (Fig. 5B) As determined by EGFP decay assay, adjacent double mismatches inhibit RGN activity. Similarly, the effects on sexuality were also significantly greater with tru-gRNA than with all three types of EGF prepared in Example 1. P was significantly greater than the similar variants of the corresponding full-length gRNA targeting the target site. (Compare the upper and lower panels of Figure 5B.) These effects appear to be site-dependent. Almost all of the double mismatched tru-gRNAs against EGFP sites #2 and #3 However, it did not show increased EGFP decay activity compared to a control gRNA lacking the complementary region. Only three of the mismatched tru-gRNA mutants for GFP site #1 remained active Furthermore, double mutations generally showed a 5'-terminal tail in the full-length gRNA. This showed a significant effect, whereas no such effect was observed with tru-gRNA. Taken together, our data suggest that tru-gRNAs are more efficient at regulating transcription than full-length gRNAs. Single and double Watson-Crick transversions at RNA-DNA junctions This suggests that it exhibits high sensitivity to mismatches.

[0139] Example 2c. tru-RGN Targeting Endogenous Genes Shows Improved Specificity in Human Cells vinegar tru-RGN expresses gRNA in human cells more efficiently than standard RGN carrying a full-length gRNA counterpart The following experiments were carried out to clarify whether the genomic off-target effects of Previous studies (see Example 1 and Fu et al., 2013; Hsu et al., 2013) (See Figure 3B) at VEGFA site 1, VEGFA site 3, and EMX1 site 1 (see Figure 3B). 13 true off-target sites for full-length gRNAs targeting the target gene (described in Having identified the sites, we then generated the corresponding full-length ribonucleotides that target these sites. gRNA and tru-gRNA were tested. (VEGFA site 2 was identified using the U6 promoter.) Gs at positions 17 and 18 of the complementary region required for efficient gRNA expression. Therefore, we tested the tru-gRNA for this target sequence from our earlier study. Notably, as determined by the T7EI assay, In addition, tru-RGN was found to be responsible for all 13 of these true off-target sites. Its mutagenic activity in human U2OS.EGFP cells was significantly reduced compared to standard RGN. It was found that the nucleotide sequence was reduced at 11 of the 13 off-target sites (Table 3A). The frequency of tru-RGN mutations was within the reliable detection limit of the T7EI assay (2–5 % (Table 3A). was tested at the same 13 off-target sites in another human cell line (FT-HEK293 cells). When tested, nearly identical results were observed (Table 3A).

[0140] To quantify the degree of specificity improvement observed in tru-RGN, we performed a low-frequency High-throughput assay, a method for detecting and quantifying mutations with higher sensitivity than the T7EI assay We measured the off-target mutation frequency using putative sequencing. 13 true sites where tru-gRNA reduced mutation rates by 7EI assay A subset of 12 off-target sites was evaluated (1 of 13 sites). In some cases, it was not possible to amplify the required short amplicons for technical reasons. ), as well as other EMX1 site 1 off-targets identified by another group 7. The off-target sites were investigated (Figure 6A). Therefore, tru-RGN showed a significantly reduced absolute frequency of mutagenesis compared to the corresponding standard RGN. It exhibits a high degree of specificity (Figure 6A and Table 3B), approximately 5000 times more specific than its standard RGN counterpart. This resulted in improved isomerism (Figure 6B). Two off-target sites (OT1-4 and O For T1-11), the absolute number of indel mutations induced by tru-RGN and and frequency are reduced to background levels or near background levels. Therefore, it is difficult to quantify the on-target to off-target ratio of tru-RGN. Therefore, in this case, the ratio of on-target to off-target is infinite. To address this issue, we instead investigated the The maximum estimated insertion / deletion frequency at a 95% confidence level is identified, and then this conservative Using the estimates, we investigated the tru- The minimum estimated specificity improvement of RGN was calculated. It is suggested that RGN results in approximately 10,000-fold or greater improvement in these areas. (Figure 6B).

[0141] To further verify the specificity of tru-RGN, we investigated whether tru-RGN is expressed in human We investigated their ability to induce off-target mutations at other closely related sites in the genome. The target regions of VEGFA site 1 and EMX1 are complementary to each other by 18 nt. For u-gRNA, a human gRNA with mismatches at one or two positions within the complementary region was used. Any other site in the genome (sites not already discussed in Table 3A above) and the present inventors The mismatches at the 5' end of the site described in Example 1, which is supported by A subset of all sites with a match was identified computationally. For tru-gRNA against VEGFA site 3 with topological complementarity, Any site with a mismatch at one of the positions where a mismatch is supported and any site with a mismatch at two of the positions where a mismatch is supported Any site with a mismatch at position (also not yet considered in Table 3A) This computer-based analysis identified a subset of VEGFA site 1, VEFG Further potential off-targeting of tru-RGN by targeting A-site 3 and EMX1 sites, respectively A total of 30, 30, and 34 target sites were obtained, respectively, and then RGN T7EI assay was performed on human U2OS.EGFP cells and HEK293 cells expressing α-glucan. mutations at these sites were assessed using the .

[0142] Notably, VEGFA site 1, VEFGA site 3, and EMX1 were identified. The three types of tru-RGNs were identified from 94 potential sites examined in human U2OS.EGFP cells. The assay did not induce detectable Cas9-mediated indel mutations at 93 of the target off-target sites. No activity was detected at any of the 94 potential off-target sites in human HEK293 cells. No potential Cas9-mediated indel mutations were induced (Table 3C). For one site found, a target with a full-length gRNA targeting VEGFA site 1 was identified. We investigated whether a comparable RGN could also mutate this same off-target site. The results showed that standard RGN was slightly less frequent than tru-RGN, but still detectable. This off-target site was not improved by shortening the gRNA (Figure 6C). The fact that this was not observed is due to the 20-nt and 18-nt sequences of the full-length gRNA and tru-gRNA. This can be understood by comparing the full-length 20 nt sequence with the The two additional bases in the nt target are both mismatched (Figure 6C ) In summary, this survey of 94 additional potential off-target sites revealed The results show that shortening the gRNA does not induce a high frequency of new off-target mutations. It seems that...

[0143] The most closely matched potential off-target site (i.e., Deep sequencing of 30 sites (i.e., sites with one or two mismatches) This was comparable to that observed at other previously identified off-target sites (Table 3B). Indeletions that were not reproducible or had very low rates of indeletions (Table 3D) were shown. The inventors found that tru-RGNs generally have one or two mismatches with the on-target site. It is thought that the mutations induce very low or undetectable levels at different sites depending on the This conclusion was supported by the fact that five mismatches resulted in high frequency at different sites. This is in contrast to standard RGN, where off-target mutations are relatively common. (See Example 1).

[0144] [Table 12]

[0145] [Table 13]

[0146] [Table 14] JPEG0007812830000130.jpg102167JPEG0007812830000131.jpg104167JPEG0007812830000132.jpg10416 7JPEG0007812830000133.jpg102167JPEG0007812830000134.jpg103167JPEG0007812830000135.jpg20167

[0147] [Table 15] JPEG0007812830000137.jpg82169

[0148] Example 2d. tru-gRNA Used with Dual Cas9 Nickases to Efficacy in Human Cells It can efficiently induce genome editing Recently described dual Cas9-nickase method for inducing indel mutations was used to To perform this test, Cas9-D10A nickase was used to The regions of the human VEGFA gene (VEGFA site 1 and what we call VEGFA site 4) U2OS.EGFP cells with two full-length gRNAs targeting the GFP-targeting gene (an additional sequence referred to as gRNA). As previously described (Ran et al., 2013), this The pair of cyclases functions cooperatively to induce high rates of indel mutations at the VEGFA target locus. Interestingly, it was only co-expressed with the gRNA targeting VEGFA site 4 (Figure 4B). The expressed Cas9-D10A nickase showed a mutation rate similar to that observed with the paired full-length gRNA. Although somewhat lower, a high rate of indel mutations was induced (Figure 4B). Using tru-gRNA instead of full-length gRNA at VEGFA site 1 also resulted in double-nicked The case method did not affect the efficiency of inducing indel mutations (Fig. 4B).

[0149] The dual nickase strategy also stimulates the introduction of specific sequence changes using ssODNs. Since it has also been used for the purpose of We also tested whether ru-gRNA could be used for this type of alteration. Paired full-length gRNAs targeting A sites 1 and 4 cooperate with the Cas9-D10A nickase. The VEGFA gene of human U2OS.EGFP cells predicted from the ssODN donor This enhanced the efficient introduction of short inserts into the locus (Figure 3A) (Figure 3C). The efficiency of ssODN-mediated sequence alteration by cloning was significantly improved by the full-length gamma-gamma targeting VEGFA site 1. The levels were maintained at a similar level to when tru-gRNA was used instead of RNA (Figure 3 C) Taken together, these results demonstrate that tru-gRNA is a promising candidate for the dual Cas9 nickase strategy. By using the gene as a gene for genome editing, insertion / deletion mutations and ss mutations can be generated without compromising the efficiency of genome editing by this method. These results demonstrate that ODN-mediated sequence changes can induce both

[0150] The use of tru-gRNA did not eliminate the on-target genome editing activity of the paired nickase. Since it was confirmed that there was no loss of DNA, the inventors then used deep sequencing to VEGFA site 1 gRN at four previously identified bona fide off-target sites From this analysis, we investigated the mutation frequency caused by A. When used together, the mutation rate at any of the four off-target sites is substantially It was revealed that the level of the oxidative stress decreased to an undetectable level (Table 4). At one of the target sites (OT1-3), tru-RGN (Table 3B), full-length Both pairs of nickases with gRNA (Table 4) completely eliminated off-target mutations. These results suggest that the use of tru-gRNA allows Eliminates off-target effects of paired Cas9 nickases without compromising the efficiency of targeted genome editing can be further reduced (and vice versa).

[0151] [Table 16]

[0152] References Cheng, A. W., Wang, H., Yang, H., Shi, L., Katz, Y. .,Theunissen,T.W.,Rangarajan,S.,Shivalil a,CS,Dadon,DB,and Jaenisch,R.Multipl exed activation of endogenous genes by C RISPR-on, an RNA-guided transcriptional a ctivator system.Cell Res 23,1163-1171.(2 013). Cho, SW, Kim, S., Kim, JM & Kim, JSTarget ed genome engineering in human cells wit h the Cas9 RNA-guided endonuclease.Nat B iotechnol 31,230-232(2013). Cong,L.et al.Multiplex genome engineerin g using CRISPR / Cas systems.Science 339,8 19-823(2013). Cradick,T.J.,Fine,E.J.,Antico,C.J.,and B ao,G.CRISPR / Cas9 systems targeting beta- globin and CCR5 genes have substantial o ff-target activity.Nucleic Acids Res.(20 13). Dicarlo,J.E.et al.Genome engineering in Saccharomyces cerevisiae using CRISPR-Ca s systems.Nucleic Acids Res(2013). Ding,Q.,Regan,S.N.,Xia,Y.,Oostrom,L.A.,C owan,C.A.,and Musunuru,K.Enhanced effici ency of human pluripotent stem cell geno me editing through replacing TALENs with CRISPRs.Cell Stem Cell 12,393-394.(2013 ). Fisher,S.,Barry,A.,Abreu,J.,Minie,B.,Nol an,J.,Delorey,T.M.,Young,G.,Fennell,T.J. ,Allen,A.,Ambrogio,L.,et al.A scalable,f ully automated process for construction of sequence-ready human exome targeted c apture libraries.Genome Biol 12,R1.(2011 ). Friedland,A.E.,Tzur,Y.B.,Esvelt,K.M.,Col aiacovo,M.P.,Church,G.M.,and Calarco,J.A .Heritable genome editing in C.elegans v ia a CRISPR-Cas9 system.Nat Methods 10,7 41-743.(2013). Fu,Y.,Foden,J.A.,Khayter,C.,Maeder,M.L., Reyon,D.,Joung,J.K.,and Sander,J.D.High- frequency off-target mutagenesis induced by CRISPR-Cas nucleases in human cells. Nat Biotechnol 31,822-826.(2013). Gabriel,R.et al.An unbiased genome-wide analysis of zinc-finger nuclease specifi city.Nat Biotechnol 29,816-823(2011). Gilbert,L.A.,Larson,M.H.,Morsut,L.,Liu,Z .,Brar,G.A.,Torres,S.E.,Stern-Ginossar,N .,Brandman,O.,Whitehead,E.H.,Doudna,J.A. ,et al.(2013).CRISPR-Mediated Modular RN A-Guided Regulation of Transcription in Eukaryotes.Cell 154,442-451. Gratz,S.J.et al.Genome engineering of Dr osophila with the CRISPR RNA-guided Cas9 nuclease.Genetics(2013). Hockemeyer,D.et al.Genetic engineering o f human pluripotent cells using TALE nuc leases.Nat Biotechnol 29,731-734(2011). Horvath,P.& Barrangou,R.CRISPR / Cas,the i mmune system of bacteria and archaea.Sci ence 327,167-170(2010). Hsu,P.D.,Scott,D.A.,Weinstein,J.A.,Ran,F .A.,Konermann,S.,Agarwala,V.,Li,Y.,Fine, E.J.,Wu,X.,Shalem,O.,et al.DNA targeting specificity of RNA-guided Cas9 nuclease s.Nat Biotechnol 31,827-832.(2013). Hwang,W.Y.et al.Efficient genome editing in zebrafish using a CRISPR-Cas system. Nat Biotechnol 31,227-229(2013). Hwang,W.Y.,Fu,Y.,Reyon,D.,Maeder,M.L.,Ka ini,P.,Sander,J.D.,Joung,J.K.,Peterson,R .T.,and Yeh,J.R.Heritable and Precise Ze brafish Genome Editing Using a CRISPR-Ca s System.PLoS One 8,e68708.(2013a). Jiang,W.,Bikard,D.,Cox,D.,Zhang,F.& Marr affini,L.A.RNA-guided editing of bacteri al genomes using CRISPR-Cas systems.Nat Biotechnol 31,233-239(2013). Jinek,M.et al.A programmable dual-RNA-gu ided DNA endonuclease in adaptive bacter ial immunity.Science 337,816-821(2012). Jinek,M.et al.RNA-programmed genome edit ing in human cells.Elife 2,e00471(2013). Li,D.,Qiu,Z.,Shao,Y.,Chen,Y.,Guan,Y.,Liu ,M.,Li,Y.,Gao,N.,Wang,L.,Lu,X.,et al.Her itable gene targeting in the mouse and r at using a CRISPR-Cas system.Nat Biotech nol 31,681-683.(2013a). Li,W.,Teng,F.,Li,T.,and Zhou,Q.Simultane ous generation and germline transmission of multiple gene mutations in rat using CRISPR-Cas systems.Nat Biotechnol 31,68 4-686.(2013b). Maeder,M.L.,Linder,S.J.,Cascio,V.M.,Fu,Y .,Ho,Q.H.,and Joung,J.K.CRISPR RNA-guide d activation of endogenous human genes.N at Methods 10,977-979.(2013). Mali,P.,Aach,J.,Stranges,P.B.,Esvelt,K.M .,Moosburner,M.,Kosuri,S.,Yang,L.,and Ch urch,G.M.CAS9 transcriptional activators for target specificity screening and pa ired nickases for cooperative genome eng ineering.Nat Biotechnol 31,833-838.(2013 a). Mali,P.,Esvelt,K.M.,and Church,G.M.Cas9 as a versatile tool for engineering biol ogy.Nat Methods 10,957-963.(2013b). Mali,P.et al.RNA-guided human genome eng ineering via Cas9.Science 339,823-826(20 13c). Pattanayak,V.,Lin,S.,Guilinger,J.P.,Ma,E .,Doudna,J.A.,and Liu,D.R.High-throughpu t profiling of off-target DNA cleavage r eveals RNA-programmed Cas9 nuclease spec ificity.Nat Biotechnol 31,839-843.(2013) . Pattanayak,V.,Ramirez,C.L.,Joung,J.K.& L iu,D.R.Revealing off-target cleavage spe cificities of zinc-finger nucleases by i n vitro selection.Nat Methods 8,765-770( 2011). Perez,E.E.et al.Establishment of HIV-1 r esistance in CD4+ T cells by genome edit ing using zinc-finger nucleases.Nat Biot echnol 26,808-816(2008). Perez-Pinera,P.,Kocak,D.D.,Vockley,C.M., Adler,A.F.,Kabadi,A.M.,Polstein,L.R.,Tha kore,P.I.,Glass,K.A.,Ousterout,D.G.,Leon g,K.W.,et al.RNA-guided gene activation by CRISPR-Cas9-based transcription facto rs.Nat Methods 10,973-976.(2013). Qi,L.S.,Larson,M.H.,Gilbert,L.A.,Doudna, J.A.,Weissman,J.S.,Arkin,A.P.,and Lim,W. A.Repurposing CRISPR as an RNA-guided pl atform for sequence-specific control of gene expression.Cell 152,1173-1183.(2013 ). Ran,F.A.,Hsu,P.D.,Lin,C.Y.,Gootenberg,J. S.,Konermann,S.,Trevino,A.E.,Scott,D.A., Inoue,A.,Matoba,S.,Zhang,Y.,et al.Double nicking by RNA-guided CRISPR Cas9 for e nhanced genome editing specificity.Cell 154,1380-1389.(2013). Reyon,D.et al.FLASH assembly of TALENs f or high-throughput genome editing.Nat Bi otech 30,460-465(2012). Sander,J.D.,Maeder,M.L.,Reyon,D.,Voytas, D.F.,Joung,J.K.,and Dobbs,D.ZiFiT(Zinc F inger Targeter):an updated zinc finger e ngineering tool.Nucleic Acids Res 38,W46 2-468.(2010). Sander,J.D.,Ramirez,C.L.,Linder,S.J.,Pat tanayak,V.,Shoresh,N.,Ku,M.,Foden,J.A.,R eyon,D.,Bernstein,B.E.,Liu,D.R.,et al.In silico abstraction of zinc finger nucle ase cleavage profiles reveals an expande d landscape of off-target sites.Nucleic Acids Res.(2013). Sander,J.D.,Zaback,P.,Joung,J.K.,Voytas, D.F.,and Dobbs,D.Zinc Finger Targeter(Zi FiT):an engineered zinc finger / target si te design tool.Nucleic Acids Res 35,W599 -605.(2007). Shen,B.et al.Generation of gene-modified mice via Cas9 / RNA-mediated gene targeti ng.Cell Res(2013). Sugimoto,N.et al.Thermodynamic parameter s to predict stability of RNA / DNA hybrid duplexes.Biochemistry 34,11211-11216(19 95). Terns,MP& Terns,RMCRISPR-based adaptation ive immune systems.Curr Opin Microbiol 1 4,321-327(2011). Wang,H.et al.One-Step Generation of Mice Carrying Mutations in Multiple Genes by CRISPR / Cas-Mediated Genome Engineering. Cell 153,910-918(2013). Wiedenheft, B., Sternberg, SH & Doudna, JA. .RNA-guided genetic silencing systems in bacteria and archaea.Nature 482,331-338 (2012). Yang, L., Guell, M., Byrne, S., Yang, J. L., De L. os Angeles, A., Mali, P., Aach, J., Kim-Kisela k,C.,Briggs,AW,Rios,X.,et al.(2013).Op timization of scarless human stem cell g enome editing.Nucleic Acids Res 41,9049- 9061.

[0153] Other embodiments While the present invention has been described in connection with its detailed description, it should be understood that the above description is for illustrative purposes only and is not intended to limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims. The invention described herein may include the following aspects. [1] 1. A method for increasing the specificity of RNA-guided genome editing in a cell, comprising contacting the cell with a guide RNA comprising a complementary region consisting of 17-18 nucleotides that is complementary to 17-18 consecutive nucleotides of the complementary strand of a selected target genome sequence. [2] 1. A method for inducing breaks in a target region of a double-stranded DNA molecule in a cell, comprising: Cas9 nuclease or Cas9 nickase, a guide RNA containing a complementary region consisting of 17 to 18 nucleotides complementary to 17 to 18 consecutive nucleotides of the complementary strand of the double-stranded DNA molecule; in the cell or introducing into the cell. [3] 1. A method for modifying a target region of a double-stranded DNA molecule in a cell, comprising: a dCas9-heterologous functional domain fusion protein (dCas9-HFD); a guide RNA comprising a complementary region consisting of 17 to 18 nucleotides complementary to 17 to 18 consecutive nucleotides of the complementary strand of the selected target genome sequence; in the cell or introducing into the cell. [4] The guide RNA is (i) a single guide RNA comprising a complementary region consisting of 17 to 18 nucleotides complementary to 17 to 18 consecutive nucleotides of the complementary strand of a selected target genome sequence; or (ii) a crRNA containing a complementary region consisting of 17 to 18 nucleotides complementary to 17 to 18 consecutive nucleotides of the complementary strand of a selected target genome sequence, and a tracrRNA. The method according to any one of the above [1] to [3], [5] The guide RNA is (X 17~18 or X 17~19 )GUUUUAGAGCUA(SEQ ID NO:2404); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 2407); or (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(SEQ ID NO:2408); (X 17~18 )GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCG(X N ) (SEQ ID NO: 1), (X 17~18 )GUUUUAGAGCUAUGCUGAAAAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N ) (SEQ ID NO: 2), (X 17~18 )GUUUUAGAGCUAUGCUGUUUUGGAAACAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N ) (SEQ ID NO: 3), (X 17~18 )GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(X N ) (SEQ ID NO: 4), (X 17~18 )GUUUAAGAGCUAGAAAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 5); (X 17~18 )GUUUUAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 6); or (X 17~18 )GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7) is or comprises a ribonucleic acid selected from the group consisting of: X 17~18 is a complementary region complementary to 17 to 18 consecutive nucleotides of the complementary strand of a selected target sequence, preferably adjacent to the 5' side of the protospacer adjacent motif (PAM), and X N is any sequence that does not interfere with the binding of the ribonucleic acid to Cas9, and N can be 0 to 200, for example, 0 to 100, 0 to 50, or 0 to 20. The method according to any one of the above [1] to [3]. [6] A guide RNA molecule having a target-complementary region of 17-18 nucleotides. [7] The gRNA according to [6] above, wherein the target-complementary region consists of 17 to 18 nucleotides. [8] The gRNA according to [6] above, wherein the target-complementary region consists of target-complementary sequences of 17 to 18 nucleotides. [9] The following array (X 17~18 or X 17~19 )GUUUUAGAGCUA(SEQ ID NO:2404); (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 2407); or (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU(SEQ ID NO:2408); (X 17~18 )GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCG(X N ) (SEQ ID NO: 1), (X 17~18 )GUUUUAGAGCUAUGCUGAAAAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N ) (SEQ ID NO: 2), (X 17~18 )GUUUUAGAGCUAUGCUGUUUUGGAAACAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUC(X N ) (SEQ ID NO: 3), (X 17~18 )GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(X N ) (SEQ ID NO: 4), (X 17~18 )GUUUAAGAGCUAGAAAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 5); (X 17~18 )GUUUUAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 6); or (X 17~18 )GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7) It consists of X 17~18 is a sequence complementary to 17 to 18 consecutive nucleotides of the complementary strand of a selected target sequence, preferably adjacent to the 5' side of the protospacer adjacent motif (PAM), and X N is any sequence that does not interfere with the binding of the ribonucleic acid to Cas9, and N can be 0 to 200, for example, 0 to 100, 0 to 50, or 0 to 20. The gRNA described in [6] above.

[10] The method according to [5] above or the ribonucleic acid according to [6] above, wherein the ribonucleic acid comprises one or more U at the 3' end of the ribonucleic acid molecule.

[11] The method according to [5] above or the ribonucleic acid according to [6] above, wherein the ribonucleic acid comprises one or more additional nucleotides at the 5' end of the RNA molecule that are not complementary to the target sequence.

[12] The method according to [5] above or the ribonucleic acid according to [6] above, wherein the ribonucleic acid comprises one, two or three additional nucleotides at the 5' end of the RNA molecule that are not complementary to the target sequence.

[13] The method according to any one of [1] to [5] above or the ribonucleic acid according to any one of [6] to

[12] above, wherein the complementary region is complementary to 17 consecutive nucleotides of a complementary strand of a selected target sequence.

[14] The method according to any one of [1] to [5] above or the ribonucleic acid according to any one of [6] to

[12] above, wherein the complementary region is complementary to 18 consecutive nucleotides of a complementary strand of a selected target sequence.

[15] A DNA molecule encoding the ribonucleic acid according to any one of [6] to

[14] above.

[16] A vector comprising the DNA molecule described in

[15] above.

[17] A host cell expressing the vector described in

[16] above.

[18] The method according to any one of [1] to [5] above or the ribonucleic acid according to any one of [6] to

[12] above, wherein the target region is within a target genome sequence.

[19] The method according to any one of [1] to [5] above or the ribonucleic acid according to any one of [6] to

[12] above, wherein the target genomic sequence is adjacent to the 5' side of a protospacer adjacent motif (PAM).

[20] the tracrRNA having the sequence GGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 8), or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2405), or an active portion thereof; AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2407) or an active portion thereof; CAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2409) or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUG (SEQ ID NO: 2410), or an active portion thereof; UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA (SEQ ID NO: 2411) or an active portion thereof; or UAGCAAGUUAAAAUAAGGCUAGUCCG (SEQ ID NO: 2412) or an active portion thereof The method according to [4] above, comprising:

[21] The cRNA is (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 2407) and the tracrRNA is GGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 8); or the cRNA is (X 17~18 or X 17~19 )GUUUUAGAGCUA (SEQ ID NO: 2404) and the tracrRNA is UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2405); or the cRNA is (X 17~18 or X 17~19 )GUUUUAGAGCUAUGCU (SEQ ID NO: 2408), and the tracrRNA is AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 2406).

[22] The method according to [3] above, wherein the dCas9-heterologous functional domain fusion protein (dCas9-HFD) comprises an HFD that modifies gene expression, histones, or DNA.

[23] The method according to

[22] above, wherein the heterologous functional domain is a transcription activation domain, an enzyme that catalyzes DNA demethylation, an enzyme that catalyzes histone modification, or a transcription silencing domain.

[24] The method according to

[23] above, wherein the transcription activation domain is derived from VP64 or NF-κB p65.

[25] The method according to

[23] above, wherein the enzyme that catalyzes histone modification is LSD1, histone methyltransferase (HNMT), histone acetyltransferase (HAT), histone deacetylase (HDAC) or histone demethylase.

[26] The method according to

[23] above, wherein the transcriptional silencing domain is derived from heterochromatin protein 1 (HP1), such as HP1α or HP1β.

[27] The method according to any one of [1] to [5] above, wherein an insertion / deletion mutation or a sequence change is generated in the selected target genome sequence.

[28] The method according to any one of [1] to [5] above, wherein the cell is a eukaryotic cell.

[29] The method according to

[28] above, wherein the cells are mammalian cells.

Claims

1. An ex vivo method for increasing the specificity of S. pyogenes CRISPR-Cas9 (Cas9) RNA-guided genome editing in cells, comprising: The method comprises contacting the cell with a guide RNA, the guide RNA comprising a complementary region at the 5' end of the guide RNA consisting of 17-18 nucleotides that is complementary to 17-18 consecutive nucleotides of a complementary strand of a selected target genomic sequence; the selected target genomic sequence is adjacent to the 5' side of a protospacer adjacent motif (PAM); the guide RNA comprises SEQ ID NO: 4, In the presence of the S. pyogenes Cas9 genome editing enzyme, the guide RNA complementary region binds to the selected target genome sequence and guides the Cas9 genome editing enzyme to the selected target genome sequence; An ex vivo method, thereby increasing the specificity of RNA-guided genome editing in cells.

2. 1. An ex vivo method for inducing breaks in a target region of a double-stranded DNA molecule in a cell, comprising: The method comprises: S. pyogenes CRISPR / Cas9 nuclease or nickase; a guide RNA comprising a complementary region at the 5' end of the guide RNA consisting of 17 to 18 nucleotides complementary to 17 to 18 consecutive nucleotides of a complementary strand of a double-stranded DNA molecule comprising a target sequence; in the cell or by introducing into the cell the target sequence is adjacent to the 5' side of a protospacer adjacent motif (PAM); the guide RNA complementary region binds to the target region of the double-stranded DNA molecule and guides the Cas9 nuclease or nickase to the target region of the double-stranded DNA molecule; the guide RNA comprises SEQ ID NO: 4, This ex vivo method induces breaks in targeted regions of double-stranded DNA molecules within cells.

3. 1. An ex vivo method for modifying a target region of a double-stranded DNA molecule in a cell, comprising: The method comprises: S. pyogenes CRISPR dCas9-heterologous function domain fusion protein (dCas9-HFD); a guide RNA comprising a complementary region at the 5' end of the guide RNA consisting of 17-18 nucleotides that is complementary to 17-18 consecutive nucleotides of the complementary strand of a selected target sequence present on a double-stranded DNA molecule; in the cell or by introducing into the cell the selected target sequence is adjacent to the 5' side of a protospacer adjacent motif (PAM); the guide RNA comprises SEQ ID NO: 4, the guide RNA complementary region binds to the selected target sequence and guides the dCas9-HFD to the selected target sequence; This provides an ex vivo method for modifying target regions of double-stranded DNA molecules within cells.

4. The ex vivo method of claim 2 or 3, wherein the target region is in a target genomic sequence.

5. 4. The ex vivo method of claim 3, wherein the dCas9-HFD comprises a heterologous functional domain (HFD) that modifies gene expression, histones, or DNA.

6. 6. The ex vivo method of claim 5, wherein the HFD is a transcriptional activation domain, an enzyme that catalyzes DNA demethylation, an enzyme that catalyzes histone modification, or a transcriptional silencing domain.

7. 7. The ex vivo method of claim 6, wherein the transcription activation domain is derived from VP64 or NF-kappa B subunit p65 (NF-κB p65).

8. 7. The ex vivo method of claim 6, wherein the enzyme that catalyzes the histone modification is lysine-specific histone demethylase 1 (LSD1), histone methyltransferase (HNMT), histone acetyltransferase (HAT), histone deacetylase (HDAC), or histone demethylase.

9. 7. The ex vivo method of claim 6, wherein the transcriptional silencing domain is derived from heterochromatin protein 1 alpha (HP1α) or heterochromatin protein 1 beta (HP1β).

10. The ex vivo method according to any one of claims 1 to 3, wherein the cell is a eukaryotic cell.

11. The ex vivo method of claim 10, wherein the cell is a mammalian cell.

12. 1. An ex vivo method for RNA-guided genome editing in a cell, comprising: The method comprises: contacting the cell with a guide RNA (gRNA) comprising a complementary region consisting of 17-18 nucleotides that is complementary to 17-18 consecutive nucleotides of a complementary strand of a target genomic sequence; the gRNA comprises SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:2404, SEQ ID NO:2407, or SEQ ID NO:2408; In the presence of S. pyogenes Cas9 nuclease, the gRNA complementary region binds to the target genomic sequence and guides the Cas9 nuclease to the target genomic sequence; The ex vivo method, wherein the Cas9 nuclease edits the target genomic sequence.

13. 1. A Streptococcus pyogenes gRNA molecule comprising: a complementary region consisting of 17 to 18 nucleotides at the 5' end of the gRNA molecule that is complementary to 17 to 18 consecutive nucleotides of the complementary strand of a target genomic sequence; the target genomic sequence is adjacent to the 5' side of the protospacer adjacent motif; the gRNA molecule is a single gRNA or a CRISPR RNA (crRNA); In the presence of S. pyogenes Cas9 nuclease, the gRNA complementary region binds to the target genomic sequence and guides the Cas9 nuclease to the target genomic sequence; A Streptococcus pyogenes gRNA molecule, wherein the Cas9 nuclease edits the target genomic sequence.

14. 14. The gRNA molecule of claim 13, wherein the complementary region of the gRNA molecule consists of 17 nucleotides.

15. 14. The gRNA molecule of claim 13, comprising a ribonucleic acid consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:2404, SEQ ID NO:2407, or SEQ ID NO:2408.

16. 14. The gRNA molecule of Claim 13, comprising a ribonucleic acid comprising one or more uracils (U) at the 3' end of the molecule.

17. 14. The gRNA molecule of Claim 13, comprising a ribonucleic acid at the 5' end of the RNA molecule comprising one or more additional nucleotides that are not complementary to the target genomic sequence.

18. 14. The gRNA molecule of Claim 13, comprising a ribonucleic acid at the 5' end of the RNA molecule comprising one, two or three additional nucleotides that are not complementary to the target genomic sequence.

19. 14. The gRNA molecule of Claim 13, wherein the region of complementarity is complementary to 17 contiguous nucleotides of the complementary strand of a target genomic sequence.

20. 14. The gRNA molecule of Claim 13, wherein the region of complementarity is complementary to 18 contiguous nucleotides of the complementary strand of a target genomic sequence.

21. A DNA molecule encoding the gRNA molecule of claim 13.

22. A vector comprising the DNA molecule of claim 21.

23. A host cell expressing the vector of claim 22.

24. The host cell of claim 23 , wherein the cell is a eukaryotic cell.

25. The host cell of claim 24 , wherein the cell is a mammalian cell.

26. 14. The gRNA molecule of claim 13, wherein the complementary region of the gRNA molecule consists of 18 nucleotides.

27. 14. The gRNA molecule of Claim 13, wherein the gRNA molecule retains the ability to form a complex with a Cas9 nuclease or a catalytically inactivated Cas9 (dCas9) nuclease.

28. A complex, Streptococcus pyogenes Cas9 nuclease; and a Streptococcus pyogenes gRNA molecule comprising a complementary region at the 5' end of the gRNA molecule consisting of 17-18 nucleotides that are complementary to 17-18 consecutive nucleotides of the complementary strand of a target genomic sequence; Including, the target genomic sequence is adjacent to the 5' side of the protospacer adjacent motif; the gRNA molecule is a single gRNA or a CRISPR RNA (crRNA); a complex, wherein in the presence of S. pyogenes Cas9 nuclease, the gRNA complementary region binds to the target genomic sequence and guides the Cas9 nuclease to the target genomic sequence, and the Cas9 nuclease edits the target genomic sequence.

29. 29. The complex of claim 28, wherein the Cas9 nuclease is a dCas9 nuclease.

30. 29. The complex of Claim 28, wherein the complementary region of the gRNA molecule consists of 17 nucleotides.

31. 29. The complex of Claim 28, wherein the gRNA molecule comprises a ribonucleic acid consisting of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:2404, SEQ ID NO:2407, or SEQ ID NO:2408.

32. 29. The complex of Claim 28, wherein the gRNA molecule comprises one or more uracils (U) at the 3' end of the molecule.

33. 29. The complex of Claim 28, wherein the gRNA molecule comprises one or more additional nucleotides at the 5' end of the gRNA molecule that are not complementary to the target genomic sequence.

34. 29. The complex of Claim 28, wherein the gRNA molecule comprises one, two, or three additional nucleotides at the 5' end of the RNA molecule that are not complementary to the target genomic sequence.

35. 29. The complex of claim 28, wherein the region of complementarity is complementary to 17 consecutive nucleotides of the complementary strand of the target genomic sequence.

36. 29. The complex of claim 28, wherein the region of complementarity is complementary to 18 consecutive nucleotides of the complementary strand of the target genomic sequence.

37. 29. The complex of Claim 28, wherein the complementary region of the gRNA molecule consists of 18 nucleotides.

38. A DNA molecule encoding the complex of claim 28.

39. A vector comprising the DNA molecule of claim 38.

40. A host cell expressing the vector of claim 39.

41. 41. The host cell of claim 40, wherein the cell is a eukaryotic cell.

42. 42. The host cell of claim 41, wherein the cell is a mammalian cell.

Citation Information

Patent Citations

  • artificial transcription factor

    JP2006513694A

  • Methods and compositions for RNA-dependent targeted DNA modification and RNA-dependent transcriptional regulation.

    JP2015523856A

  • Methods of transcription activator like effector assembly

    WO2013012674A1

  • Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription

    WO2013176772A1

  • Crispr-based genome modification and regulation

    WO2014089290A1