Genetically engineered CRISPR-Cas9 nucleases
Mutating Cas9 proteins with specific amino acid changes enhances their specificity, reducing off-target effects and ensuring precise genome editing.
Patent Information
- Application Number
- JP2023125901
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-02-04
- Filing Date
- 2023-08-02
- Publication Date
- 2025-09-01
- Estimated Expiration
- 2036-08-26
AI Technical Summary
Existing CRISPR-Cas9 nucleases exhibit low specificity, leading to numerous off-target effects at imperfectly matched or mismatched DNA sites, which can cause catastrophic effects.
Engineering the Cas9 protein with specific mutations, such as N497A, R661A, Q695A, and Q926A, to reduce its binding affinity and enhance target specificity, thereby minimizing off-target effects.
The mutated Cas9 proteins demonstrate significantly reduced off-target activity, maintaining high on-target efficiency while minimizing unintended genomic alterations.
Smart Images

Figure 0007731943000029 
Figure 0007731943000030 
Figure 0007731943000031
Abstract
Description
[Technical Field]
[0001] Priority claims This application is filed on August 28, 2015 under 35 U.S.C. § 119(e). Patent Application No. 62 / 211,553; filed September 9, 2015 No. 216,033; No. 62 / 258,28 filed November 20, 2015 No. 0; No. 62 / 271,938 filed December 28, 2015; and U.S. Patent Application Publication No. 15 / 015,947, filed February 4, 2016. Priority is claimed to the following document, the entire contents of which are incorporated herein by reference: .
[0002] Sequence Listing This application has been filed electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy contains a sequence listing incorporated herein by reference on August 26, 2016. It is created and named SEQUENCE LISTING.txt and has a size of 129.9 It is 55 bytes.
[0003] Federally funded research or development This invention was made possible by a grant from the National Institutes of Health Government under grant numbers DP1GM105378 and R01GM088040. The Government has certain rights in this invention.
[0004] The present invention relates, at least in part, to genetically engineered proteins with altered and improved target specificity. Clustered Regularly Interspaced Short Palindromic Repeats terspaced Short Palindromic Repeats)(CRI SPR / CRISPR-associated protein 9 (Cas9) nuclease and genomic inheritance Reproduction, epigenomic gene manipulation, genome targeting, genome editing, and in vitro diagnostics Concerning its use in [Background technology]
[0005] CRISPR-Cas9 nucleases are efficient in a wide range of organisms and cell types enables efficient genome editing (Sander & Joung, Nat Biotech hnol 32,347-355(2014);Hsu et al.,Cell 15 7,1262-1278(2014);Doudna&Charpentier,Sci ence 346,1258096(2014);Barrangou&May,Exp ert Opin Biol Ther 15,311-314(2015)). Cas Target site recognition by 9 is achieved by the use of a chimeric single molecule encoding a sequence complementary to the target protospacer. It is programmed by guide RNA (sgRNA) (Jinek et al. ,Science 337,816-821(2012)), also requires recognition of short adjacent PAMs. (Mojica et al., Microbiology 155, 733-7 40(2009);Shah et al.,RNA Biol 10,891-899 (2013);Jiang et al.,Nat Biotechnol 31,23 3-239(2013);Jinek et al.,Science 337,816 -821(2012);Sternberg et al.,Nature 507,6 2-67(2014)). Summary of the Invention [Means for solving the problem]
[0006] As described herein, the Cas9 protein theoretically acts as a Ca2+ receptor for DNA. Engineering s9 to exhibit increased specificity by reducing its binding affinity Therefore, it has increased specificity compared to the wild-type protein (i.e. i.e., substantially fewer off-targets at imperfectly matched or mismatched DNA sites. Numerous Cas9 variants (which induce a catastrophic effect) and methods of using them are described herein. It is described in.
[0007] In a first aspect, the present invention provides a method for the preparation of a medicament for the treatment of a medicament comprising administering to a patient a medicament for the treatment of ... 1, Q695, Q926, and / or D1135E 1, 2, 3, 4, 5, 6, or has mutations at all seven positions, e.g., L169A, Y450, N4 97, R661, Q695, Q926, D1135E 1, 2, 3, 4, 5, 6 or 7 A sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1 with a mutation in one of the and optionally a nuclear localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag. Isolates of Streptococcus pyogenes containing one or more of the following: s) Cas9 (SpCas9) protein. The mutation replaces the amino acid with a natural amino acid. (e.g., 497 is anything other than N). In a preferred embodiment, the mutation changes an amino acid to something other than the natural one, arginine or lysine. to any amino acid in the group consisting of:
[0008] In some embodiments, the variant SpCas9 protein comprises one of the following: N497, R Mutations in one, two, three, or all four of 661, Q695, and Q926, e.g. For example, one of the following mutations: N497A, R661A, Q695A, and Q926A; Includes 2, 3, or all four.
[0009] In some embodiments, the variant SpCas9 protein comprises Q695 and / or or Q926 and optionally L169, Y450, N497, R661, and D Mutations in one, two, three, four, or all five of 1135E, e.g., Although not specified, Y450A / Q695A, L169A / Q695A, Q695 A / Q926A, Q695A / D1135E, Q926A / D1135E, Y450A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q92 6A, Y450A / Q695A / Q926A, R661A / Q695A / Q926A, N 497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y450 A / Q926A / D1135E, Q695A / Q926A / D1135E, L169A / Y450A / Q695A / Q926A, L169A / R661A / Q695A / Q926 A, Y450A / R661A / Q695A / Q926A, N497A / Q695A / Q9 26A / D1135E, R661A / Q695A / Q926A / D1135E, and Y Includes 450A / Q695A / Q926A / D1135E.
[0010] In some embodiments, the variant SpCas9 protein has the structure N14;S15;S 55;R63;R78;H160;K163;R165;L169;R403;N407 ;Y450;M495;N497;K510;Y515;W659;R661;M694 ;Q695;H698;A728;S730;K775;S777;R778;R780 ;K782;R783;K789;K797;Q805;N808;K810;R832 ;Q844;S845;K848;S851;K855;R859;K862;K890 ;Q920;Q926;K961;S964;K968;K974;R976;N980 ;H982;K1003;K1014;S1040;N1041;N1044;K104 7;K1059;R1060;K1107;E1108;S1109;K1113;R1 114;S1116;K1118;D1135;S1136;K1153;K1155; K1158;K1200;Q1221;H1241;Q1254;Q1256;K128 9;K1296;K1297;R1298;K1300;H1311;K1325;K1 334; containing mutations at T1337 and / or S1216.
[0011] In some embodiments, the variant SpCas9 protein comprises the following mutation: N 14A;S15A;S55A;R63A;R78A;R165A;R403A;N407 A;N497A;Y450A;K510A;Y515A;R661A;Q695A;S7 30A;K775A;S777A;R778A;R780A;K782A;R783A; K789A;K797A;Q805A;N808A;K810A;R832A;Q844 A;S845A;K848A;S851A;K855A;R859A;K862A;K8 90A;Q920A;Q926A;K961A;S964A;K968A;K974A; R976A;N980A;H982A;K1003A;K1014A;S1040A;N 1041A;N1044A;K1047A;K1059A;R1060A;K1107A ;E1108A;S1109A;K1113A;R1114A;S1116A;K111 8A;D1135A;S1136A;K1153A;K1155A;K1158A;K1 200A;Q1221A;H1241A;Q1254A;Q1256A;K1289A; K1296A;K1297A;R1298A;K1300A;H1311A;K1325 Also contains one or more of A;K1334A;T1337A and / or S1216A. In an embodiment, the variant protein is HF1(N497A / R661A / Q69 5A / Q926A)+K810A, HF1+K848A, HF1+K855A, HF1+ H982A, HF1+K848A / K1003A, HF1+K848A / R1060A, HF1+K855A / K1003A, HF1+K855A / R1060A, HF1+H9 82A / K1003A, HF1+H982A / R1060A, HF1+K1003A / R 1060A, HF1+K810A / K1003A / R1060A, HF1+K848A / In some embodiments, the variant protein comprises K1003A / R1060A. , HF1+K848A / K1003A, HF1+K848A / R1060A, HF1+K 855A / K1003A, HF1+K855A / R1060A, HF1+K1003A / Includes R1060A, HF1+K848A / K1003A / R1060A. In this state, the variant protein is Q695A / Q926A / R780A, Q695 A / Q926A / R976A, Q695A / Q926A / H982A, Q695A / Q9 26A / K855A, Q695A / Q926A / K848A / K1003A, Q695A / Q926A / K848A / K855A, Q695A / Q926A / K848A / H98 2A, Q695A / Q926A / K1003A / R1060A, Q695A / Q926A / K848A / R1060A, Q695A / Q926A / K855A / H982A, Q6 95A / Q926A / K855A / K1003A, Q695A / Q926A / K855A / R1060A, Q695A / Q926A / H982A / K1003A, Q695A / Q 926A / H982A / R1060A, Q695A / Q926A / K1003A / R10 60A, Q695A / Q926A / K810A / K1003A / R1060A, Q695 A / Q926A / K848A / K1003A / R1060A. The variants are N497A / R661A / Q695A / Q926A / K810A, N497A / R661A / Q695A / Q926A / K848A, N497A / R661 A / Q695A / Q926A / K855A, N497A / R661A / Q695A / Q9 26A / R780A, N497A / R661A / Q695A / Q926A / K968A, N497A / R661A / Q695A / Q926A / H982A, N497A / R661 A / Q695A / Q926A / K1003A, N497A / R661A / Q695A / Q 926A / K1014A, N497A / R661A / Q695A / Q926A / K104 7A, N497A / R661A / Q695A / Q926A / R1060A, N497A / <h2 style=";text-align:left;direction:ltr">R661A / Q695A / Q926A / K810A / K968A、N497A / R661<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / Q695A / Q926A / K810A / K848A、N497A / R661A / Q6<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 95A / Q926A / K810A / K1003A、N497A / R661A / Q695A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / Q926A / K810A / R1060A、N497A / R661A / Q695A / Q9<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 26A / K848A / K1003A、N497A / R661A / Q695A / Q926A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / K848A / R1060A、N497A / R661A / Q695A / Q926A / K8<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 55A / K1003A, N497A / R661A / Q695A / Q926A / K855A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / R1060A、N497A / R661A / Q695A / Q926A / K968A / K1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 003A, N497A / R661A / Q695A / Q926A / H982A / K1003<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A、N497A / R661A / Q695A / Q926A / H982A / R1060A、N<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 497A / R661A / Q695A / Q926A / K1003A / R1060A、N49<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 7A / R661A / Q695A / Q926A / K810A / K1003A / R1060A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 、N497A / R661A / Q695A / Q926A / K848A / K1003A / R1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 060A、Q695A / Q926A / R780A、Q695A / Q926A / K810A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q695A / Q926A / R832A, Q695A / Q926A / K848A, Q69<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5A / Q926A / K855A、Q695A / Q926A / K968A、Q695A / Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 926A / R976A, Q695A / Q926A / H982A, Q695A / Q926A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / K1003A、Q695A / Q926A / K1014A、Q695A / Q926A / K<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1047A、Q695A / Q926A / R1060A、Q695A / Q926A / K84 8A / K968A, Q695A / Q926A / R976A, Q695A / Q926A / H 982A, Q695A / Q926A / K855A, Q695A / Q926A / K848A / K1003A, Q695A / Q926A / K848A / K855A, Q695A / Q9 26A / K848A / H982A, Q695A / Q926A / K1003A / R1060 A, Q695A / Q926A / R832A / R1060A, Q695A / Q926A / K 968A / K1003A, Q695A / Q926A / K968A / R1060A, Q69 5A / Q926A / K848A / R1060A, Q695A / Q926A / K855A / H982A, Q695A / Q926A / K855A / K1003A, Q695A / Q92 6A / K855A / R1060A, Q695A / Q926A / H982A / K1003A , Q695A / Q926A / H982A / R1060A, Q695A / Q926A / K1 003A / R1060A, Q695A / Q926A / K810A / K1003A / R10 60A, Q695A / Q926A / K1003A / K1047A / R1060A, Q69 5A / Q926A / K968A / K1003A / R1060A, Q695A / Q926A / R832A / K1003A / R1060A, or Q695A / Q926A / K848 Includes A / K1003A / R1060A.
[0012] Mutations of amino acids other than alanine are also included and may be made in the present methods and compositions. and can be used.
[0013] In some embodiments, the variant SpCas9 protein comprises the following additional mutations: Different: R63A, R66A, R69A, R70A, R71A, Y72A, R74A, R75 A, K76A, N77A, R78A, R115A, H160A, K163A, R165A , L169A, R403A, T404A, F405A, N407A, R447A, N49 7A, I448A, Y450A, S460A, M495A, K510A, Y515A, R 661A, M694A, Q695A, H698A, Y1013A, V1015A, R11 22A, K1123A, K1124A, K1158A, K1185A, K1200A, S 1216A, Q1221A, K1289A, R1298A, K1300A, K1325A , R1333A, K1334A, R1335A, and T1337A.
[0014] In some embodiments, the variant SpCas9 protein comprises multiple substitution mutations. :N497 / R661 / Q695 / Q926 (quadruple variant mutant);Q695 / Q926 (double mutant); R661 / Q695 / Q926 and N497 / Q695 In some embodiments, L169, Y450 and Q926 (triple mutant). and / or additional substitution mutations at D1135 in their double, triple, and quadruple The mutant can be added or carry a substitution at Q695 or Q926 In some embodiments, the mutation can be in addition to a single mutation that In some embodiments, the mutant has an alanine in place of the wild-type amino acid. It has any amino acid other than arginine or lysine (or any naturally occurring amino acid).
[0015] In some embodiments, the variant SpCas9 protein is D10, E762, mutations at D839, H983, or D986; and H840 or N863 The nuclease activity of the nuclease may also be reduced by one or more mutations selected from the group consisting of: In some embodiments, the mutations are (i) D10A or D10N, and (ii) H840A, H840N, or H840Y.
[0016] In some embodiments, the SpCas9 variant comprises the following set of mutations: D1 135V / R1335Q / T1337R (VQR variant); D1135E / R133 5Q / T1337R (EQR variant); D1135V / G1218R / R1335Q / T1337R (VRQR variant); or D1135V / G1218R / R133 It may also contain one of the 5E / T1337R (VRER variants).
[0017] As used herein, the following positions: Y211, Y212, W229, Y230, R245 , T392, N419, Y651, or R654 have mutations in the following positions: Y211, Y212, W229, Y230, 1, 2, 3, or 4 of R245, T392, N419, Y651, or R654, or 2 with 5 or 6 mutations and at least 80% sequences that are identical and optionally a nuclear localization sequence, a cell penetrating peptide sequence, and / or or affinity tags. Also provided are S. aureus Cas9 (SaCas9) proteins. In this regard, the SaCas9 variants described herein have the following positions: Y211, Y212, W229, Y230, R245, T392, N419, Y651 and / or R654 The amino acid sequence of SEQ ID NO: 2 having mutations in one, two, three, four, five, six or more of In some embodiments, the variant comprises the following mutations: Y211A, Y212 A, W229, Y230A, R245A, T392A, N419A, Y651, and / or containing one or more of R654A.
[0018] In some embodiments, the variant SaCas9 protein comprises N419 and / or or R654 and optionally additional mutations Y211, Y212 , W229, Y230, R245 and T392, preferably one, two, three, four or more of N419A / R654A, Y211A / R654A, Y211A / Y212A, Y211 A / Y230A, Y211A / R245A, Y212A / Y230A, Y212A / R2 45A, Y230A / R245A, W229A / R654A, Y211A / Y212A / Y230A, Y211A / Y212A / R245A, Y211A / Y212A / Y651 A, Y211A / Y230A / R245A, Y211A / Y230A / Y651A, Y2 11A / R245A / Y651A, Y211A / R245A / R654A, Y211A / R245A / N419A, Y211A / N419A / R654A, Y212A / Y230 A / R245A, Y212A / Y230A / Y651A, Y212A / R245A / Y6 51A, Y230A / R245A / Y651A, R245A / N419A / R654A, T392A / N419A / R654A, R245A / T392A / N419A / R654 A, Y211A / R245A / N419A / R654A, W229A / R245A / N4 19A / R654A, Y211A / R245A / T392A / N419A / R654A, or containing Y211A / W229A / R245A / N419A / R654A.
[0019] In some embodiments, the variant SaCas9 protein has the amino acid sequence Y211;Y212 ;W229;Y230;R245;T392;N419;L446;Q488;N492 ;Q495;R497;N498;R499;Q500;K518;K523;K525 ;H557;R561;K572;R634;Y651;R654;G655;N658 ;S662;N667;R686;K692;R694;H700;K751;D786 ;T787;Y789;T882;K886;N888;889;L909;N985; N986;R991;R1015;N44;R45;R51;R55;R59;R60; R116;R165;N169;R208;R209;Y211;T238;Y239; K248;Y256;R314;N394;Q414;K57;R61;H111;K1 14;V164;R165;L788;S790;R792;N804;Y868;K8 70;K878;K879;K881;Y897;R901;and / or K906 This includes mutations in
[0020] In some embodiments, the variant SaCas9 protein comprises the following mutation: Y 211A;Y212A;W229A;Y230A;R245A;T392A;N419A ;L446A;Q488A;N492A;Q495A;R497A;N498A;R49 9A;Q500A;K518A;K523A;K525A;H557A;R561A;K 572A;R634A;Y651A;R654A;G655A;N658A;S662A ;N667A;R686A;K692A;R694A;H700A;K751A;D78 6A;T787A;Y789A;T882A;K886A;N888A;A889A;L 909A;N985A;N986A;R991A;R1015A;N44A;R45A; R51A;R55A;R59A;R60A;R116A;R165A;N169A;R2 08A;R209A;T238A;Y239A;K248A;Y256A;R314A; N394A;Q414A;K57A;R61A;H111A;K114A;V164A; R165A;L788A;S790A;R792A;N804A;Y868A;K870 One of A;K878A;K879A;K881A;Y897A;R901A;K906A Including the above.
[0021] In some embodiments, the variant SaCas9 protein contains the following additional mutations: Different: Y211A, W229A, Y230A, R245A, T392A, N419A, L4 46A, Y651A, R654A, D786A, T787A, Y789A, T882A, K886A, N888A, A889A, L909A, N985A, N986A, R991 A, R1015A, N44A, R45A, R51A, R55A, R59A, R60A, R 116A, R165A, N169A, R208A, R209A, T238A, Y239A , K248A, Y256A, R314A, N394A, Q414A, K57A, R61A , H111A, K114A, V164A, R165A, L788A, S790A, R79 2A, N804A, Y868A, K870A, K878A, K879A, K881A, Y Contains one or more of 897A, R901A, K906A.
[0022] In some embodiments, the variant SaCas9 protein comprises multiple substitution mutations. :R245 / T392 / N419 / R654 and Y221 / R245 / N419 / R6 54 (quadruple variant mutant); N419 / R654, R245 / R654, Y22 1 / R654, and Y221 / N419 (double mutant); R245 / N419 / R 654, Y211 / N419 / R654, and T392 / N419 / R654 (triple thrust) In some embodiments, mutants include a nucleotide sequence that replaces the wild-type amino acid. Contains ranin.
[0023] In some embodiments, the variant SaCas9 protein is selected from the group consisting of D10, E477, Mutations at D556, H701, or D704; and H557 or N580 The nuclease activity of the nuclease may also be reduced by one or more mutations selected from the group consisting of: In some embodiments, the mutations are: (i) D10A or D10N; (ii) H5 57A, H557N, or H557Y, (iii) N580A, and / or (i v) D556A.
[0024] In some embodiments, the variant SaCas9 protein comprises the following mutation: E Contains one or more of 782K, K929R, N968K, or R1015H. Specifically: , E782K / N968K / R1015H(KKH variant); E782K / K929 R / R1015H (KRH variant); or E782K / K929R / N968K / R1015H (KRKH variant).
[0025] In some embodiments, variant Cas9 proteins are used to increase specificity. It contains mutations in one or more of the following regions:
[0026] [Table 1]
[0027] As used herein, the present invention relates to a method for producing a heterologous functional domain fused to a heterologous functional domain by an optional intervening linker. A fusion protein comprising an isolated variant Cas9 protein as described herein, Fusion proteins are also provided in which the Car does not interfere with the activity of the fusion protein. In the present invention, the heterologous functional domain is a functional domain that binds to DNA or protein, e.g., chromatin. In some embodiments, the heterologous functional domain is a transcription activation domain. In some embodiments, the transcription activation domain is from VP64 or NF-κB p65. In some embodiments, the heterologous functional domain is a transcriptional silencer or a transcription factor. In some embodiments, the transcriptional repression domain is a Krüppel-associated domain. KRAB domain, ERF repressor domain (ERD), or mSin In some embodiments, the transcriptional silencer is a 3A interaction domain (SID). Heterochromatin protein 1 (HP1), for example, HP1α or HP1β. In some embodiments, the heterologous functional domain is an enzyme that modifies the methylation status of DNA. In some embodiments, the enzyme that modifies the methylation state of DNA is a DNA methyltransferase. transferase (DNMT) or TET protein whole or dioxygenase domains, e.g., a cysteine-rich extension and seven highly conserved exons. The 2OGFeDO domain encoded by, for example, T et1 catalytic domain, Tet2 including amino acids 1290–1905 and amino acid 966 In some embodiments, the catalytic module comprises Tet3 comprising TE3-1678. The T protein or TET-derived dioxygenase domain is from TET1 In some embodiments, the heterologous functional domain is an enzyme that modifies a histone subunit. In some embodiments, the enzyme that modifies the histone subunit is a histone a Cetyltransferase (HAT), histone deacetylase (HDAC), histone methyltransferases (HMTs), or histone demethylases. In some embodiments, the heterologous functional domain is a biological tether. The biological tether is MS2, Csy4, or lambda N protein. In this case, the heterologous functional domain is FokI.
[0028] As used herein, nucleic acids encoding the variant Cas9 proteins described herein , isolated nucleic acids, and optionally expression of variant Cas9 proteins described herein. a vector comprising the isolated nucleic acid operably linked to one or more regulatory domains for Also provided herein are nucleic acids comprising the nucleic acids described herein and optionally the Host cells expressing the variant Cas9 proteins described herein, such as bacteria, yeast, Insect or mammalian host cells or transgenic animals (e.g., mice) are also provided. do.
[0029] Provided herein are isolated nucleic acids encoding Cas9 variants and optionally their and its isolation, operably linked to one or more regulatory domains for expression of the variant. A vector comprising the nucleic acid and a protein comprising the nucleic acid and optionally a variant thereof Host cells, eg, mammalian host cells, that express the polypeptide are also provided.
[0030] As used herein, a variant Cas9 protein or a variant Cas9 protein described herein is used in a cell. fusion protein and cells with optimal nucleotide spacing at the genomic target site expressing at least one guide RNA having a region complementary to a selected portion of the genome of by contacting cells with them, Also provided are methods for altering the Cas9 gene expression in a cell, e.g., by transforming a cell with a Cas9 gene in a single vector. contacting the cells with nucleic acids encoding the protein and guide RNA; The nucleic acid encoding the Cas9 protein and the nucleic acid encoding the guide RNA in the vector and contacting the cells with purified Cas9 protein and synthetic or purified gR. In some embodiments, the cells are contacted with a complex of g Stably express RNA or variant / fusion proteins or both. For example, the cells may be transfected or introduced with other elements as described herein. The method comprises: stably expressing the variant or fusion protein of the present invention; contacting the cells with synthetic gRNA, purified recombinantly produced gRNA, or nucleic acid encoding gRNA In some embodiments, the variant protein or fusion protein may comprise , a nuclear localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag.
[0031] As used herein, dsDNA refers to purified variant proteins or a fusion protein and a guide RNA having a region complementary to a selected portion of a dsDNA molecule; Altering isolated dsDNA molecules in vitro by contacting, e.g., selectively A method for modifying is also provided.
[0032] Unless otherwise defined, all technical and scientific terms used herein are defined by the present invention. The term "method" has the same meaning as commonly understood by a person skilled in the art to which the present invention pertains. and materials described herein for use in the present invention and known in the art. Other suitable methods and materials may also be used. Materials, methods, and examples are provided in the accompanying drawings. All publications cited herein are for illustrative purposes only and are not intended to be limiting. Patent applications, patents, sequences, database entries, and other references are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control.
[0033] Other features and advantages of the invention will become apparent from the following detailed description and drawings, as well as the claims. This is clear from the range. [Brief explanation of the drawings]
[0034] [Figure 1-1] Figure 1A: Identification and characterization of SpCas9 variants carrying mutations in residues that form nonspecific DNA contacts. Schematic showing wild-type SpCas9 recognition of the target DNA:sgRNA duplex based on PDB 4OOG and 4UN3 (sources, refs. 31 and 32, respectively). [Figure 1-2]Figure 1B: Identification and characterization of SpCas9 variants bearing mutations in residues that form nonspecific DNA contacts. Characterization of SpCas9 variants containing alanine substitutions at positions that form hydrogen bonds to the DNA backbone. Wild-type SpCas9 and variants were evaluated using the EGFP decay assay in human cells programmed with a perfect match sgRNA or four other sgRNAs encoding mismatches to the target site. Error bars represent s.e.m. for n=3; the average level of background EGFP loss is represented by the red dotted line (for this panel and panel C). Figure 1C: Identification and characterization of SpCas9 variants bearing mutations in residues that form nonspecific DNA contacts. On-target activity of wild-type SpCas9 and SpCas9-HF1 across 24 sites assessed by EGFP decay assay. Error bars represent s.e.m. for n=3. [Figure 1-3] Figure 1D: Identification and characterization of SpCas9 variants bearing mutations in residues that form nonspecific DNA contacts. On-target activity of wild-type SpCas9 and SpCas9-HF1 across 13 endogenous sites assessed by the T7E1 assay. Error bars represent sem for n=3. Figure 1E: Identification and characterization of SpCas9 variants bearing mutations in residues that form nonspecific DNA contacts. Ratio of the on-target activity of SpCas9-HF1 to that of wild-type SpCas9 (from panels C and D). [Figure 2-1] Figure 2A: Genome-wide specificity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs for standard target sites. Off-target sites for wild-type SpCas9 and SpCas9-HF1 using eight sgRNAs targeted to endogenous human genes, as determined by GUIDE-seq. Read counts represent a measure of cleavage frequency at a given site; mismatch positions within the spacer or PAM are highlighted by color. [Figure 2-2]Figure 2B: Genome-wide specificity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs for standard target sites. Summary of the total number of genome-wide off-target sites identified by GUIDE-seq for wild-type SpCas9 and SpCas9-HF1 from the eight sgRNAs used in panel A. Figure 2C: Genome-wide specificity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs for standard target sites. Identified off-target sites for wild-type SpCas9 and SpCas9-HF1 for the eight sgRNAs, binned according to the total number of mismatches (within the protospacer and PAM) relative to on-target activity. [Figure 3-1] Figure 3A: Validation of improved SpCas9-HF1 specificity by targeted deep sequencing of off-target sites identified by GUIDE-seq. Average percent on-target modification determined by deep sequencing for wild-type SpCas9 and SpCas9-HF1 using the six sgRNAs from Figure 2. Error bars represent s.e.m. for n=3. [Figure 3-2]Figure 3B: Validation of improved SpCas9-HF1 specificity by targeted deep sequencing of off-target sites identified by GUIDE-seq. Percentage of deep-sequenced on-target sites containing indel mutations and GUIDE-seq-detected off-target sites. Triplicate experiments are plotted for wild-type SpCas9, SpCas9-HF1, and control conditions. Black circles below the x-axis represent replicates in which no insertion or deletion mutations were observed. Off-target sites that could not be amplified by PCR are shown in red text with an asterisk. Hypothesis testing using a one-sided Fisher's exact test with pooled read counts found significant differences between SpCas9-HF1 and the control condition only for EMX1-1 and FANCF-3 off-targets (p<0.05 after adjusting for multiple comparisons using the Benjamini-Hochberg method). Significant differences were also found between wild-type SpCas9 and SpCas9-HF1 at all off-target sites, and between wild-type SpCas9 and the control condition at all off-target sites except RUNX1-1 off-target 2. [Figure 3-3] (C) Validation of improved SpCas9-HF1 specificity by targeted deep sequencing of off-target sites identified by GUIDE-seq. Scatter plot of the correlation between GUIDE-seq read counts (from Figure 2A) and the mean percent modification determined by deep sequencing at on- and off-target cleavage sites with wild-type SpCas9. [Figure 4-1]Figure 4A: Genome-wide specificity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs for non-canonical repeat sites. GUIDE-seq specificity profiles of wild-type SpCas9 and SpCas9-HF1 using two sgRNAs known to cleave multiple off-target sites (Fu et al., Nat Biotechnol 31, 822-826 (2013); Tsai et al., Nat Biotechnol 33, 187-197 (2015)). GUIDE-seq read counts represent a measure of cleavage efficiency at a given site; mismatch positions within the spacer or PAM are highlighted by color; red circles indicate sites at the sgRNA-DNA interface that are likely to have the indicated bulge (Lin et al., Nucleic Acids Res 42, 7473-7485 (2014)); and blue circles indicate sites that may have alternative gapped alignments to those shown (see Figure 8). [Figure 4-2] Figure 4B: Genome-wide specificity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs for non-canonical repeat sites. Summary of the total number of genome-wide off-target sites identified by GUIDE-seq for wild-type SpCas9 and SpCas9-HF1 from the two sgRNAs used in panel A. Figure 4C: Genome-wide specificity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs for non-canonical repeat sites. Off-target sites identified using wild-type SpCas9 or SpCas9-HF1 for VEGFA sites 2 and 3, binned according to the total number of mismatches (within the protospacer and PAM) relative to the on-target site. Off-target sites marked with red circles in panel A are not included in these counts; sites marked with blue circles in panel A are counted along with the number of mismatches in the ungapped alignment. [Figure 5-1]Figure 5A: Activity of SpCas9-HF1 derivatives carrying additional substitutions. Human cell EGFP-disruption activity of wild-type SpCas9, SpCas9-HF1, and SpCas9-HF1 derivative variants using eight sgRNAs. SpCas9-HF1 carries N497A, R661A, Q695, and Q926A mutations; HF2 = HF1 + D1135E; HF3 = HF1 + L169A; HF4 = HF1 + Y450A. Error bars represent s.e.m. for n = 3; the average level of background EGFP loss is represented by the red dotted line. Figure 5B: Activity of SpCas9-HF1 derivatives carrying additional substitutions. Summary of on-target activity using SpCas9-HF1 variants compared to wild-type SpCas9 using the eight sgRNAs from panel a. Median and interquartile ranges are shown; intervals showing >70% of wild-type activity are highlighted in green. Figure 5C: Activity of SpCas9-HF1 derivatives carrying additional substitutions. Average percent modification by SpCas9 and HF variants at the FANCF site 2 and VEGFA site 3 on-target sites, as well as off-target sites from Figures 2A and 4A that were resistant to the effects of SpCas9-HF1. Percent modification was determined by T7E1 assay; background indel rates were subtracted for all experiments. Error bars represent s.e.m. for n=3. Figure 5D: Activity of SpCas9-HF1 derivatives carrying additional substitutions. Specificity ratios of wild-type SpCas9 and HF variants using FANCF site 2 or VEGFA site 3 sgRNAs, plotted as the ratio of on-target activity to off-target activity (from panel C). [Figure 5-2]Figure 5E: Genome-wide specificity of SpCas9-HF1, -HF2, and -HF4 using sgRNAs with off-target sites resistant to the effects of SpCas9-HF1. Average GUIDE-seq tag integration at the intended on-target site for the GUIDE-seq experiment in panel F. SpCas9-HF1 = N497A / R661A / Q695A / Q926A; HF2 = HF1 + D1135E; HF4 = HF1 + Y450A. Error bars represent s.e.m. for n = 3. [Figure 5-3] Figure 5F: Genome-wide specificity of SpCas9-HF1, -HF2, and -HF4 using sgRNAs with off-target sites resistant to the effects of SpCas9-HF1. GUIDE-seq-identified off-target sites for SpCas9-HF1, -HF2, or -HF4 using either FANCF site 2 or VEGFA site 3 sgRNA. Read counts represent a measure of cleavage frequency at a given site; mismatch positions within the spacer or PAM are highlighted in color. Fold improvement in off-target discrimination was calculated by normalizing off-target read counts for SpCas9-HF variants to read counts at the on-target site before comparison between SpCas9-HF variants. [Figure 6-1] Figure 6A: SpCas9 interaction with sgRNA and target DNA. Schematic illustrating the SpCas9:sgRNA complex along with base pairing between the sgRNA and target DNA. [Figure 6-2] Figure 6B: SpCas9 interaction with sgRNA and target DNA. Structural representation of the SpCas9:sgRNA complex bound to target DNA from PDB:4UN3 (Reference 32). The four residues that form hydrogen-bonding contacts to the target strand DNA backbone are highlighted in blue; the HNH domain is hidden for visualization purposes. [Figure 7]Figure 7A: Comparison of on-target activity of wild-type and SpCas9-HF1 using various sgRNAs used in GUIDE-seq experiments. Average GUIDE-seq tag integration at the intended on-target site for the GUIDE-seq experiment shown in Figure 2A, as quantified by restriction fragment length polymorphism assay. Error bars represent sem for n=3. Figure 7B: Comparison of on-target activity of wild-type and SpCas9-HF1 using various sgRNAs used in GUIDE-seq experiments. Average percent modification at the intended on-target site for the GUIDE-seq experiment shown in Figure 2A, as detected by T7E1 assay. Error bars represent sem for n=3. Figure 7C: Comparison of on-target activity of wild-type and SpCas9-HF1 using various sgRNAs used in GUIDE-seq experiments. Average GUIDE-seq tag integration at the intended on-target site for the GUIDE-seq experiment shown in Figure 4A, as quantified by restriction fragment length polymorphism assay. Error bars represent sem for n=3. Figure 7D: Comparison of on-target activity of wild-type and SpCas9-HF1 with the various sgRNAs used in GUIDE-seq experiments. Average percent modification at the intended on-target site for the GUIDE-seq experiments shown in Figure 4A, as detected by the T7E1 assay. Error bars represent sem for n=3. [Figure 8] Potential alternative alignments for VEGFA site 2 off-target sites. Ten VEGFA site 2 off-target sites (Lin et al., Nucleic Acids Res 42, 7473-7485 (2014)) identified by GUIDE-seq (left) that could potentially be recognized as off-target sites containing single nucleotide gaps, aligned using Geneious (Kearse et al., Bioinformatics 28, 1647-1649 (2012)) version 8.1.6 (right). [Figure 9]Activity of wild-type SpCas9 and SpCas9-HF1 with truncated sgRNA14. EGFP-decay activity of wild-type SpCas9 and SpCas9-HF1 using full-length or truncated sgRNAs targeted to four sites in EGFP. Error bars represent sem for n=3; the average level of background EGFP loss in control experiments is represented by the red dotted line. [Figure 10] Figure 1 shows wild-type SpCas9 and SpCas9-HF1 activity using sgRNAs bearing a 5'-mismatched guanine base. Figure 2 shows EGFP decay activity of wild-type SpCas9 and SpCas9-HF1 using sgRNAs targeted to four different sites. For each target site, the sgRNA contains either a matching non-guanine 5'-base or an intentionally mismatched 5'-guanine. [Figure 11] Figure 1 shows titration of the amount of wild-type SpCas9 and SpCas9-HF1 expression plasmid. Human cell EGFP decay activity from transfection with varying amounts of wild-type and SpCas9-HF1 expression plasmids. For all transfections, the amount of sgRNA-containing plasmid was fixed at 250 ng. Two sgRNAs targeting distinct sites were used; error bars represent sem for n=3; the average level of background EGFP loss in the negative control is represented by the red dotted line. [Figure 12-1]Figure 12A: Altered PAM recognition specificity of SpCas9-HF1. Comparison of the mean percent modification of on-target endogenous human sites by SpCas9-VQR (reference 15) and the improved SpCas9-VRQR using eight sgRNAs, as quantified by the T7E1 assay. Both variants were engineered to recognize the NGAN PAM. Error bars represent s.e.m. for n=2 or 3. Figure 12B: Altered PAM recognition specificity of SpCas9-HF1. On-target EGFP decay activity of SpCas9-VQR and SpCas9-VRQR using eight sgRNAs compared to their -HF1 counterparts. Error bars represent s.e.m. for n=3; the mean level of background EGFP loss in the negative control is represented by the red dotted line. [Figure 12-2] Figure 12C: Altered PAM recognition specificity of SpCas9-HF1. Comparison of the mean percent on-target modification by SpCas9-VQR and SpCas9-VRQR compared to their -HF1 variants at eight endogenous human gene sites quantified by the T7E1 assay. Error bars represent sem for n=3; ND is not detected. Figure 12D: Altered PAM recognition specificity of SpCas9-HF1. Summary of fold changes in on-target activity using SpCas9-VQR or SpCas9-VRQR compared to their corresponding -HF1 variants (from panels B and C). Median and interquartile ranges are shown; intervals showing >70% of wild-type activity are highlighted in green. [Figure 13]Figure 13A: Activity of wild-type SpCas9, SpCas9-HF1, and wild-type SpCas9 derivatives bearing one or more alanine substitutions at positions that could potentially contact non-target DNA strands. Nuclease activity was assessed using an EGFP decay assay using sgRNAs that perfectly match sites in the EGFP gene and sgRNAs intentionally mismatched at positions 11 and 12. Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the red dotted line represents background levels of EGFP decay; HF1 = SpCas9 with N497A / R661A / Q695A / Q926A substitutions. Figure 13B: Activity of wild-type SpCas9, SpCas9-HF1, and wild-type SpCas9 derivatives bearing one or more alanine substitutions at positions that could potentially contact non-target DNA strands. Nuclease activity was assessed using an EGFP decay assay with an sgRNA that perfectly matches a site in the EGFP gene and an sgRNA intentionally mismatched at positions 9 and 10 (panel B). Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the red dotted line represents background levels of EGFP decay; HF1 = SpCas9 with N497A / R661A / Q695A / Q926A substitutions. [Figure 14]Figure 14A: Activity of wild-type SpCas9, SpCas9-HF1, and SpCas9-HF1 derivatives bearing one or more alanine substitutions at positions that could potentially contact non-target DNA strands. Nuclease activity was assessed using an EGFP decay assay using sgRNAs that perfectly match sites in the EGFP gene and sgRNAs intentionally mismatched at positions 11 and 12. Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the red dotted line represents the background level of EGFP decay; HF1 = SpCas9 with N497A / R661A / Q695A / Q926A substitutions. Figure 14B: Activity of wild-type SpCas9, SpCas9-HF1, and SpCas9-HF1 derivatives bearing one or more alanine substitutions at positions that could potentially contact non-target DNA strands. Nuclease activity was assessed using an EGFP decay assay with an sgRNA that perfectly matches a site in the EGFP gene and an sgRNA intentionally mismatched at positions 9 and 10 (panel B). Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the red dotted line represents background levels of EGFP decay; HF1 = SpCas9 with N497A / R661A / Q695A / Q926A substitutions. [Figure 15] Figure 1 shows the activity of wild-type SpCas9, SpCas9-HF1, and the SpCas9(Q695A / Q926A) derivative, which carries one or more alanine substitutions at positions that could potentially contact the non-target DNA strand. Nuclease activity was assessed using an EGFP decay assay using sgRNAs that perfectly match sites in the EGFP gene and sgRNAs intentionally mismatched at positions 11 and 12. Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the dotted red line represents the background level of EGFP decay; HF1 = SpCas9 with N497A / R661A / Q695A / Q926A substitutions; Dbl = SpCas9 with Q695A / Q926A substitutions. [Figure 16]Figure 1 shows the activity of wild-type SpCas9, SpCas9-HF1, and eSpCas9-1.1 using matched sgRNAs and sgRNAs with single mismatches at each position in the spacer. Nuclease activity was assessed using an EGFP decay assay with sgRNAs that perfectly match sites in the EGFP gene ("match") and sgRNAs intentionally mismatched at the positions indicated. Mismatch positions are numbered, with position 20 being the most distal position from the PAM. SpCas9-HF1 = N497A / R661A / Q695A / Q926A, and eSP1.1 = K848A / K1003A / R1060A. [Figure 17]Figure 17A: Activity of wild-type SpCas9 and variants using matched sgRNAs and sgRNAs with single mismatches at various positions in the spacer. The activity of SpCas9 nucleases containing combinations of alanine substitutions (directed at positions that could potentially contact the target or non-target DNA strand) was assessed using an EGFP decay assay with sgRNAs that perfectly match sites in the EGFP gene ("match") and sgRNAs intentionally mismatched at the indicated spacer positions. Mismatch positions are numbered, with position 20 being the position most distal to the PAM; mm = mismatch. WT = wild-type, Db = Q695A / Q926A, HF1 = N497A / R661A / Q695A / Q926A, 1.0 = K810A / K1003A / R1060A, and 1.1 = K848A / K1003A / R1060A. Figure 17B: Activity of wild-type SpCas9 and variants using matched sgRNAs and sgRNAs with single mismatches at various positions in the spacer. A subset of these nucleases from (a) was tested using all remaining possible single-mismatched sgRNAs for match-on-target sites. Mismatch positions are numbered, with position 20 being the most distal position from the PAM; mm = mismatch. WT = wild-type, Db = Q695A / Q926A, HF1 = N497A / R661A / Q695A / Q926A, 1.0 = K810A / K1003A / R1060A, and 1.1 = K848A / K1003A / R1060A. [Figure 18]Activity of wild-type SpCas9 and variants using matched sgRNAs and sgRNAs with mismatches at various individual positions in the spacer. The activity of SpCas9 nucleases containing combinations of alanine substitutions (directed at positions that could potentially contact the target or non-target DNA strand) was assessed using an EGFP decay assay with sgRNAs that perfectly match sites in the EGFP gene ("match") and sgRNAs intentionally mismatched at the positions indicated. Db = Q695A / Q926A, and HF1 = N497A / R661A / Q695A / Q926A. [Figure 19-1] Figure 19A: Activity of wild-type SpCas9 and variants using matched sgRNAs and sgRNAs with mismatches at various individual positions in the spacer. The on-target activity of SpCas9 nucleases containing combinations of alanine substitutions (targeting positions that could potentially contact the target or non-target DNA strand) was assessed using an EGFP decay assay with two sgRNAs that perfectly match sites in the EGFP gene. Db = Q695A / Q926A, and HF1 = N497A / R661A / Q695A / Q926A. [Figure 19-2] Figure 19B: Activity of wild-type SpCas9 and variants using matched sgRNAs and sgRNAs with mismatches at various individual positions in the spacer. A subset of these nucleases from (a) was tested with sgRNAs containing mismatches at positions 12, 14, 16, or 18 (of sgRNA "site 1") in the spacer sequence to determine whether these substitutions conferred mismatch intolerance. Db=Q695A / Q926A, and HF1=N497A / R661A / Q695A / Q926A. [Figure 20] Structural comparison of SpCas9 (top) and SaCas9 (bottom) illustrating the similarities between the positions of the mutations in the quadruple mutant constructs (shown as yellow globules). Other residues that contact the DNA backbone are also shown as pink globules. [Figure 21]Figure 21A: Activity of wild-type SaCas9 and SaCas9 derivatives bearing one or more alanine substitutions. SaCas9 substitutions were targeted to positions that could potentially contact the target DNA strand. Nuclease activity was assessed using an EGFP decay assay with sgRNAs that perfectly match sites in the EGFP gene and sgRNAs intentionally mismatched at positions 11 and 12. Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the red dotted line represents background levels of EGFP decay. Figure 21B: Activity of wild-type SaCas9 and SaCas9 derivatives bearing one or more alanine substitutions. SaCas9 substitutions were targeted to positions previously shown to affect PAM specificity. Nuclease activity was assessed using an EGFP decay assay with sgRNAs that perfectly match sites in the EGFP gene and sgRNAs intentionally mismatched at positions 11 and 12. Mismatch positions are numbered, with position 20 being the most distal position from the PAM; the red dotted line represents the background level of EGFP decay. [Figure 22-1] Figure 22A: Activity of wild-type (WT) SaCas9 and SaCas9 derivatives bearing one or more alanine substitutions at residues that could potentially contact the target DNA strand. Nuclease activity was assessed using an EGFP disruption assay with an sgRNA that perfectly matches a site in the EGFP gene ("match") and an sgRNA intentionally mismatched at positions 19 and 20. Mismatch positions are numbered, with position 20 being the most distal position from the PAM. [Figure 22-2] Figure 22B: Activity of wild-type (WT) SaCas9 and SaCas9 derivatives bearing one or more alanine substitutions at residues that could potentially contact the target DNA strand. Nuclease activity was assessed using an EGFP disruption assay with an sgRNA that perfectly matches a site in the EGFP gene ("match") and an sgRNA intentionally mismatched at positions 19 and 20. Mismatch positions are numbered, with position 20 being the most distal position from the PAM. [Figure 23]The activity of wild-type (WT) SaCas9 and SaCas9 variants carrying triple combinations of alanine substitutions at residues potentially contacting the target DNA strand was assessed using an EGFP decay assay. Four different sgRNAs (matches 1-4) were used, and each of the four target sites was tested with a mismatch sgRNA known to be efficiently used by wild-type SaCas9. The mismatch sgRNA for each site is shown to the right of the respective match sgRNA (e.g., for match site 3, the only mismatch sgRNAs are mm11 and 12). Mismatch positions are numbered, with position 21 being the most distal position from the PAM; mm is the mismatch. [Figure 24]Figure 24A: Activity of wild-type (WT) SaCas9 and SaCas9 derivatives bearing one or more alanine substitutions at residues that could potentially contact the target DNA strand. SaCas9 variants bearing dual combination substitutions were evaluated using the T7E1 assay against matched and single-mismatched endogenous human gene target sites. Matched "on-target" sites are named according to the sgRNA number of their gene target sites from Kleinstiver et al., Nature Biotechnology 2015. Mismatched sgRNAs are numbered, with the mismatch occurring at position 21, the most distal position from the PAM; the mismatched sgRNA is derived from the matched-on-target site listed to the left of the mismatched sgRNA. Figure 24B: Activity of wild-type (WT) SaCas9 and SaCas9 derivatives bearing one or more alanine substitutions at residues that could potentially contact the target DNA strand. SaCas9 variants carrying triple combinatorial substitutions were evaluated against matched and single-mismatched endogenous human gene target sites using the T7E1 assay. Matched "on-target" sites are named according to the sgRNA number of their gene target sites from Kleinstiver et al., Nature Biotechnology 2015. Mismatched sgRNAs are numbered, with the mismatch occurring at position 21, the most distal position from the PAM; the mismatched sgRNA is derived from the matched on-target site listed to the left of the mismatched sgRNA. DETAILED DESCRIPTION OF THE INVENTION
[0035] Restriction of CRISPR-Cas9 nuclease at imperfectly matched target sites their potential to induce unwanted "off-target" mutations (e.g., Tsa i et al., Nat Biotechnol. 2015), and in some cases, is comparable to that observed at the intended on-target site (Fu et al., Nat Biotechnol. 2013). Using CRISPR-Cas9 nuclease Previous studies have focused on the sequence between the guide RNA (gRNA) and the spacer region of the target site. By reducing the number of specific interactions, off-target sites of cleavage in human cells are avoided. This suggests that mutations may reduce the effects of mutations in the Biotechnol. 2014).
[0036] This is achieved by truncating the gRNA by 2 or 3 nt at its 5' end. The mechanism of this increased specificity is achieved earlier through the interaction of the gRNA / Cas9 complex. The reduction in energy required for cleaving the on-target site is This balances the energy required to generate a mismatch in the target DNA site. To cleave off-target sites, which are assumed to have an energy penalty due to It was hypothesized that the probability of having sufficient energy for / Brochure No. 099850).
[0037] Off-target effects of SpCas9 (incomplete transcription of the intended target site for the guide RNA) (at all matched or mismatched DNA sites) is a non-specific interaction with its target DNA site. It was hypothesized that this could be minimized by reducing heterologous interactions. The gRNA complex contains the NGG PAM sequence (recognized by SpCas9) (Deltc heva,E.et al.Nature 471,602-607(2011);Ji nek,M.et al.Science 337,816-821(2012);Ji ang,W.,et al.,Nat Biotechnol 31,233-239( 2013);Sternberg, SH, et al., Nature 507,6 2-67(2014)) and an adjacent 20 bp protospacer sequence (5' end of sgRNA (Jinek, M. et al. Science 337, 816- 821(2012);Jinek,M.et al.Elife 2,e00471(2 013);Mali,P.et al.,Science 339,823-826(2 013);Cong,L.et al.,Science 339,819-823(2 013)) to cleave the target site. may have greater energy than is required for recognition of the intended target DNA site, thereby It has been theorized that this allows for cleavage of mismatched off-target sites (Fu ,Y.,et al.,Nat Biotechnol 32,279-284(201 4) This property may be advantageous for the intended role of Cas9 in adaptive bacterial immunity. This gives it the ability to cleave foreign sequences that can be mutated. The excess energy model was achieved by decreasing the SpCas9 concentration (Hsu, PD .et al.Nat Biotechnol 31,827-832(2013);P attanayak,V.et al.Nat Biotechnol 31,839- 843 (2013)), or by reducing the complementarity length of the sgRNA (Fu, Y.,et al.,Nat Biotechnol 32,279-284(2014 ), which can reduce (but not eliminate) off-target effects This effect is supported by previous research demonstrating that It has been proposed (Josephs, EA et al. Nucleic Acids Res 43,8924-8941(2015);Sternberg,SH,et al.Nature 527,110-113(2015);Kiani,S.et al. Nat Methods 12, 1051-1054 (2015)). Structural data The SpCas9-sgRNA-target DNA complex was then bound to four SpCas9 residues (N4 97, R661, Q695, Q926) to the phosphate backbone of the target DNA strand It can be stabilized by several SpCas9-mediated DNA contacts, including direct hydrogen bonds. This suggests that (Nishimasu, H. et al. Cell 156, 935-94 9(2014);Anders,C.,et al.Nature 513,569-5 73 (2014)) (Fig. 1a and Fig. 6a and 6b). Disruption of one or more of these is just sufficient to retain robust on-target activity. However, SpCa at levels that reduce its ability to cleave mismatched off-target sites We hypothesized that the s9-sgRNA complex could be energetically balanced.
[0038] As described herein, the Cas9 protein theoretically acts as a Ca2+ receptor for DNA. Engineering s9 to exhibit increased specificity by reducing its binding affinity It can be widely used against Streptococcus pyogenes (Streptococcus pyogenes) Several variants of SpCas9 (SpCas9) have been identified based on structural information, bacterial selection vectors, and Interacting with phosphates on the DNA backbone using source-directed evolution and combinatorial design Individual alanine substitutions were made at various residues in SpCas9 that could be predicted to A robust E. coli-based script was introduced to engineer The variants were further tested for cellular activity using a screening assay to identify their variants. In this bacterial system, cell viability is assessed by the toxic gyrase toxin cc The gene for dB and the 23 bases targeted by gRNA and SpCas9 Retention or loss of activity depends on the cleavage and subsequent disruption of the selection plasmid containing the complement sequence. This led to the identification of residues associated with the loss of ATP. Furthermore, it showed improved target specificity in human cells. We identified and characterized another SpCas9 variant.
[0039] Furthermore, single alanine substitution mutations of SpCas9 were evaluated in a bacterial cell-based system. The activity of the isomers is generally robust, with 50-100% viability indicating robust cleavage, whereas 0% viability indicates robust cleavage. The enzyme functionally attenuated. R69A, R70A, R71A, Y72A, R74A, R75A, K76A, N77A, R78A, R115A, H160A, K163A, R165A, L169A, R403A ,T404A, F405A, N407A, R447A, N497A, I448A, Y45 0A, S460A, M495A, K510A, Y515A, R661A, M694A, Q 695A, H698A, Y1013A, V1015A, R1122A, K1123A, K 1124A, K1158A, K1185A, K1200A, S1216A, Q1221A , K1289A, R1298A, K1300A, K1325A, R1333A, K133 Additional mutations of SpCas9, including 4A, R1335A, and T1337A, were introduced into bacteria. Two mutants (R69A and F) with <5% viability in bacteria were identified. All of these additional single mutations were on-target for SpCas9, except for 405A. It appeared to have little effect on activity (>70% in bacterial screens). survival rate).
[0040] Cas9 variants identified in a bacterial screen function efficiently in human cells To further determine whether EGFP decay was mediated by ATP, a human U2OS cell-based EGFP decay assay was used. Various alanine-substituted Cas9 mutants were tested in this assay. Identification of target sites within the coding sequence of an integrated constitutively expressed EGFP gene The favorable cleavage was associated with the induction of indel mutations and the subsequent cleavage of ribosomal proteins, which was quantitatively assessed by flow cytometry. and resulted in a collapse of EGFP activity (e.g., Reyon et al., Nat B iotechnol.2012 May;30(5):460-5).
[0041] These experiments demonstrated that results obtained in bacterial cell-based assays were consistent with those obtained in human cells. The results show that the enzyme activity is well correlated with the enzyme activity of the α-glucanase, and that these genetic manipulation strategies can be applied to other species and different These findings suggest that the present invention may be extended to Cas9 from other cells. SpCas9, collectively referred to in the specification as "variants" or "variants thereof" and provides support for SaCas9 variants.
[0042] All of the variants described herein can be generated by, for example, simple site-directed mutagenesis. They can be rapidly incorporated into existing widely used vectors, and they contain few mutations. Since only requiring a single gene, the variants are compatible with the previously described SpCas9 platform. Other improvements that have been made (e.g., truncated sgRNA (Tsai et al., Nat B iotechnol 33,187-197(2015);Fu et al.,Nat Biotechnol 32,279-284(2014)), nickase mutation ( Mali et al., Nat Biotechnol 31,833-838(20 13);Ran et al.,Cell 154,1380-1389(2013)) , FokI-dCas9 fusion (Guilinger et al., Nat Biote chnol 32,577-582(2014);Tsai et al.,Nat B iotechnol 32,569-576(2014);International Publication No. 20141442 88); and engineered CRISPR-Ca with altered PAM specificity s9 nuclease (Kleinstiver et al., Nature. 2015 Jul 23;523(7561):481-5).
[0043] Thus, as used herein, Cas9 variants, e.g., SpCas9 variants, The wild-type sequence of SpCas9 is as follows: [ka] [ka]
[0044] The SpCas9 variants described herein have the following positions: N497, R661, Q6 95, mutations at one or more of Q926 (or positions similar thereto) (i.e. , substitution of a natural amino acid with a different amino acid, e.g., alanine, glycine, or serine. In some embodiments, the amino acid sequence of SEQ ID NO: 1 may include the amino acid sequence of SEQ ID NO: 1 with the Sp Cas9 may have at least the amino acid sequence of SEQ ID NO:1 in addition to the mutations described herein. At least 80%, e.g., at least 85%, 90%, or 95% identical, e.g., conserved For example, up to 5%, 10%, or even more of the residues of SEQ ID NO: 1 have been replaced by selective mutations. In a preferred embodiment, the variant has a difference of 15%, or 20%. , the desired activity of the parent, e.g., nuclease activity (if the parent is a nickase or dead Cas9) interacting with the target DNA and / or the guide RNA (excluding cases where the target DNA is present) maintain the ability to
[0045] To determine percent identity of two nucleic acid sequences, the sequences are aligned for optimal comparison purposes. Align (e.g., for optimal alignment, gaps should be between the first and second amino acids) The non-correlated sequences may be introduced into either or both of the acid or nucleic acid sequences for comparison purposes. (The sequence can be ignored.) The length of the reference sequence to be aligned for comparison purposes is At least 80% of the length of the reference sequence, and in some embodiments, at least 90% or 100%. Then, the corresponding amino acid or nucleotide position The nucleotides are compared. A position in the first sequence is the same as the corresponding position in the second sequence. If the position is occupied by a single nucleotide, then the molecules are identical at that position (as defined herein). As used herein, nucleic acid "identity" is equivalent to nucleic acid "homology." The percent identity is a function of the number of identical positions shared by the sequences, the number of gaps that need to be introduced for optimal alignment of the two sequences, and The length of each gap is taken into account. The invention can be accomplished by various means within the skill of the art, for example, by extracting publicly available components. Computer software, e.g., Smith Waterman Alignment ( Smith, TFand MS Waterman (1981) J Mol Bi ol 147:195-7); incorporated into GeneMatcher Plus™ "BestFit" (Smith and Waterman, Advances in n Applied Mathematics,482-489(1981)), Sch. warz and Dayhof(1979) Atlas of Protein Se Quence and Structure,Dayhof,MO,Ed,pp 3 53-358; BLAST program (Basic Local Alignment Search Tool;(Altschul,SF,W.Gish,et al. (1990) J Mol Biol 215:403-10), BLAST-2, BLA ST-P, BLAST-N, BLAST-X, WU-BLAST-2, ALIGN, AL IGN-2, CLUSTAL, or Megalign (DNASTAR) software Furthermore, those skilled in the art will be able to determine the appropriate parameters for measuring alignment. data required to achieve maximal alignment over the length of the sequences being compared, e.g. Any algorithm required can be determined. Generally, The comparison length can be any length less than the total length (for example, 5%, 10%, 20%, 30% , 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100%) For purposes of the present compositions and methods, at least 80% of the entire length of the sequences are aligned. do.
[0046] For purposes of the present invention, the comparison of sequences and determination of percent identity between two sequences is carried out using the method of lossum 62 scoring matrix with gap penalty 12 and gap extension penalty This can be achieved using a gap penalty of 4, and a frameshift gap penalty of 5.
[0047] Conservative substitutions typically include amino acids from the following groups: glycine, alanine; valine, isoleucine Synthin, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, These include substitutions within threonine, lysine, arginine, and phenylalanine, and tyrosine. can be.
[0048] In some embodiments, the SpCas9 variant comprises the following set of mutations: N4 97A / R661A / Q695 / Q926A (quadruple alanine mutant); Q695A / Q926A (double alanine mutant); R661A / Q695A / Q926A and N Some of the mutants contain one of the 497A / Q695A / Q926A (triple alanine mutants). In some embodiments, additional substitution mutations at L169 and / or Y450 are or Q695 or Q It can be added to a single mutant carrying a substitution at 926. In some embodiments, the mutant has an alanine in place of the wild-type amino acid. In this case, the mutant is a nucleotide sequence of any amino acid other than arginine or lysine (or any naturally occurring amino acid). It has amino acids.
[0049] In some embodiments, the SpCas9 variant comprises a nuclease portion of the protein. Reduce or destroy the nuclease activity of Cas9 to render it catalytically inactive The following mutations: D10, E762, D839, H983, or D986 and H84 0 or one of N863, e.g., D10A / D10N and H840A / H840N / H840Y; substitutions at those positions include alanine (Nishimasu et al. ., Cell 156, 935-949 (2014)), or other residues , e.g., glutamine, asparagine, tyrosine, serine, or aspartic acid, e.g., For example, E762Q, H983N, H983Y, D986N, N863D, N863S, or The compound may be N863H (see International Publication No. WO 2014 / 152432). In some embodiments, the variant is D10A or H840A (single-chain nickase Mutations in the nuclease gene (which produces a nuclease activity) or D10A and H840A (which deactivate the nuclease activity) This mutant is known as dead Cas9 or dCas9. This includes mutations that
[0050] SpCas9N497A / R661A / Q695A / R926A mutations are responsible for the yellow grape Staphylococcus aureus Cas9 (SaCas9) See, for example, Figure 20. Residues that contact the DNA or RNA backbone These mutations increase the specificity of SaCas9, as observed for SpCas9. Accordingly, SaCas9 variants are also provided herein. do.
[0051] The SaCas9 wild-type sequence is as follows: [ka]
[0052] The SaCas9 variants described herein have the following positions: Y211, W229, R2 1, 2, 3, 4, 5, or 6 of 45, T392, N419, and / or R654 2 with mutations at all of the following positions: 1 or 2 of Y211, W229, R245, T392, N419, and / or R654 2 with at least one mutation in position 1, 3, 4, 5 or 6 of the amino acid sequence of SEQ ID NO: 2 also contain sequences that are 80% identical.
[0053] In some embodiments, the variant SaCas9 protein comprises the following mutation: Y 211A;W229A;Y230A;R245A;T392A;N419A;L446A ;Y651A;R654A;D786A;T787A;Y789A;T882A;K88 6A;N888A;A889A;L909A;N985A;N986A;R991A;R 1015A;N44A;R45A;R51A;R55A;R59A;R60A;R116 A;R165A;N169A;R208A;R209A;Y211A;T238A;Y2 39A;K248A;Y256A;R314A;N394A;Q414A;K57A;R 61A;H111A;K114A;V164A;R165A;L788A;S790A; R792A;N804A;Y868A;K870A;K878A;K879A;K881 Also contains one or more of A;Y897A;R901A;K906A.
[0054] In some embodiments, the variant SaCas9 protein contains the following additional mutations: Different: Y211A, W229A, Y230A, R245A, T392A, N419A, L4 46A, Y651A, R654A, D786A, T787A, Y789A, T882A, K886A, N888A, A889A, L909A, N985A, N986A, R991 A, R1015A, N44A, R45A, R51A, R55A, R59A, R60A, R 116A, R165A, N169A, R208A, R209A, Y211A, T238A , Y239A, K248A, Y256A, R314A, N394A, Q414A, K57 A, R61A, H111A, K114A, V164A, R165A, L788A, S79 0A, R792A, N804A, Y868A, K870A, K878A, K879A, K Contains one or more of 881A, Y897A, R901A, K906A.
[0055] In some embodiments, the variant SaCas9 protein comprises multiple substitution mutations. :R245 / T392 / N419 / R654 and Y221 / R245 / N419 / R6 54 (quadruple variant mutant); N419 / R654, R245 / R654, Y22 1 / R654, and Y221 / N419 (double mutant); R245 / N419 / R 654, Y211 / N419 / R654, and T392 / N419 / R654 (triple thrust) In some embodiments, mutants include a nucleotide sequence that replaces the wild-type amino acid. Contains ranin.
[0056] In some embodiments, the variant SaCas9 protein is E782K, K92 9R, N968K, and / or R1015H mutations. H variant (E782K / N968K / R1015H), KRH variant (E782 K / K929R / R1015H), or KRKH variant (E782K / K929R / N968K / R1015H).
[0057] In some embodiments, the variant SaCas9 protein is selected from the group consisting of D10, E477, Mutations at D556, H701, or D704; and H557 or N580 The nuclease activity of the nuclease may also be reduced by one or more mutations selected from the group consisting of:
[0058] In some embodiments, the mutation is: (i) D10A or D10N; (ii) H 557A, H557N, or H557Y, (iii) N580A, and / or (i v) D556A.
[0059] As used herein, optionally, one or more preparations for the expression of a variant protein are Isolated nucleic acid encoding a Cas9 variant operably linked to a knot domain, isolated Vectors containing the nucleic acid, and proteins containing the nucleic acid and optionally variant proteins Host cells, eg, mammalian host cells, expressing the proteins are also provided.
[0060] The variants described herein can be used to modify the genome of a cell; The methods generally involve inserting a variant protein in a cell into a target gene that is complementary to a selected portion of the cell's genome. This involves expressing a gene along with a guide RNA that has a specific region. Methods for modifying the solubility of cellulose are known in the art and are described, for example, in U.S. Pat. No. 8,993,233. No. 20140186958; U.S. Patent Application Publication No. 9,023 ,649 specification; WO 2014 / 099744 pamphlet; WO 20 14 / 089290 Brochure; International Publication No. 2014 / 144592 Brochure ;International Publication No. 144288 Brochure;International Publication No. 2014 / 204578 Brochure Lett.; International Publication No. 2014 / 152432; International Publication No. 2115 / 09 No. 9850; U.S. Pat. No. 8,697,359; U.S. Pat. App. Pub. No. 20160024529; U.S. Patent Application Publication No. 20160024524 ;U.S. Patent Application Publication No. 20160024523;U.S. Patent Application Publication No. 20160 024510; U.S. Patent Application Publication No. 20160017366; U.S. Patent Publication No. 20160017301; Publication No. 2015037665 2; U.S. Patent Application Publication No. 20150356239; U.S. Patent Application Publication No. No. 20150315576; U.S. Patent Application Publication No. 20150291965 ;U.S. Patent Application Publication No. 20150252358;U.S. Patent Application Publication No. 20150 247150; U.S. Patent Application Publication No. 20150232883; U.S. Patent Publication No. 20150232882; U.S. Patent Application Publication No. 2015020387 2; U.S. Patent Application Publication No. 20150191744; U.S. Patent Application Publication No. No. 20150184139; U.S. Patent Application Publication No. 20150176064 ;U.S. Patent Application Publication No. 20150167000;U.S. Patent Application Publication No. 20150 166969; U.S. Patent Application Publication No. 20150159175; U.S. Patent Publication No. 20150159174; U.S. Patent Application Publication No. 2015009347 3; U.S. Patent Application Publication No. 20150079681; U.S. Patent Application Publication No. No. 20150067922; U.S. Patent Application Publication No. 20150056629 ;U.S. Patent Application Publication No. 20150044772;U.S. Patent Application Publication No. 20150 024500; U.S. Patent Application Publication No. 20150024499; U.S. Patent Publication No. 20150020223; Publication No. 2014035686 7; U.S. Patent Application Publication No. 20140295557; U.S. Patent Application Publication No. No. 20140273235; U.S. Patent Application Publication No. 20140273226 ;U.S. Patent Application Publication No. 20140273037;U.S. Patent Application Publication No. 20140 No. 189896; U.S. Patent Application Publication No. 20140113376; U.S. Patent Publication No. 20140093941; Publication No. 2013033077 8; U.S. Patent Application Publication No. 20130288251; U.S. Patent Application Publication No. No. 20120088676; U.S. Patent Application Publication No. 20110300538 ;U.S. Patent Application Publication No. 20110236530;U.S. Patent Application Publication No. 20110 217739; U.S. Patent Application Publication No. 20110002889; U.S. Patent Publication No. 20100076057; U.S. Patent Application Publication No. 2011018977 6; U.S. Patent Application Publication No. 20110223638; U.S. Patent Application Publication No. No. 20130130248; U.S. Patent Application Publication No. 20150050699 ;U.S. Patent Application Publication No. 20150071899;U.S. Patent Application Publication No. 20150 050699; U.S. Patent Application Publication No. 20150045546; U.S. Patent Publication No. 20150031134; U.S. Patent Application Publication No. 2015002450 0; U.S. Patent Application Publication No. 20140377868; U.S. Patent Application Publication No. Specification No. 20140357530; WO 2008 / 108989; International Publication No. 2010 / 054108; International Publication No. 2012 / 164565 No. brochure; International Publication No. 2013 / 098244 brochure; International Publication No. 201 3 / 176772 Brochure; International Publication No. 20150071899 Brochure; U.S. Patent Application Publication No. 20140349400; U.S. Patent Application Publication No. 201403 35620; U.S. Patent Application Publication No. 20140335063; U.S. Patent No. Publication No. 20140315985; Publication No. 20140310830 No. 20140310828; U.S. Patent Application Publication ... 0140309487; U.S. Patent Application Publication No. 20140304853; U.S. Patent Application Publication No. 20140298547; U.S. Patent Application Publication No. 201402 95556; U.S. Patent Application Publication No. 20140294773; U.S. Patent No. Publication No. 20140287938; Publication No. 20140273234 No. 20140273232; U.S. Patent Application Publication ... 0140273231; U.S. Patent Application Publication No. 20140273230; U.S. Patent Application Publication No. 20140271987; U.S. Patent Application Publication No. 201402 56046; U.S. Patent Application Publication No. 20140248702; U.S. Patent No. Publication No. 20140242702; Publication No. 20140242700 No. 20140242699; U.S. Patent Application Publication No. 20140242699; U.S. Patent Application Publication No. 0140242664; U.S. Patent Application Publication No. 20140234972; U.S. Patent Application Publication No. 20140227787; U.S. Patent Application Publication No. 201402 12869; U.S. Patent Application Publication No. 20140201857; U.S. Patent No. Publication No. 20140199767; Publication No. 20140189896 No. 20140186958; U.S. Patent Application Publication ... 0140186919; U.S. Patent Application Publication No. 20140186843; U.S. Patent Application Publication No. 20140179770; U.S. Patent Application Publication No. 201401 79006; U.S. Patent Application Publication No. 20140170753; Makar ova et al., “Evolution and classification of the CRISPR-Cas systems”9(6)Nature Re views Microbiology 467-477(1-23)(Jun.201 1);Wiedenheft et al.,“RNA-guided genetic silencing systems in bacteria and archaea of”482 Nature 331-338(Feb.16,2012);Gasiu nas et al.,“Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria”109(39). )Proceedings of the National Academy of Sciences Science USA E2579-E2586(Sep.4,2012);Jin ek et al.,“A Programmable Dual-RNA Guide d DNA Endonuclease in Adaptive Bacterial Immunity.”337 Science 816–821(Aug.17,201). 2);Carroll,“A CRISPR Approach to Gene Taking rgeting”20(9)Molecular Therapy 1658-1660 (Sep.2012);Appealed on May 25, 2012, under National Assessment Act No. 61 / 652 ,086 Instrumentation System;Al-Attar et al.,Clustered Regul arly Interspaced Short Palindromic Repeat ts(CRISPRs):The Hallmark of an Ingenious Antiviral Defense Mechanisms in Prokaryotes tes,Biol Chem.(2011)vol.392,Issue 4,pp.2 77-289;Hale et al.,Essential Features an d Rational Design of CRISPR RNAs That Fu nction With the Cas RAMP Module Complex to Cleave RNAs,Molecular Cell,(2012)vol. 45, Issue 3, 292-302.
[0061] The variant proteins described herein are similar to the Cas9 proteins described in the above references. Alternatively or in addition to any of the mutations described above, or in combination with the mutations described above. Furthermore, the variants described herein can be used in Known wild-type Cas9 or other Cas9 mutants (e.g., dCas9 or C, as described above) as9 nickase), fusion proteins, e.g., U.S. Pat. No. 8,993,233 Specification; U.S. Patent Application Publication No. 20140186958 Specification; U.S. Patent No. 9,023, 649; WO 2014 / 099744; WO 201 4 / 089290; WO 2014 / 144592; International Publication No. 144288 Brochure; International Publication No. 2014 / 204578 Brochure International Publication No. 2014 / 152432; International Publication No. 2115 / 099 No. 850; U.S. Pat. No. 8,697,359; U.S. Pat. App. Pub. No. 2 US Patent Application Publication No. 2011 / 0189776 Publication No. 2011 / 0223638; Publication No. 201 3 / 0130248 specification; WO 2008 / 108989 pamphlet; International Publication No. 2010 / 054108 Pamphlet; International Publication No. 2012 / 164565 Pamphlet Brochure; International Publication No. 2013 / 098244 Brochure; International Publication No. 2013 / No. 176772; U.S. Patent Application Publication No. 20150050699; U.S. Patent Application Publication No. 20150071899 and International Publication No. 2014 / 124 284 brochure. For example, variants, preferably those with one or more nuclease-reduced, altered or variants containing lethal mutations are inserted into the N- or C-terminus of Cas9, along with the transcriptional activation domain. or other heterologous functional domains (e.g., transcriptional repressors (e.g., KRAB, ERD, SID, e.g., ets2 repressor factor (ERF) repressor domain (ER D) amino acids 473 to 530, amino acids 1 to 97 of the KRAB domain of KOX1, is amino acids 1–36 of the Mad mSIN3 interaction domain (SID); Beerli et al., PNAS USA 95:14628-14633 (1998)) or silencers, e.g., heterochromatin protein 1 (HP1, also known as swi6) known), e.g., HP1α or HP1β; long non-codon-binding domains fused to fixed RNA binding sequences; Proteins or peptides that can recruit incRNAs, such as M to S2 coat protein, endoribonuclease Csy4, or lambda N protein DNA methylation is known in the art. Enzymes that modify the methylation state (e.g., DNA methyltransferases (DNMTs) or T ET proteins); or enzymes that modify histone subunits (e.g., histone ase Histone deacetylases (HATs), histone deacetylases (HDACs), histone metabolites methyltransferases (e.g., for methylation of lysine or arginine residues) or histone demethylases (e.g., for demethylation of lysine or arginine residues) The sequence of many such domains, e.g., methyltransferases in DNA, can also be used. Domains that catalyze the hydroxylation of hydroxylated cytosine are known in the art. Typical proteins include ten-eleven translocation proteins (Ten-Eleven- Translocation Translocation (TET) 1-3 family, 5-methylcytosine translocation (5-methylcytosine translocation) in DNA Enzymes that convert 5-hydroxymethylcytosine (5-mC) to 5-hydroxymethylcytosine (5-hmC) include can be done.
[0062] Human TET1-3 sequences are known in the art and are shown in the table below.
[0063] [Table 2]
[0064] In some embodiments, all or part of the full-length sequence of the catalytic domain, e.g., 2OGFe is encoded by an in-rich extension and seven highly conserved exons DO domains, e.g., the Tet1 catalytic domain containing amino acids 1580-2052; Tet2, which contains amino acids 1290 to 1905, and Tet3, which contains amino acids 966 to 1678 The catalytic module containing the key catalysts in all three Tet proteins can be included. For alignments accounting for catalytic residues and supporting material for full-length sequences (f tp site ftp.ncbi.nih.gov / pub / aravind / DONS / s supplementary_material_DONS.html) is an example For example, Iyer et al.,Cell Cycle.2009 Jun 1;8(1 1):1698-710. Epub 2009 Jun 27, see Figure 1 (e.g. See, e.g., seq 2c); in some embodiments, the sequence is amino acid 14 of Tet1 18 to 2136 or the corresponding region in Tet2 / 3.
[0065] Another catalytic module is the protein identified in Iyer et al., 2009. It may be from a protein.
[0066] In some embodiments, the heterologous functional domain is a biological tether and protein, endoribonuclease Csy4, or lambda N protein in whole or in part These proteins contain specific RNA molecules containing stem-loop structures are defined by dCas9 gRNA targeting sequences It can be used to recruit to locations where the MS2 coat is used. dCas9 bait fused to protein, endoribonuclease Csy4, or lambda N Ant is a long non-coding RNA that binds to Csy4, MS2 or lambda N-binding sequences. (lncRNA), e.g., used to recruit XIST or HOTAIR See, for example, Keryer-Bibens et al., Biol. Ce ll 100:125-138(2008). Alternatively, see Csy4, MS 2 or lambda N protein binding sequences are described, for example, in Keryer-Bibens et al. al., supra, can be conjugated to another protein, as described herein. The methods and compositions are used to target proteins to dCas9 variant binding sites. In some embodiments, Csy4 is catalytically inactive. In embodiments, the Cas9 variant, preferably a dCas9 variant, is described in U.S. Pat. No. 8,993,233; U.S. Patent Application Publication No. 20140186958 ;U.S. Patent No. 9,023,649;WO 2014 / 099744 Lett; International Publication No. 2014 / 089290 Brochure; International Publication No. 2014 / 14 No. 4592; International Publication No. 144288; International Publication No. 2014 / 204578 Brochure; International Publication No. 2014 / 152432 Brochure; Country International Publication No. 2115 / 099850; U.S. Patent No. 8,697,359 Publication No. 2010 / 0076057; Publication No. 201 No. 1 / 0189776; U.S. Patent Application Publication No. 2011 / 0223638; U.S. Patent Application Publication No. 2013 / 0130248; WO 2008 / 1089 Pamphlet No. 89; Pamphlet No. WO 2010 / 054108; Pamphlet No. WO 2010 / 054108 Pamphlet No. 012 / 164565; Pamphlet No. WO 2013 / 098244 International Publication No. 2013 / 176772; U.S. Patent Application Publication No. 20150 050699; U.S. Patent Application Publication No. 20150071899; and As described in the pamphlet of International Publication No. 2014 / 204578, do.
[0067] In some embodiments, the fusion protein comprises a dCas9 variant and a heterologous functional domain. These fusion proteins (or fusion proteins with linked structures) contain linkers between the domains. The linker that can be used in the intersubstrate (intercellular) synthesis may be any sequence that does not interfere with the function of the fusion protein. In a preferred embodiment, the linker is a short chain, e.g., 2-20 amino acids, which are typically flexible (i.e., amino acids with a high degree of freedom) In some embodiments, The linker may be one or more of GGGS (SEQ ID NO: 3) or GGGGS (SEQ ID NO: 4). units, for example, 2, 3 GGGS (SEQ ID NO: 5) or GGGGS (SEQ ID NO: 6) units , 4, or more repeats. Other linker sequences can also be used. .
[0068] In some embodiments, the variant protein facilitates delivery to the intracellular space. Cell-penetrating peptide sequences, such as the HIV-derived TAT peptide, penetratin, transposon, These include tans, or hCT-derived cell-penetrating peptides, e.g., Caron et al. ,(2001)Mol Ther.3(3):310-8;Langel,Cell-P enetrating Peptides:Processes and Applic ations(CRC Press,Boca Raton FL 2002);El- Andaloussi et al.,(2005)Curr Pharm Des.1 1(28):3597-611; and Deshayes et al., (2005) See Cell Mol Life Sci. 62(16):1839-49.
[0069] Cell-penetrating peptides (CPPs) are peptides that penetrate the cell membrane and enter the cytoplasm or other organelles, e.g. For example, it is a short peptide that facilitates the translocation of a wide range of biomolecules into the mitochondria and nucleus. Examples of molecules that can be delivered by CPPs include therapeutic agents, plasmid DNA, Oligonucleotides, siRNA, peptide nucleic acids (PNAs), proteins, peptides, nucleic acids CPPs generally contain 30 amino acids or less and include nanoparticles, nanoparticles, and liposomes. Positively charged amino acids with high relative abundance derived from natural or non-naturally occurring proteins or chimeric sequences amino acids, e.g., lysine or arginine, or an alternating pattern of polar and nonpolar amino acids Commonly used CPPs in the art include Tat (Fr ankel et al.,(1988)Cell.55:1189-1193,Viv es et al.,(1997)J.Biol.Chem.272:16010-16 017), penetratin (Derossi et al., (1994) J. Biol. Chem. 269:10444-10450), polyarginine peptide sequence (Wend er et al.,(2000)Proc.Natl.Acad.Sci.USA 9 7:13003-13008,Futaki et al.,(2001)J.Biol .Chem.276:5836-5840), and transportan (Pooga e t al., (1998) Nat. Biotechnol. 16:857-861) It can be obtained.
[0070] A CPP can be attached to its cargo via covalent or non-covalent strategies. Methods for covalently linking CPPs and their cargoes are known in the art, e.g. , chemical crosslinking (Stetsenko et al., (2000) J.Org.Chem .65:4900-4909,Gait et al.(2003)Cell.Mol. Life.Sci.60:844-853) or cloning of fusion proteins (Na Gahara et al.,(1998) Nat.Med.4:1449-1453) The non-covalent interaction between the cargo and the short-chain amphiphilic CPP containing polar and non-polar domains Coupling is established via electrostatic and hydrophobic interactions.
[0071] CPPs are used in the art for the delivery of potentially therapeutic biomolecules into cells. Examples include cyclosporin-conjugated polyarginine for immunosuppression. Phosphorus (Rothbard et al., (2000) Nature Medicine 6(11):1253-1257), a CPP called MPG for the inhibition of tumorigenesis. siRNA against cyclin B1 bound to cyclin B1 (Crombez et al., 2007)Biochem Soc.Trans.35:44-46), reduces cancer cell growth. Tumor suppressor p53 peptide (Takenobu) binding to CPPs to reduce et al.,(2002)Mol.Cancer Ther.1(12):1043- 1049,Snyder et al.,(2004)PLoS Biol.2:E36 ), and Ras or phosphoinositol fused to Tat for treating asthma. A dominant-negative form of PI3K (Myou et al., 200 3) J. Immunol. 171:4399-4405).
[0072] CPPs are widely used in the art to produce contrast agents for imaging and biosensing applications. For example, green fluorescent protein attached to Tat GFP has been used to label cancer cells (Shokolenko et al. al., (2005) DNA Repair 4(4):511-518). Tat conjugated to ribosomal RNA successfully penetrates the blood-brain barrier for visualization in rat brain. It has been used to filter out (Santra et al., (2005) Chem. Commun. 3144-3146). CPP uses magnetic resonance imaging techniques for cell imaging. It can also be combined with other techniques (Liu et al., (2006) Biochem. and Biophys.Res.Comm.347(1):133-140). RAM sey and Flynn,Pharmacol Ther.2015 Jul 22 See also .pii:S0163-7258(15)00141-2.
[0073] Alternatively or additionally, the variant protein may contain a nuclear localization sequence, e.g., SV40 LAMBDA. genomic T antigen NLS (PKKKRRV (SEQ ID NO: 7)) and nucleoplasmin NLS ( Other NLSs may be used in the art. are known in the art; see, for example, Cokol et al., EMBO Rep. 20 00 Nov 15;1(5):411-415;Freitas and Cunha ,Curr Genomics.2009 Dec;10(8):550-557 I want to be done that.
[0074] In some embodiments, the variant comprises a moiety that has a high affinity for the ligand. Such affinity tags include, for example, GST, FLAG, or hexahistidine sequences. can facilitate the purification of recombinant variant proteins.
[0075] Any method known in the art for delivering variant proteins to cells can be used. For example, in vitro translation from nucleic acid encoding a variant protein using the method of or by expression in a suitable host cell; the protein can be produced by Numerous methods are known in the art for production. For example, proteins can be produced by enzymes. Mother, E. coli, insect cell lines, plants, transgenic animals, or cultures It can be produced in and purified from mammalian cells; e.g., Palomares et al., “Production of Recombinant Protei ns:Challenges and Solutions,”Methods Mol Biol. 2004;267:15-52. The protein is optionally attached to the cell using a linker that is cleaved when the protein is present in the cell. It can be attached to a moiety that facilitates its transfer into the vesicle, for example, a lipid nanoparticle. LaFountaine et al.,Int J Pharm.2015 Au See g 13;494(1):180-194.
[0076] Expression system In order to use the Cas9 variants described herein, they must be isolated from the nucleic acids that encode them. It may be desirable to express these. This can be done in a variety of ways. For example, nucleic acids encoding Cas9 variants can be incorporated into vectors for replication and / or expression. It can be cloned into intermediate vectors for transformation into prokaryotic or eukaryotic cells. Intermediate vectors typically contain a Cas9 barrier for the production of Cas9 variants. Prokaryotic vectors, e.g., plasmids, for the storage or manipulation of nucleic acids encoding the compounds. , or a shuttle vector, or an insect vector. The nucleic acid can be expressed in a plant cell, an animal cell, preferably a mammalian or human cell, a fungal cell, or a fungal cell. Cloning into an expression vector for administration to cells, bacteria, or protozoan cells It is also possible.
[0077] To obtain expression, the sequence encoding the Cas9 variant is typically inserted into a transcriptional target. The vector is subcloned into an expression vector containing a promoter for expression in a suitable bacterium or and eukaryotic promoters are well known in the art, see, e.g., Sambroo k et al.,Molecular Cloning,A Laboratory Manual(3d ed.2001);Kriegler, Gene Transfe r and Expression:A Laboratory Manual(199 0); and Current Protocols in Molecular Bio The text is described in the logy (Ausubel et al., eds., 2010). Bacterial expression systems for expressing engineered proteins include, for example, E. coli. , Bacillus sp., and Salmonella spp. la) (Palva et al., 1983, Gene 22 Kits for such expression systems are commercially available. Eukaryotic expression systems for bacteria, yeast, and insect cells are also well known in the art, They are also commercially available.
[0078] The promoter used to direct expression of a nucleic acid will depend on the particular application. For example, a strong constitutive promoter is typically used for expression and purification of a fusion protein. In contrast, when Cas9 variants are to be administered in vivo for gene regulation, Either constitutive or inducible promoters depending on the specific use of the Cas9 variant Additionally, preferred promoters for administration of Cas9 variants include: A weak promoter, such as HSV TK or a promoter with similar activity. The promoter may contain an element that is responsive to transactivation, e.g., a hypoxia response. response element, Gal4 response element, lac repressor response element, and Small molecule regulatory systems may also be included, such as the tetracycline regulatory system and the RU-486 system (e.g., Gossen & Bujard, 1992, Proc. Natl. Acad. Sci. USA,89:5547;Oligino et al.,1998,Gene The r.,5:491-496;Wang et al.,1997,Gene Ther. ,4:432-441;Neering et al.,1996,Blood,88: 1147-55; and Rendahl et al., 1998, Nat. Biote chnol.,16:757-761).
[0079] In addition to a promoter, expression vectors typically include a promoter, which may be used in either prokaryotic or eukaryotic organisms. Any transcription unit containing all additional elements required for expression of the nucleic acid in a host cell. Thus, a typical expression cassette may contain, for example, a promoter operably linked to a nucleic acid sequence encoding an s9 variant; and For example, efficient polyadenylation of transcripts, transcription termination, ribosome binding sites, or translation It contains any signals required for termination. Additional elements of the cassette include, for example, Examples include enhancers and heterologous splicing intron signals. do.
[0080] The specific expression vectors used to transport genetic information into cells are Cas9 vectors. The intended use of the invention, e.g., expression in plants, animals, bacteria, fungi, protozoa, etc. Standard bacterial expression vectors include pBR322-based plasmids, Plasmids such as pSKF and pET23D, as well as commercially available tagged fusion expression systems, e.g., G Examples include ST and LacZ.
[0081] For eukaryotic expression vectors, expression vectors containing regulatory elements from eukaryotic viruses vectors, such as SV40 vectors, papillomavirus vectors, and Epstein-Barr virus vectors. Other exemplary eukaryotic vectors include vectors derived from viruses. pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo-5, and Vacuum pDSVE, and SV40 early promoter, SV40 late promoter, Metallothionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus The promoters are either the nucleotide sequence of ... and any other promoter that allows expression of the protein under the direction of another promoter as indicated. vectors.
[0082] The vectors for expressing Cas9 variants contain R, which drives the expression of guide RNA. NA PolIII promoter, e.g., H1, U6, or 7SK promoter These human promoters are capable of activating Cas in mammalian cells after plasmid transfection. It allows the expression of 9 variants.
[0083] Some expression systems include a marker for selection of stably transfected cell lines, e.g., Myridine kinase, hygromycin B phosphotransferase, and dihydrofolate receptor High-yield expression systems, such as baculovirus vectors in insect cells, under the direction of the polyhedrin promoter or other strong baculovirus promoter Also suitable are those for use with gRNA coding sequences.
[0084] Elements typically included in an expression vector include: Antibiotic resistance allows for selection of bacteria harboring a functional replicon and recombinant plasmid and a non-essential region of the plasmid to allow for the insertion of recombinant sequences. Also included are unique restriction sites within the region.
[0085] Bacterial, mammalian, yeast or other organisms that express large amounts of protein using standard transfection methods. The protein is then purified using standard techniques (e.g., Colley et al., 1989, J. Biol. Chem., 264:17 619-22;Guide to Protein Purification,in Methods in Enzymology,vol.182(Deutscher, Transformation of eukaryotic and prokaryotic cells is carried out according to standard techniques. (e.g., Morrison, 1977, J. Bacteriol. 132:34 9-351;Clark-Curtiss&Curtiss,Methods in E nzymology 101:347-362 (Wu et al., eds, 1983 reference).
[0086] Any of the known procedures for introducing foreign nucleotide sequences into host cells can be used. Such methods include calcium phosphate transfection, polybrene, protoplasts, and fusion, electroporation, nucleofection, liposomes, microinjection injection, naked DNA, plasmid vectors, viral vectors (episomes) (both endothelial and integrated) and cloned genomic DNA, cDNA, synthetic DNA A or other well-known method of introducing foreign genetic material into a host cell may be used. (See, e.g., Sambrook et al., supra). All that is required is In some cases, the particular genetic engineering procedure used may result in the generation of a small number of Cas9 variants in the host cell that can express the Cas9 variant. The key is to be able to successfully transfer at least one gene.
[0087] This method involves injecting purified Cas9 protein into cells along with gRNA to form ribonucleoproteins (RNPs). By introducing the gRNA and Cas9 protein as a complex, It may also involve modifying the gDNA by introducing RNA. The nucleic acid may be a nucleic acid (e.g., in an expression vector) encoding a NA or a guide RNA.
[0088] The present invention also includes vectors and cells containing the vectors. [Example]
[0089] The invention is further described in the following examples, which are included in the appended claims. The descriptions are not intended to limit the scope of the invention.
[0090] method A bacterial-based positive selection assay for evolving SpCas9 variants Competent Escherichia coli (E. coli) containing a positive selection plasmid (with the target site embedded) coli) BW25141(λDE3) 23 into the Cas9 / sgRNA-encoding plasmid After 60 minutes of recovery in SOB medium, chloramphenicol (non- containing either chloramphenicol + 10 mM arabinose (selection) or chloramphenicol + 10 mM arabinose (selection) The transformants were plated on LB medium containing 1000 kJ / ml of PBS.
[0091] To identify additional positions that may be important for genome-wide target specificity, homing A bacterial selection system (hereafter referred to as positive) has already been used to study the properties of guanine nuclease. This is called sexual selection. (Chen & Zhao, Nucleic Acids Res 33 ,e154(2005);Doyon et al.,J Am Chem Soc 1 28, 2477-2484 (2006)) was adapted.
[0092] In this adaptation of the system, the Cas gene of a positive selection plasmid encoding an inducible toxic gene was 9-mediated cleavage allows cell survival due to the subsequent degradation and loss of the linearized plasmid. After establishing that SpCas9 can function in a positive selection system, we performed the selection of wild-type and variant Both the cleavage sites were then cleaved using a selection plasmid carrying a target site selected from the known human genome. Cells carrying a positive selection plasmid containing the target site were tested for their ability to These variants were introduced into bacteria and plated on selective media. Cleavage of the midi was performed by counting the survival frequency: colonies on selective plates / colonies on non-selective plates. The estimation was made by calculating (see Figures 1, 5-6).
[0093] [Table 3]
[0094] Human cell culture and transfection Constitutively expressed EGFP-PEST reporter gene 15 Single integrated core U2OS.EGFP cells carrying the PI were cultured in 10% FBS, 2 mM GlutaMax (Life Technologies), penicillin / streptomycin, and 4 Advanced DMEM medium (Life Technologies) supplemented with 0.000 μg / ml G418 The cells were cultured in a Lonz (London Technologies) at 37°C with 5% CO2. a) 4D-nucleofector DN-100 program according to the manufacturer's protocol Use 750 ng of Cas9 plasmid and 250 ng of sgRNA plasmid according to the protocol. Cells were co-transfected with the empty U6 promoter and the smid (unless otherwise noted). For all human cell experiments, the Cas9 plasmid was transfected along with the plasmid. was used as a negative control (see Figures 2, 7-10).
[0095] Human cell EGFP decay assay EGFP decay experiments were performed as previously described 16 Transfection for EGFP expression Approximately 52 hours after transfection, the transfected cells were analyzed using a Fortessa flow cytometer (BD Background EGFP loss was analyzed using a cytochrome P4500 (Cell Biosciences). Experiments were gated at approximately 2.5% (see Figures 2 and 7).
[0096] T7E1 Assay for Quantifying Nuclease-Induced Mutation Rates, Targeted DeepSeq Sensing, and GUIDE-seq The T7E1 assay was performed as previously described for human cells (Kleinstiver, BP et al., Nature 523, 481-485 (2015)) For U2OS.EGFP human cells, Agencourt DNAdvance G genomic DNA Isolation Kit(Beckman Coulter) Transfected cells were transfected with genomic DNA approximately 72 hours after transfection using a genomics Approximately 200 ng of purified PCR product was denatured, annealed, and then purified to T7E1( Digestion was performed using a 500-kDa ELISA kit (New England BioLabs) as previously described for human cells. As stated (Kleinstiver et al., Nature 523, 481-4 85(2015);Reyon et al,.Nat Biotechnol 30, 460-465(2012)), Qiaxcel capillary electrophoresis instrument (QIage n) was used to quantify the mutagenesis frequency.
[0097] GUIDE-seq experiments were performed as previously described (Tsai et al., 2014). at Biotechnol 33,187-197(2015)). Briefly, Phosphorylated phosphorothioate-modified double-stranded oligodeoxynucleotides (dsODN) were used. Cas9 nuclease was used together with the Cas9 and sgRNA expression plasmids as described below. Both were transfected into U2OS cells. dsODN-specific amplification and high-throughput sequencing Sequencing and mapping were performed to identify genomic intervals containing DSB activity. Turn on off-target read counting for double or quadruple mutant variant experiments Normalization to target read counts corrected for differences in sequencing depth between samples. The normalized ratios for wild-type and variant SpCas9 were then compared to determine off-target activity. The fold change in activity at the target site was calculated. and variant samples have similar oligotag integration rates at the target site. To determine whether Phu By amplifying the target locus with Hot-Start Flex Restriction fragment length polymorphism (RFLP) assays were performed. Approximately 150 ng of PCR product was diluted with 20 U Digestion was performed with NdeI (New England BioLabs) at 37°C for 3 hours. Then cleaned up using the Agencourt Ampure XP kit RFLP fragments were analyzed using a Qiaxcel capillary electrophoresis system (Qiagen). The results were quantified to estimate the oligo-tag integration rate. It was also conducted for the same purpose.
[0098] Example 1 One to address targeting specificity of CRISPR-Cas9 RNA-guided gene editing A potential solution to this problem is to engineer Cas9 variants with novel mutations. be.
[0099] Based on these previous results, the specificity of CRISPR-Cas9 nuclease is Nonspecific binding affinity of Cas9 for DNA mediated by binding to phosphate groups on A by reducing the affinity or hydrophobic or base-stacking interactions with DNA. It was assumed (without wishing to be bound by theory) that the The approach involves the use of gRNA / Cas9 complexes, such as the previously described truncated gRNA approach. The advantage of this method is that it does not reduce the length of the target site recognized by the binding. The nonspecific binding affinity of all Cas9s is determined by the amino acid residues that contact the phosphate groups on the target DNA. It was speculated that this could be reduced by mutating the gene.
[0100] Similar approaches for generating variants of non-Cas9 nucleases, e.g., TALENs, are also available. approaches have been used (e.g., Guilinger et al., Nat. Met. thods.11:429(2014)).
[0101] In a first test of this hypothesis, we investigated the role of phosphates in the DNA backbone. Introducing individual alanine substitutions into various residues in SpCas9 that can be predicted to By using the widely used Streptococcus pyogenes (S. pyogenes) Cas9 (S We attempted to engineer a reduced affinity variant of pCas9. The activity of these variants was assessed using an oli)-based screening assay. (Kleinstiver et al.,Nature.2015 Jul 23;5 23(7561):481-5). In this bacterial system, cell viability is determined by the toxic gyrase toxin. The gene for ccdB and its targeting by gRNA and SpCas9 This experiment relied on the cleavage (and subsequent disruption) of a selection plasmid containing a three base pair sequence. The results of the experiment identified residues that retained or lost activity (Table 1).
[0102] [Table 4]
[0103] A survival rate of 50–100% usually indicates robust cleavage, while a survival rate of 0% indicates the enzyme is ineffective. Additional studies assayed in bacteria (but not shown in the table above) demonstrated that the Additional mutations include R69A, R71A, Y72A, R75A, K76A, and N77A. , R115A, H160A, K163A, L169A, T404A, F405A, R44 7A, I448A, Y450A, S460A, M495A, M694A, H698A, Y 1013A, V1015A, R1122A, K1123A, and K1124A. With the exception of R69A and F405A (which had a survival rate of <5% in bacteria), All of the additional single mutations had little effect on the on-target activity of SpCas9. were considered to be non-viral (>70% viability in bacterial screen).
[0104] All possible mutations: N497A, R661A, Q695A, and Q926A Fifteen different SpCas9 variants carrying single, double, triple, and quadruple combinations were analyzed. The contacts made by these residues are then displaceable for on-target activity. For these experiments, we tested whether a single integrated EG Cleavage within the FP reporter gene and insertion or deletion by non-homologous end joining (NHEJ)-mediated repair The induction of a deletion mutation (indel) leads to a loss of cell fluorescence in previously described human cells. A cell-based assay was used (Reyon, D. et al., Nat Biotech Nol. 30, 460-465, 2012). When combined with wild-type SpCas9, EGFP-targeted sgRNAs have been shown to efficiently disrupt EGFP expression in mouse cells. When RNA is used (Fu, Y. et al., Nat Biotechnol 31, 822-826 (2013), all 15 SpCas9 variants were wild-type SpCas 9 (Fig. 1b, gray bar). Substitution of one or all of the residues in the sgRNA allows for the synthesis of SpCas9 using this EGFP-targeting sgRNA. did not reduce the on-target cleavage efficiency of
[0105] Next, we compared the relative activity of all 15 SpCas9 variants at mismatched target sites. To do this, we conducted experiments to evaluate the Positions 7 and 18, and 18 and 19 (the bases most proximal to the PAM) The numbering begins with 20 for the base most distal to the PAM; Figure 1b) A derivative of the EGFP-targeting sgRNA used in previous experiments containing the base pair This analysis revealed that one of the triple mutants (R 661A / Q695A / Q926A) and quadruple mutant (N497A / R661A / Both the mismatched sgRNAs (Q695A / Q926A) were detected in the background using all four of the mismatched sgRNAs. It was revealed that the level of EGFP decay was equivalent to that of the round (Fig. 1b, In particular, among the 15 variants, the one with the lowest activity using mismatched sgRNAs All of the individuals with the mutations carried the Q695A and Q926A. Based on similar data from experiments using sgRNAs for different EGFP target sites, The quadruple mutant (N497A / R661A / Q695A / Q926A) was further analyzed. and cloned it into SpCas9-HF1 (high fidelity variant #1). ) and named it.
[0106] On-target activity of SpCas9-HF1 How robustly SpCas9-HF1 functions at a greater number of on-target sites To determine whether this variant and wild-type SpCas can be expressed using additional sgRNAs, A direct comparison between the 9 sgRNAs was performed. A total of 37 different sgRNAs were tested: 24 13 targeted to EGFP (assayed using an EGFP decay assay) Targeted to the target gene (T7 endonuclease I (T7EI) mismatch assay) (The assay was performed using the EGFP decay assay.) 24 sgRNs tested using the EGFP decay assay 20 in A (Fig. 1c) and 13 sgRNs tested against endogenous human gene sites. Twelve of the A (Fig. 1d) genes were transfected with SpCas9-HF1 using the same sgRNA. The activity of SpCas9 was at least 70% of that of live SpCas9 (Fig. 1e). Cas9-HF1 showed activity highly comparable to wild-type SpCas9 with most sgRNAs. Three of the 37 sgRNAs tested showed Sp Cas9-HF1 showed essentially no activity, and testing of these target sites revealed high activity. did not suggest any obvious differences in their sequence characteristics compared to those seen in Overall, SpCas9-HF1 was effective in 86% (32 / 37) of the sgRNAs tested. ) had comparable activity (>70% of wild-type SpCas9 activity).
[0107] [Table 5]
[0108] [Table 6]
[0109] [Table 7]
[0110] [Table 8]
[0111] [Table 9]
[0112] [Table 10]
[0113] Genome-wide specificity of SpCas9-HF1 We tested whether SpCas9-HF1 exhibited reduced off-target effects in human cells. To test this, we performed a sequencing (GUIDE-seq) method to detect double-strand breaks. GUIDE-seq uses genome-wide unbiased identification of adjacent genome sequences. The double-stranded oligodeoxynucleotides are split into short double-stranded fragments to allow for width and sequencing. Integration of nucleotide (dsODN) tags at any given site is used. The number of tag integrations that occur provides a quantitative measure of cleavage efficiency (Tsai, SQ et al, Nat Biotechnol 33, 187-197(2015)). Endogenous human EMX1, FANCF, RUNX1, and Z were identified using GUIDE-seq. Wild-type mice were infected with eight different sgRNAs targeted to various sites in the SCAN2 gene. The range of off-target effects induced by SpCas9 and SpCas9-HF1 The sequences targeted by these sgRNAs were unique and compared with the reference human genome. The eight sgRNAs have predicted mismatch sites of various lengths in the genome (Table 2). Evaluation of on-target dsODN tag integration (restriction fragment length polymorphism (RFLP)) ) assay) and indel formation (by T7EI assay) Equivalent on-target activity was demonstrated using Cas9 and SpCas9-HF1. The GUIDE-seq experiment identified seven of the eight sgRNAs (Figures 7a and 7b, respectively). However, wild-type SpCas9 was used to target multiple genome-wide off-target sites (e.g., sgRNAs). Induce cleavage in the 8th sgRNA (FANCF domain) for position 4) did not produce any detectable off-target sites ( However, wild-type SpCas9 was used to induce indels. Six of the sgRNAs were GUIDE-seq detectable using SpCas9-HF1. The remaining seventh target events were notably absent (Figures 2a and 2b). The sgRNA (for FANCF site 2) is located at one site within the protospacer seed sequence. Single detectable genome-wide off-target cleavage events at sites harboring mismatches In total, only venting was induced when SpCas9-HF1 was used (Figure 2a). The off-target sites that were not identified were those in the protospacer and / or PAM sequences. The wild-type SpCas9 harbored ~6 mismatches (Fig. 2c). NA (for FANCF site 4) was not observed when tested with SpCas9-HF1. It did not result in any detectable off-target cleavage events (Fig. 2a).
[0114] Using targeted amplicon sequencing to confirm GUIDE-seq findings NHEJ-mediated injury induced by wild-type SpCas9 and SpCas9-HF1 For these experiments, we measured the frequency of del mutations more directly. A- and Cas9-encoding plasmids only (i.e., without GUIDE-seq tags) Next-generation sequencing was then used to conduct GUIDE-seq experiments. 40 off-target genes identified using wild-type SpCas9 for six sgRNAs in Thirty-six of the target sites were tested (four of the 40 sites were isolated from genomic DNA). (These could not be tested because they could not be specifically amplified.) Deep sequencing experiments were performed using (1) wild-type SpCas9 and SpCas9-HF1 induces indels at equal frequencies at each of the six sgRNA on-target sites (2) wild-type SpCas9, as expected, 36 at frequencies that correlate well with GUIDE-seq read counts for the site Statistically significant evidence of indel mutations at 35 off-target sites (Fig. 3b) (3) S at 34 of 36 off-target sites (Fig. 3c); The frequency of indels induced by pCas9-HF1 was significantly higher than that observed in samples from control transfections. The results showed that the levels of indels were indistinguishable from the background levels of indels expected (Figure 3b). Statistically significant mutation frequencies were observed using SpCas9-HF1 compared to the negative control. For the two potential off-target sites, the average indel frequency was 0.049%. and 0.037%, at which level they are sequencing / PCR errors. It is difficult to determine whether these are due to nuclease-induced indels or genuine nuclease-induced indels. Based on these results, SpCas9-HF1 was able to induce different frequencies of cytotoxicity compared with wild-type SpCas9. Reduce off-target mutations, which occur over a range of frequencies, to undetectable levels completely or nearly so. It was concluded that the effect could be almost completely reduced.
[0115] Next, we performed genome-wide analysis of sgRNAs targeting atypical homopolymeric or repetitive sequences. We evaluated the ability of SpCas9-HF1 to reduce the off-target effect. On-target sites with characteristics that are often due to a relative lack of orthogonality to the genome However, SpCas9-HF1 is able to bypass these difficult targets. It may be desirable to investigate whether target indels can be reduced. Therefore, the cytosine-rich homopolymer sequence or multiple TG linkers in the human VEGFA gene either of the PEAT-containing sequences (VEGFA site 2 and VEGFA site 3, respectively) A previously characterized sgRNA targeting otechnol 31,Tsai,SQet al.,Nat Biotechn ol 33, 187-197 (2015) was used (Table 2). Each of these sgRNAs inhibits both wild-type SpCas9 and SpCas9-HF1. Comparable levels of GUIDE-seq dsODN tag incorporation (Figure 7c) and SpCas9-HF1 induced a stochastic mutation (Fig. 7d) in either of these sgRNAs. We demonstrated that on-target activity was not impaired using either of these. Importantly, GUID E-seq experiments revealed that SpCas9-HF1 targeted the off-target genes of their sgRNAs. It was found to be highly effective in reducing VEGFA site 2. 123 / 144 sites and 31 / 32 sites for VEGFA site 3 were examined. Not detected using SpCas9-HF1 (Figures 4a and 4b). Examination of these off-target sites revealed that they were associated with their protospacer and Total series of mismatches within the sgRNA and PAM sequences: 2 to 7 for VEGFA site 2 sgRNA The sgRNAs had one mismatch and the VEGFA site had one to four mismatches. Furthermore, their off-target activity for VEGFA site 2 was shown to be significantly different from that of the control (Fig. 4c). Nine of these are potential bulge bases at the sgRNA-DNA interface (Lin, Y. et al. al,.Nucleic Acids Res 42,7473-7485(2014) (Figure 4a and Figure 8). Sites not detected using SpCas9-HF1 2 to 6 mismatches for VEGFA site 2 sgRNA and VEGFA site There are two mismatches in a single site for the 3sgRNA (Figure 4c), and the VEGFA region Three off-target sites for position 2 sgRNA again had potential bulges ( Taken together, these results demonstrate that SpCas9-HF1 targets simple repeat sequences. can be highly effective in reducing off-target effects of the sgRNA targeted to it, and We demonstrated that the primer sequence can also have a significant effect on the sgRNA targeted to it.
[0116] [Table 11]
[0117] [Table 12]
[0118] [Table 13]
[0119] [Table 14]
[0120] [Table 15]
[0121] [Table 16]
[0122] Improving the specificity of SpCas9-HF1 The transfection was carried out according to previously described methods, for example, truncated gRNA (Fu, Y. et al., Na t Biotechnol 32,279-284(2014)) and SpCas9- D1135E variant (Kleinstiver, BP et al., Nature e 523,481-485(2015)) partially suppressed SpCas9 off-target effects. The present inventors have combined them with SpCas9-HF1 to We considered whether the genome-wide specificity could be further improved. Match full-length and truncated sgR targeted to four sites in the P decay assay Testing SpCas9-HF1 with sgRNAs revealed that shortening the sgRNA complementarity length significantly increased the It was found that the addition of D1 SpCas9-HF1 (referred to herein as SpCas9-HF) carrying the 135E mutation The variants (referred to as 2) were tested using a human cell-based EGFP decay assay. Six of the eight sgRNAs retained more than 70% of the activity of wild-type SpCas9 (Figure 5a and 5b). The side chains on the proximal end of the PAM interact with the target DNA via hydrophobic nonspecific interactions. S carrying the L169A or Y450A mutation, respectively, at the position mediating the effect The pCas9-HF3 and SpCas9-HF4 variants were also generated (Nishima su,H.et al.,Cell 156,935-949(2014);Jiang ,F.,et al.,Science 348,1477-1481(2015)). SpCas9-HF3 and SpCas9-HF4 contain eight EGFP-targeting sgRNAs The same six of these showed over 70% of the activity observed with wild-type SpCas9. was maintained (Figs. 5a and 5b).
[0123] SpCas9-HF2, -HF3, and -HF4 were resistant to SpCas9-HF1. Two off-target sites (FANCF site 2 and VEGFA site 3) were identified in the sgRNA. Further experiments will be performed to determine whether or not the indel frequency in FANCF site 2 carrying a single mismatch in the seed sequence of the protospacer Off-target activity of SpCas9-HF4 was assessed by T7EI assay. As shown, the indel mutation frequency was reduced to near background levels, while the This also beneficially increased targeting activity (Fig. 5c), resulting in the greatest increase in specificity among the three variants. Two protospacer mismatches (one in the seed sequence and one in the VEGFA site 3 off carrying one at the most distal nucleotide from the PAM sequence For the target site, SpCas9-HF2 showed the greatest reduction in indel formation, whereas , which showed only a small effect on on-target mutation frequency (Fig. 5c), and The results showed that the 5-fold increase in specificity between the two variants was the greatest (Fig. 5d). The results suggest that additional residues at other residues may mediate nonspecific DNA contacts or alter PAM recognition. By introducing additional mutations, off-target genes that are resistant to SpCas9-HF1 can be identified. Demonstrate the potential for effect reduction.
[0124] FANCF site 2 and VEGFA site 3 sgRNA off-targets, respectively SpCas9-HF4 and SpCas9-HF2 were found to be highly potent against SpCas9-HF1. To generalize the findings of the T7E1 assay described above, which demonstrates improved discrimination, The genome-wide specificity of these variants was tested using UIDE-seq. Using the LP assay, SpCas9-HF4 and SpCas9-HF2 were isolated using the GUI. SpCas9-H as assayed by DE-seq tag integration rate It was determined to have similar on-target activity to F1 (Figure 5E). When analyzing eq data, for SpCas9-HF2 or SpCas9-HF4 No new off-target sites were identified (Figure 5F). Compared to SpCas9-HF1 As a result, off-target activity at all sites was undetectable by GUIDE-seq. In contrast to SpCas9-HF1, SpCas9-HF4 The single FANCF site 2 on the cytoplasm remains resistant to improving the specificity of SpCas9-HF1. SpCas9 had approximately 26-fold better specificity for the target site (Figure 5F). HF2 suppresses SpCas9-HF1 against the frequent VEGFA site 3 off-target had nearly four-fold improved specificity at GU1 at other low-frequency off-target sites, while IDE-seq detectable events were also significantly reduced (>38-fold) or eliminated. Remarkably, three of the rare sites identified for SpCas9-HF1 were genomic DNA fragments. The cytoplasmic position is adjacent to a previously characterized breakpoint hotspot in background U2OS cells. Taken together, these results demonstrate the efficacy of SpCas9-HF2 and SpCas9-HF4. These results suggest that the variants may improve the genome-wide specificity of SpCas9-HF1.
[0125] SpCas9 when using sgRNAs designed against standard, non-repetitive target sequences -HF1 robustly and consistently reduced off-target mutations. The two off-target sites most tolerant to F1 were 1 and 2 in the protospacer. Taken together, these observations support the use of SpCas9-HF1 to closely related clones carrying one or two mismatches anywhere in the genome. By targeting non-repetitive sequences that do not have a site (existing publicly available software) Program (Bae, S., et al., Bioinformatics 30, 14 73-1475(2014)), which can be easily achieved using This suggests that the nucleotide mutations can be minimized to undetectable levels. One parameter is the mismatch of SpCas9-HF1 to the protospacer sequence. This may not be compatible with the common approach of using a G at the 5' end of the gRNA. Testing four sgRNAs carrying mismatched 5'Gs to the target site yielded four cytoplasmic RNAs. Three of these showed reduced activity with SpCas9-HF1 compared to wild-type SpCas9. (Figure 10) and SpCas9-HF1, which better discriminates between partial match sites. It may reflect ability.
[0126] Further biochemical testing will confirm that SpCas9-HF1 achieves its high genome-wide specificity. The precise mechanism by which the four mutations introduced increase SpCa expression in cells can be confirmed or clarified. It is not apparent that this alters the stability or steady-state expression level of s9. Titration experiments using decreasing concentrations of the expression plasmid demonstrated that wild-type SpCas9 and This suggests that SpCas9-HF1 and SpCas9-HF1 behave comparably as their concentrations decrease. Instead, the simplest mechanistic explanation is that these mutations just enough to retain target activity but not to cleave off-target sites inefficiently or The energy of the complex is so low that it is non-existent. This mechanism is based on structural data. This is consistent with the nonspecific interactions observed between the mutated residues in the target DNA and the phosphate backbone. (Nishimasu, H. et al., Cell 156, 935-949(2 014);Anders,C et.Al.,Nature 513,569-573( 2014). A somewhat similar mechanism is reported for transcriptional activators carrying substitutions at positively charged residues. has been proposed to explain the increased specificity of activator-like effector nucleases ( Guilinger,JPet al.,Nat Methods 11,429- 435(2014)).
[0127] SpCas9-HF1 was combined with other mutations shown to alter Cas9 function. For example, Sp carrying three amino acid substitutions Cas9 mutant (D1135V / R1335Q / T1337R, SpCas9-VQ R variant) is a NGAN PAM (NGAG>NGAT=NGAA>N (Kleinstiver, BP et al, Nature 523, 481-485 (2015)), recently identified The quadruple SpCas9 mutant (D1135V / G1218R / R1335Q / T1 337R, referred to as the SpCas9-VRQR variant) is a NGAH (H = A, C, or T) have improved activity against VQR variants at sites with PAM (Fig. 12a). SpCas9-HF1 to SpCas9-VQR and SpCas9- Four mutations (N497A / R661A / Q695A / Q926A) in VRQR The transfections were SpCas9-VQR-HF1 and SpCas9-VRQR-HF1, respectively. Both of these HF versions of nucleases express an EGFP reporter gene. Using five of the eight sgRNAs targeted to the chromosome and to an endogenous human gene site Seven of the eight sgRNAs synthesized were used to generate sgRNAs equivalent to their non-HF counterparts (i.e., 70 % or more) of on-target activity (Figures 12b to 12d).
[0128] Considered more broadly, these results support the development of additional high fidelity CRISPR-associated nucleases. General guidelines for genetic manipulation of diversity variants are described. Adding additional mutations in the gene fragments allows for the elution of surviving polarized regions using SpCas9-HF1. This further reduced the small number of remaining off-target sites. For example, SpCas9-HF2, SpCas9-HF3, SpCas9-HF4, etc. , can be utilized in a customized manner depending on the nature of the off-target sequence. Successful genetic engineering of high-fidelity variants of as9 has demonstrated the ability to inhibit nonspecific DNA contacts. The mutagenic approach has been applied to other naturally occurring and engineered Cas9 orthologues (Ran ,FAet al.,Nature 520,186-191(2015),Esv elt, KMet al., Nat Methods 10, 1116-1121( 2013);Hou, Z. et al., Proc Natl Acad Sci US A(2013);Fonfara,I.et al.,Nucleic Acids R es 42,2577-2590(2014);Kleinstiver,BPet al, Nat Biotechnol (2015) and found with increasing frequency and Newer CRISPR-associated nucleases have been characterized (Zetsche, Be t al.,Cell 163,759-771(2015);Shmakov,Se al., Molecular Cell 60, 385-397) This suggests that...
[0129] Example 2 Alanine substitutions in residues that contact target strand DNA, e.g., N497A, Q695A, SpCas9 variants having R661A and Q926A are described herein. In addition to these residues, the inventors have also investigated their variants, e.g., SpCas9-HF The specificity of one variant (N497A / R661A / Q695A / Q926A) was confirmed by comparing the non-targeted Positively charged SpCas9 residues that likely contact the DNA strand: R780, K810, R83 2, K848, K855, K968, R976, H982, K1003, K1014, K 1047, and / or R1060 (Slaymaker et al., Science ce.2016 Jan 1;351(6268):84-8) We sought to determine whether further improvements could be made by
[0130] Wild-type SpC carrying single alanine substitutions at these positions and their combinations The activity of as9 derivatives was evaluated using a perfect match sgRNP designed against a site in the EGFP gene. A (to assess on-target activity) and positions 11 and 12 (position 1 is the most PA The same sgRNA (O) carrying an intentional mismatch in the base proximal to M (to assess activity at mismatch sites found in the target site) The triple-substituted K810 was first tested using an EGFP decay assay (Figure 13A). Carries A / K1003A / R1060A or K848A / K1003A / R1060A The corresponding derivatives are designated eSpCas9(1.0) and eSpCas9(1.1), respectively. (Note that this is identical to a recently described variant known in the literature; see reference 1). As expected, wild-type SpCas9 exhibits robust on-target and mismatch properties. As a control, SpCas9-HF1 was also tested in this experiment. As expected, this maintains on-target activity while reducing mismatch target activity. We found that the 1-terminal nucleotide sequence at a position that could potentially contact the non-target DNA strand was significantly different from the 1-terminal nucleotide sequence (Figure 13A). All of the wild-type SpCas9 derivatives carrying one or more alanine substitutions were significantly different from wild-type SpCas 9 showed comparable on-target activity (Figure 13A). A portion of the mismatches (11 / 12) were significantly lower than the activity observed with wild-type SpCas9. We also demonstrate reduced cleavage using sgRNAs, with a subset of substitutions in those derivatives being significantly lower than wild-type. This result shows that SpCas9 confers improved specificity for this mismatch site. However, these single substitutions and combinations of substitutions resulted in 11 / 12 mistakes. This was not sufficient to completely eliminate the activity observed with the matched sgRNA. Wild-type SpCas was generated using an additional sgRNA carrying a mismatch at positions 10 and 11. 9, SpCas9-HF1, and their identical wild-type SpCas9 derivatives were tested. In the case of β-actin (Figure 13B), only minimal changes in mismatch target activity were observed for most derivatives. Again, this is due to the presence of single, double, or double cleavage sites at these potential non-target strand contact residues. or even triple substitutions (already described eSpCas9(1.0) and (1.1) variants). (equivalent to a variant) is insufficient to eliminate activity at mismatched DNA sites. Taken together, these data demonstrate that wild-type SpCas9 variants Retaining on-target activity with matched sgRNA and (wild-type SpCa (Regarding s9) The substitutions contained in themselves and their derivatives are two different mismatches. This demonstrates that the nuclease activity against the target DNA site is not sufficient to eliminate the nuclease activity (Figure 1). 13A and 13B).
[0131] Considering these results, one or more additions at residues that can contact the non-target DNA strand SpCas9-HF1 derivatives carrying the amino acid substitutions are expressed as parent SpCas9-HF1 proteins. It was hypothesized that the specificity for proteins could be further improved. Various SpCas9-HF1 derivatives carrying combinations of triple alanine substitutions were analyzed using the complete Matched sgRNA (to test for on-target activity) and at positions 11 and 12 The same sgRNA carrying a mismatch at the off-target site (the mismatch found for the off-target site) Human cell-based EGFP disruption using Smamatch (to assess activity at target sites) These sgRNAs were tested in the assay. These sgRNAs were the same as those used in Figures 13A-B. This experiment demonstrated that most of the SpCas9-HF1 derivative variants tested was similar to that observed with both wild-type SpCas9 and SpCas9-HF1. It was revealed that the 11 / 12 mismatches exhibited on-target activity (Fig. 14A). Using sgRNAs, some of the tested SpCas9-HF1 derivatives (e.g., SpC as9-HF1+R832A and SpCas9-HF1+K1014A) Using the sgRNA did not show any discernible change in cleavage. Most of the SpCas9-HF1 derivatives were synthesized using the 11 / 12 mismatched sgRNA. SpCas9-HF1, eSpCas9(1.0), or eSpCas9(1.1) A set of these new variants had significantly lower activity than that observed with The combination has reduced mismatch targeting activity and therefore improved specificity. This suggests that mismatch targeting activity was enhanced using 11 / 12 mismatched sgRNAs (Figure 14A). Of the 16 SpCas9-HF1 derivatives that reduced expression to near background levels, Nine were deemed to have only a minimal effect on on-target activity (perfect match sgRNAs were used (Figure 14A). These SpCas9-H in EGFP decay assays using sgRNAs Further testing of a subset of F1 derivatives (Figure 14B) confirmed that these variants were Using matched sgRNAs, SpCas9-HF1 (Figure 14b), eSpCas9 (1 .1) (Figure 13A), or the same substitution added to the wild-type SpCas9 nuclease (Figure 13B) were also found to have lower activity than that observed with either Importantly, five variants were identified in this approach using 9 / 10 mismatched sgRNAs. The antibody showed background levels of off-target activity in the assay.
[0132] These alanine substitutions in the non-target strand were then used to identify the target site of interest in our SpCas9-HF1 barrier. SpCas9 variants containing only the Q695A and Q926A substitutions from the control (this Here, we tested whether it could be combined with the "double" variant. Many of the tested HF1 derivatives showed observable (and undesirable) reductions in on-target activity. Therefore, only the two most significant substitutions from SpCas9-HF1 (Q695A and Q92 6A; see Figure 1B) with one or more non-target strand contact substitutions. These substitutions can rescue the activity of the SpCas9-HF1 variant, but adding them to the SpCas9-HF1 variant It was hypothesized that the increased specificity observed when Sets of single, double, or triple alanine substitutions at potential non-target DNA strand interaction positions Various SpCas9(Q695A / Q926A) derivatives carrying the combination are shown in Figures 13A-B The same perfect match sgRNA (on-target) targeted to EGFP as above was used in (to test the cloning activity) and the same clone carrying mismatches at positions 11 and 12. sgRNA (activity at mismatched target sites found at off-target sites) This was tested in a human cell-based EGFP decay assay using Experiments showed that most of the tested SpCas9(Q695A / Q926A) derivative variants Most of these results were similar to those observed with both wild-type SpCas9 and SpCas9-HF1. These results revealed that Sp1000 and Sp2000 exhibited comparable on-target activity (Figure 15). Many of the Cas9-HF1 derivatives were engineered to express SpCa using 11 / 12 mismatched sgRNAs. Observation was performed using s9-HF1, eSpCas9(1.0), or eSpCas9(1.1). Some combinations of these new variants have significantly lower activity compared to previously observed suggest that the α-glucan-1-phosphate dehydrogenase (GAD) has reduced mismatch targeting activity and therefore improved specificity. Mismatch targeting activity was approximately 11 / 12 mismatched sgRNAs (Figure 15). 13 SpCas9 (Q695A / Q926A) reduced to background levels Of the derivatives, only one appeared to have a significant effect on on-target activity. (Assessed using perfect match sgRNAs; Figure 15).
[0133] Overall, these data support the use of SpCas9- in positions that can contact non-target DNA strands. One, two, or three alanines to HF1 or SpCas9(Q695A / Q926A) The addition of the substitutions was performed using either the parental clone or the recently described eSpCas9(1.0) or (vs. 1.1)) improved ability to discriminate between mismatched off-target sites. Importantly, we demonstrate that wild-type SpCa can be used to generate new variants with the same phenotype. These same substitutions for s9 are not expected to provide any substantial specificity benefit. I can't get it.
[0134] SpCas9-HF1 against mismatches at the sgRNA-target DNA complementarity interface To better define and compare the tolerance of eSpCas9-1.1 and eSpCas9-1.2, we used spacers - sgRNAs containing single mismatches at all possible positions in the complementary region Their activity was tested using SpCas9-HF1 and eSPCas9-1.1. Both variants have the most single-mismatched sgRNAs compared to wild-type SpCas9 with some exceptions, SpCas9-HF1 has similar activity to eSpCas9- exceeded 1.1 (Figure 16).
[0135] Next, we developed a double mutant (Db = Q695A / Q926A), SpCas9-HF1(N 497A / R661A / Q695A / Q926A), eSpCas9-1.0(1.0= K810A / K1003A / R1060A), or eSpCas9-1.1 (1.1 = Amino acid substitutions from either of the target strands (K848A / K1003A / R1060A) and target strand D Additional alanine substitutions in residues that contact NA or potentially contact non-target strand DNA. The single nucleotide mismatch tolerance of some variants containing the combination of Figure 17A-B). Perfect match sgRNAs were used to assess on-target activity, while Carrying such mismatches at positions 4, 8, 12, or 16 in the spacer sequence Single-nucleotide mismatch tolerance was assessed using sgRNAs containing the nucleotides (Figure 17A). A number of these variants maintained on-target activity and were successfully treated with mismatched sgRNAs. The observed activity was significantly reduced. Three of these variants (Q695A / K848A / Q926A / K1003A / R1060A, N497A / R661A / Q695A / K 855A / Q926A / R1060A, and N497A / R661A / Q695A / Q 926A / H982A / R1060A) to the remaining single-mismatched sgRNAs (1–3, containing mismatches at positions 5-7, 9-11, 13-15, and 17-20) These variants were further tested using sg compared to eSpCas9-1.1. We demonstrate robust intolerance to single nucleotide substitutions in RNA and identify novel The improved specificity profile of the variants was demonstrated (Figure 17B). Additional variant nucleases containing substitutions were added at positions 5, 7, and 9 in the spacer. The test was performed using an sgRNA containing a mismatch at position 1 (the conventional variant These specific mismatches were expected to be tolerated. (Figure 18). Many of these nucleases target mismatched sites. The on-target activity was only slightly reduced ( Figure 18).
[0136] To further determine whether additional combinations of mutations could confer improved specificity, A greatly expanded panel of nuclease variants using two additional matched sgRNAs We tested the EGFP-degrading activity of the GFP-antigens by using the ELISA kit (Fig. 19A). Many of these variants maintain robust on-target activity, and their These results suggest that this may be useful in generating further improvements to specificity (Figure 19B). A number of these variants were identified as containing a single substitution at position 12, 14, 16, or 18. Test using sgRNAs with different positions to determine whether improved specificity was observed. It was found that α- and β-nucleotides exhibit a greater intolerance to single nucleotide mismatches in (Figure 19B).
[0137] Example 3 Similar to what was done with SpCas9, By adopting a strategy of using SaCas9 (SaCas9) , residues known to contact the target DNA strand (Figures 20 and 21A), residues known to contact non-target DNA strands (Figures 21B and 21C), The residues that can be contacted (experiments in progress) and that we have already shown can affect PAM specificity We improved the specificity of SaCas9 by introducing alanine substitutions in the selected residues (Figure 21B). The residues that can contact the DNA backbone of the target strand include Y211, Y212, W229, Y230, R245, T392, N419, L446, Y651, and R654; residues that may contact non-target strand DNA include Q848, N4 92, Q495, R497, N498, R499, Q500, K518, K523, K5 25, H557, R561, K572, R634, R654, G655, N658, S6 62, N668, R686, K692, R694, H700, K751; PA Residues that contact M include E782, D786, T787, Y789, T882, and K8. 86, N888, A889, L909, K929, N985, N986, R991, and In preliminary experiments, the target strand DNA contact residues or PAM contact residues were identified. Single alanine substitutions (or some combinations thereof) in residues (Figures 21A and B, respectively) ) have variable effects on on-target EGFP decay activity (perfect match sgRNA (using 11 and 12), off-target cleavage could not be excluded (mismatches at positions 11 and 12). Interestingly, SpCas9 mutations in HF1 completely abolish off-target activity using similar mismatched target / sgRNA pairs and contain a combination of target / non-target strand displacement (as observed with SpCas9). We suggest that variants that enhance specificity at such sites may be necessary. suggested.
[0138] Further strategies for improving SaCas9 specificity by mutating potential target strand DNA contacts are discussed. To evaluate the effectiveness of the mismatch-tolerant mutations at positions 19 and 20 in the sgRNA, The potential of different single, double, triple, and quadruple combinations was tested (Figures 22A and B). These combinations allow alanine substitutions at Y230 and R245 to be combined with other substitutions. When combined, the specificity is enhanced as judged by the ability to better discriminate between mismatched sites. It has been shown that it can increase sexual activity.
[0139] Next, two of these triple alanine substitution variants (Y211A / Y230A / R24 5A and Y212A / Y230A / R245A) to determine the on-target gene disruption activity of Four on-target sites in EGFP were tested (match sites #1-4; Figure 23). These variants maintain robust on-target activity for match sites 1 and 2. However, sites 3 and 4 showed approximately 60-70% loss of on-target activity. Both of these triple alanine substitution variants were substituted with various nucleotides in the spacer of target sites 1-4. As determined by using an sgRNA carrying a double mismatch at position This significantly improved specificity over wild-type SaCas9 (Figure 23).
[0140] These double and triple combinations of alanine substitutions (Figures 24A and B, respectively) were carried SaCas9 variants were tested against six endogenous sites for on-target activity. However, position 21 (the most distal position from the PAM) is predicted to be a mismatch that is difficult to distinguish. ) was used to assess improved specificity. In some examples, on-target activity using the matched sgRNA is maintained using the variant. While the sgRNA mismatch at position 21 was used, the "off-target" activity was not observed. In other cases, only a small amount of activity was detected using the matched sgRNA (Figure 24A and B). Losses to complete losses were observed.
[0141] References 1.Sander,JD&Joung,JKCRISPR-Cas syst ems for editing, regulating and targeting genomes.Nat Biotechnol 32,347-355(2014) . 2. Hsu, PD, Lander, ES & Zhang, F. Developm ent and applications of CRISPR-Cas9 for genome engineering.Cell 157,1262-1278(20 14). 3.Doudna,JA&Charpentier,E.Genome edit ing.The new frontier of genome engineeri ng with CRISPR-Cas9.Science 346,1258096( 2014). 4.Barrangou,R.&May,A.P.Unraveling the p otential of CRISPR-Cas9 for gene therapy .Expert Opin Biol Ther 15,311-314(2015). 5.Jinek,M.et al.A programmable dual-RNA -guided DNA endonuclease in adaptive bac terial immunity.Science 337,816-821(2012 ). 6.Sternberg,S.H.,Redding,S.,Jinek,M.,Gr eene,E.C.&Doudna,J.A.DNA interrogation b y the CRISPR RNA-guided endonuclease Cas 9.Nature 507,62-67(2014). 7.Hsu,P.D.et al.DNA targeting specifici ty of RNA-guided Cas9 nucleases.Nat Biot echnol 31,827-832(2013). 8.Tsai,S.Q.et al.GUIDE-seq enables geno me-wide profiling of off-target cleavage by CRISPR-Cas nucleases.Nat Biotechnol 33,187-197(2015). 9.Hou,Z.et al.Efficient genome engineer ing in human pluripotent stem cells usin g Cas9 from Neisseria meningitidis.Proc Natl Acad Sci USA(2013). 10.Fonfara,I.et al.Phylogeny of Cas9 de termines functional exchangeability of d ual-RNA and Cas9 among orthologous type II CRISPR-Cas systems.Nucleic Acids Res 42,2577-2590(2014). 11.Esvelt,K.M.et al.Orthogonal Cas9 pro teins for RNA-guided gene regulation and editing.Nat Methods 10,1116-1121(2013). 12.Cong,L.et al.Multiplex genome engine ering using CRISPR / Cas systems.Science 3 39,819-823(2013). 13.Horvath,P.et al.Diversity,activity,a nd evolution of CRISPR loci in Streptoco ccus thermophilus.J Bacteriol 190,1401-1 412(2008). 14.Anders,C.,Niewoehner,O.,Duerst,A.&Ji nek,M.Structural basis of PAM-dependent target DNA recognition by the Cas9 endon uclease.Nature 513,569-573(2014). 15.Reyon,D.et al.FLASH assembly of TALE Ns for high-throughput genome editing.Na t Biotechnol 30,460-465(2012). 16.Fu,Y.et al.High-frequency off-target mutagenesis induced by CRISPR-Cas nucle ases in human cells.Nat Biotechnol 31,82 2-826(2013). 17.Chen,Z.&Zhao,H.A highly sensitive se lection method for directed evolution of homing endonucleases.Nucleic Acids Res 33,e154(2005). 18.Doyon,J.B.,Pattanayak,V.,Meyer,C.B.& Liu,D.R.Directed evolution and substrate specificity profile of homing endonucle ase I-SceI.J Am Chem Soc 128,2477-2484(2 006). 19.Jiang,W.,Bikard,D.,Cox,D.,Zhang,F.&M arraffini,L.A.RNA-guided editing of bact erial genomes using CRISPR-Cas systems.N at Biotechnol 31,233-239(2013). 20.Mali,P.et al.RNA-guided human genome engineering via Cas9.Science 339,823-82 6(2013). 21.Hwang,W.Y.et al.Efficient genome edi ting in zebrafish using a CRISPR-Cas sys tem.Nat Biotechnol 31,227-229(2013). 22.Chylinski,K.,Le Rhun,A.&Charpentier, E.The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems.RNA Biol 10,726-737(2013). 23.Kleinstiver,B.P.,Fernandes,A.D.,Gloo r,G.B.&Edgell,D.R.A unified genetic,comp utational and experimental framework ide ntifies functionally relevant residues o f the homing endonuclease I-BmoI.Nucleic Acids Res 38,2411-2427(2010). 24.Gagnon,J.A.et al.Efficient mutagenes is by Cas9 protein-mediated oligonucleot ide insertion and large-scale assessment of single-guide RNAs.PLoS One 9,e98186( 2014).
[0142] array SEQ ID NO: 271 - JDS246:CMV-T7-human SpCas9-NLS-3xF LAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, N LS, double underline, 3xFLAG tag, bold: [ka]
[0143] SEQ ID NO: 272 - VP12:CMV-T7-humanSpCas9-HF1(N497A , R661A, Q695A, Q926A)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0144] SEQ ID NO: 273 - MSP2135:CMV-T7-human SpCas9-HF2(N4 97A, R661A, Q695A, Q926A, D1135E)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0145] SEQ ID NO: 274 - MSP2133:CMV-T7-human SpCas9-HF4(Y4 50A, N497A, R661A, Q695A, Q926A)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0146] SEQ ID NO: 275 - MSP469:CMV-T7-human SpCas9-VQR(D11 35V, R1335Q, T1337R)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0147] SEQ ID NO: 276 - MSP2440:CMV-T7-HumanSpCas9-VQR-HF 1(N497A, R661A, Q695A, Q926A, D1135V, R1335Q, T1337R)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0148] SEQ ID NO: 277 - BPK2797: CMV-T7-HumanSpCas9-VRQR(D 1135V, G1218R, R1335Q, T1337R)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0149] SEQ ID NO: 278 - MSP2443:CMV-T7-HumanSpCas9-VRQR-H F1(N497A, R661A, Q695A, Q926A, D1135V, G1218R , R1335Q, T1337R)-NLS-3xFLAG Human codon-optimized Streptococcus pyogenes (S. pyogenes) Cas9, regular font, modified Variant codon, lowercase, NLS, double underline, 3xFLAG tag, bold: [ka]
[0150] SEQ ID NO: 279 - BPK1520: U6-BsmBI cassette-Sp-sgRNA U6 promoter, normal font, BsmBI site, italics, Streptococcus pyogenes (S. pyogenes) sgRNA, lowercase, U6 terminator, double underline: [ka]
[0151] Other embodiments While the present invention has been described with reference to its detailed description, the above description is not intended to be limiting unless the appended claims are incorporated herein by reference. This is illustrative, but not limiting, of the scope of the invention, which is defined by the range of It should be understood that other aspects, advantages, and modifications are within the scope of the following claims. be. Examples of inventions based on the above disclosure include the following. [1] At the following locations: L169, Y450, N497, R661, Q695, Q926, and and / or mutations in 1, 2, 3, 4, 5, 6, or all 7 of D1135 and preferably at the following positions: L169, Y450, N497, R661, Q695, Mutations at 1, 2, 3, 4, 5, 6, or 7 of Q926, D1135 A sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1 and optionally a nucleic acid sequence isolated suppurative cells comprising one or more of a localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag. Streptococcus pyogenesCas9(SpCas9 )protein. [2] One, two, three, or four of the following: N497, R661, Q695, and Q926 All mutations, preferably the following: N497A, R661A, Q69 5A, and Q926A, and one, two, three, or all four of the isolated protein according to [1]. Quality. [3] Q695 and / or Q926 and optionally L169 , Y450, N497, R661, and D1135, one, two, three, four, or all five of them Mutations in the nucleotide sequence Y450A / Q695A, ... 695A / Q926A, Q695A / D1135E, Q926A / D1135E, Y45 0A / D1135E, L169A / Y450A / Q695A, L169A / Q695A / Q926A, Y450A / Q695A / Q926A, R661A / Q695A / Q926 A, N497A / Q695A / Q926A, Y450A / Q695A / D1135E, Y 450A / Q926A / D1135E, Q695A / Q926A / D1135E, L16 9A / Y450A / Q695A / Q926A, L169A / R661A / Q695A / Q 926A, Y450A / R661A / Q695A / Q926A, N497A / Q695A / Q926A / D1135E, R661A / Q695A / Q926A / D1135E, or or Y450A / Q695A / Q926A / D1135E, Plagiarism. [4] N14;S15;S55;R63;R78;H160;K163;R165;L1 69;R403;N407;Y450;M495;N497;K510;Y515;W6 59;R661;M694;Q695;H698;A728;S730;K775;S7 77;R778;R780;K782;R783;K789;K797;Q805;N8 08;K810;R832;Q844;S845;K848;S851;K855;R8 59;K862;K890;Q920;Q926;K961;S964;K968;K9 74;R976;N980;H982;K1003;Y1013;K1014;V101 5;S1040;N1041;N1044;K1047;K1059;R1060;K1 107;E1108;S1109;K1113;R1114;S1116;K1118; R1122;K1123;K1124;D1135;S1136;K1153;K115 5;K1158;K1200;Q1221;H1241;Q1254;Q1256;K1 289;K1296;K1297;R1298;K1300;H1311;K1325; Mutations at K1334; T1337 and / or S1216, preferably N4 97A / R661A / Q695A / Q926A / K810A, N497A / R661A / <h2 style=";text-align:left;direction:ltr">Q695A / Q926A / K848A、N497A / R661A / Q695A / Q926<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / K855A、N497A / R661A / Q695A / Q926A / R780A、N4<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 97A / R661A / Q695A / Q926A / K968A、N497A / R661A / <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q695A / Q926A / H982A、N497A / R661A / Q695A / Q926<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / K1003A、N497A / R661A / Q695A / Q926A / K1014A、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> N497A / R661A / Q695A / Q926A / K1047A、N497A / R66<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1A / Q695A / Q926A / R1060A、N497A / R661A / Q695A / <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q926A / K810A / K968A、N497A / R661A / Q695A / Q926<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / K810A / K848A、N497A / R661A / Q695A / Q926A / K8<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 10A / K1003A、N497A / R661A / Q695A / Q926A / K810A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / R1060A、N497A / R661A / Q695A / Q926A / K848A / K1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 003A、N497A / R661A / Q695A / Q926A / K848A / R1060<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A、N497A / R661A / Q695A / Q926A / K855A / K1003A、N<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 497A / R661A / Q695A / Q926A / K855A / R1060A、N497<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / R661A / Q695A / Q926A / K968A / K1003A、N497A / R<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 661A / Q695A / Q926A / H982A / K1003A、N497A / R661<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / Q695A / Q926A / H982A / R1060A、N497A / R661A / Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 695A / Q926A / K1003A / R1060A、N497A / R661A / Q69<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5A / Q926A / K810A / K1003A / R1060A、N497A / R661A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / Q695A / Q926A / K848A / K1003A / R1060A、Q695A / Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 926A / R780A、Q695A / Q926A / K810A、Q695A / Q926A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> / R832A、Q695A / Q926A / K848A、Q695A / Q926A / K85<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5A, Q695A / Q926A / K968A, Q695A / Q926A / R976A, Q<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 695A / Q926A / H982A、Q695A / Q926A / K1003A、Q695<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / Q926A / K1014A、Q695A / Q926A / K1047A、Q695A / <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q926A / R1060A、Q695A / Q926A / K848A / K968A、Q69<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 5A / Q926A / K848A / K1003A、Q695A / Q926A / K848A / <h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> K855A, Q695A / Q926A, K848A, H982A, Q695A / Q926<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / K1003A / R1060A、Q695A / Q926A / R832A / R1060A<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q695A / Q926A / K968A / K1003A Q695A / Q926A / K9<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 68A / R1060A、Q695A / Q926A / K848A / R1060A、Q695<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / Q926A / K855A / H982A、Q695A / Q926A / K855A / K1<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 003A、Q695A / Q926A / K855A / R1060A、Q695A / Q926<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> A / H982A / K1003A、Q695A / Q926A / H982A / R1060A、<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> Q695A / Q926A / K1003A / R1060A、Q695A / Q926A / K8<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 10A / K1003A / R1060A、Q695A / Q926A / K1003A / K10<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 47A / R1060A, Q695A / Q926A / K968A / K1003A / R106<h2 style=";text-align:left;direction:ltr"> 0A, Q695A / Q926A / R832A / K1003A / R1060A, or Q6 95A / Q926A / K848A / K1003A / R1060A, as described in [1] The isolated protein is described. [5] The following mutations: D1135E; D1135V; D1135V / R1335Q / T 1337R (VQR variant); D1135E / R1335Q / T1337R (EQR Variant); D1135V / G1218R / R1335Q / T1337R (VRQR variant) riant); or D1135V / G1218R / R1335E / T1337R(VRE The isolated protein according to [1], further comprising one or more of: [6] D10, E762, D839, H983, or D986; and H840 or 1, which reduces nuclease activity, selected from the group consisting of mutations at N863 and N863. The isolated protein described in [1], further comprising one or more mutations. [7] The mutation that reduces nuclease activity is (i) D10A or D10N, and (ii) H840A, H840N, or H840Y [6] An isolated protein according to [6]. [8] At the following locations: Y211, Y212, W229, Y230, R245, T392, N 419, Y651, R654, or one or more mutations; Preferably, at the following positions: Y211, Y212, W229, Y230, R245, T39 2. Sudden onset of 1, 2, 3, 4, or 5, 6 or more of N419, Y651, R654 A sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1 with a mutation, as well as any Optionally, one or more of a nuclear localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag Staphylococcus aureus Cas9 (S) isolates aCas9) protein. [9] The following mutations: Y211A, Y212A, W229, Y230A, R245A, containing one or more of T392A, N419A, Y651, and / or R654A.[8] 2. The isolated protein according to claim 1 .
[10] Mutations at N419 and / or R654 and optionally additional 1, 2, for mutations Y211, Y212, W229, Y230, R245, and T392 3, 4 or more, preferably N419A / R654A, Y211A / R654A, Y21 1A / Y212A, Y211A / Y230A, Y211A / R245A, Y212A / Y 230A, Y212A / R245A, Y230A / R245A, W229A / R654A , Y211A / Y212A / Y230A, Y211A / Y212A / R245A, Y21 1A / Y212A / Y651A, Y211A / Y230A / R245A, Y211A / Y 230A / Y651A, Y211A / R245A / Y651A, Y211A / R245A / R654A, Y211A / R245A / N419A, Y211A / N419A / R65 4A, Y212A / Y230A / R245A, Y212A / Y230A / Y651A, Y 212A / R245A / Y651A, Y230A / R245A / Y651A, R245A / N419A / R654A, T392A / N419A / R654A, R245A / T39 2A / N419A / R654A, Y211A / R245A / N419A / R654A, W 229A / R245A / N419A / R654A, Y211A / R245A / T392A / N419A / R654A, or Y211A / W229A / R245A / N419A / The isolated protein according to [8], which contains R654A.
[11] Y211;Y212;W229;Y230;R245;T392;N419;L 446;Q488A;N492A;Q495A;R497A;N498A;R499;Q 500;K518;K523;K525;H557;R561;K572;R634;Y 651;R654;G655;N658;S662;N667;R686;K692;R 694;H700;K751;D786;T787;Y789;T882;K886;N 888;889;L909;N985;N986;R991;R1015;N44;R4 5;R51;R55;R59;R60;R116;R165;N169;R208;R2 09;Y211;T238;Y239;K248;Y256;R314;N394;Q4 14;K57;R61;H111;K114;V164;R165;L788;S790 ;R792;N804;Y868;K870;K878;K879;K881;Y897 and / or K906. protein.
[12] of the following mutations: E782K, K929R, N968K, or R1015H One or more, specifically E782K / N968K / R1015H (KKH variant); E782K / K929R / R1015H (KRH variant); or E782K / K9 [8], further comprising 29R / N968K / R1015H (KRKH variant). Isolated proteins.
[13] D10, E477, D556, H701, or D704; and H557 or is selected from the group consisting of a mutation at N580, which reduces nuclease activity The isolated protein according to [8], further comprising one or more mutations.
[14] The mutation is: (i) D10A or D10N, and / or (ii) H557A, H557N, or H557Y, and / or (iii) N580A, and / or (iv) D556A
[13] The isolated protein according to
[13] ,
[15] fused to a heterologous functional domain via an optional intervening linker [1]-
[14] 10. A fusion protein comprising the isolated protein of any one of claims 1 to 9, wherein the linker A fusion protein that does not interfere with the activity of said fusion protein.
[16] The fusion protein according to
[15] , wherein the heterologous functional domain is a transcription activation domain. Quality.
[17] The transcriptional activation domain is from VP64 or NF-κB p65. The fusion protein according to
[16] .
[18] The fusion protein described in
[0015] , wherein the heterologous functional domain is a transcriptional silencer or transcriptional repression domain.
[19] The transcriptional repression domain is a Kruppel-associated box (KRAB) domain, an ER domain, F repressor domain (ERD), or mSin3A interaction domain (SID) The fusion protein described in
[18] .
[20] The transcriptional silencer is heterochromatin protein 1 (HP1), preferably The fusion protein according to
[18] , which is HP1α or HP1β.
[21] The heterologous functional domain is an enzyme that modifies the methylation status of DNA.
[15] The fusion protein according to claim 1.
[22] The enzyme that modifies the methylation state of the DNA is a DNA methyltransferase. The fusion protein according to
[21] , which is a DNMT or TET protein.
[23] The fusion protein according to
[22] , wherein the TET protein is TET1.
[24] The heterologous functional domain is an enzyme that modifies a histone subunit.
[15] The fusion protein according to claim 1.
[25] The enzyme that modifies histone subunits is histone acetyltransferase. Histone deacetylases (HATs), histone deacetylases (HDACs), histone methyltransferases The fusion protein according to
[15] , which is a histone demethylase (HMT).
[26] The fusion protein according to
[15] , wherein the heterologous functional domain is a biological tether. .
[27] The fusion protein described in
[0026] , wherein the biological tether is MS2, Csy4 or lambda N protein.
[28] The fusion protein according to
[26] , wherein the heterologous functional domain is FokI.
[29] An isolated nucleic acid encoding the protein according to any one of [1] to
[14] .
[30] Optionally, a method for expressing the protein according to any one of [1] to
[24] The isolated nucleic acid of
[29] is operably linked to one or more regulatory domains for Vector.
[31] The nucleic acid according to
[29] , and optionally the nucleic acid according to any one of [1] to
[14] . A host cell, preferably a mammalian host cell, that expresses the protein.
[32] A method for modifying the genome or epigenome of a cell, comprising:
[14] The isolated protein according to any one of
[14] and the selected portion of the genome of the cell. expressing or contacting the cell with a guide RNA having a complementary region; The method includes:
[33] An isolated nucleic acid encoding the protein described in
[15] .
[34] Optionally, one or more regulatory domains for expressing the protein of
[15] . A vector comprising the isolated nucleic acid of
[33] operably linked to a domain.
[35] A method for producing a nucleic acid according to
[33] , which optionally expresses a protein according to
[15] . A host cell, preferably a mammalian host cell.
[36] A method for modifying the genome or epigenome of a cell, comprising:
[14] The isolated protein according to any one of
[14] and the selected portion of the genome of the cell. expressing or contacting the cell with a guide RNA having a complementary region; The method includes:
[37] A method for modifying the genome or epigenome of a cell, comprising:
[28] The isolated fusion protein according to any one of
[28] to
[28] , and the selection of the genome of the cell or infecting the cells with a guide RNA having a region complementary to the region of the target gene. A method that involves touching.
[38] The isolated protein or fusion protein may contain a nuclear localization sequence, a cell-penetrating peptide, The method of
[36] or
[37] , comprising one or more of a sequence, and / or affinity tag.
[39] The cells are stem cells, preferably embryonic stem cells, mesenchymal stem cells, or induced pluripotent stem cells. The method according to
[36] or
[0037] , wherein the stem cells are potent stem cells, present in a living animal, or present in an embryo.
[40] A method for modifying a double-stranded DNA (dsDNA) molecule, comprising: The molecule is the isolated protein according to any one of [1] to
[14] and the dsDNA molecule. contacting the selected portion with a guide RNA having a region complementary to the selected portion.
[41] The method of
[40] , wherein the dsDNA molecule exists in vitro.
[42] A method for modifying a double-stranded DNA (dsDNA) molecule, comprising: The molecule is a fusion protein according to
[15] and a region complementary to a selected portion of the dsDNA molecule. A method comprising contacting a nucleic acid sequence with a guide RNA having a target sequence.
[43] The method of
[42] , wherein the dsDNA molecule exists in vitro.
Claims
1. A sequence having at least 90% sequence identity with the amino acid sequence of SEQ ID NO:2, comprising the following mutations: Y211A / R654A, Y211A / Y212A, Y211A / Y230A, Y211A / R245A, Y212A / Y230A, Y2 12A / R245A, Y230A / R245A, W229A / R654A, Y211A / Y212A / Y230A, Y211A / Y212A / R245A, Y211A / Y212A / Y651A, Y211A / Y230A / R245A, Y211A / Y230A / Y651A, Y 211A / R245A / Y651A, Y211A / R245A / R654A, Y211A / R245A / N419A, Y211A / N419 A / R654A, Y212A / Y230A / R245A, Y212A / Y230A / Y651A, Y212A / R245A / Y651A, Y230A / R245A / Y651A, R245A / N419A / R654A, T392A / N419A / R654A, R245A / T39 2A / N419A / R654A, Y211A / R245A / N419A / R654A, W229A / R245A / N419A / R654A , Y211A / R245A / T392A / N419A / R654A, or Y211A / W229A / R245A / N419A / R654A and has improved off-target activity compared to wild-type SaCas9.
2. The following mutations: E782K; K929R; N968K; R1015H; E782K / N968K / R1015H (KKH mutant); E782K / K929R / R1015H (KRH mutant); or E782K / K929R / N968K / R1015H (KRKH mutant).
2. The isolated protein of claim 1, further comprising one or more of:
3. 2. The isolated protein of claim 1, further comprising a mutation that reduces nuclease activity, said mutation selected from the group consisting of mutations at D10, E477, D556, H701 and D704, and mutations at H557 and N580.
4. the mutation in D10 is D10A or D10N, the mutation at D556 is D556A, the mutation in H557 is H557A, H557N, or H557Y; or the mutation at N580 is N580A; The isolated protein of claim 3.
5. 2. The isolated protein of claim 1, wherein the SaCas9 protein is fused to one or more of a nuclear localization sequence, a cell-penetrating peptide sequence, and / or an affinity tag.
6. A fusion protein comprising the isolated protein of claim 1 , fused to a heterologous functional domain by an optional intervening linker, wherein the linker does not interfere with the activity of the fusion protein.
7. The fusion protein of claim 6 , wherein the heterologous functional domain is a transcription activation domain.
8. The fusion protein of claim 7, wherein the transcription activation domain is from VP64 or NF-κB p65.
9. The fusion protein of claim 6, wherein the heterologous functional domain is a transcriptional silencer or transcriptional repression domain.
10. The fusion protein of claim 9, wherein the transcriptional repression domain is a Kruppel-associated box (KRAB) domain, an ERF repressor domain (ERD), or an mSin3A-interacting domain (SID).
11. The fusion protein of claim 9, wherein the transcriptional silencer is heterochromatin protein 1 (HP1).
12. The fusion protein of claim 11, wherein the HP1 is HP1α or HP1β.
13. The fusion protein of claim 6 , wherein the heterologous functional domain is an enzyme that modifies the methylation status of DNA.
14. 14. The fusion protein of claim 13, wherein the enzyme that modifies the methylation status of DNA is a DNA methyltransferase (DNMT) or a ten-eleven translocation (TET) protein.
15. 15. The fusion protein of claim 14, wherein the TET protein is TET1.
16. The fusion protein of claim 6 , wherein the heterologous functional domain is an enzyme that modifies a histone subunit.
17. 17. The fusion protein of claim 16, wherein the enzyme that modifies a histone subunit is a histone acetyltransferase (HAT), a histone deacetylase (HDAC), a histone methyltransferase (HMT), or a histone demethylase.
18. The fusion protein of claim 6 , wherein the heterologous functional domain is a biological tether.
19. 19. The fusion protein of claim 18, wherein the biological tether is MS2, Csy4, or lambda N protein.
20. The fusion protein of claim 6 , wherein the heterologous functional domain is FokI.
21. An isolated nucleic acid encoding the isolated protein of claim 1.
22. 22. A vector comprising the isolated nucleic acid of claim 21, optionally operably linked to one or more regulatory domains.
23. 23. The vector of claim 22, wherein the vector is a prokaryotic vector.
24. 24. The vector of claim 23, wherein the prokaryotic vector is a plasmid or a shuttle vector.
25. The vector of claim 22, wherein the vector is an expression vector.
26. 22. A host cell comprising the isolated nucleic acid of claim 21.
27. 1. An in vitro method for altering the genome or epigenome of a cell, comprising expressing in or contacting the cell with the isolated protein of claim 1 and a guide RNA having a region complementary to a selected portion of the genome of the cell.
28. 10. An in vitro method for altering the genome or epigenome of a cell, comprising expressing in or contacting the cell with a fusion protein of claim 6 and a guide RNA having a region complementary to a selected portion of the genome of the cell.
29. 10. An in vitro method for altering a double-stranded DNA (dsDNA) molecule, comprising contacting the dsDNA molecule with the isolated protein of claim 1 and a guide RNA having a region complementary to a selected portion of the dsDNA molecule.
30. 10. An in vitro method for altering a double-stranded DNA (dsDNA) molecule, comprising contacting the dsDNA molecule with the isolated protein of claim 6 and a guide RNA having a region complementary to a selected portion of the dsDNA molecule.
Citation Information
Patent Citations
JPP6799586B
Using truncated guide rnas (tru-grnas) to increase specificity for RNA-guided genome editing
WO2014144592A2