Engineered target-specific nuclease

By mutation of the DNA binding domain and cleavage domain of nucleases, especially changing the charge state of key amino acid residues and the position close to the DNA backbone, and using different proportions of nuclease complexes, the problem of high off-target cleavage activity when targeting genomic sequences is solved, achieving higher target specificity and genome editing accuracy.

CN120248076APending Publication Date: 2025-07-04SANGAMO THERAPEUTICS INC
View PDF 143 Cites 0 Cited by

Patent Information

Application Number
CN202510129840.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-01-09
Filing Date
2017-08-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing engineered nucleases have the problem of high off-target cleavage activity when targeting genomic sequences, which affects their specificity and efficiency.

Method used

By mutation of the DNA binding domain and cleavage domain of nucleases, especially changing the charge state of certain key amino acid residues and the position close to the DNA backbone, reducing non-specific interactions, and using different proportions of nuclease complexes, improving target specificity and reducing off-target activity.

Benefits of technology

It significantly improves the specific cleavage of the expected target by nucleases, reduces the cleavage of off-target sites, and enhances the accuracy and safety of genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120248076A_ABST
    Figure CN120248076A_ABST
Patent Text Reader

Abstract

Described herein are engineered nuclease enzymes comprising mutations in a cleavage domain (e.g., FokI or a homolog thereof) and / or a DNA binding domain (zinc finger protein, TALE, single guide RNA) such that target specificity is increased.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for "Engineered Target-Specific Nucleases" with the application date of August 24, 2017, application number 201780051010.X (International Application Number PCT / US2017 / 048409).

[0002] Cross-Reference to Related Applications

[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 378,978, filed on August 24, 2016, and U.S. Provisional Application No. 62 / 443,981, filed on January 9, 2017, the disclosures of which are hereby incorporated by reference in their entireties.

[0004] Claims of the Invention

[0005] Made under Federal Government Support

[0006] Not applicable. Technical Field

[0007] The present disclosure relates to the fields of polypeptide and genome engineering and homologous recombination. Background Art

[0008] Engineered nucleases such as engineered zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), CRISPR / Cas systems with engineered crRNA / tracrRNA (‘single guide RNAs’) (also referred to as RNA-guided nucleases), and / or Argonaute system-based nucleases (e.g., from Thermus thermophilus, referred to as ‘TtAgo’ (Swarts et al. (2014) Nature 507(7491):258-261)) include a DNA-binding domain (nucleotide or polypeptide) associated or operably linked to a cleavage domain and have been used for targeted alteration of genomic sequences. For example, nucleases have been used to insert exogenous sequences, inactivate one or more endogenous genes, generate organisms (e.g., crops) and cell lines with altered gene expression patterns, and the like. See, e.g., U.S. Patent Nos. 9,255,250; 9,200,266; 9,045,763; 9,005,973; 8,956,828; 8,945,868; 8,703,489; 8,586,526; 6,534,261; 6,599,692; 6,503,717; 6,689,558; 7,067,317; 7,262,054; 7,888,121; 7,972,854; 7,914,796; 7,951,925; 8,110,379; 8,409,861; U.S. Patent Publications 20030232410; 20050208489; 20050026157; 20050064474; 20060063231; 20080159996; 201000218264; 20120017290; 20110265198; 20130137104; 20130122591; 20130177983 and 20130177960 and 20150056705. For example, nuclease pairs (e.g., zinc finger nucleases, TALENs, dCas-Fok fusions) can be used to cleave genomic sequences. Each member of the pair typically includes an engineered (non-naturally occurring) DNA-binding protein linked to one or more cleavage domains (or half-domains) of a nuclease. When the DNA-binding proteins bind to their target sites, the cleavage domains linked to those DNA-binding proteins are positioned such that dimerization and subsequent cleavage of the genome can occur.

[0009] Generally, intermolecular ion pairs (salt bridges) are required for many DNA-protein interactions. Typically, charged amino acid side chains (i.e., –NH3+, =NH2+) interact with the negatively charged phosphate groups of the DNA backbone to form salt bridges. These ion pairs can be very dynamic and can alternate between direct pairing of the two ions and pairing as a ‘solvent-separated ion pair’ when a solvent (e.g., a water molecule) inserts between the two ions (Chen et al. (2015) J Phys Chem Lett 6:2733-2737).

[0010] Regarding zinc finger proteins, the specificity of ZFPs for target DNA sequences depends on sequence-specific contacts between the zinc finger domains and specific DNA bases. In addition, the zinc finger domains also contain amino acid residues involved in non-specific ion pair interactions with the phosphates of the DNA backbone. Elrod-Erickson et al. ((1996) Structure 4:1171) demonstrated by co-crystallization of a zinc finger protein and its cognate DNA target that there are specific amino acids capable of interacting with the phosphates on the DNA backbone by forming hydrogen bonds. Zinc finger proteins with the well-known Zif268 backbone typically have arginine as the amino-terminal residue of their second β-strand, which is also the second position carboxyl-terminal to the second invariant cysteine (see Figure 5A ). This position can be designated as (-5) within each zinc finger domain because it is the 5th residue before the start of the α-helix ( Figure 5A ). The arginine at this position can interact with the phosphates on the DNA backbone via a charged hydrogen bond formed with its side chain guanidine group. Zinc finger proteins in the Zif268 backbone typically have lysine at the position of 4 residues amino-terminal to the first invariant cysteine. This position can be designated as (-14) within each finger because it is the 14th residue before the start of the α-helix of the zinc finger, with two residues between the zinc-coordinating cysteine residues ( Figure 5A ). Lysine can interact with the phosphates on the DNA backbone via a water-mediated charged hydrogen bond formed with its side chain amino group. Since phosphate groups are found along the DNA backbone, this type of interaction between zinc fingers and DNA molecules is generally considered non-sequence-specific (J. Miller, Massachusetts Institute of Technology Ph.D. Thesis, 2002).

[0011] Recent studies have hypothesized that non-specific phosphate contacting side chains in some nucleases may also contribute to a certain amount of non-specific cleavage activity of those nucleases (Kleinstiver et al., (2016) Nature 529(7587):490-5; Guilinger et al (2014) Nat Meth:429-435). Researchers have proposed that these nucleases may have 'excess DNA binding energy', meaning that the affinity of the nuclease for its DNA target may be greater than that required to substantially bind and cleave the target site. Thus, attempts have been made to reduce the cationic charge in the TALE DNA binding domain (Guilinger, ibid) or the Cas9 DNA binding domain (Kleinstiver, ibid) to reduce the DNA binding energy of these nucleases, thereby producing increased in vitro cleavage specificity. However, additional studies (Sternberg et al (2015) Nature 527(7576):110-113) have also shown the role of some cationic amino acids in the proper folding and activation of the Cas9 nuclease domain, which were mutated in the Kleinstiver study of the Cas9 DNA binding domain. Thus, the exact role of these amino acids in Cas9 activity remains unclear.

[0012] For optimal cleavage specificity by sequence-selective (engineered) nucleases, it is desirable to arrange conditions such that on-target binding and activity are not saturated. Under saturated conditions - by definition - an excess of nuclease is used rather than that required to achieve full on-target activity. This excess does not provide an on-target benefit but can still lead to increased cleavage at off-target sites. For monomeric nucleases, saturated conditions can be easily avoided by performing a simple dose-response study to identify and avoid the saturated plateau phase on the titration curve. However, for dimeric nucleases such as ZFNs, TALENs or dCas-Fok, if the binding affinities of the individual monomers are different, identifying and avoiding saturated conditions may be more complex. In such cases, a dose-response study using a simple 1:1 nuclease ratio will only reveal the saturation point of the weaker-binding monomer. In this case, if, for example, the monomer affinities differ by 10-fold, at the saturation point identified in a 1:1 titration study, the higher-affinity monomer will be present at a concentration 10-fold higher than that required. The resulting excess of the higher-affinity monomer can in turn lead to increased off-target activity without providing any beneficial increase in cleavage at the intended target, potentially resulting in a decrease in the overall specificity of any given nuclease pair.

[0013] To reduce off-target cleavage events, engineered obligate heterodimeric cleavage half-domains have been developed. See, for example, U.S. Patent Nos. 7,914,796; 8,034,598; 8,961,281 and 8,623,618; U.S. Patent Publication Nos. 20080131962 and 20120040398. These obligate heterodimers dimerize and cleave their target only when different engineered cleavage domains are targeted to an appropriate target site by a ZFP, thereby reducing and / or eliminating monomer off-target cleavage.

[0014] However, there remains a need for additional methods and compositions for engineered nuclease cleavage systems to reduce off-target cleavage activity. SUMMARY OF THE INVENTION

[0015] The present disclosure provides methods and compositions for increasing the specificity of nucleases (e.g., nuclease pairs) for their intended targets relative to other unintended cleavage sites (also referred to as off-target sites). Thus, artificial nucleases (e.g., zinc finger nucleases (ZFNs), TALENs, CRISPR / Cas nucleases) are described herein that contain one or more mutations in one or more DNA binding domain regions (e.g., the backbone of a zinc finger protein or TALE) and / or one or more mutations in a FokI nuclease cleavage domain or cleavage half-domain. In addition, methods for increasing the specificity of cleavage activity by using these novel nucleases (e.g., ZFNs, TALENs, etc.) and / or by independently titrating engineered cleavage half-domain partners of a nuclease complex are described herein. When used alone or in combination, the methods and compositions of the present invention provide a surprising and unexpected increase in targeting specificity through a reduction in off-target cleavage activity. The present disclosure also provides methods for using these compositions for targeting cleavage of cellular chromatin in a target region and / or integrating a transgene via targeted integration at a predetermined target region in a cell.

[0016] Accordingly, in one aspect, an engineered nuclease cleavage half-domain is described herein that contains one or more mutations compared to a parental (e.g., wild-type) cleavage domain from which these mutants are derived. In certain embodiments, the one or more mutations are one or more of the mutations shown in any of the schedules and figures, including any combination of these mutants with each other and with other mutants (such as dimerization and / or catalytic domain mutants and nickase mutants). Mutations as described herein include, but are not limited to, mutations that alter the charge of the cleavage domain, such as mutations of positively charged residues to non-positively charged residues (e.g., mutations of K and R residues (e.g., to S); mutations of N residues (e.g., to D) and mutations of Q residues (e.g., to E); mutations of residues predicted by molecular modeling to be close to the DNA backbone and showing variation in FokI homologs ( Figure 1 and17 ); and / or mutations at other residues (e.g., U.S. Patent No. 8,623,618 and Guo et al., (2010) J. Mol. Biol. 400(1):96-107).

[0017] The most promising mutations are identified using a second criterion. When FokI binds to DNA, initially promising mutations are positively charged residues predicted to be close to the DNA backbone. The cleavage domain described herein can include one, two, three, four, five or more mutations described herein and can also include additional known mutations. Thus, when used alone, the mutations of the present invention do not include the specific mutations disclosed in U.S. Patent No. 8,623,618 (e.g., N527D, S418P, K448M, Q531R, etc.); however, novel mutants are provided herein that can be used in combination with the mutants of U.S. Patent No. 8,623,618. Nickase mutants in which one of the catalytic nuclease domains in a dimer pair contains one or more mutations that inactivate its catalysis (see U.S. Patent Nos. 8,703,489; 9,200,266; and 9,631,186) can also be used in combination with any of the mutants described herein. The nickase can be a ZFN nickase, a TALEN nickase, and a CRISPR / dCas system.

[0018] In certain embodiments, the engineered cleavage half-domain is derived from FokI or a FokI homolog and contains a mutation in one or more of amino acid residues 416, 422, 447, 448, and / or 525 numbered relative to wild-type full-length FokI as shown in SEQ ID NO:1 or the corresponding residues in a FokI homolog (see Figure 17 ). In other embodiments, the cleavage half-domain derived from FokI contains one or more of amino acid residues 414-426, 443-450, 467-488, 501-502, and / or 521-531, including one or more of 387, 393, 394, 398, 400, 416, 418, 422, 427, 434, 439, 441, 442, 444, 446, 448, 472, 473, 476, 478, 479, 480, 481, 487, 495, 497, 506, 516, 523, 525, 527, 529, 534, 559, 569, 570, and / or 571. The mutations can include mutations of residues found in native restriction enzymes homologous to FokI at the corresponding positions ( Figure 17)。In certain embodiments, the mutation is a substitution, e.g., with any different amino acid, such as alanine (A), cysteine (C), aspartic acid (D), glutamic acid (E), histidine (H), phenylalanine (F), glycine (G), asparagine (N), serine (S), or threonine (T) to replace the wild-type residue. Any combination of mutants is contemplated, including but not limited to those shown in the tables and figures. In certain embodiments, the FokI nuclease domain comprises a mutation at one or more of 416, 422, 447, 479, and / or 525 (numbered relative to wild-type, SEQ ID NO:1). The nuclease domain may also comprise one or more mutations at positions 418, 432, 441, 448, 476, 481, 483, 486, 487, 490, 496, 499, 523, 527, 537, 538, and 559, including but not limited to ELD, KKR, ELE, KKS. See, e.g., U.S. Patent No. 8,623,618. In other embodiments, the cleavage domain comprises one or more of the residues shown in Table 15 (e.g., 419, 420, 425, 446, 447, 470, 471, 472, 475, 478, 480, 492, 500, 502, 521, 523, 526, 530, 536, 540, 545, 573, and / or 574) at the mutation. In certain embodiments, the variant cleavage domain described herein comprises a mutation of a residue involved in nuclease dimerization (dimerization domain mutation), and one or more additional mutations; e.g., a mutation of a phosphate contact residue: e.g., a combination of a dimerization mutant (such as ELD, KKR, ELE, KKS, etc.) and one, two, three, four, five, six, or more mutations at amino acid positions outside the dimerization domain (e.g., in amino acid residues that may be involved in phosphate contact). In a preferred embodiment, the mutation at positions 416, 422, 447, 448, and / or 525 comprises replacing a positively charged amino acid with an uncharged or negatively charged amino acid. In other embodiments, mutations are made at positions 446, 472, and / or 478 (and optionally additional residues in, e.g., the dimerization or catalytic domain).

[0019] In other embodiments, in addition to the mutations described herein, the engineered cleavage half-domain further comprises mutations in the dimerization domain, such as mutations in amino acid residues 490, 537, 538, 499, 496, and 486. In a preferred embodiment, the present invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, wherein in addition to one or more of the mutations described herein, the wild-type Gln (Q) residue at position 486 is replaced with a Glu (E) residue, the wild-type Ile (I) residue at position 499 is replaced with a Leu (L) residue, and the wild-type Asn (N) residue at position 496 is replaced with an Asp (D) or Glu (E) residue ("ELD" or "ELE"). In another embodiment, the engineered cleavage half-domain is derived from wild-type FokI or a FokI homologous cleavage half-domain and further comprises mutations in amino acid residues 490, 538, and 537 relative to the numbering of wild-type FokI (SEQ ID NO:1) in addition to one or more mutations at amino acid residues 416, 422, 447, 448, or 525. In a preferred embodiment, the present invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, wherein in addition to one or more of the mutations described herein, the wild-type Glu (E) residue at position 490 is replaced with a Lys (K) residue, the wild-type Ile (I) residue at position 538 is replaced with a Lys (K) residue, and the wild-type His (H) residue at position 537 is replaced with a Lys (K) residue or an Arg (R) residue ("KKK" or "KKR") (see U.S. Patent 8,962,281, which is incorporated herein by reference).

[0020] In another embodiment, the engineered cleavage half-domain is derived from a wild-type FokI cleavage half-domain or homolog thereof and contains mutations at amino acid residues 490 and 538 relative to wild-type FokI numbering in addition to one or more mutations at amino acid residues 416, 422, 447, 448 or 525. In a preferred embodiment, the invention provides a fusion protein wherein the engineered cleavage half-domain comprises a polypeptide wherein, in addition to one or more mutations at position 416, 422, 447, 448 or 525, the wild-type Glu (E) residue at position 490 is replaced by a Lys (K) residue and the wild-type Ile (I) residue at position 538 is replaced by a Lys (K) residue (“KK”). In a preferred embodiment, the invention provides a fusion protein wherein the engineered cleavage half-domain comprises a polypeptide wherein, in addition to one or more mutations at position 416, 422, 447, 448 or 525, the wild-type Gln (Q) residue at position 486 is replaced by a Glu (E) residue and the wild-type Ile (I) residue at position 499 is replaced by a Leu (L) residue (“EL”) (see U.S. Patent 8,034,598, which is incorporated herein by reference).

[0021] In one aspect, the invention provides fusion molecules wherein the engineered cleavage half-domain comprises a polypeptide wherein a wild-type amino acid residue at one or more of positions 387, 393, 394, 398, 400, 402, 416, 422, 427, 434, 439, 441, 446, 447, 448, 469, 472, 478, 487, 495, 497, 506, 516, 525, 529, 534, 559, 569, 570, 571 in the FokI catalytic domain is mutated. In some embodiments, the one or more mutations change the wild-type amino acid from a positively charged residue to a neutral or negatively charged residue. In any of these embodiments, the described mutants can also be prepared in FokI domains containing one or more additional mutations. In preferred embodiments, these additional mutations are located in the dimerization domain, e.g., at positions 499, 496, 486, 490, 538 and 537. Mutations include substitution, insertion and / or deletion of one or more amino acid residues.

[0022] In another aspect, any of the above-described engineered cleavage half-domains can be incorporated into artificial nucleases, including but not limited to zinc finger nucleases, TALENs, CRISPR / Cas nucleases, etc., for example by associating them with a DNA-binding domain. The zinc finger proteins of zinc finger nucleases can contain non-canonical zinc-coordinating residues (e.g., CCHC instead of the canonical C2H2 configuration, see U.S. Patent 9,234,187).

[0023] In another aspect, there are provided fusion molecules that produce an artificial nuclease, the fusion molecules comprising a DNA binding domain as described herein and an engineered cleavage half-domain of FokI or its homolog. In certain embodiments, the DNA binding domain of the fusion molecule is a zinc finger binding domain (e.g., an engineered zinc finger binding domain). In other embodiments, the DNA binding domain is a TALE DNA binding domain. In other embodiments, the DNA binding domain comprises a DNA binding molecule (e.g., a guide RNA) and a catalytically inactive Cas9 or Cfp1 protein (dCas9 or dCfp1). In some embodiments, the engineered fusion molecule forms a nuclease complex with a catalytically inactive engineered cleavage half-domain such that the dimeric nuclease is capable of cleaving only one strand of a double-stranded DNA molecule, thereby forming a nickase (see U.S. Patent 9,200,266).

[0024] The methods and compositions of the invention also include mutations of one or more amino acids within the DNA binding domain outside of the residues of the nucleotides that recognize the target sequence, the residues that can interact non-specifically with the phosphate groups on the DNA backbone. Thus, in certain embodiments, the invention includes mutations of cationic amino acid residues in the ZFP backbone that are not required for nucleotide target specificity. In some embodiments, these mutations in the ZFP backbone include mutating the cationic amino acid residues to neutral or anionic amino acid residues. In some embodiments, these mutations in the ZFP backbone include mutating polar amino acid residues to neutral or non-polar amino acid residues. In preferred embodiments, the mutations are made at positions (-5), (-9) and / or position (-14) relative to the DNA binding helix. In some embodiments, the zinc finger may comprise one or more mutations at (-5), (-9) and / or (-14). In other embodiments, one or more zinc fingers in a multi-fingered zinc finger protein may comprise mutations at (-5), (-9) and / or (-14). In some embodiments, the amino acids (e.g., arginine (R) or lysine (K)) at (-5), (-9) and / or (-14) are mutated to alanine (A), leucine (L), Ser (S), Asp (N), Glu (E), Tyr (Y) and / or glutamine (Q).

[0025] In another aspect, there are provided polynucleotides that encode any of the engineered cleavage half-domains or fusion proteins as described herein.

[0026] In yet another aspect, cells are also provided, the cells comprising any nuclease, polypeptide (e.g., a fusion molecule or fusion polypeptide), and / or polynucleotide as described herein. In one embodiment, the cell comprises a pair of fusion polypeptides, one fusion polypeptide comprising an ELD or ELE cleavage half-domain in addition to one or more mutations in one or more of amino acid residues 393, 394, 398, 416, 421, 422, 442, 444, 447, 448, 473, 480, 530, and / or 525, and one fusion polypeptide comprising a KKK or KKR cleavage half-domain in addition to one or more mutations at residues 393, 394, 398, 416, 421, 422, 442, 444, 446, 447, 448, 472, 473, 478, 480, 530, and / or 525 (see U.S. Patent 8,962,281).

[0027] In any of these fusion polypeptides described herein, the ZFP moiety may also comprise mutations in the zinc finger DNA-binding domain at positions (-5), (-9), and / or (-14). In some embodiments, the Arg (R) at position -5 is changed to Tyr (Y), Asp (N), Glu (E), Leu (L), Gln (Q), or Ala (A). In other embodiments, the Arg (R) at position (-9) is replaced by Ser (S), Asp (N), or Glu (E). In other embodiments, the Arg (R) at position (-14) is replaced by Ser (S) or Gln (Q). In other embodiments, the fusion polypeptide may comprise mutations in the zinc finger DNA-binding domain, wherein the amino acids at positions (-5), (-9), and / or (-14) are changed to any one of the amino acids listed above in any combination.

[0028] Cells modified by the polypeptides and / or polynucleotides of the invention are also provided herein. In some embodiments, the cell comprises an insertion of a nuclease-mediated transgene, or a nuclease-mediated gene knockout. The modified cells and any cells derived from the modified cells do not necessarily transiently contain only the nuclease of the invention, but the genomic modifications mediated by such nucleases remain.

[0029] In yet another aspect, methods are provided for targeting cleavage of cellular chromatin in a target region; methods for causing homologous recombination to occur in a cell; methods for treating an infection; and / or methods for treating a disease. These methods can be practiced in vitro, ex vivo, in vivo, or in combination thereof. The methods include cleaving cellular chromatin in a predetermined target region in a cell by expressing a pair of fusion polypeptides as described herein (i.e., a pair of fusion polypeptides, wherein one or both of the fusion polypeptides comprises an engineered cleavage half-domain as described herein). In certain embodiments, targeted cleavage at the target site is increased by at least 50% to 200% (or any value therebetween) or more, including 50%-60% (or any value therebetween), 60%-70% (or any value therebetween), 70%-80% (or any value therebetween), 80%-90% (or any value therebetween), 90% to 200% (or any value therebetween), as compared to a cleavage domain that does not have the mutations described herein. Similarly, using the methods and compositions described herein, off-target site cleavage is reduced 1- to 100-fold or more, including but not limited to 1- to 50-fold (or any value therebetween).

[0030] The engineered cleavage half-domains described herein can be used in methods for targeting cleavage of cellular chromatin in a target region and / or for performing homologous recombination in a predetermined target region in a cell. Cells include cultured cells, cell lines, cells in an organism, cells that have been removed from an organism for treatment in cases where the cells and / or their progeny will be returned to the treated organism, and cells that have been removed from an organism, modified using the fusion molecules of the invention, and then returned to the organism in a therapeutic method (cell therapy). The target region in cellular chromatin can be, for example, a genomic sequence or a portion thereof. Compositions comprise a fusion molecule or a polynucleotide encoding a fusion molecule, the fusion molecule comprising a DNA-binding molecule (e.g., an engineered zinc finger or TALE binding domain or an engineered CRISPR guide RNA) and a cleavage half-domain as described.

[0031] The fusion molecule can be expressed in a cell, for example, by delivering the fusion molecule as a polypeptide to the cell, or by delivering a polynucleotide encoding the fusion molecule to the cell, wherein the polynucleotide (if DNA) is transcribed and translated to produce the fusion molecule. Additionally, if the polynucleotide is an mRNA encoding the fusion molecule, then after delivery of the mRNA to the cell, the mRNA is translated, thereby producing the fusion molecule.

[0032] In other aspects of the invention, methods and compositions for enhancing the specificity of engineered nucleases are provided. In one aspect, methods for enhancing overall on-target cleavage specificity by reducing off-target cleavage activity are provided. In some embodiments, the engineered cleavage half-domain partners of an engineered nuclease complex are used to contact a cell, wherein each partner of the complex is provided in a ratio other than 1:1 relative to the other partner. In some embodiments, the ratio of the two partners (half-cleavage domains) is provided at a ratio of 1:2, 1:3, 1:4, 1:5, 1:6, 1:8, 1:9, 1:10, or 1:20 or any value therebetween. In other embodiments, the ratio of the two partners is greater than 1:30. In other embodiments, the two partners are deployed in a ratio selected to be different from 1:1. In some aspects, each partner is delivered to the cell as mRNA or in a viral or non-viral vector, wherein different amounts of the mRNA or vector encoding each partner are delivered. In other embodiments, each partner of the nuclease complex can be included on a single viral or non-viral vector, but is intentionally expressed such that one partner is expressed at a value higher or lower than the other partner, thereby ultimately delivering a ratio of cleavage half-domains other than 1:1 to the cell. In some embodiments, each cleavage half-domain is expressed using different promoters with different expression efficiencies. In other embodiments, two cleavage domains are delivered to the cell using a viral or non-viral vector, wherein both are expressed from the same open reading frame, but the genes encoding the two partners are separated by a sequence (e.g., a self-cleaving 2A sequence or an IRES) that results in the 3' partner being expressed at a lower rate such that the ratio of the two partners is 1:2, 1:3, 1:4, 1:5, 1:6, 1:8, 1:9, 1:10, or 1:20 or any value therebetween. In other embodiments, the two partners are deployed in a ratio selected to be different from 1:1.

[0033] Also provided are methods for reducing off-target nuclease activity when using two or more nuclease complexes. For example, the present invention provides methods for altering the ratio of DNA-binding molecules when using two or more nuclease complexes. In some embodiments, the DNA-binding molecule is a polypeptide DNA-binding domain (e.g., ZFN, TALEN, dCas-Fok, megaTAL, meganuclease), while in other embodiments, the DNA-binding molecule is a guide RNA used with an RNA-guided nuclease. In preferred embodiments, the ratio of two or more DNA-binding molecules is a 1:2, 1:3, 1:4, 1:5, 1:6, 1:8, 1:9, 1:10, or 1:20 ratio, or any value therebetween. In other embodiments, two DNA-binding molecules are deployed in a ratio selected to be different from 1:1. In some aspects, a non-1:1 ratio is achieved by altering the ratio of guide RNAs used to transfect cells. In other aspects, the ratio is altered by changing the ratio of each Cas9 protein-guide RNA complex used to treat target cells. In another aspect, an altered ratio is achieved by treating cells with different ratios of DNA (viral or non-viral) encoding the guide RNA, or by using promoters with different expression intensities to differentially express the DNA-binding molecule inside the cell. Off-target events can be reduced 2 to 1000-fold (or any amount therebetween) or more, including but not limited to reduced by at least 10, 50, 60, 70, 80, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000-fold (or any value therebetween) or more.

[0034] Thus, in another aspect, a method for cleaving cellular chromatin in a target region can include (a) selecting a first sequence in the target region; (b) engineering a first DNA-binding molecule to specifically bind to the first sequence; (c) expressing a first fusion molecule in a cell, the first fusion molecule comprising the first DNA-binding domain (e.g., zinc finger, TALE, sgRNA) and a cleavage domain (or half-domain); and (d) expressing a second fusion protein in the cell, the second fusion molecule comprising a second DNA-binding domain and a second cleavage domain (or half-domain), wherein at least one of the fusion molecules comprises a linker as described herein, and further wherein the first fusion molecule binds to the first sequence and the second fusion molecule binds to a second sequence located between 2 and 50 nucleotides from the first sequence such that an active nuclease complex can be formed and cellular chromatin can be cleaved in the target region. In certain embodiments, both fusion molecules comprise a linker as described herein between the DNA-binding domain and the catalytic nuclease domain.

[0035] Methods are also provided for altering a cellular chromatin region, such as to introduce a targeted mutation. In certain embodiments, methods of altering cellular chromatin include introducing into the cell one or more targeted nucleases to create a double-strand break in the cellular chromatin at a predetermined site; and a donor polynucleotide that has homology to the nucleotide sequence of the cellular chromatin in the region of the break. The cellular DNA repair process is activated by the presence of the double-strand break, and the donor polynucleotide serves as a template for break repair, such that all or part of the nucleotide sequence of the donor is introduced into the cellular chromatin. Thus, the sequence in the cellular chromatin can be altered and, in certain embodiments, converted to the sequence present in the donor polynucleotide.

[0036] Targeted alterations include, but are not limited to, point mutations (i.e., conversion of a single base pair to a different base pair), substitutions (i.e., conversion of multiple base pairs to a different sequence of the same length), insertions of one or more base pairs, deletions of one or more base pairs, and any combination of the foregoing sequence alterations. Alterations can also include conversion of base pairs that are part of a coding sequence such that the encoded amino acid is changed.

[0037] The donor polynucleotide can be DNA or RNA, can be linear or circular, and can be single-stranded or double-stranded. It can be delivered as a naked nucleic acid, as a complex with one or more delivery agents (e.g., liposomes, nanoparticles, poloxamers), or contained within a viral delivery vehicle (e.g., an adenovirus, lentivirus, or adeno-associated virus (AAV)) to the cell. The length of the donor sequence can range from 10 to 1,000 nucleotides (or any integer value of nucleotides therebetween) or longer. In some embodiments, the donor comprises a full-length gene flanked by homologous regions having a targeted cleavage site. In some embodiments, the donor lacks homologous regions and integrates into the target locus by a mechanism independent of homology (i.e., NHEJ). In other embodiments, the donor comprises a smaller fragment of nucleic acid flanked by homologous regions for the cell (i.e., for gene correction). In some embodiments, the donor comprises a gene encoding a functional or structural component such as shRNA, RNAi, miRNA, etc. In other embodiments, the donor comprises a sequence encoding a regulatory element that binds to the target gene and / or regulates the expression of the target gene. In other embodiments, the donor is a target regulatory protein (e.g., ZFP TF, TALE TF, or CRISPR / Cas TF) that binds to the target gene and / or regulates the expression of the target gene.

[0038] For any of the above methods, the cellular chromatin can be in chromosomes, episomes, or organelle genomes. The cellular chromatin can be present in any type of cell, including but not limited to prokaryotic and eukaryotic cells, fungal cells, plant cells, animal cells, mammalian cells, primate cells, and human cells.

[0039] In yet another aspect, provided are cells comprising any polypeptide (e.g., a fusion molecule) and / or polynucleotide as described herein. In one embodiment, the cell comprises a pair of fusion molecules, each fusion molecule comprising a cleavage domain as disclosed herein. Cells include cultured cells, cells in an organism, and cells that have been removed from an organism for treatment in cases where the cells and / or their progeny will be returned to the treated organism. A target region in cellular chromatin can be, for example, a genomic sequence or a portion thereof.

[0040] In another aspect, described herein is a kit that includes a fusion protein as described herein or a polynucleotide encoding one or more zinc finger proteins, cleavage domains, and / or fusion proteins as described herein; ancillary reagents; and optionally instructions and a suitable container. The kit can also include one or more nucleases or polynucleotides encoding such nucleases.

[0041] In view of the overall disclosure, these and other aspects will be apparent to those of skill in the art.

[0042] Specifically, the present invention includes, but is not limited to, the following:

[0043] 1. An engineered FokI cleavage half-domain, wherein the engineered cleavage half-domain comprises one or more mutations in residues 393, 394, 398, 416, 421, 422, 442, 444, 472, 473, 478, 480, 525, or 530, or wherein the wild-type residue at position 418 is replaced with a Glu (E) or Asp (D) residue, the wild-type residue at position 446 is replaced with an Asp (D) residue, the wild-type residue at position 448 is replaced with an Ala (A) residue (K448A), the wild-type residue at position 479 is replaced with a Gln (Q) or Thr (T) residue (I479Q or I479T), the wild-type residue at position 481 is replaced with an Ala (A), Asn (N), or Glu (E) residue (Q481A, Q481N, Q481E), or the wild-type residue at position 523 is replaced with a Phe (F) residue, wherein the amino acid residues are numbered relative to the full-length FokI wild-type cleavage domain as shown in SEQ ID NO:1.

[0044] 2. The engineered cleavage half-domain as described in item 1, which comprises mutations at residues 416 and 422, the mutation at position 416 and K448A, K448A and I479Q, K448A and Q481A and / or K448A, and the mutation at position 525.

[0045] 3. The engineered cleavage half-domain as described in item 2, wherein the wild-type residue at position 416 is replaced by a Glu (E) residue (R416E), the wild-type residue at position 422 is replaced by a His (H) residue (R422H), and the wild-type residue at position 525 is replaced by an Ala (A) residue.

[0046] 4. The engineered cleavage half-domain as described in any one of items 1 to 3, which further comprises additional amino acid mutations at one or more of positions 432, 441, 483, 486, 487, 490, 496, 499, 527, 537, 538 and 559.

[0047] 5. A heterodimer, which comprises a first engineered cleavage half-domain and a second cleavage half-domain as described in any one of items 1 to 4.

[0048] 6. An artificial nuclease, which comprises an engineered cleavage half-domain and a DNA-binding domain as described in any one of items 1 to 4.

[0049] 7. The artificial nuclease as described in item 6, wherein the DNA-binding domain comprises a zinc finger protein, a TALE effector domain or a single guide RNA (sgRNA).

[0050] 8. The artificial nuclease as described in item 7, wherein the zinc finger protein comprises one or more mutations at positions -14, -9 and -5.

[0051] 9. A polynucleotide, which encodes an engineered cleavage half-domain as described in any one of items 1 to 4, a first engineered cleavage domain and a second engineered cleavage domain as described in item 5, or an artificial nuclease as described in any one of items 6 to 8.

[0052] 10. An isolated cell, which comprises the polynucleotide as described in item 9.

[0053] 11. A method for cleaving genomic cell chromatin in a target region, the method comprising:

[0054] expressing an artificial nuclease comprising an engineered cleavage domain as described in any one of items 6 to 8 in a cell,

[0055] wherein the nuclease site-specifically cleaves the nucleotide sequence in the target region of the genomic cell chromatin.

[0056] 12. The method according to item 11, further comprising contacting the cell with a donor polynucleotide; wherein lysis of the cell chromatin promotes homologous recombination between the donor polypeptide and the cell chromatin.

[0057] 13. A method of lysing at least two target sites in genomic cell chromatin, the method comprising:

[0058] Lysing at least a first target site and a second target site in genomic cell chromatin, wherein each target site is lysed using a composition comprising an artificial nuclease according to any one of items 6 to 8.

[0059] 14. An isolated cell or cell line comprising at least one site-specific genomic modification generated by an artificial nuclease according to any one of items 6 to 8. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 The amino acid sequence (SEQ ID NO:1) and nucleotide sequence (SEQ ID NO:2) depicting a portion of the wild-type FokI nuclease are shown. The sequences show the FokI catalytic nuclease domain, and the numbering is relative to the wild-type FokI protein (amino acid Q of the nuclease domain starting at 384) used to generate the crystal structures 1FOK.pdb and 2FOK.pdb (Wah et al., supra). The boxed positions indicate possible mutation sites.

[0061] Figures 2A to 2C is a schematic diagram showing a model of the FokI domain interacting with a DNA molecule. Figure 2A Indicates the positions of amino acids R422, R416, and K525. Figure 2B Indicates the positions of amino acids R447, K448, and R422. Figure 2C is an illustration showing a subset of different types of ZFNs that can be prepared, the ZFNs incorporating 1, 2, or 3 (1x, 2x, or 3x, respectively) mutations (R->Q or R->L) in the zinc finger backbone. The black arrows indicate the positions of the mutations.

[0062] Figure 3A and Figure 3B show the activity of the BCL11A-specific ZFN carrying the novel FokI mutations described herein. Figure 3AShows targeted modification of BCL11A-specific ZFN SBS#51857-ELD / SBS#51949-KKR in CD34+ cells for the BCL11A homologous target (represented by the unique 'license' identifier PRJIYLFN, SEQ ID NO:13) and two off-target sites also identified by their 'license' identifiers NIFMAEVG (SEQ ID NO:14) and PEVYOHIU (SEQ ID NO:20). The ZFP is described in PCT / US2016 / 032049. All experiments were performed with 2 μg of each ZFN mRNA for nuclease delivery, and the values represent the percentage of sequence reads containing insertions and deletions (indel %) consistent with nuclease activity. Figure 3A Shows the results when serine residues are substituted into positions 416, 422, 447, 448, and 525 of the FokI domain in one or both ZFNs. Figure 3B Shows a similar data set, except that the heterodimer dimerization domain FokI backbone has been converted, i.e., Figure 3A Shows the results using the mutations in the SBS#51857-ELD / SBS#51949-KKR pair, while Figure 3B Shows the results using the mutations in the SBS#51857-KKR / SBS#51949-ELD pair.

[0063] Figure 4 Is a graph depicting on-target and off-target activities of specific ZFN FokI variants (PCT publication WO2017106528) of many TCRA (the targeted constant region, also known as TRAC). Except for two repeats of the parental ZFN pair, the FokI domain in one of the two ZFNs carries a mutation at a positively charged residue. The distance between the alpha carbon of the mutated residue in FokI of the ZFN-DNA molecular model and the closest phosphate oxygen in the DNA backbone was calculated (Miller et al. (2007) Nat Biotech 25(7):778-785), and the data points are color-coded based on this calculated distance (<10 angstroms: gray; >10 angstroms: black). Each data point represents the on-target activity and combined off-target activity of a different ZFN pair carrying an FokI mutation on one of the ZFNs in the pair. The data points representing the parental pair are indicated.

[0064] Figure 5A and Figure 5B Is a schematic diagram depicting the backbone region of the zinc finger. Figure 5A(SEQ ID NO:3) shows the amino acids in the second finger of the Zif268 protein, where the β-sheet and α-helix structures are indicated. Also shown are the positions of the amino acids (-1 to 6) involved in specific DNA base recognition. Positively charged residues with the potential to interact with the phosphate backbone on DNA are represented by squares. The invariant cysteine residues involved in zinc coordination are underlined. Figure 5B is a close-up view of a single finger in its three-dimensional state (coordinated zinc ions are represented by solid spheres), and indicates how the different regions of each zinc finger tend to interact with DNA. DNA is represented by a diagram where the phosphate is indicated by the letter P, and the DNA bases are represented by boxes with rounded corners. The grey arrows indicate the approximate positions of the residue positions indicated in the boxes, and the black arrows indicate the interactions between the zinc finger protein and DNA.

[0065] Figure 6 (SEQ ID NO:4 - 6) depicts the conservation of the amino acids at each position within the zinc finger. The first few lines show the alignment of the amino acid sequences from well-known zinc fingers from Zif268 and Sp1 (finger 2 from Zif268 (SEQ ID NO:4), finger 3 from Zif268 (SEQ ID NO:5), and finger 2 from Sp1 (SEQ ID NO:6)). The zinc-coordinating cysteine and histidine residues are boxed, and the recognition helix is also boxed. The positively charged residues of arginine (R) and lysine (K) that contact the DNA backbone phosphate are also indicated in the box. The numbers below the first three lines are the frequencies of each amino acid at each position, where 4867 different naturally occurring zinc fingers were analyzed. The letters on the left side of the figure are the single-letter codes corresponding to the amino acid residues whose frequencies are given in the table. Three uncharged amino acids, alanine, leucine, and glutamine (indicated by ovals) occur at low but non-zero frequencies at the phosphate contact positions.

[0066] Figure 7A and Figure 7B (SEQ ID NO:7 and 8) depict diagrams of the ZFP backbone, which comprises modules of a six-finger zinc finger protein ( Figure 7A , SEQ ID NO:7) or a five-finger zinc finger protein ( Figure 7B , SEQ ID NO:8). The letters above some of the boxed positions indicate the mutations tested at the specified positions. The identity of each finger is given by the labels F1 to F6. Each of these proteins consists of three different "module assemblies" designated as "module A", "module B", and "module C". Mutations at positions -14, -9, and -5 of the N-terminal finger in each module can be made by altering the sequences of the PCR primers used during the assembly process.

[0067] Figures 8A to 8CA figure depicting the on-target and off-target cleavage activities of TCRA (TRAC)-specific ZFNs (PCT Publication WO2017106528) containing the novel zinc finger backbone mutations of the present invention. The TCRA (TRAC)-specific ZFNs each contain 6 zinc finger repeats, and for ease of experimentation, the mutation at position -5 was introduced only into the N-terminal finger of each module (e.g., F1, F3, or F5 in the full-length ZFN). Thus, each individual ZFN can have 0, 1, 2, or 3 mutations, and the entire ZFN pair can have up to 6 mutations in total (e.g., 0, 1, 2, 3, 4, 5, or 6 mutations). The plotted values indicate the average of all tested ZFN pairs having the specified number and type of mutations at position -5. Error bars represent the standard error of the mean. For each ZFN pair, three off-targets indicated in Table 3 were measured; the off-target values averaged to produce the plotted values include the activity scores of the parental TCRA (TRAC) ZFNs for each of these three off-targets of each construct. Figure 8A Shows the activity scores of the parental TCRA (TRAC) ZFNs, where the data set shows the change in on-target (black bars) or off-target (gray bars) activity of the indicated amino acid substitutions at position -5 in one or more zinc finger repeats of only one of the two ZFNs in the pair. Figure 8B Shows the activity scores from the simultaneous arginine-to-alanine substitutions indicated in one or both ZFN partners. Figure 8B The left half of represents ZFN pairs where the indicated number of mutations appears in only one of the two ZFNs in the pair (and corresponds to Figure 8A the left third of ), while Figure 8B the right half of represents ZFN pairs where the same number of mutations were made in both ZFNs in the pair (e.g., 2 represents one mutation in each ZFN in the pair, 4 represents two mutations in each ZFN in the pair, and 6 represents three mutations in each ZFN in the pair). Figure 8A and 8B The experiments conducted in were performed on CD34+ cells at a dose of 6 μg per experiment. Figure 8C Shows Figure 8A similar data for the right two-thirds of, where the dose of RNA was 2 μg per experiment.

[0068] Figure 9Figure depicting on-target (black bars) and off-target (gray bars) cleavage activity of a BCL11A-specific ZFN containing a novel zinc finger backbone mutation of the invention at the off-target site NIFMAEVG (in this case, the three-letter abbreviation was used and the number of arginine residues mutated to the specified residue at position -5 was indicated; e.g., "6Gln" indicates that 6 arginines (3 per ZFN) were mutated to 6 glutamines in the ZFN pair). Error bars represent standard error. Experiments were performed in CD34+ cells at a dose of 2 μg mRNA per ZFN per experiment.

[0069] Figure 10 Figure depicting Western blot for detection of BCL11A-specific ZFN expression in CD34+ cells after transfection with either mRNA encoding two ZFN partners encoded on a single polynucleotide linked by a 2A sequence (51857- 2a -51949), compared to mRNA encoding the ZFN alone; or a mixture of two mRNAs encoding each partner at the indicated doses. Proteins were detected by anti-Flag antibody and the amount of protein expressed after mRNA transfection was shown to be consistent with the amount of mRNA used. As expected, the 2a construct produced a greater amount of the 5' ZFN SBS#51857 compared to the 3' ZFN, SBS#51949.

[0070] Figure 11 Figure depicting titration of administration of two BCl11A-specific ZFN partners 51949 and 51857 against the on-target position (BCL11A, left panel) or against the off-target position NIFMAEVG (right panel). Results show that changing the ratio of ZFN partners maintains on-target activity while reducing off-target activity (compare on-target activity of 85.92% indel at the BCL11A target at 60 μg of each mRNA, or 86.42% on-target activity at 60 μg 51949, 6.6 μg 51857 with reduced off-target (27.34% off-target activity at 60 μg of each mRNA compared to 4.21% indel when 60 μg 51949 was used with 6.6 μg 51857)).

[0071] Figure 12Table listing the on-target and off-target cleavage activities of a BCL11A-specific ZFN when CD34+ cells are treated with a single mRNA encoding two ZFN monomers (51857 / 51949 2a) as described above or with titrated doses of ZFN monomers, where one monomer (51949) contains a FokI R->S mutation at position 416. The 'license plate' identifiers are shown in Table 1 of Example 2 (SEQ ID NO:13-53). Data corresponding to PRJIYLFN represent the fraction of sequence reads containing indels at the expected target in BCL11A consistent with ZFN activity. Data corresponding to all other 'license plate' identifiers listed in the leftmost column correspond to confirmed or suspected off-target loci for the 51857 / 51949 ZFN pair. The ratios shown in the right column represent the activity in samples treated with 51857 / 51949 2a divided by the activity in samples treated with titrated 51857 / 51949 R416S at the designated locus.

[0072] Figure 13 Shows the results of an unbiased capture assay comparing two ZFN pairs. The left panel ("parental ZFN pair") shows the results using the SBS51857 and SBS51949 pair, and the right panel ("variant ZFN pair") shows the results using SBS63014 and SBS65721, which pair contains the parental pair as well as additional ZFP backbone mutations as described herein and a FokI R416S mutation on the SBS65721 construct. In particular, in the variant pair, each ZFN in the pair contains three R->Q mutations in the fingers, and the SBS65721 construct also contains a FokI R416S mutation. The data show that the mutations reduce the number of unique capture events from 21 positions for the parental pair to 4 for the variant. In addition, when the monomers in the ZFN pair are given in non-equal amounts, the capture events are also reduced. For the parental pair, the capture events decrease from 21 (equal dosing) to 13 (unequal dosing) positions (28% to 3.4% overall off-target, respectively), and for the variant pair, the capture events decrease from 4 to 2 (0.26% to 0.08% overall off-target cleavage, respectively). The combination of these two methods results in a reduction from 21 positions overall for the parental to 2 positions for the variant with unequal monomer concentration dosing, for an overall reduction in off-target events from 28% for the parental to 0.08% for the variant.

[0073] Figure 14 Shows the results of using a ZFN as described herein that exhibits reduced off-target cleavage events under large-scale manufacturing conditions. The ZFN pair used contains SBS63014 and SBS65722.

[0074] Figures 15A to 15DResults showing reduced off - target cleavage events using ZFN mutants (targeting AAVS1) as described herein. Figure 15A Depicts the activity results from parental ZFN 30035 / 30054. Figure 15B Depicts the ratios of on - target and on - target / off - target cleavage activities for three sets of FokI mutants: ELD FokI mutants with additional single mutations (left - most data set); KKR FokI mutants with additional single mutations (middle data set); and ELD and KKR FokI mutants with the same additional single mutation (right - most data set). Figure 15C Shows a grid of on - target activities where the ELD or KKR FokI domain contains two mutations, and Figure 15D Shows Figure 15C The on - target / off - target ratios of the data shown in

[0075] Figure 16A and Figure 16B Results showing reduced off - target cleavage events using exemplary AAVS1 - targeting ZFN mutants as described herein. Figure 16A Shows mutants in the ELD and KKR backgrounds, and Figure 16B Shows mutants in the ELD - KKR background.

[0076] Figure 17 Shows the alignment of FokI and FokI homologs (SEQ ID NO: 54 to 64). Shading indicates the degree of conservation. Numbering is according to the wild - type FokI domain (SEQ ID NO: 1).

[0077] Figure 18 Shows exemplary mutations where the positions correspond to Figure 17 the FokI or FokI homologs shown in Detailed Description

[0078] Methods and compositions for increasing the specificity of on - target engineered nuclease cleavage by differentially reducing off - target cleavage are disclosed herein. The methods involve reducing non - specific interactions between the FokI cleavage domain and DNA, reducing non - specific interactions between the zinc - finger backbone and DNA, and altering the relative ratio of each half - cleavage domain partner away from the default ratio of 1:1.

[0079] General Principles

[0080] Unless otherwise indicated, the practice of the methods, and the preparation and use of the compositions disclosed herein, employ conventional techniques in molecular biology, biochemistry, chromatin structure and analysis, computational chemistry, cell culture, recombinant DNA, and related fields within the skill of the art. These techniques are well explained in the literature. See, for example, Sambrook et al., MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed., Cold Spring Harbor Laboratory Press, 1989 and 3rd ed., 2001; Ausubel et al., CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, New York, 1987 and periodic updates; the series METHODS IN ENZYMOLOGY, Academic Press, San Diego; Wolffe, CHROMATIN STRUCTURE AND FUNCTION, 3rd ed., Academic Press, San Diego, 1998; METHODS IN ENZYMOLOGY, Vol. 304, "Chromatin" (P.M. Wassarman and A.P. Wolffe, eds.), Academic Press, San Diego, 1999; and METHODS IN MOLECULAR BIOLOGY, Vol. 119, "Chromatin Protocols" (P.B. Becker, ed.) Humana Press, Totowa, 1999.

[0081] Definitions

[0082] The terms "nucleic acid", "polynucleotide", and "oligonucleotide" are used interchangeably and refer to a polymer of deoxyribonucleotides or ribonucleotides in a linear or circular conformation, and in single-stranded or double-stranded form. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of the polymer. The terms can encompass known analogs of natural nucleotides, as well as nucleotides that are modified in the bases, sugars, and / or phosphate moieties (e.g., phosphorothioate backbones). In general, analogs of a particular nucleotide have the same base-pairing specificity; i.e., an analog of A will base-pair with T.

[0083] The terms "polypeptide", "peptide", and "protein" are used interchangeably to refer to a polymer of amino acid residues. The terms also apply to amino acid polymers in which one or more amino acids are chemical analogs or modified derivatives of the corresponding naturally occurring amino acids.

[0084] "Binding" refers to sequence-specific, non-covalent interactions between macromolecules (e.g., between a protein and a nucleic acid). Not all components of the binding interaction need to be sequence-specific (e.g., contacts with phosphate residues in the DNA backbone), as long as the interaction is sequence-specific overall. Such interactions are typically characterized by a dissociation constant (K -6 M -1 ) of 10 d or lower. "Affinity" refers to the strength of binding: increased binding affinity correlates with a lower K d . "Non-specific binding" refers to non-covalent interactions that occur between any target molecule (e.g., an engineered nuclease) and a macromolecule (e.g., DNA) that is independent of the target sequence.

[0085] "Binding protein" is a protein that can non-covalently bind to another molecule. A binding protein can bind to, for example, a DNA molecule (DNA-binding protein), an RNA molecule (RNA-binding protein), and / or a protein molecule (protein-binding protein). In the case of a protein-binding protein, it can bind to itself (to form a homodimer, homotrimer, etc.), and / or it can bind to one or more molecules of one or more different proteins. A binding protein can have more than one type of binding activity. For example, a zinc finger protein has DNA-binding, RNA-binding, and protein-binding activities. In the case of an RNA-guided nuclease system, the RNA guidance is heterologous to the nuclease component (Cas9 or Cfp1), and both can be engineered.

[0086] "DNA-binding molecule" is a molecule that can bind to DNA. Such DNA-binding molecules can be polypeptides, domains of proteins, domains within larger proteins, or polynucleotides. In some embodiments, the polynucleotide is DNA, while in other embodiments, the polynucleotide is RNA. In some embodiments, the DNA-binding molecule is the protein domain of a nuclease (e.g., the FokI domain), while in other embodiments, the DNA-binding molecule is the guide RNA component of an RNA-guided nuclease (e.g., Cas9 or Cfp1).

[0087] "DNA-binding protein" (or binding domain) is a protein or a domain within a larger protein that binds to DNA in a sequence-specific manner, for example, through one or more zinc fingers or through interactions with one or more RVDs in a zinc finger protein or TALE. The term zinc finger DNA-binding protein is often abbreviated as zinc finger protein or ZFP.

[0088] A "zinc finger DNA binding protein" (or binding domain) is a protein or a domain within a larger protein that binds DNA in a sequence-specific manner through one or more zinc fingers, where the one or more zinc fingers are regions of amino acid sequence within a binding domain whose structure is stabilized by zinc ion coordination. The term zinc finger DNA binding protein is often abbreviated as zinc finger protein or ZFP.

[0089] A "TALE DNA binding domain" or "TALE" is a polypeptide that contains one or more TALE repeat domains / units. The repeat domains are involved in the binding of the TALE to its cognate target DNA sequence. A single "repeat unit" (also referred to as "repeat sequence") is typically 33-35 amino acids in length and exhibits at least some sequence homology to other TALE repeat sequences within a naturally occurring TALE protein. See, e.g., U.S. Patent No. 8,586,526, which is incorporated herein by reference in its entirety.

[0090] Zinc finger and TALE DNA binding domains can be "engineered" to bind to a predetermined nucleotide sequence, e.g., via engineering of the recognition helix region of a naturally occurring zinc finger protein (altering one or more amino acids) or by engineering of the amino acids involved in DNA binding ("repeat variable diresidue" or RVD region). Thus, engineered zinc finger proteins or TALE proteins are non-naturally occurring proteins. Non-limiting examples of methods for engineering zinc finger proteins and TALEs are design and selection. Designed proteins are proteins that do not exist in nature and whose design / composition is primarily the result of rational criteria. Rational criteria for design include substitution rules and the application of computer algorithms for processing database information that stores information on existing ZFP and / or TALE designs and binding data. See, e.g., U.S. Patent No. 8,586,526; 6,140,081; 6,453,242; and 6,534,261; also see WO 98 / 53058; WO 98 / 53059; WO 98 / 53060; WO 02 / 016536 and WO 03 / 016496.

[0091] The "selected" zinc finger protein, TALE protein or CRISPR / Cas system is a protein not found in nature and the production of such protein is primarily the result of empirical processes such as phage display, interaction trap, rational design or hybrid selection. See, e.g., US 5,789,538; US 5,925,523; US 6,007,988; US 6,013,453; US 6,200,759; WO 95 / 19431; WO 96 / 06166; WO 98 / 53057; WO 98 / 54311; WO 00 / 27878; WO 01 / 60970; WO 01 / 88197 and WO 02 / 099084.

[0092] "TtAgo" is a prokaryotic Argonaute protein thought to be involved in gene silencing. TtAgo is derived from the bacterium Thermus thermophilus. See, e.g., Swarts et al., supra; G. Sheng et al., (2013) Proc. Natl. Acad. Sci. U.S.A. 111, 652). The "TtAgo system" is all of the components required, including, for example, the guide DNA for cleavage by the TtAgo enzyme.

[0093] "Recombination" refers to the process of exchange of genetic information between two polynucleotides, including but not limited to capture by non-homologous end joining (NHEJ) and homologous recombination. For the purposes of this disclosure, "homologous recombination (HR)" refers to a particular form of such exchange that occurs, for example, during double-strand break repair in cells via a homologous-directed repair mechanism. This process requires nucleotide sequence homology, uses a "donor" molecule as a repair template for the "target" molecule (i.e., the one that has undergone the double-strand break), and this process is referred to as "non-crossover gene conversion" or "short tract gene conversion", respectively, because it results in the transfer of genetic information from the donor to the target. Without wishing to be bound by any particular theory, such transfer may involve mismatch correction of heteroduplex DNA formed between the broken target and the donor, and / or "synthesis-dependent strand annealing", where the donor is used to resynthesize genetic information that will become part of the target, and / or related processes. This particular HR often results in a change in the sequence of the target molecule such that part or all of the sequence of the donor polynucleotide is incorporated into the target polynucleotide.

[0094] In certain methods of the present disclosure, one or more targeted nucleases as described herein create double-strand breaks (DSBs) in a target sequence (e.g., cellular chromatin) at a pre-determined site (e.g., a target gene or locus). The DSBs mediate the integration of a construct (e.g., a donor) as described herein. Optionally, the construct is homologous to the nucleotide sequence in the break region. The expression construct can be physically integrated or, alternatively, the expression cassette is used as a template for repairing the break via homologous recombination, resulting in the introduction of all or part of the nucleic acid sequence in the expression cassette into the cellular chromatin. Thus, the first sequence in the cellular chromatin can be altered and, in certain embodiments, can be converted to the sequence present in the expression cassette. Thus, the use of the term “replace” or “replacement” can be understood to mean the replacement of one nucleotide sequence with another nucleotide sequence (i.e., replacement of the sequence in an informational sense), and does not necessarily require the physical or chemical replacement of one polynucleotide with another polynucleotide.

[0095] In any of the methods described herein, additional engineered nucleases can be used for additional double-strand cleavage of additional target sites within the cell.

[0096] In embodiments of methods for targeted recombination and / or replacement and / or alteration of a sequence in a target region in cellular chromatin, the chromosomal sequence is altered by homologous recombination with an exogenous “donor” nucleotide sequence. If a sequence homologous to the break region exists, then such homologous recombination can be stimulated by the presence of a double-strand break in the cellular chromatin.

[0097] In any of the methods described herein, a first nucleotide sequence (“donor sequence”) can contain a sequence that is homologous but different from the genomic sequence in the target region, thereby stimulating homologous recombination to insert a non-identical sequence in the target region. Thus, in certain embodiments, the portion of the donor sequence that is homologous to the sequence in the target region exhibits sequence identity of between about 80% and 99% (or any integer therebetween) with the genomic sequence being replaced. In other embodiments, the homology between the donor and genomic sequences is higher than 99%, e.g., if there is only 1 nucleotide difference between a donor having more than 100 contiguous base pairs and the genomic sequence. In certain cases, the non-homologous portion of the donor sequence can contain a sequence that is not present in the region of interest, such that a new sequence is introduced into the region of interest. In these cases, the non-homologous sequence is generally flanked by 50 to 1,000 base pairs (or any integer value therebetween) or any number of base pairs greater than 1,000 that are homologous or identical to the sequence in the region of interest. In other embodiments, the donor sequence is non-homologous to the first sequence and the donor sequence is inserted into the genome via a non-homologous recombination mechanism.

[0098] Any method described herein can be used for partial or complete inactivation of one or more target sequences in a cell by targeted integration of a donor sequence that disrupts expression of one or more target genes or via cleavage of one or more target sequences that disrupt expression of one or more target genes, followed by error-prone NHEJ-mediated repair. Cell lines with partially or completely inactivated genes are also provided.

[0099] In addition, the targeted integration methods described herein can also be used to integrate one or more exogenous sequences. Exogenous nucleic acid sequences can include, for example, one or more genes or cDNA molecules, or any type of coding or non-coding sequence, as well as one or more control elements (e.g., promoters). Additionally, the exogenous nucleic acid sequences can produce one or more RNA molecules (e.g., short hairpin RNA (shRNA), inhibitory RNA (RNAi), microRNA (miRNA), etc.).

[0100] "Cleavage" refers to the breakage of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods, including but not limited to enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand cleavage and double-strand cleavage are possible, and double-strand cleavage can occur as a result of two distinct single-strand cleavage events. DNA cleavage can result in the generation of blunt ends or staggered ends. In certain embodiments, a fusion polypeptide is used to target double-strand DNA cleavage.

[0101] "Cleavage half-domain" is a polypeptide sequence that, when linked to a second polypeptide (identical or different), forms a complex having cleavage activity (preferably double-strand cleavage activity). The terms "first and second cleavage half-domains"; "+ and - cleavage half-domains" and "right and left cleavage half-domains" can be used interchangeably to refer to a dimerized pair of cleavage half-domains. The term "cleavage domain" can be used interchangeably with the term "cleavage half-domain". The term "FokI cleavage domain" includes the FokI sequence as shown in SEQ ID NO:1 and any FokI homologs, including but not limited to Figure 17 the sequences shown in

[0102] "Engineered cleavage half-domain" is a cleavage half-domain that has been modified to form an obligate heterodimer with another cleavage half-domain (e.g., another engineered cleavage half-domain).

[0103] The term "sequence" refers to a nucleotide sequence of any length, which can be DNA or RNA; it can be linear, circular or branched, and can be single-stranded or double-stranded. The term "transgene" refers to a nucleotide sequence inserted into the genome. A transgene can have any length, for example, a length between 2 and 100,000,000 nucleotides (or any integer value therebetween or thereon), preferably a length between about 100 and 100,000 nucleotides (or any integer therebetween), more preferably a length between about 2000 and 20,000 nucleotides (or any value therebetween) and even more preferably between about 5 and 15 kb (or any value therebetween).

[0104] A "chromosome" is a chromatin complex that contains all or a portion of the genome of a cell. The genome of a cell is typically characterized by its karyotype, which is the collection of all the chromosomes that contain the genome of the cell. The genome of a cell can contain one or more chromosomes.

[0105] An "episome" is a replicative nucleic acid, nucleoprotein complex or other structure that contains nucleic acids that are not part of the chromosomal karyotype of a cell. Examples of episomes include plasmids, minicircles and certain viral genomes. The liver-specific constructs described herein can be episomally maintained or can be stably integrated into the cell.

[0106] An "exogenous" molecule is a molecule that is not normally present in a cell but can be introduced into the cell by one or more genetic, biochemical or other means. "Normally present in a cell" is determined relative to the specific developmental stage and environmental conditions of the cell. Thus, for example, a molecule that is present only during embryonic muscle development is an exogenous molecule relative to an adult muscle cell. Similarly, a molecule induced by heat shock is an exogenous molecule relative to a non-heat-shocked cell. Exogenous molecules can include, for example, an operating form of an endogenous molecule that is not functioning properly or a non-operating form of an endogenous molecule that is functioning properly.

[0107] In other cases, the exogenous molecule can be a small molecule such as those generated by combinatorial chemistry processes, or a macromolecule such as a protein, nucleic acid, carbohydrate, lipid, glycoprotein, lipoprotein, polysaccharide, any modified derivative of the foregoing molecules, or any complex comprising one or more of the foregoing molecules. Nucleic acids include DNA and RNA, can be single-stranded or double-stranded; can be linear, branched, or circular; and can have any length. Nucleic acids include those capable of forming duplexes as well as triplexes. See, for example, U.S. Patent Nos. 5,176,996 and 5,422,251. Proteins include, but are not limited to, DNA-binding proteins, transcription factors, chromatin remodeling factors, methylated DNA-binding proteins, polymerases, methylases, demethylases, acetyltransferases, deacetylases, kinases, phosphatases, ligases, deubiquitinating enzymes, integrases, recombinases, ligases, topoisomerases, gyrases, and helicases.

[0108] The exogenous molecule can be a molecule of the same type as an endogenous molecule, such as an exogenous protein or nucleic acid. For example, exogenous nucleic acids can include infectious viral genomes, plasmids or episomes introduced into cells, or chromosomes that are not normally present in the cell. Methods for introducing exogenous molecules into cells are known to those of skill in the art and include, but are not limited to, lipid-mediated transfer (i.e., liposomes, including neutral and cationic lipids), electroporation, direct injection, cell fusion, particle bombardment, calcium phosphate co-precipitation, DEAE-dextran-mediated transfer, and virus vector-mediated transfer. The exogenous molecule can also be a molecule of the same type as an endogenous molecule, but derived from a species different from the species from which the cell is produced. For example, a human nucleic acid sequence can be introduced into a cell line originally derived from a mouse or a hamster. Methods for introducing exogenous molecules into plant cells are known to those of skill in the art and include, but are not limited to, protoplast transformation, silicon carbide (e.g., WHISKERS TM ), Agrobacterium-mediated transformation, lipid-mediated delivery (i.e., liposomes, including neutral and cationic lipids), electroporation, direct injection, cell fusion, particle bombardment (e.g., using a "gene gun"), calcium phosphate co-precipitation, DEAE-dextran-mediated transfer, and virus vector-mediated transfer.

[0109] In contrast, an "endogenous" molecule is a molecule that is normally present in a particular cell under particular environmental conditions at a particular stage of development. For example, endogenous nucleic acids can include chromosomes, the genomes of mitochondria, chloroplasts, or other organelles, or naturally occurring episomal nucleic acids. Additional endogenous molecules can include proteins, such as transcription factors and enzymes.

[0110] As used herein, the term "product of an exogenous nucleic acid" includes polynucleotide and polypeptide products, such as transcriptional products (polynucleotides, such as RNA) and translational products (polypeptides).

[0111] A "fusion" molecule is a molecule in which two or more subunit molecules are preferably covalently linked. The subunit molecules can be molecules of the same chemical type or can be molecules of different chemical types. Examples of fusion molecules include, but are not limited to, fusion proteins (e.g., a fusion between a protein DNA-binding domain and a cleavage domain), a fusion between a polynucleotide DNA-binding domain (e.g., sgRNA) operably associated with a cleavage domain, and fusion nucleic acids (e.g., a nucleic acid encoding a fusion protein).

[0112] Expression of a fusion protein in a cell can be caused by delivery of the fusion protein to the cell or by delivery of a polynucleotide encoding the fusion protein to the cell, wherein the polynucleotide is transcribed and the transcript is translated to produce the fusion protein. Trans-splicing, polypeptide cleavage, and polypeptide ligation can also be involved in protein expression in a cell. Methods for delivering polynucleotides and polypeptides to cells are presented elsewhere in the present disclosure.

[0113] For the purposes of the present disclosure, a "gene" includes the DNA region encoding a gene product (see below), as well as all DNA regions that regulate the production of the gene product, whether or not the regulatory sequences are adjacent to the coding and / or transcriptional sequences. Thus, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions.

[0114] "Gene expression" refers to the conversion of the information contained in a gene into a gene product. The gene product can be the direct transcriptional product of the gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, structural RNA, or any other type of RNA) or a protein produced by translation of the mRNA. Gene products also include RNAs modified by methods such as capping, polyadenylation, methylation, and editing, as well as proteins modified by, for example, methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristoylation, and glycosylation.

[0115] "Regulation" of gene expression refers to a change in gene activity. Expression regulation can include, but is not limited to, gene activation and gene repression. Gene editing (e.g., cleavage, alteration, inactivation, random mutagenesis) can be used to regulate expression. Gene inactivation refers to any decrease in gene expression compared to a cell that does not include a ZFP, TALE, or CRISPR / Cas system as described herein. Thus, gene inactivation can be partial or complete.

[0116] "Target region" is any chromatin region of a cell, such as a gene or non-coding sequence within or near a gene, where it is desirable to bind an exogenous molecule. The binding can be for the purpose of targeted DNA cleavage and / or targeted recombination. Target regions can be present, for example, in chromosomes, episomes, organelle genomes (e.g., mitochondria, chloroplasts), or infecting viral genomes. A target region can be within a gene coding region, within a transcribed non-coding region such as, for example, a leader sequence, a trailing sequence, or an intron, or within a non-transcribed region upstream or downstream of the coding region. The length of a target region can be as small as a single nucleotide pair, or as large as 2,000 nucleotide pairs, or any integer value of nucleotide pairs.

[0117] A "safe harbor" locus is a locus within a genome where a gene can be inserted without any deleterious effects on the host cell. Most preferably, a safe harbor locus is one where the expression of the inserted gene sequence is not disrupted by any read-through expression from adjacent genes. Non-limiting examples of safe harbor loci targeted by one or more nucleases include CCR5, HPRT, AAVS1, Rosa, and albumin. See, for example, U.S. Patent Nos. 7,951,925; 8,771,985; 8,110,379; 7,951,925; U.S. Publication Nos. 20100218264; 20110265198; 20130137104; 20130122591; 20130177983; 20130177960; 20150056705; and 20150159172.

[0118] A "reporter gene" or "reporter sequence" refers to any sequence that produces a protein product that is preferably, but not necessarily, easily measurable in a conventional assay. Suitable reporter genes include, but are not limited to, sequences encoding proteins that mediate antibiotic resistance (e.g., ampicillin resistance, neomycin resistance, G418 resistance, puromycin resistance), sequences encoding colored or fluorescent or luminescent proteins (e.g., green fluorescent protein, enhanced green fluorescent protein, red fluorescent protein, luciferase), and proteins that mediate enhanced cell growth and / or gene amplification (e.g., dihydrofolate reductase). Epitope tags include, for example, one or more copies of FLAG, His, myc, Tap, HA, or any detectable amino acid sequence. An "expression tag" includes a sequence encoding a reporter gene that can be operably linked to a desired gene sequence in order to monitor the expression of the target gene.

[0119] "Eukaryotic" cells include, but are not limited to, fungal cells (such as yeast), plant cells, animal cells, mammalian cells, and human cells (e.g., T cells), including stem cells (pluripotent and multipotent cells).

[0120] The terms "operative linkage" and "operatively linked" (or "operably linked") are used interchangeably with respect to two or more juxtaposed components, such as sequence elements, where the components are arranged so that the components function in a normal manner and allow the possibility that at least one component may mediate a function exerted on at least one other component. By way of illustration, if a transcriptional regulatory sequence controls the level of transcription of a coding sequence in response to the presence or absence of one or more transcriptional regulatory factors, the transcriptional regulatory sequence (such as a promoter) is operably linked to the coding sequence. Transcriptional regulatory sequences are generally operably linked to the coding sequence in cis, but need not be directly adjacent to the coding sequence. For example, enhancers are transcriptional regulatory sequences that are operably linked to the coding sequence, although they are not contiguous.

[0121] A "functional fragment" of a protein, polypeptide or nucleic acid is a protein, polypeptide or nucleic acid whose sequence is different from the full-length protein, polypeptide or nucleic acid, but still retains the same function as the full-length protein, polypeptide or nucleic acid. A functional fragment may have more, fewer or the same number of residues compared to the corresponding native molecule, and / or may contain one or more amino acid or nucleic acid substitutions. Methods for determining nucleic acid function (e.g., coding function, ability to hybridize to another nucleic acid) are well known in the art.

[0122] A polynucleotide "vector" or "construct" is capable of transferring a gene sequence to a target cell. Generally, "vector construct", "expression vector", "expression construct", "expression cassette" and "gene transfer vector" refer to any nucleic acid construct capable of directing the expression of a target gene and transferring the gene sequence to a target cell. Thus, the terms include cloning, and expression mediators, as well as integration vectors.

[0123] The terms "subject" and "patient" are used interchangeably and refer to mammals such as human patients and non-human primates as well as laboratory animals such as rabbits, dogs, cats, rats, mice, and other animals. Thus, as used herein, the term "subject" or "patient" means any mammalian patient or subject to whom the expression cassette of the present invention can be administered. Subjects of the present invention include subjects suffering from a disorder.

[0124] As used herein, the terms "treating" and "treatment" refer to reducing the severity and / or frequency of symptoms, eliminating symptoms and / or underlying causes, preventing the occurrence of symptoms and / or their underlying causes, and improving or repairing damage. Cancer, monogenic diseases and graft-versus-host disease are non-limiting examples of conditions that can be treated using the compositions and methods described herein.

[0125] "Chromatin" is the nucleoprotein structure that contains the cell's genome. Cellular chromatin includes nucleic acids (primarily DNA) and proteins, and the proteins include histone and non-histone chromosomal proteins. Most eukaryotic cell chromatin exists in the form of nucleosomes, where the nucleosome core includes approximately 150 DNA base pairs associated with an octamer, and the octamer includes two each of histones H2A, H2B, H3, and H4; and linker DNA (of variable length depending on the organism) extends between the nucleosome cores. Histone H1 molecules are generally associated with the linker DNA. For the purposes of this disclosure, the term "chromatin" is intended to encompass all types of prokaryotic and eukaryotic nuclear proteins. Cellular chromatin includes chromosomal chromatin and episomal chromatin.

[0126] "Accessible region" is a site in cellular chromatin at which a target site present in a nucleic acid can be bound by an exogenous molecule that recognizes the target site. Without wishing to be bound by any particular theory, it is believed that accessible regions are regions that are not packaged into nucleosome structures. The unique structure of accessible regions can often be detected by their sensitivity to chemical and enzymatic probes (such as nucleases).

[0127] "Target site" or "target sequence" is a nucleic acid sequence that defines the portion of a nucleic acid to which a binding molecule will bind (provided that sufficient conditions for binding exist). For example, the sequence 5'-GAATTC-3' is a target site for the Eco RI restriction endonuclease. An "intended" or "on-target" sequence is the sequence to which a binding molecule is intended to bind, and an "unintended" or "off-target" sequence includes any sequence bound by a binding molecule that is not the intended target.

[0128] DNA binding molecule / domain

[0129] Compositions are described herein that include a DNA binding molecule / domain that specifically binds to a target site in any target gene or locus. Any DNA-binding molecule / domain can be used in the compositions and methods disclosed herein, including but not limited to zinc finger DNA-binding domains, TALE DNA-binding domains, the DNA-binding portion (guide or sgRNA) of a CRISPR / Cas nuclease, or a DNA-binding domain from a meganuclease.

[0130] In certain embodiments, the DNA binding domain comprises a zinc finger protein. Preferably, the zinc finger protein is non-naturally occurring as it has been engineered to bind to a selected target site. See, for example, Beerli et al. (2002) Nature Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nature Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416; U.S. Patent Nos. 6,453,242; 6,534,261; 6,599,692; 6,503,717; 6,689,558; 7,030,215; 6,794,136; 7,067,317; 7,262,054; 7,070,934; 7,361,635; 7,253,273; and U.S. Patent Publication Nos. 2005 / 0064474; 2007 / 0218528; 2005 / 0267061, which are hereby incorporated by reference in their entirety. In certain embodiments, the DNA binding domain comprises a zinc finger protein disclosed in U.S. Patent Publication No. 2012 / 0060230 (e.g., Table 1), which is hereby incorporated by reference.

[0131] Engineered zinc finger binding domains can have novel binding specificities compared to naturally occurring zinc finger proteins. Engineering methods include, but are not limited to, rational design and various types of selection. Rational design includes, for example, using databases that contain triplet (or quadruplet) nucleotide sequences and individual zinc finger amino acid sequences, where each triplet or quadruplet nucleotide sequence is associated with the amino acid sequence of one or more zinc fingers that bind to the specific triplet or quadruplet sequence. See, for example, U.S. Patent 6,453,242 and 6,534,261, which are hereby incorporated by reference.

[0132] Exemplary selection methods, including phage display and two-hybrid systems, are disclosed in U.S. Pat. Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759; and 6,242,568; as well as WO98 / 37186; WO 98 / 53057; WO 00 / 27878; WO 01 / 88197 and GB 2,338,237. In addition, enhancement of the binding specificity of zinc finger binding domains has been described, for example, in U.S. Patent No. 6,794,136.

[0133] In addition, as disclosed in these and other references, zinc finger domains and / or multi-fingered zinc finger proteins can be joined together using any suitable linker sequence (including, for example, linkers of 5 or more amino acids in length). For exemplary linker sequences of 6 or more amino acids in length, also see U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949. The proteins described herein can include any combination of suitable linkers between individual protein zinc fingers. In addition, enhancement of the binding specificity of zinc finger binding domains has been described, for example, in U.S. Patent No. 6,794,136.

[0134] Selection of target sites; ZFPs and methods for designing and constructing fusion proteins (and polynucleotides encoding said fusion proteins) are known to those of skill in the art and are described in detail in U.S. Patent Nos. 6,140,081; 5,789,538; 6,453,242; 6,534,261; 5,925,523; 6,007,988; 6,013,453; 6,200,759; WO 95 / 19431; WO96 / 06166; WO 98 / 53057; WO 98 / 54311; WO 00 / 27878; WO 01 / 60970WO 01 / 88197; WO 02 / 099084; WO 98 / 53058; WO 98 / 53059; WO 98 / 53060; WO 02 / 016536 and WO 03 / 016496.

[0135] In addition, as disclosed in these and other references, zinc finger domains and / or multi-fingered zinc finger proteins can be joined together using any suitable linker sequence (including, for example, linkers of 5 or more amino acids in length). For exemplary linker sequences of 6 or more amino acids in length, also see U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949. The proteins described herein can include any combination of suitable linkers between individual protein zinc fingers.

[0136] Generally, the ZFP contains at least three fingers. Some ZFPs contain four, five, or six fingers. A ZFP containing three fingers generally recognizes a target site containing 9 or 10 nucleotides; a ZFP containing four fingers generally recognizes a target site containing 12 to 14 nucleotides; and a ZFP with six fingers can recognize a target site containing 18 to 21 nucleotides. The ZFP can also be a fusion protein containing one or more regulatory domains, which can be transcriptional activation or repression domains.

[0137] In some embodiments, the DNA-binding domain can be derived from a nuclease. For example, the recognition sequences of homing endonucleases and meganucleases such as I-SceI, I-CeuI, PI-PspI, PI-Sce, I-SceIV, I-CsmI, I-PanI, I-SceII, I-PpoI, I-SceIII, I-CreI, I-TevI, I-TevII, and I-TevIII are known. See also U.S. Patent No. 5,420,032; U.S. Patent No. 6,833,252; Belfort et al. (1997) Nucleic Acids Res. 25:3379–3388; Dujon et al. (1989) Gene 82:115–118; Perler et al. (1994) Nucleic Acids Res. 22, 1125–1127; Jasin (1996) Trends Genet. 12:224–228; Gimble et al. (1996) J. Mol. Biol. 263:163–180; Argast et al. (1998) J. Mol. Biol. 280:345–353 and the New England Biolabs catalog. Additionally, the DNA-binding specificities of homing endonucleases and meganucleases can be engineered to bind non-native target sites. See, for example, Chevalier et al. (2002) Molec. Cell 10:895-905; Epinat et al. (2003) Nucleic Acids Res. 31:2952-2962; Ashworth et al. (2006) Nature 441:656-659; Paques et al. (2007) Current Gene Therapy 7:49-66; U.S. Patent Publication No. 20070117128.

[0138] In certain embodiments, the zinc finger proteins used with the mutant cleavage domains described herein comprise mutations (substitutions, deletions, and / or insertions) in the backbone region (e.g., the region outside of the 7 - amino acid recognition helix region numbered - 1 to 6), such as one or more mutations at one or more of positions - 14, - 9, and / or - 5 (see, for example Figure 5A ). The wild - type residues at one or more of these positions may be deleted, replaced with any amino acid residue, and / or contain one or more additional residues. In some embodiments, the Arg (R) at position - 5 is changed to Tyr (Y), Asp (N), Glu (E), Leu (L), Gln (Q), or Ala (A). In other embodiments, the Arg (R) at position (- 9) is replaced with Ser (S), Asp (N), or Glu (E). In other embodiments, the Arg (R) at position (- 14) is replaced with Ser (S) or Gln (Q). In other embodiments, the fusion polypeptide may comprise mutations in the zinc finger DNA - binding domain, wherein the amino acids at positions (- 5), (- 9), and / or (- 14) are changed to any of the amino acids listed above in any combination.

[0139] In other embodiments, the DNA binding domain comprises an engineered domain from a transcription activator-like (TAL) effector (TALE), which is similar to those derived from the plant pathogen Xanthomonas (see Boch et al., (2009) Science 326:1509-1512 and Moscou and Bogdanove, (2009) Science 326:1501) and Ralstonia (see Heuer et al. (2007) Applied and Environmental Microbiology 73(13):4379-4384); those of U.S. Patent Publication Nos. 20110301073 and 20110145940. Plant pathogens of the genus Xanthomonas are known to cause many diseases in important crops. The pathogenicity of Xanthomonas depends on the conserved type III secretion (T3S) system, which injects more than 25 different effector proteins into plant cells. Among these injected proteins are the transcription activator-like effectors (TALEs), which mimic plant transcription activators and manipulate the plant transcriptome (see Kay et al. (2007) Science 318:648-651). These proteins contain a DNA binding domain and a transcription activation domain. One of the best-characterized TALE effectors is AvrBs3 from Xanthomonas campestgris pv. Vesicatoria (see Bonas et al. (1989) Mol Gen Genet 218:127-136 and WO2010079430). TALEs contain a central domain of tandem repeats, each repeat containing approximately 34 amino acids, which are critical for the DNA binding specificity of these proteins. Additionally, it contains a nuclear localization sequence and an acidic transcription activation domain (for a review, see Schornack S, et al. (2006) J Plant Physiol 163(3):256-272). Additionally, in the plant pathogenic bacterium Ralstonia solanacearum, two genes designated brg11 and hpx17 have been found, which are homologous to the AvrBs3 family of Xanthomonas in Ralstonia solanacearum biovar 1 strain GMI1000 and biovar 4 strain RS1000 (see Heuer et al. (2007) Appl and Envir Micro 73(13):4379-4384). These genes are 98.9% identical to each other in nucleotide sequence but differ in that there is a 1,575 base pair deletion in the repeat domain of hpx17. However, the two gene products have less than 40% sequence identity with the AvrBs3 family proteins of Xanthomonas.

[0140] The specificity of these TAL effectors depends on the sequences found in the tandem repeats. The repeats contain approximately 102 base pairs and the repeats are typically 91-100% homologous to each other (Bonas et al., supra). The polymorphism of the repeats is typically located at positions 12 and 13 and there appears to be a one-to-one correspondence between the identity of the hypervariable di-residues (repeat variable di-residues or RVD region) at positions 12 and 13 and the identity of the adjacent nucleotides in the target sequence of the TAL-effector (see Moscou and Bogdanove, (2009) Science 326:1501 and Boch et al. (2009) Science 326:1509-1512). Experimentally, the natural code for DNA recognition by these TAL-effectors has been determined such that the HD sequence (repeat variable di-residue or RVD) at positions 12 and 13 results in binding to cytosine (C), NG binds to T, NI binds to A, C, G or T, NN binds to A or G, and ING binds to T. These DNA-binding repeats have been assembled into proteins in new combinations and multiple repeats to produce artificial transcription factors that are capable of interacting with new sequences and activating the expression of non-endogenous reporter genes in plant cells (Boch et al., supra). Engineered TAL proteins have been linked to the FokI cleavage half-domain to produce TAL effector domain nuclease fusions (TALENs), including TALENs with non-canonical RVDs. See, e.g., U.S. Patent No. 8,586,526.

[0141] In some embodiments, the TALEN comprises an endonuclease (e.g., FokI) cleavage domain or cleavage half-domain. In other embodiments, the TALE-nuclease is a mega TAL. These mega TAL nucleases are fusion proteins that comprise a TALE DNA-binding domain and a meganuclease cleavage domain. The meganuclease cleavage domain functions as a monomer and does not require dimerization to obtain activity. (See Boissel et al., (2013) Nucl Acid Res:1-13, doi:10.1093 / nar / gkt1224).

[0142] In a further embodiment, the nuclease comprises a compact TALEN. These nucleases are single-chain fusion proteins that link a TALE DNA-binding domain to a TevI nuclease domain. The fusion protein can act as a nickase targeted by the TALE region, or can generate a double-strand break, depending on the position of the TALE DNA-binding domain relative to the TevI nuclease domain (see Beurdeley et al. (2013) Nat Comm: 1-8 DOI: 10.1038 / ncomms2782). Additionally, the nuclease domain can also exhibit DNA-binding function. Any TALEN can be used in combination with another TALEN having one or more mega-TALEs (e.g., one or more TALENs (cTALEN or FokI-TALEN)).

[0143] Additionally, as disclosed in these and other references, zinc finger domains and / or multi-fingered zinc finger proteins or TALEs can be linked together using any suitable linker sequence (including, for example, linkers of 5 or more amino acids in length). For exemplary linker sequences of 6 or more amino acids in length, also see U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949. The proteins described herein can include any combination of suitable linkers between individual protein zinc fingers. Additionally, enhancement of the binding specificity for zinc finger binding domains has been described, for example, in U.S. Patent No. 6,794,136. In certain embodiments, the DNA-binding domain is part of a CRISPR / Cas nuclease system, including a single-guide RNA (sgRNA) DNA-binding molecule that binds to DNA. See, for example, U.S. Patent No. 8,697,359 and U.S. Patent Publication Nos. 20150056705 and 20150159172. The CRISPR (clustered regularly interspaced short palindromic repeats) locus encoding the RNA component of the system and the cas (CRISPR-associated) locus encoding the protein (Jansen et al., 2002. Mol. Microbiol. 43:1565-1575; Makarova et al., 2002. Nucleic Acids Res. 30:482-496; Makarova et al., 2006. Biol. Direct 1:7; Haft et al., 2005. PLoS Comput. Biol. 1:e60) make up the gene sequences of the CRISPR / Cas nuclease system. The CRISPR locus in a microbial host contains a combination of CRISPR-associated (Cas) genes and non-coding RNA elements capable of programming the specificity of CRISPR-mediated nucleic acid cleavage.

[0144] In some embodiments, the DNA binding domain is part of the TtAgo system (see Swarts et al., supra; Sheng et al., supra). In eukaryotes, gene silencing is mediated by the Argonaute (Ago) protein family. In this paradigm, Ago binds small (19 - 31 nt) RNAs. This protein - RNA silencing complex recognizes target RNAs via Watson - Crick base - pairing between the small RNA and the target, and cleaves the target RNA endonucleolytically (Vogel (2014) Science 344:972 - 973). In contrast, prokaryotic Ago proteins bind to small single - stranded DNA fragments and may be used to detect and remove foreign (often viral) DNA (Yuan et al., (2005) Mol. Cell 19, 405; Olovnikov, et al. (2013) Mol. Cell 51, 594; Swarts et al., supra). Exemplary prokaryotic Ago proteins include those from Aquifex aeolicus, Rhodobacter sphaeroides, and Thermus thermophilus.

[0145] One of the best - characterized prokaryotic Ago proteins is the one from Thermus thermophilus (TtAgo; Swarts et al., supra). TtAgo associates with 15 - nt or 13 - 25 - nt single - stranded DNA fragments that have a 5' phosphate group. This "guide DNA" bound by TtAgo is used to direct the protein - DNA complex to bind to a Watson - Crick complementary DNA sequence in a third - party DNA molecule. Once the sequence information in these guide DNAs allows for the recognition of the target DNA, the TtAgo - guide DNA complex cleaves the target DNA. This mechanism is also supported by the structure of the TtAgo - guide DNA complex when it binds its target DNA (G. Sheng et al., supra). The Ago from Rhodobacter sphaeroides (RsAgo) has similar properties (Olivnikov et al., supra).

[0146] An exogenous guide DNA having any DNA sequence can be loaded onto the TtAgo protein (Swarts et al., supra). Since the specificity of TtAgo cleavage is guided by the guide DNA, the TtAgo-DNA complex formed with exogenous, researcher-specified guide DNA will thus direct TtAgo target DNA cleavage to complementary, researcher-specified target DNA. In this way, targeted double-strand breaks can be generated in DNA. The use of the TtAgo guide DNA system (or orthologous Ago-guide DNA systems from other organisms) allows for the targeted cleavage of genomic DNA within cells. Such cleavage can be single-stranded or double-stranded. For the cleavage of mammalian genomic DNA, a codon-optimized form of TtAgo for expression in mammalian cells will preferably be used. Additionally, cells can preferably be treated with an in vitro-formed TtAgo-DNA complex in which the TtAgo protein is fused to a cell-penetrating peptide. Additionally, a form of the TtAgo protein that has been altered by mutagenesis to have improved activity at 37°C can preferably be used. Ago-RNA-mediated DNA cleavage can be used to effect a series of outcomes including, but not limited to, gene knockout, targeted gene addition, gene correction, and targeted gene deletion, using the art-standard techniques for developing DNA breaks.

[0147] Accordingly, any DNA binding molecule / domain can be used.

[0148] Fusion molecule

[0149] Also provided are fusion molecules that comprise a DNA binding domain as described herein (e.g., a ZFP or TALE, a CRISPR / Cas component such as a single guide RNA) and a heterologous regulatory (functional) domain (or a functional fragment thereof). Common domains include, for example, transcription factor domains (activators, repressors, co-activators, co-repressors), silencers, oncogenes (e.g., myc, jun, fos, myb, max, mad, rel, ets, bcl, myb, mos family members, etc.); DNA repair enzymes and their associated factors and modifiers; DNA rearrangement enzymes and their associated factors and modifiers; chromatin-associated proteins and their modifiers (e.g., kinases, acetyltransferases, and deacetylases); and DNA modifying enzymes (e.g., methyltransferases, topoisomerases, helicases, ligases, kinases, phosphatases, polymerases, endonucleases) and their associated factors and modifiers. For details regarding fusions of DNA binding domains and nuclease cleavage domains, see U.S. Patent Publications No. 20050064474; 20060188987, and 2007 / 0218528, which are incorporated herein by reference in their entireties.

[0150] Suitable domains for achieving activation include the HSV VP16 activation domain (see, e.g., Hagmann et al., J. Virol. 71, 5952 - 5962 (1997)), nuclear hormone receptors (see, e.g., Torchia et al., Curr. Opin. Cell Biol. 10:373 - 383 (1998)); the p65 subunit of nuclear factor κB (Bitko & Barik, J. Virol. 72:5610 - 5618 (1998) and Doyle & Hunt, Neuroreport 8:2937 - 2942 (1997)); Liu et al., Cancer Gene Ther. 5:3 - 28 (1998)) or artificial chimeric functional domains such as VP64 (Beerli et al., (1998) Proc. Natl. Acad. Sci. USA 95:14623 - 33) and degrons (Molinari et al., (1999) EMBO J. 18, 6439 - 6447). Additional exemplary activation domains include Oct1, Oct - 2A, Sp1, AP - 2, and CTF1 (Seipel et al., EMBO J. 11, 4961 - 4968 (1992) as well as p300, CBP, PCAF, SRC1 PvALF, AtHD2A, and ERF - 2. See, e.g., Robyr et al. (2000) Mol. Endocrinol. 14:329 - 347; Collingwood et al. (1999) J. Mol. Endocrinol. 23:255 - 275; Leo et al. (2000) Gene 245:1 - 11; Manteuffel - Cymborowska (1999) Acta Biochim. Pol. 46:77 - 89; McKenna et al. (1999) J. Steroid Biochem. Mol. Biol. 69:3 - 12; Malik et al. (2000) Trends Biochem. Sci. 25:277 - 283; and Lemon et al. (1999) Curr. Opin. Genet. Dev. 9:499 - 504. Additional exemplary activation domains include but are not limited to OsGAI, HALF - 1, C1, AP1, ARF - 5, - 6, - 7, and - 8, CPRF1, CPRF4, MYC - RP / GP, and TRAB1.See, for example, Ogawa et al. (2000) Gene 245:21-29; Okanami et al. (1996) Genes Cells 1:87-99; Goff et al. (1991) Genes Dev. 5:298-309; Cho et al. (1999) Plant Mol. Biol. 40:419-429; Ulmason et al. (1999) Proc. Natl. Acad. Sci. USA 96:5844-5849; Sprenger-Haussels et al. (2000) Plant J. 22:1-8; Gong et al. (1999) Plant Mol. Biol. 41:33-44; and Hobo et al. (1999) Proc. Natl. Acad. Sci. USA 96:15,348-15,353.

[0151] It will be clear to those skilled in the art that in forming a fusion protein (or nucleic acid encoding said fusion protein) between a DNA binding domain and a functional domain, an activation domain or a molecule that interacts with an activation domain is suitable as a functional domain. Basically any molecule capable of recruiting an activation complex and / or activation activity (such as histone acetylation) to a target gene can be used as the activation domain of a fusion protein. Insulator domains, localization domains, and chromatin remodeling proteins, such as ISWI-containing domains and / or methyl-binding domain proteins suitable as functional domains in fusion molecules, are described in, for example, U.S. Patent Publication 2002 / 0115215 and 2003 / 0082552 and WO 02 / 44376.

[0152] Exemplary repressor domains include, but are not limited to, KRAB A / B, KOX, TGF-β-inducible early gene (TIEG), v-erbA, SID, MBD2, MBD3, members of the DNMT family (e.g., DNMT1, DNMT3A, DNMT3B, Rb, and MeCP2. See, for example, Bird et al. (1999) Cell 99:451-454; Tyler et al. (1999) Cell 99:443-446; Knoepfler et al. (1999) Cell 99:447-450; and Robertson et al. (2000) Nature Genet. 25:338-342. Additional exemplary repressor domains include, but are not limited to, ROM2 and AtHD2A. See, for example, Chem et al. (1996) Plant Cell 8:305-321; and Wu et al. (2000) Plant J. 22:19-27.

[0153] Fusion molecules are constructed by cloning and biochemical conjugation methods well known to those skilled in the art. The fusion molecule comprises a DNA binding domain and a functional domain (e.g., a transcriptional activation or repression domain). The fusion molecule also optionally comprises a nuclear localization signal (e.g., the signal from the SV40 large T antigen) and an epitope tag (e.g., FLAG and hemagglutinin). The fusion protein (and the nucleic acid encoding the fusion protein) is designed such that the translational reading frame is maintained among the fused components.

[0154] On the one hand, a fusion between the polypeptide component (or a functional fragment thereof) of the functional domain and on the other hand, a non-protein DNA binding domain (e.g., an antibiotic, an intercalator, a minor groove binder, a nucleic acid) is constructed by biochemical conjugation methods known to those skilled in the art. See, e.g., the Pierce Chemical Company (Rockford, IL) catalog. Methods and compositions for preparing fusions between minor groove binders and polypeptides have been described. Mapp et al. (2000) Proc. Natl. Acad. Sci. USA 97:3930 - 3935. In addition, a single guide RNA of the CRISPR / Cas system associates with the functional domain to form an active transcriptional regulator and nuclease.

[0155] In certain embodiments, the target site is present in an accessible region of the cell chromatin. Accessible regions can be determined as described, for example, in U.S. Patent Nos. 7,217,509 and 7,923,542. If the target site is not present in an accessible region of the cell chromatin, one or more accessible regions can be generated as described in U.S. Patent Nos. 7,785,792 and 8,071,370. In additional embodiments, the DNA binding domain of the fusion molecule is capable of binding to cell chromatin regardless of whether its target site is in an accessible region. For example, such DNA binding domains are capable of binding to linker DNA and / or nucleosomal DNA. Examples of this type of "pioneer" DNA binding domain are found in certain steroid receptors and hepatocyte nuclear factor 3 (HNF3) (Cordingley et al. (1987) Cell 48:261 - 270; Pina et al. (1990) Cell 60:719 - 731; and Cirillo et al. (1998) EMBO J. 17:244 - 254).

[0156] As known to those skilled in the art, the fusion molecule can be formulated with a pharmaceutically acceptable carrier. See, e.g., Remington's Pharmaceutical Sciences, 17th ed., 1985; and U.S. Patent Nos. 6,453,242 and 6,534,261.

[0157] The functional component / domain of the fusion molecule can be selected from any of a variety of different components that are capable of affecting gene transcription once the fusion molecule binds to the target sequence via its DNA-binding domain. Thus, the functional component can include, but is not limited to, various transcription factor domains such as activators, repressors, co-activators, co-repressors, and silencers.

[0158] Additional exemplary functional domains are disclosed, for example, in U.S. Patent Nos. 6,534,261 and 6,933,113.

[0159] Functional domains that are regulated by exogenous small molecules or ligands can also be selected. For example, techniques can be used where the functional domain assumes its active conformation only in the presence of an external RheoChem TM ligand (see, for example, US20090136465). Thus, the ZFP can be operably linked to a regulatable functional domain, where the resulting activity of the ZFP-TF is controlled by an external ligand.

[0160] Nuclease

[0161] In certain embodiments, the fusion protein comprises a DNA-binding domain and a cleavage (nuclease) domain. Thus, nucleases (e.g., engineered nucleases) can be used to effect gene modification. Engineered nuclease technologies are based on the engineering of naturally occurring DNA-binding proteins. For example, the engineering of homing endonucleases with customized DNA-binding specificities has been described. Chames et al. (2005) Nucleic Acids Res 33(20):e178; Arnould et al. (2006) J. Mol. Biol. 355:443-458. In addition, the engineering of ZFPs has also been described. See, for example, U.S. Patent Nos. 6,534,261; 6,607,882; 6,824,978; 6,979,539; 6,933,113; 7,163,824; and 7,013,219.

[0162] In addition, ZFPs and / or TALEs have been fused to nuclease domains to create ZFNs and TALENs – functional entities capable of recognizing their intended nucleic acid target via their engineered (ZFP or TALE) DNA-binding domain and causing DNA cleavage near the DNA-binding site via nuclease activity. See, e.g., Kim et al. (1996) Proc Nat'l Acad Sci USA 93(3):1156-1160. More recently, such nucleases have been used for genome modification in a variety of organisms. See, e.g., U.S. Patent Publication 20030232410; 20050208489; 20050026157; 20050064474; 20060188987; 20060063231; and International Publication WO 07 / 014275.

[0163] Accordingly, the methods and compositions described herein are widely applicable and may relate to any target nuclease. Non-limiting examples of nucleases include meganucleases, TALENs, and zinc finger nucleases. A nuclease may comprise a heterologous DNA-binding domain and a cleavage domain (e.g., zinc finger nuclease; meganuclease DNA-binding domain with a heterologous cleavage domain), or alternatively, the DNA-binding domain of a naturally occurring nuclease may be altered to bind to a selected target site (e.g., a meganuclease that has been engineered to bind to a site different from its cognate binding site).

[0164] In any of the nucleases described herein, the nuclease may comprise an engineered TALE DNA-binding domain and a nuclease domain (e.g., an endonuclease and / or meganuclease domain), also referred to as a TALEN. Methods and compositions for engineering these TALEN proteins for robust site-specific interaction with a user-selected target sequence have been disclosed (see U.S. Patent No. 8,586,526). In some embodiments, the TALEN comprises an endonuclease (e.g., FokI) cleavage domain or cleavage half-domain. In other embodiments, the TALE-nuclease is a megaTAL. These megaTAL nucleases are fusion proteins comprising a TALE DNA-binding domain and a meganuclease cleavage domain. The meganuclease cleavage domain functions as a monomer and does not require dimerization to obtain activity. (See Boissel et al., (2013) Nucl Acid Res:1-13, doi:10.1093 / nar / gkt1224). Additionally, the nuclease domain may also exhibit DNA-binding function.

[0165] In other embodiments, the nuclease comprises a compact TALEN (cTALEN). These nucleases are single-chain fusion proteins that link a TALE DNA-binding domain to a TevI nuclease domain. The fusion protein can act as a nickase targeted by the TALE region, or can generate a double-strand break, depending on the location of the TALE DNA-binding domain relative to the TevI nuclease domain (see Beurdeley et al. (2013) Nat Comm: 1-8 DOI: 10.1038 / ncomms2782). Any TALEN can be used in combination with another TALEN (e.g., one or more TALENs (cTALEN or FokI-TALEN) having one or more mega-TALs) or other DNA-cleaving enzymes.

[0166] In certain embodiments, the nuclease comprises a meganuclease (homing endonuclease) or a portion thereof that exhibits cleavage activity. Naturally occurring meganucleases recognize 15-40 base pair cleavage sites and are generally divided into four families: the LAGLIDADG family, the GIY-YIG family, the His-Cyst box family, and the HNH family. Exemplary homing endonucleases include I-SceI, I-CeuI, PI-PspI, PI-Sce, I-SceIV, I-CsmI, I-PanI, I-SceII, I-PpoI, I-SceIII, I-CreI, I-TevI, I-TevII, and I-TevIII. Their recognition sequences are known. See also U.S. Patent No. 5,420,032; U.S. Patent No. 6,833,252; Belfort et al. (1997) Nucleic Acids Res. 25:3379–3388; Dujon et al. (1989) Gene 82:115–118; Perler et al. (1994) Nucleic Acids Res. 22, 1125–1127; Jasin (1996) Trends Genet. 12:224–228; Gimble et al. (1996) J. Mol. Biol. 263:163–180; Argast et al. (1998) J. Mol. Biol. 280:345–353 and New England Biolabs catalog.

[0167] DNA binding domains from naturally occurring homing endonucleases (primarily from the LAGLIDADG family) have been used to facilitate site-specific genome modification in plants, yeast, Drosophila, mammalian cells, and mice, but this approach has been limited to homologous genes that preserve the homing endonuclease recognition sequence (Monet et al. (1999), Biochem. Biophysics. Res. Common. 255:88-93) or genomes in which the recognition sequence has been pre-engineered (Route et al. (1994), Mol. Cell. Biol. 14:8096-106; Chilton et al. (2003), Plant Physiology. 133:956-65; Puchta et al. (1996), Proc. Natl. Acad. Sci. USA 93:5055-60; Rong et al. (2002), Genes Dev. 16:1568-81; Gouble et al. (2006), J. Gene Med. 8(5):616-622). Thus, attempts have been made to engineer homing endonucleases to exhibit novel binding specificities at medically or biotechnologically relevant sites (Porteus et al. (2005), Nat. Biotechnol. 23:967-73; Sussman et al. (2004), J. Mol. Biol. 342:31-41; Epinat et al. (2003), Nucleic Acids Res. 31:2952-62; Chevalier et al. (2002) Molec. Cell 10:895-905; Epinat et al. (2003) Nucleic Acids Res. 31:2952-2962; Ashworth et al. (2006) Nature 441:656-659; Paques et al. (2007) Current Gene Therapy 7:49-66; U.S. Patent Publication Nos. 20070117128; 20060206949; 20060153826; 20060078552; and 20040002092). In addition, naturally occurring or engineered DNA binding domains from homing endonucleases can be operably linked to cleavage domains from heterologous endonucleases (e.g., FokI), and / or cleavage domains from homing endonucleases can be operably linked to heterologous DNA binding domains (e.g., ZFP or TALE).

[0168] In other embodiments, the nuclease is a zinc finger nuclease (ZFN) or a TALE DNA binding domain-nuclease fusion (TALEN). ZFNs and TALENs comprise a DNA binding domain (zinc finger protein or TALE DNA binding domain) that has been engineered to bind to a selected gene and a cleavage domain or cleavage half-domain (e.g., from a restriction and / or meganuclease as described herein) at a target site.

[0169] As described in detail above, zinc finger binding domains and TALE DNA binding domains can be engineered to bind to selected sequences. See, e.g., Beerli et al. (2002) Nature Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nature Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416. Engineered zinc finger binding domains or TALE proteins can have novel binding specificities compared to naturally occurring proteins. Engineering methods include, but are not limited to, rational design and various types of selection. Rational design includes, for example, using databases that contain triplet (or quadruplet) nucleotide sequences and individual zinc finger or TALE amino acid sequences, where each triplet or quadruplet nucleotide sequence is associated with one or more amino acid sequences of a zinc finger or TALE repeat unit that binds to the specific triplet or quadruplet sequence. See, e.g., U.S. Pat. Nos. 6,453,242 and 6,534,261, which are incorporated herein by reference in their entirety.

[0170] The selection of target sites; methods for designing and constructing fusion proteins (and polynucleotides encoding the fusion proteins) are known to those of skill in the art and are described in detail in U.S. Patent Nos. 7,888,121 and 8,409,861, which are incorporated herein by reference in their entirety.

[0171] In addition, as disclosed in these and other references, zinc finger domains, TALEs, and / or multi-fingered zinc finger proteins can be linked together using any suitable linker sequence (including, for example, linkers of 5 or more amino acids in length). (For example, TGEKP (SEQ ID NO:9), TGGQRP (SEQ ID NO:10), TGQKP (SEQ ID NO:11), and / or TGSQKP (SEQ ID NO:12). For exemplary linker sequences of 6 or more amino acids in length, see, for example, U.S. Patent Nos. 6,479,626; 6,903,185; and 7,153,949. The proteins described herein can include any combination of suitable linkers between individual protein zinc fingers. See also U.S. Patent No. 8,772,453.)

[0172] Thus, nucleases such as ZFNs, TALENs, and / or meganucleases can comprise any DNA-binding domain and any nuclease (cleavage) domain (cleavage domain, cleavage half-domain). As mentioned above, the cleavage domain can be heterologous to the DNA-binding domain, e.g., a zinc finger or TAL effector DNA-binding domain and a cleavage domain from a nuclease or a meganuclease DNA-binding domain and a cleavage domain from a different nuclease. Heterologous cleavage domains can be obtained from any endonuclease or exonuclease. Exemplary endonucleases from which the cleavage domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, for example, the 2002-2003 Catalogue, New England Biolabs, Beverly, MA; and Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Additional enzymes that cleave DNA are known (e.g., S1 nuclease; mung bean nuclease; pancreatic DNase I; micrococcal nuclease; yeast HO endonuclease; see also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993). One or more of these enzymes (or functional fragments thereof) can be used as a source of cleavage domains and cleavage half-domains.

[0173] Similarly, the cleavage half-domains can be derived from any nuclease or a portion thereof that requires dimerization for cleavage activity, as described above. In general, if the fusion protein contains a cleavage half-domain, cleavage requires two fusion proteins. Alternatively, a single protein containing two cleavage half-domains can be used. The two cleavage half-domains can be derived from the same endonuclease (or functional fragments thereof), or each cleavage half-domain can be derived from a different endonuclease (or functional fragments thereof). Additionally, the target sites of the two fusion proteins are preferably arranged relative to each other such that binding of the two fusion proteins to their respective target sites positions the cleavage half-domains in a spatial relationship that permits the cleavage half-domains to form a functional cleavage domain, e.g., by dimerization. Thus, in certain embodiments, the proximal edges of the target sites are separated by 5 - 10 nucleotides or by 15 - 18 nucleotides. However, any integer number of nucleotides or nucleotide pairs can be between the two target sites (e.g., 2 to 50 nucleotide pairs or more). In general, the site of cleavage is between the target sites.

[0174] Restriction endonucleases (restriction enzymes) are present in many species and are capable of binding to DNA (at the recognition site) in a sequence-specific manner and cleaving the DNA at or near the binding site. Certain restriction enzymes (e.g., type IIS) cleave DNA at sites distant from the recognition site and have separable binding and cleavage domains. For example, the type IIS enzyme FokI catalyzes double-stranded cleavage of DNA 9 nucleotides from its recognition site on one strand and 13 nucleotides from its recognition site on the other strand. See, e.g., U.S. Pat. Nos. 5,356,802; 5,436,150 and 5,487,994; and Li et al. (1992) Proc. Natl. Acad. Sci. USA 89:4275 - 4279; Li et al. (1993) Proc. Natl. Acad. Sci. USA 90:2764 - 2768; Kim et al. (1994a) Proc. Natl. Acad. Sci. USA 91:883 - 887; Kim et al. (1994b) J. Biol. Chem. 269:31,978 - 31,982. Thus, in one embodiment, the fusion protein contains a cleavage domain (or cleavage half-domain) from at least one type IIS restriction enzyme and one or more zinc finger binding domains, which may or may not be engineered.

[0175] An exemplary type IIS restriction enzyme with a cleavage domain separated from the binding domain is Fok I. This particular enzyme is active as a dimer. Bitinaite et al. (1998) Proc. Natl. Acad. Sci. USA 95:10,570-10,575. Thus, for the purposes of the present disclosure, the portion of the Fok I enzyme used in the disclosed fusion proteins is considered a cleavage half-domain. Accordingly, to use a zinc finger-Fok I fusion for targeted double-strand cleavage and / or targeted replacement of a cellular sequence, two fusion proteins each containing a FokI cleavage half-domain can be used to reconstitute the catalytically active cleavage domain. Alternatively, a single polypeptide molecule containing a zinc finger binding domain and two Fok I cleavage half-domains can also be used. The present disclosure provides elsewhere parameters for targeted cleavage and targeted sequence alteration using zinc finger-Fok I fusions.

[0176] The cleavage domain or cleavage half-domain can be any portion of a protein that retains cleavage activity or retains the ability to multimerize (e.g., dimerize) to form a functional cleavage domain.

[0177] Exemplary type IIS restriction enzymes are described in International Publication WO 07 / 014275, which is incorporated herein by reference in its entirety. Additional restriction enzymes also contain separable binding and cleavage domains, and these enzymes are encompassed by the present disclosure. See, e.g., Roberts et al. (2003) Nucleic Acids Res. 31:418-420.

[0178] In certain embodiments, the cleavage domain comprises the FokI cleavage domain used to generate crystal structures 1FOK.pdb and 2FOK.pdb (see Wah et al. (1997) Nature 388:97-100), which has the sequence shown below:

[0179] Wild-type FokI cleavage half-domain (SEQ ID NO:1)

[0180]

[0181] The cleavage half-domain derived from FokI can contain mutations in one or more of the amino acid residues shown in SEQ ID NO:1. Mutations include substitutions of (wild-type amino acid residues with different residues), insertions of (one or more amino acid residues), and / or deletions of (one or more amino acid residues). In certain embodiments, one or more of residues 414-426, 443-450, 467-488, 501-502, and / or 521-531 (relative to SEQ ID NO:1 and Figure 17The sequence numbers shown in are mutated because these residues are located near the DNA backbone in the molecular model of the ZFN that binds to its target site as described in Miller et al. ((2007) Nat Biotechnol 25:778-784). In certain embodiments, one or more residues at positions 416, 422, 447, 448, and / or 525 are mutated. In certain embodiments, the mutation comprises substituting the wild-type residue with any different residue, such as an alanine (A) residue, a cysteine (C) residue, an aspartic acid (D) residue, a glutamic acid (E) residue, a histidine (H) residue, a phenylalanine (F) residue, a glycine (G) residue, an asparagine (N) residue, a serine (S) residue, or a threonine (T) residue. In other embodiments, the wild-type residue at one or more of positions 416, 418, 422, 446, 448, 476, 479, 480, 481, and / or 525 is replaced with any other residue, including but not limited to, R416D, R416E, S418E, S418D, R422H, S446D, K448A, N476D, I479Q, I479T, G480D, Q481A, Q481E, K525S, K525A, N527D, R416E+R422H, R416D+R422H, R416E+K448A, R416D+R422H, K448A+I479Q, K448A+Q481A, K448A+K525A.

[0182] In certain embodiments, the cleavage domain comprises one or more engineered cleavage half-domains (also referred to as dimerization domain mutants) that minimize digestion or prevent homodimerization, as described, for example, in U.S. Patent Nos. 7,914,796; 8,034,598; and 8,623,618; and U.S. Patent Publication No. 20110201055, the entire disclosures of which are incorporated herein by reference in their entirety. The amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of Fok I (relative to SEQ ID NO:1 and Figure 17 the sequence numbers shown in are all targets for affecting the dimerization of the Fok I cleavage half-domain. The mutations can include mutations of residues found in natural restriction enzymes homologous to FokI. In a preferred embodiment, at positions 416, 422, 447, 448, and / or 525 (relative to SEQ ID NO:1 and Figure 17The mutations (in the sequence numbers shown in ) include substitution of a positively charged amino acid with an uncharged or negatively charged amino acid. In another embodiment, in addition to the mutations in one or more of amino acid residues 416, 422, 447, 448, or 525, the engineered cleavage half-domain also includes mutations in amino acid residues 499, 496, and 486, all relative to SEQ ID NO:1 or Figure 17 the sequence numbers shown in .

[0183] In certain embodiments, the compositions described herein include an engineered cleavage half-domain of Fok I that forms an obligate heterodimer, as described, for example, in U.S. Patent Nos. 7,914,796; 8,034,598; 8,961,281; and 8,623,618; and U.S. Patent Publication Nos. 20080131962 and 20120040398. Thus, in a preferred embodiment, the invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, wherein in addition to one or more mutations at positions 416, 422, 447, 448, or 525 (relative to SEQ ID NO:1 and Figure 17 the sequence numbers shown in ), the wild-type Gln (Q) residue at position 486 is replaced with a Glu (E) residue, the wild-type Ile (I) residue at position 499 is replaced with a Leu (L) residue, and the wild-type Asn (N) residue at position 496 is replaced with an Asp (D) or Glu (E) residue ("ELD" or "ELE"). In another embodiment, the engineered cleavage half-domain is derived from the wild-type FokI cleavage half-domain and, in addition to one or more mutations at amino acid residues 416, 422, 447, 448, or 525, also includes relative to wild-type FokI (SEQ ID NO:1 and Figure 17Mutations in amino acid residues 490, 538, and 537 of the sequence shown in . In a preferred embodiment, the present invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, wherein in addition to one or more mutations at positions 416, 422, 447, 448, or 525, the wild-type Glu (E) residue at position 490 is replaced with a Lys (K) residue, the wild-type Ile (I) residue at position 538 is replaced with a Lys (K) residue, and the wild-type His (H) residue at position 537 is replaced with a Lys (K) residue or an Arg (R) residue ("KKK" or "KKR") (see U.S. 8,962,281, which is incorporated herein by reference). See, for example, U.S. Patent No. 7,914,796; 8,034,598, and 8,623,618, the disclosures of which are incorporated herein by reference in their entirety for all purposes. In other embodiments, the engineered cleavage half-domain comprises "Sharkey" and / or "Sharkey" mutations (see Guo et al., (2010) J. Mol. Biol. 400(1):96-107).

[0184] In another embodiment, the engineered cleavage half-domain is derived from the wild-type FokI cleavage half-domain and comprises mutations in amino acid residues 490 and 538 relative to the wild-type FokI or FokI homolog numbering in addition to one or more mutations at amino acid residues 416, 422, 447, 448, or 525. In a preferred embodiment, the present invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, wherein in addition to one or more mutations at positions 416, 422, 447, 448, or 525, the wild-type Glu (E) residue at position 490 is replaced with a Lys (K) residue, and the wild-type Ile (I) residue at position 538 is replaced with a Lys (K) residue ("KK"). In a preferred embodiment, the present invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, wherein in addition to one or more mutations at positions 416, 422, 447, 448, or 525, the wild-type Gln (Q) residue at position 486 is replaced with a Glu (E) residue, and the wild-type Ile (I) residue at position 499 is replaced with a Leu (L) residue ("EL") (see U.S. 8,034,598, which is incorporated herein by reference).

[0185] In one aspect, the present invention provides a fusion protein, wherein the engineered cleavage half-domain comprises a polypeptide, and wherein one or more of the wild-type amino acid residues at positions 387, 393, 394, 398, 400, 402, 416, 422, 427, 434, 439, 441, 447, 448, 469, 487, 495, 497, 506, 516, 525, 529, 534, 559, 569, 570, 571 in the FokI catalytic domain are mutated. A nuclease domain comprising one or more mutations as shown in any of the attached tables and figures is provided. In some embodiments, the one or more mutations change the wild-type amino acid from a positively charged residue to a neutral residue or a negatively charged residue. In any of these embodiments, the described mutants can also be prepared in a FokI domain comprising one or more additional mutations. In a preferred embodiment, these additional mutations are located in the dimerization domain, for example, at positions 418, 432, 441, 481, 483, 486, 487, 490, 496, 499, 523, 527, 537, 538 and / or 559. Non-limiting examples of mutations include mutating (e.g., substituting) the wild-type residue of any cleavage domain (e.g., FokI or a homolog of FokI) at positions 393, 394, 398, 416, 421, 422, 442, 444, 472, 473, 478, 480, 525 or 530 with any amino acid residue (e.g., K393X, K394X, R398X, R416S, D421X, R422X, K444X, S472X, G473X, S472, P478X, G480X, K525X and A530X, where the first residue depicts the wild-type and X refers to the amino acid substituting the wild-type residue). In some embodiments, X is E, D, H, A, K, S, T, D or N. Other exemplary mutations include S418E, S418D, S446D, K448A, I479Q, I479T, Q481A, Q481N, Q481E, A530E and / or A530K, where the amino acid residues are numbered relative to the full-length FokI wild-type cleavage domain and its homologs ( Figure 17)。In certain embodiments, the combination may include 416 and 422, mutations at positions 416 and K448A, K448A and I479Q, K448A and Q481A and / or K448A, and a mutation at position 525. In one embodiment, the wild-type residue at position 416 may be replaced with a Glu (E) residue (R416E), the wild-type residue at position 422 may be replaced with a His (H) residue (R422H), and the wild-type residue at position 525 may be replaced with an Ala (A) residue. The cleavage domain as described herein may also contain additional mutations, including but not limited to positions 432, 441, 483, 486, 487, 490, 496, 499, 527, 537, 538, and / or 559, such as dimerization domain mutants (e.g., ELD, KKR) and / or nickase mutants (mutations to the catalytic domain). The cleavage half-domains having the mutations described herein form heterodimers as known in the art.

[0186] Alternatively, the so-called "split enzyme" technique can be used in vivo to assemble nucleases at a nucleic acid target site (see, for example, U.S. Patent Publication No. 20090068164). Components of such split enzymes can be expressed on separate expression constructs or can be linked in a single open reading frame by separate components such as a self-cleaving 2A peptide or an IRES sequence. The components can be individual zinc finger binding domains or domains having a meganuclease nucleic acid binding domain.

[0187] The activity of nucleases (e.g., ZFNs and / or TALENs) can be screened in a yeast-based chromosomal system as described, for example, in U.S. Patent No. 8,563,314, prior to use.

[0188] In certain embodiments, the nuclease comprises a CRISPR / Cas system. The CRISPR (clustered regularly interspaced short palindromic repeats) locus encoding the RNA component of the system and the Cas (CRISPR-associated) locus encoding the protein (Jansen et al., 2002. Mol. Microbiol. 43:1565 - 1575; Makarova et al., 2002. Nucleic Acids Res. 30:482 - 496; Makarova et al., 2006. Biol. Direct 1:7; Haft et al., 2005. PLoS Comput. Biol. 1:e60) constitute the gene sequences of the CRISPR / Cas nuclease system. The CRISPR locus in a microbial host contains a combination of CRISPR-associated (Cas) genes and non-coding RNA elements capable of programming the specificity of CRISPR-mediated nucleic acid cleavage.

[0189] Type II CRISPR is one of the best characterized systems and makes a targeted DNA double-strand break in four successive steps. First, two non-coding RNAs (the pre-crRNA array and tracrRNA) are transcribed from the CRISPR locus. Second, tracrRNA hybridizes to the repeat regions of the pre-crRNA and mediates the processing of the pre-crRNA into mature crRNAs containing individual spacer sequences. Third, the mature crRNA:tracrRNA complex guides Cas9 to the target DNA via Watson-Crick base pairing between the spacer on the crRNA and the protospacer near the protospacer adjacent motif (PAM) on the target DNA, which is an additional requirement for target recognition. Finally, Cas9 mediates cleavage of the target DNA to create a double-strand break within the protospacer. The activities of the CRISPR / Cas system include three steps: (i) insertion of foreign DNA sequences into the CRISPR array in a process called 'adaptation' to prevent future attacks, (ii) expression of the associated proteins, and expression and processing of the array, followed by (iii) RNA-mediated interference of foreign nucleic acids. Thus, in bacterial cells, several so-called 'Cas' proteins are involved in the native function of the CRISPR / Cas system and play roles in functions such as insertion of foreign DNA.

[0190] In some embodiments, the CRISPR-Cpf1 system is used. The CRISPR-Cpf1 system identified in Francisella spp is a type II CRISPR-Cas system that mediates robust DNA interference in human cells. Although functionally conserved, Cpf1 and Cas9 differ in many respects, including in their guide RNA and substrate specificities (see Fagerlund et al., (2015) Genom Bio16:251). The major difference between the Cas9 and Cpf1 proteins is that Cpf1 does not utilize a tracrRNA and thus requires only a crRNA. The FnCpf1 crRNA is 42-44 nucleotides long (a 19-nucleotide repeat sequence and a 23-25-nucleotide spacer) and contains a single stem loop that tolerates sequence variations that retain the secondary structure. In addition, the Cpf1 crRNA is significantly shorter than the engineered sgRNA of approximately 100 nucleotides required by Cas9, and the PAM requirement for FnCpfl is 5'-TTN-3' and 5'-CTA-3' on the displaced strand. Although both Cas9 and Cpf1 produce double-strand breaks in the target DNA, Cas9 uses its RuvC- and HNH-like domains to produce blunt-ended cuts within the seed sequence of the guide RNA, while Cpf1 uses the RuvC-like domain to produce staggered cuts outside the seed. Since Cpf1 produces staggered cuts away from the critical seed region, NHEJ will not disrupt the target site, thus ensuring that Cpf1 can continue to cut the same site until the desired HDR recombination event occurs. Thus, in the methods and compositions described herein, it should be understood that the term "Cas" includes both Cas9 protein and Cfp1 protein. Thus, as used herein, "CRISPR / Cas system" refers to both CRISPR / Cas and / or CRISPR / Cfp1 systems, including both nuclease and / or transcription factor systems.

[0191] In certain embodiments, the Cas protein can be a "functional derivative" of a naturally occurring Cas protein. A "functional derivative" of a native sequence polypeptide is a compound having the same qualitative biological properties as the native sequence polypeptide. "Functional derivatives" include, but are not limited to, fragments of the native sequence and derivatives of the native sequence polypeptide and its fragments, provided that they have the same biological activity as the corresponding native sequence polypeptide. The biological activity covered herein is the ability of the functional derivative to hydrolyze a DNA substrate into fragments. The term "derivative" encompasses amino acid sequence variants of the polypeptide, its covalent modifications, and fusions, such as derivative Cas proteins. Suitable derivatives of the Cas polypeptide or its fragment include, but are not limited to, mutants, fusions, and covalent modifications of the Cas protein or its fragment. The Cas protein (which includes the Cas protein or its fragment and the derivative of the Cas protein or its fragment) can be obtained from cells or chemically synthesized or obtained by a combination of these two procedures. The cells can be cells that naturally produce the Cas protein, or cells that naturally produce the Cas protein and are genetically engineered to produce the endogenous Cas protein at a higher expression level or produce the Cas protein from an exogenously introduced nucleic acid that encodes a Cas that is the same as or different from the endogenous Cas. In some cases, the cells do not naturally produce the Cas protein and are genetically engineered to produce the Cas protein. In some embodiments, the Cas protein is a small Cas9 ortholog for delivery via an AAV vector (Ran et al. (2015) Nature 510, p. 186).

[0192] One or more nucleases can create one or more double-stranded and / or single-stranded nicks at the target site. In certain embodiments, the nuclease comprises a catalytically inactivated cleavage domain (e.g., FokI and / or Cas protein). See, for example, U.S. Patent No. 9,200,266; 8,703,489 and Guillinger et al. (2014) Nature Biotech. 32(6):577-582. The catalytically inactivated cleavage domain can be combined with a catalytically active domain to act as a nickase to create a single-stranded nick. Thus, two nickases can be used in combination to create a double-stranded nick in a specific region. Additional nickases are also known in the art, such as McCaffery et al. (2016) Nucleic Acids Res. 44(2):e11.doi:10.1093 / nar / gkv878. Epub October 19, 2015.

[0193] Delivery

[0194] The protein (e.g., nuclease), polynucleotide, and / or composition comprising the protein and / or polynucleotide described herein can be delivered to a target cell by any suitable means, including, for example, by injecting the protein and / or mRNA components.

[0195] Suitable cells include, but are not limited to, eukaryotic and prokaryotic cells and / or cell lines. Non-limiting examples of such cells or cell lines derived from such cells include T cells, COS, CHO (e.g., CHO-S, CHO-K1, CHO-DG44, CHO-DUXB11, CHO-DUKX, CHOK1SV), VERO, MDCK, WI38, V79, B14AF28-G3, BHK, HaK, NS0, SP2 / 0-Ag14, HeLa, HEK293 (e.g., HEK293-F, HEK293-H, HEK293-T), and perC6 cells, as well as insect cells such as Spodoptera fugiperda (Sf), or fungal cells such as Saccharomyces, Pichia, and Schizosaccharomyces. In certain embodiments, the cell line is a CHO-K1, MDCK, or HEK293 cell line. Suitable cells also include stem cells, such as, by way of example, embryonic stem cells, induced pluripotent stem cells (iPS cells), hematopoietic stem cells, neural stem cells, and mesenchymal stem cells.

[0196] Methods for delivering DNA-binding domains as described herein are described, for example, in U.S. Patent Nos. 6,453,242; 6,503,717; 6,534,261; 6,599,692; 6,607,882; 6,689,558; 6,824,978; 6,933,113; 6,979,539; 7,013,219; and 7,163,824, the entire disclosures of which are incorporated herein by reference in their entirety.

[0197] Vectors containing sequences encoding one or more DNA-binding proteins can also be used to deliver DNA-binding domains and fusion proteins comprising such DNA-binding domains as described herein. Additionally, additional nucleic acids (e.g., donors) can also be delivered via such vectors. Any vector system can be used, including but not limited to plasmid vectors, retroviral vectors, lentiviral vectors, adenoviral vectors, poxviral vectors; herpes viral vectors, and adeno-associated viral vectors, etc. See also U.S. Patent Nos. 6,534,261; 6,607,882; 6,824,978; 6,933,113; 6,979,539; 7,013,219; and 7,163,824, which are incorporated herein by reference in their entirety. In addition, it will be apparent that any of these vectors can suitably contain one or more DNA-binding protein coding sequences and / or additional nucleic acids. Thus, when introducing one or more DNA-binding proteins as described herein, and optionally additional DNA, the DNA-binding proteins can be carried on the same or different vectors. When multiple vectors are used, each vector can contain sequences encoding one or more DNA-binding proteins and additional nucleic acids as needed.

[0198] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids encoding engineered DNA-binding proteins into cells (e.g., mammalian cells) and target tissues and to co-introduce additional nucleotide sequences as needed. Such methods can be used to administer nucleic acids (e.g., encoding DNA-binding proteins and / or donors) in vitro to cells. In certain embodiments, nucleic acids used for in vivo or ex vivo gene therapy are administered. Non-viral vector delivery systems include DNA plasmids, naked nucleic acids, and nucleic acids complexed with delivery vehicles such as liposomes or poloxamers. Viral vector delivery systems include DNA and RNA viruses that have episomal genes or integrated genomes after delivery to cells. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Felgner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and Immunology Doerfler and (Eds.)(1995); and Yu et al., Gene Therapy 1:13-26 (1994).

[0199] Non-viral delivery methods of nucleic acids include electroporation, lipofection, microinjection, gene gun, virosomes, liposomes, immunoliposomes, polycations or lipid:nucleic acid conjugates, naked DNA, mRNA, artificial virus particles, and reagent-enhanced uptake of DNA. Sono-poration using, for example, the Sonitron 2000 system (Rich-Mar) can also be used for delivery of nucleic acids. In a preferred embodiment, one or more nucleic acids are delivered as mRNA. It is also preferred to use capped mRNA to increase translation efficiency and / or mRNA stability. Particularly preferred is the ARCA (anti-reverse cap analog) cap or variants thereof. See U.S. Patent Nos. 7,074,596 and 8,153,773, which are incorporated herein by reference in their entirety.

[0200] Additional exemplary nucleic acid delivery systems include those provided by Amaxa Biosystems (Cologne, Germany), Maxcyte, Inc. (Rockville, Maryland), BTX Molecular Delivery Systems (Holliston, MA), and Copernicus Therapeutics Inc. (see Example of US 6008336). Lipofection is described, for example, in US 5,049,386, US 4,946,787, and US 4,897,355), and lipofection reagents are commercially available (e.g., Transfectam TM , Lipofectin TM and Lipofectamine TM RNAiMAX). Cationic and neutral lipids useful for efficient receptor recognition lipofection of polynucleotides include those of Felgner, WO 91 / 17424, WO 91 / 16024. Delivery can be to cells (ex vivo administration) or target tissues (in vivo administration).

[0201] The preparation of lipid:nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known to those skilled in the art (see, for example, Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0202] Additional delivery methods include encapsulating the nucleic acid to be delivered into an EnGeneIC delivery vehicle (EDV). Bispecific antibodies are used to specifically deliver these EDVs to the target tissue, where one arm of the antibody has specificity for the target tissue and the other arm has specificity for the EDV. The antibody brings the EDV to the surface of the target cell and then takes the EDV into the cell by endocytosis. Once inside the cell, the contents are released (see MacDiarmid et al. (2009) Nature Biotechnology 27(7) page 643).

[0203] RNA- or DNA virus-based systems are used as needed to deliver nucleic acids encoding engineered DNA-binding proteins and / or donors (such as CAR or ACTR) using highly advanced methods for targeting the virus to specific cells in the body and for migrating the viral payload to the nucleus. Viral vectors can be administered directly to the patient (in vivo) or they can be used to treat cells ex vivo and then administer the modified cells to the patient (ex vivo). Conventional virus-based systems for delivering nucleic acids include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated, vaccinia, and herpes simplex virus vectors for gene transfer. With retroviral, lentiviral, and adeno-associated virus gene transfer methods, integration into the host genome is possible, thus often resulting in long-term expression of the inserted transgene. In addition, high transduction efficiencies have been observed in many different cell types and target tissues.

[0204] The tropism of retroviruses can be altered by incorporating foreign envelope proteins, thereby expanding the potential target population of the target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and usually produce high viral titers. The choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats that have the ability to package foreign sequences of up to 6 - 10 kb. The minimal cis-acting LTR is sufficient for vector replication and encapsidation, which is then used to integrate the therapeutic gene into the target cells to provide persistent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Virol. 66:2731 - 2739 (1992); Johann et al., J. Virol. 66:1635 - 1640 (1992); Sommerfelt et al., Virol. 176:58 - 59 (1990); Wilson et al., J. Virol. 63:2374 - 2378 (1989); Miller et al., J. Virol. 65:2220 - 2224 (1991); PCT / US94 / 05700).

[0205] In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors are capable of extremely high transduction efficiencies in many cell types and do not require cell division. In the case of such vectors, high titers and high levels of expression have been achieved. Such vectors can be produced in large quantities in relatively simple systems. Adeno-associated virus ("AAV") vectors are also used to transduce cells with target nucleic acids, e.g., for the production of nucleic acids and peptides in vitro, and for in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO 93 / 24641; Kotin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors is described in many publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS USA 81:6466-6470 (1984); and Samulski et al., J. Virol. 63:03822-3828 (1989).

[0206] At least six viral vector methods are currently available for gene transfer in clinical trials, which utilize methods involving complementing defective vectors by inserting genes into helper cell lines to produce transducing agents.

[0207] pLASN and MFG-S are examples of retroviral vectors that have been used in clinical trials (Dunbar et al., Blood 85:3048-305 (1995); Kohn et al., Nat. Med. 1:1017-102 (1995); Malech et al., PNAS USA 94:2212133-12138 (1997)). PA317 / pLASN was the first therapeutic vector used in gene therapy trials. (Blaese et al., Science 270:475-480 (1995)). Transduction efficiencies of 50% or greater have been observed for the MFG-S packaging vector. (Ellem et al., Immunol Immunother. 44(1):10-20 (1997); Dranoff et al., Hum. Gene Ther. 1:111-2 (1997).

[0208] Recombinant adeno-associated virus vectors (rAAV) are promising alternative gene delivery systems based on the defective and non-pathogenic parvovirus adeno-associated virus type 2. All vectors are derived from plasmids that retain only the 145 bp inverted terminal repeats of AAV flanking the transgene expression cassette. Efficient gene transfer and stable transgene delivery due to integration into the genome of the transduced cells are key features of this vector system. (Wagner et al., Lancet 351:9117 1702-3 (1998), Kearns et al., Gene Ther. 9:748-55 (1996)). Other AAV serotypes, including AAV1, AAV3, AAV4, AAV5, AAV6, AAV8, AAV8.2, AAV9 and AAVrh10 and pseudotyped AAVs such as AAV2 / 8, AAV2 / 5 and AAV2 / 6 can also be used according to the present invention.

[0209] Replication-deficient recombinant adenovirus vectors (Ad) can be produced at high titers and easily infect many different cell types. Most adenovirus vectors are engineered such that the transgene replaces the Ad E1a, E1b and / or E3 genes; the replication-deficient vectors are then propagated in human 293 cells that provide the missing gene functions in trans. Ad vectors can transduce multiple types of tissues in vivo, including non-dividing differentiated cells such as those found in liver, kidney and muscle tissues. Conventional Ad vectors have a large carrying capacity. Examples of the use of Ad vectors in clinical trials involve polynucleotide therapies for anti-tumor immunization by intramuscular injection (Sterman et al., Hum. Gene Ther. 7:1083-9 (1998)). Additional examples of the use of adenovirus vectors to transfer genes in clinical trials include Rosenecker et al., Infection 24:1 5-10 (1996); Sterman et al., Hum. Gene Ther. 9:7 1083-1089 (1998); Welsh et al., Hum. Gene Ther. 2:205-18 (1995); Alvarez et al., Hum. Gene Ther. 5:597-613 (1997); Topf et al., Gene Ther. 5:507-513 (1998); Sterman et al., Hum. Gene Ther. 7:1083-1089 (1998).

[0210] Packaging cells are used to form virus particles capable of infecting host cells. Such cells include 293 cells, which package adenovirus, and ψ2 cells or PA317 cells, which package retrovirus. Viral vectors for gene therapy are typically produced by a producer cell line that packages a nucleic acid vector into virus particles. The vector usually contains reduced viral sequences (if applicable) required for packaging and subsequent integration into the host, and other viral sequences are replaced by an expression cassette encoding the protein to be expressed. The missing viral functions are provided in trans by the packaging cell line. For example, AAV vectors for gene therapy typically have only the inverted terminal repeat (ITR) sequences from the AAV genome, which are required for packaging and integration into the host genome. The viral DNA is packaged in a cell line that contains a helper plasmid encoding the other AAV genes, i.e., rep and cap, but lacking the ITR sequences. An adenovirus, also used as a helper virus, is used to infect the cell line. The helper virus facilitates the replication of the AAV vector and the expression of the AAV genes from the helper plasmid. Since the helper plasmid lacks the ITR sequences, it is not packaged in large amounts. Contamination by adenovirus can be reduced, for example, by heat treatment to which adenovirus is more sensitive than AAV.

[0211] In many gene therapy applications, it is desirable for gene therapy vectors to deliver with a high degree of specificity to a particular tissue type. Thus, viral vectors are often modified to be specific for a given cell type by expressing a ligand on the outer surface of the virus as a fusion protein with a viral coat protein. Ligands with an affinity for receptors known to be present on the target cell type are selected. For example, Han et al., (Proc. Natl. Acad. Sci. USA 92:9747-9751 (1995)) reported that Moloney murine leukemia virus can be modified to express heregulin fused to gp70, and the recombinant virus infects certain human breast cancer cells expressing the human epidermal growth factor receptor. This principle can be extended to other virus-target cell pairs, where the target cell expresses a receptor and the virus expresses a fusion protein containing a ligand for the cell surface receptor. For example, filamentous phage can be engineered to display antibody fragments (e.g., FAB or Fv) with specific binding affinities for virtually any selected cell receptor. Although the above description applies mainly to viral vectors, the same principle can be applied to non-viral vectors. Such vectors can be engineered to contain specific uptake sequences that facilitate uptake by specific target cells.

[0212] Delivery methods for CRISPR / Cas systems can include those described above. For example, in animal models, in vitro transcribed Cas or recombinant Cas proteins encoding mRNA can be directly injected into single-cell stage embryos using glass needles for genome editing of animals. To express Cas and guide RNA in cells in vitro, plasmids encoding them are typically transfected into cells via lipid transfection or electroporation. In addition, recombinant Cas proteins can be complexed with in vitro transcribed guide RNA, where the Cas-guide RNA ribonucleoprotein is taken up by the template cells (Kim et al. (2014) Genome Res 24(6):1012). For therapeutic purposes, Cas and guide RNA can be delivered by a combination of viral and non-viral techniques. For example, mRNA encoding Cas can be delivered via nanoparticle delivery, while guide RNA and any desired transgene or repair template are delivered via AAV (Yin et al. (2016) Nat Biotechnol 34(3) p. 328).

[0213] Gene therapy vectors can be delivered in vivo by administering to an individual patient, typically by systemic administration (e.g., intravenous, intraperitoneal, intramuscular, subcutaneous, or intracranial infusion) or local administration as described below. Alternatively, the vector can be delivered to ex vivo cells, such as cells transplanted from an individual patient (e.g., lymphocytes, bone marrow aspirates, tissue biopsies) or universal donor hematopoietic stem cells, and the cells are then typically re-implanted into the patient after selecting the cells that have incorporated the vector.

[0214] Ex vivo transfection of cells for diagnostic, research, transplantation, or for gene therapy (e.g., via re-infusion of transfected cells into a host organism) is well known to those skilled in the art. In a preferred embodiment, cells are isolated from a subject organism, transfected with DNA-binding protein nucleic acids (genes or cDNAs), and re-infused into the subject organism (e.g., a patient). A variety of cell types suitable for ex vivo transfection are well known to those skilled in the art (for a discussion of how to isolate and culture cells from a patient, see, e.g., Freshney et al., Culture of Animal Cells,A Manual of Basic Technique (3rd ed 1994)) and references cited therein).

[0215] In one embodiment, stem cells are used in ex vivo procedures for cell transfection and gene therapy. The advantage of using stem cells is that they can differentiate into other cell types in vitro or can be introduced into a mammal, such as the donor of the cells, in which they will migrate into the bone marrow. Methods for differentiating CD34+ cells into clinically important immune cell types in vitro using cytokines such as GM-CSF, IFN-γ, and TNF-α are known (see Inaba et al., J. Exp. Med. 176:1693-1702 (1992)).

[0216] Stem cells are isolated for transduction and differentiation using known methods. For example, stem cells can be isolated from bone marrow cells by panning the bone marrow cells with antibodies that bind unwanted cells, such as CD4+ and CD8+ (T cells), CD45+ (pan B cells), GR-1 (granulocytes), and Iad (differentiated antigen-presenting cells) (see Inaba et al., J. Exp. Med. 176:1693-1702 (1992)).

[0217] In some embodiments, stem cells that have been modified can also be used. For example, neuronal stem cells that have been made resistant to apoptosis can be used as a therapeutic composition, wherein the stem cells also contain the ZFP TF of the present invention. For example, resistance to apoptosis can be generated, for example, by using BAX- or BAK-specific ZFNs in the stem cells (see U.S. Patent No. 8,597,912) or, for example, by knocking out BAX and / or BAK again using caspase-6-specific ZFNs in those caspases that have been disrupted.

[0218] Vectors containing therapeutic DNA-binding proteins (or nucleic acids encoding these proteins), such as retroviruses, adenoviruses, liposomes, etc., can also be directly administered to an organism for in vivo transduction of cells. Alternatively, naked DNA can be administered. Administration is by any route commonly used to bring a molecule into ultimate contact with blood or tissue cells, including but not limited to injection, infusion, topical administration, and electroporation. Suitable methods for administering such nucleic acids are available and well known to those of skill in the art, and while more than one route can be used to administer a particular composition, a particular route often provides a more direct and effective response than another route.

[0219] Methods for introducing DNA into hematopoietic stem cells are disclosed, for example, in U.S. Patent No. 5,928,638. Vectors that can be used to introduce a transgene into hematopoietic stem cells (such as CD34+ cells) include adenovirus type 35.

[0220] Vectors suitable for introducing transgenes into immune cells (such as T cells) include non-integrating lentiviral vectors. See, for example, Ory et al. (1996) Proc. Natl. Acad. Sci. USA 93:11382-11388; Dull et al. (1998) J. Virol. 72:8463-8471; Zuffery et al. (1998) J. Virol. 72:9873-9880; Follenzi et al. (2000) Nature Genetics 25:217-222.

[0221] A pharmaceutically acceptable carrier is determined in part by the particular composition being administered and by the particular method used to administer the composition. Thus, as described below, there are a wide variety of suitable formulations of pharmaceutical compositions available (see, for example, Remington’s Pharmaceutical Sciences, 17th ed., 1989).

[0222] As described above, the disclosed methods and compositions can be used with any type of cell, including but not limited to prokaryotic cells, fungal cells, archaeal cells, plant cells, insect cells, animal cells, vertebrate cells, mammalian cells, and human cells, including any type of T cell and stem cells. Suitable cell lines for protein expression are known to those of skill in the art and include but are not limited to COS, CHO (e.g., CHO-S, CHO-K1, CHO-DG44, CHO-DUXB11), VERO, MDCK, WI38, V79, B14AF28-G3, BHK, HaK, NS0, SP2 / 0-Ag14, HeLa, HEK293 (e.g., HEK293-F, HEK293-H, HEK293-T), perC6, insect cells such as Spodoptera frugiperda (Sf), and fungal cells such as Saccharomyces, Pichia, and Schizosaccharomyces. Progeny, variants, and derivatives of these cell lines can also be used.

[0223] Applications

[0224] The use of engineered nucleases to treat and prevent disease is expected to be one of the most important developments in medicine in the coming years. The methods and compositions described herein are used to increase the specificity of these novel tools to ensure that the desired target site will be the primary location of cleavage. For all in vitro, in vivo, and ex vivo applications, minimizing or eliminating off-target cleavage will be required to realize the full potential of this technology.

[0225] Exemplary genetic diseases include, but are not limited to: achondroplasia, achromatopsia, acid maltase deficiency, adenosine deaminase deficiency (OMIM No. 102700), adrenoleukodystrophy, Aicardi syndrome, alpha-1 antitrypsin deficiency, alpha thalassemia, androgen insensitivity syndrome, Apert syndrome, arrhythmogenic right ventricular dysplasia, ataxia telangiectasia, Barth syndrome, beta-thalassemia, blue rubber bleb nevus syndrome, Canavan disease, chronic granulomatous disease (CGD), cri du chat syndrome, cystic fibrosis, Dercum's disease, ectodermal dysplasia, Fanconi anemia, fibrodysplasia ossificans progressive, fragile X syndrome, galactosemia, Gaucher's disease, generalized gangliosidosis (e.g., GM1), hemochromatosis, hemoglobin C mutation in the 6 thbeta-globin codon, HbC), hemophilia, Huntington’s disease, Hurler Syndrome, hypophosphatasia, Klinefleter syndrome, Krabbes Disease, Langer-Giedion Syndrome, leukocyte adhesion deficiency (LAD, OMIM No.116920), leukodystrophy, long QT syndrome, Marfan syndrome, Moebius syndrome, mucopolysaccharidosis (MPS), nail-patella syndrome, nephrogenic diabetes insipidus, neurofibroma, Neimann-Pick disease, osteogenesis imperfecta, phenylketonuria (PKU), porphyria, Prader-Willi syndrome, progeria, Proteus syndrome, retinoblastoma, Rett syndrome, Rubinstein-Taybi syndrome, Sanfilippo syndrome, severe combined immunodeficiency (SCID), Shwachman syndrome, sickle cell disease (sickle cell anemia), Smith-Magenis syndrome, Stickler syndrome, Tay-Sachs disease, Thrombocytopenia Absent Radius (TAR) syndrome, Treacher Collins syndrome, trisomy, tuberous sclerosis, Turner's syndrome, urea cycle disorders, von Hippel-Landau disease, Waardenburg syndrome, Williams syndrome, Wilson's disease, Wiskott-Aldrich syndrome, X-linked lymphoproliferative syndrome (XLP, OMIM number 308240).

[0226] Additional exemplary diseases treatable by targeted DNA cleavage and / or homologous recombination include acquired immunodeficiency, lysosomal storage diseases (e.g., Gaucher's disease, GM1, Fabry disease, and Tay-Sachs disease), mucopolysaccharidoses (e.g., Hunter's disease, Hurler's disease), hemoglobinopathies (e.g., sickle cell disease, HbC, alpha-thalassemia, beta-thalassemia), and hemophilia.

[0227] Such methods also allow for the treatment of genetic diseases by treating an infection (viral or bacterial) in a host (e.g., by blocking the expression of viral or bacterial receptors, thereby preventing infection and / or transmission in the host organism).

[0228] Targeted cleavage of an infected or integrated viral genome can be used to treat viral infections in a host. Additionally, targeted cleavage of genes encoding viral receptors can be used to block the expression of such receptors, thereby preventing viral infection and / or viral transmission in the host organism. Targeted mutagenesis of genes encoding viral receptors (e.g., the CCR5 and CXCR4 receptors of HIV) can be used to render the receptors unable to bind the virus, thereby preventing new infections and blocking the spread of existing infections. See U.S. Patent Publication No. 2008 / 015996. Non-limiting examples of targetable viruses or viral receptors include herpes simplex virus (HSV), such as HSV-1 and HSV-2, varicella zoster virus (VZV), Epstein-Barr virus (EBV), and cytomegalovirus (CMV), HHV6, and HHV7. The hepatitis virus family includes hepatitis A virus (HAV), hepatitis B virus (HBV), hepatitis C virus (HCV), delta hepatitis virus (HDV), hepatitis E virus (HEV), and hepatitis G virus (HGV). Other targetable viruses or their receptors include, but are not limited to, Picornaviridae (e.g., poliovirus, etc.); Caliciviridae; Togaviridae (e.g., rubella virus, dengue virus, etc.); Flaviviridae; Coronaviridae; Reoviridae; Birnaviridae; Rhabdoviridae (e.g., rabies virus, etc.); Filoviridae; Paramyxoviridae (e.g., mumps virus, measles virus, respiratory syncytial virus, etc.); Orthomyxoviridae (e.g., influenza A, B, and C viruses, etc.); Bunyaviridae; Arenaviridae; Retroviradae; Lentivirus (e.g., HTLV-I; HTLV-II; HIV-1 (also known as HTLV-III, LAV, ARV, hTLR, etc.) HIV-II); simian immunodeficiency virus (SIV), human papillomavirus (HPV), influenza virus, and tick-borne encephalitis virus. For descriptions of these and other viruses, see, e.g., Virology, 3rd Edition (W.K. Joklik ed. 1988); Fundamental Virology, 2nd Edition (B.N. Fields and D.M. Knipe eds. 1991). Receptors for HIV include, for example, CCR-5 and CXCR-4.

[0229] Thus, the heterodimeric cleavage domain variants described herein provide broad utility for enhancing the specificity of ZFNs in gene modification applications. These variant cleavage domains can be readily incorporated into any existing ZFN by site-directed mutagenesis or subcloning to enhance the in vivo specificity of any ZFN dimer.

[0230] As described above, the compositions and methods described herein can be used for gene modification, gene correction, and gene disruption. Non-limiting examples of gene modification include targeted integration based on homologous directed repair (HDR); gene correction based on HDR; gene modification based on HDR; gene disruption based on HDR; gene disruption based on NHEJ and / or combinations of HDR, NHEJ, and / or single-strand annealing (SSA). Single-strand annealing (SSA) refers to the repair of double-strand breaks between two repeat sequences occurring in the same orientation by resection of the DSB by a 5'-3' exonuclease to expose two complementary regions. The single strands encoding the two direct repeats then anneal to each other, and the annealed intermediate can be processed such that the single-stranded tails (the portions of single-stranded DNA not annealed to any sequence) are digested away, the gaps are filled in by DNA polymerase, and the DNA ends are religated. This results in the deletion of the sequence located between the direct repeats.

[0231] Compositions comprising a cleavage domain (e.g., ZFN, TALEN, CRISPR / Cas system) and the methods described herein can also be used to treat various genetic and / or infectious diseases.

[0232] The compositions and methods can also be applied to stem cell-based therapies, including but not limited to: correction of somatic mutations by short patch gene conversion or targeted integration for monogenic gene therapy; disruption of dominant negative alleles; disruption of genes required for pathogen entry or productive infection of cells; enhancement of tissue engineering, e.g., by modifying gene activity to promote the differentiation or formation of functional tissue; and / or disruption of gene activity to promote the differentiation or formation of functional tissue; blocking or inducing differentiation, e.g., by disrupting genes that block differentiation to promote the differentiation of stem cells along a specific lineage pathway, targeted insertion of genes or siRNA expression cassettes that can stimulate stem cell differentiation, targeted insertion of genes or siRNA expression that can block stem cell differentiation and allow for better expansion and maintenance of pluripotency, and / or targeted insertion of a reporter gene in-frame with an endogenous gene that is a marker of pluripotency or differentiation status, which will allow for easy scoring of how changes in the differentiation status of stem cells and the expression of media, cytokines, growth conditions, genes, siRNA, shRNA, or miRNA molecules, the exposure of cell surface markers by antibodies, or drugs alter this status; somatic cell nuclear transfer, e.g., somatic cells of a patient can be isolated, the intended target gene can be modified in a suitable manner, cell clones can be generated (and quality controlled to ensure genomic safety), and the nuclei from these cells can be isolated and transferred into unfertilized eggs to generate patient-specific hES cells, which can be directly injected or differentiated before implantation into the patient, thereby reducing or eliminating tissue rejection; universal stem cells obtained by knocking out MHC receptors (e.g., to generate cells with reduced or completely eliminated immune identity). Cell types used for such procedures include but are not limited to T cells, B cells, hematopoietic stem cells, and embryonic stem cells. Additionally, induced pluripotent stem cells (iPSCs), which can also be generated from the somatic cells of a patient, can be used. Thus, these stem cells or their derivatives (differentiated cell types or tissues) can potentially be implanted into anyone, regardless of their origin or histocompatibility.

[0233] The compositions and methods can also be used for somatic cell therapy, thereby allowing for the generation of stocks of cells that have been modified to enhance their biological properties. Such cells can be infused into various patients, independent of the donor source of the cells and their histocompatibility with the recipient.

[0234] In addition to therapeutic applications, when used in engineered nucleases, the increased specificity provided by the variants described herein can be used for crop engineering, cell line engineering, and the construction of disease models. The obligate heterodimer cleavage half-domains provide a direct means for improving nuclease properties.

[0235] The engineered cleavage half-domains described herein can also be used in gene modification schemes that require simultaneous cleavage at multiple targets to delete intervening regions or to immediately alter two specific loci. Cleavage at two targets would require cellular expression of four ZFNs or TALENs, which could generate ten different active ZFN or TALEN combinations. For such applications, replacing the wild-type nuclease domain with these novel variants would eliminate the activity of the undesired combinations and reduce the chance of off-target cleavage. If cleavage at a particular desired DNA target requires nuclease activity against A + B and simultaneous cleavage at a second desired DNA target requires nuclease activity against X + Y, then the mutations described herein prevent pairing of A with A, A with X, A with Y, etc. Thus, these FokI mutations reduce non-specific cleavage activity due to "illegal" pair formation and allow for the generation of more effective orthogonal nuclease mutant pairs (see co-owned U.S. Patent Publications No. 20080131962 and 20090305346).

[0236] Examples

[0237] Example 1: Preparation of ZFNs

[0238] ZFNs targeting sites in the BCL11A and TCRA (targeting the constant region, also known as TRAC) genes were designed and incorporated into plasmid vectors, substantially as described in Urnov et al. (2005) Nature 435(7042):646-651, Perez et al. (2008) Nature Biotechnology 26(7):808-816, and PCT Patent Publications No. WO 2016183298 and PCT Publication No. WO2017106528. ZFNs targeting AAVS1 as described in U.S. Publication No. 20150110762 were also used.

[0239] Example 2: Mutants in FokI Residues Targeting Interaction with Phosphate

[0240] Using a model of the FokI cleavage domain (Miller et al. (2007) Nat Biotech 25(7):778-85), positively charged arginine or lysine amino acid residues potentially interacting with phosphate on the DNA backbone were identified ( Figure 1 ).

[0241] Then the identified positions (amino acids 416, 422, 447, 448, and 525) in the FokI domain were specifically altered (mutated) to serine residues to eliminate the interaction between the original positive amino acid on the DNA and the negatively charged phosphate (see Figure 2A and Figure 2B)。When two ZFN monomers contain these mutations, many different combinations can be generated (see the illustration in Figure 2C ). These new FokI mutants were prepared in the 'KKR' FokI monomer of the ELD / KKR heterodimer (see U.S. Patent No. 8,962,281), and then ligated to a ZFN pair specific for the BCL11A enhancer region (SBS#51857ELD / SBS#51949KKR or SBS#51857_KKR / SBS#51949_ELD, the 'parent' proteins highlighted in gray in Figure 3). Mutations were made in each monomer, and then the cleavage activity against the native BCL11A target was tested in CD34+ T cells in various combinations as shown in Figure 3. Off-target sites were previously identified by unbiased capture analysis (PCT Patent Publication No. WO 2016 / 183298). The off-target sites are listed in Table 1 below, where each site is identified with a unique and randomly generated 'license plate' letter identifier, with the license plate PRJIYLFN indicating the expected target BCL11A sequence. In the table below, the locus of each site is also indicated (coordinates are listed in accordance with the hg38 assembly of the U.C. Santa Cruz Human Genome Browser sequence database (Kent et al. (2002), Genome Res. 12(6):996-1006)).

[0242] Table 1: Cleavage sites identified by SBS#51857 / SBS#51949

[0243]

[0244]

[0245] The data presented in Figure 3 show that certain mutations reduce the activity of the protein against the cognate BCL11A target site, and the off-target cleavage activity correspondingly decreases (e.g., see 51857-ELD / 51949-KKR_R447S: the on-target activity decreased to 11.60% indel compared to 80.59% of the parent; the activity at NIFMAEVG off-target also decreased to 0.05% compared to the value of 9.04% of the parent, and the activity at PEVYOHIU off-target decreased from the value of 0.65% of the parent to 0.03% ( Figure 3AHowever, for other mutations, on-target cleavage activity remained robust while activity at two measured off-target sites was significantly reduced. For example, 51857-ELD / 51949-KKR_R416S had on-target activity against BCL11A very similar to the parental protein (at a 2 μg dose (20 μg / mL), 80.63% indel for the mutant pair vs. 80.59% for the parent, while activity at both off-targets was significantly reduced (at off-target site NIFMAEVG, 0.75% for the mutant pair vs. 9.04% for the parent, and at off-target site PEVYOHIU, 0.08% for the mutant pair vs. 0.65% for the parent)( Figure 3A ).

[0246] Mutant proteins were also assembled using the heterodimeric FokI domain in the reverse orientation, i.e., Figure 3A showing results using 51857-ELD / 51949-KKR, while Figure 3B depicting results using 51857-KKR / 51949-ELD. Again, there were pairs (such as 51857-KKR / 51949-ELD_K448S) that had some on-target activity remaining robust and similar to the parental pair (83.02% vs. 88.28% respectively), but showed reduced activity at off-target positions (for NIFMAEVG, 1.00% vs. 9.26%; for PEVYOHIU, 0.33% vs. 0.87%( Figure 3B ).

[0247] Experiments were also conducted using a pair of TCRA (TRAC)-specific ZFNs: SBS#52742_ELD / SBS#52774_KKR (U.S. Patent Publication No. US-2017-0211075-A1). These experiments were performed in K562 cells, where the cells were treated with 100 or 400 ng of mRNA encoding each ZFN. The ZFN pairs consisted of one mutant and one non-mutant partner as disclosed in Table 2 below. In these experiments, Figure 1 all positively charged amino acids identified were mutated to serine (S). Briefly, 2x10e5 cells were used per transfection, where the mRNA was delivered to the cells via use of the Amaxa 96-well shuttle system. The transfected cells were harvested on day 3 post-transfection and processed by standard methods for MiSeq (Illumina) analysis. The data are shown in Table 2 below and demonstrate that some mutations maintained robust on-target activity while other mutations (such as 52742ELD_K469S, identified as an active site residue in Figure 1 ) knocked out cleavage activity.

[0248] Table 2: Targeting Results of TCRA (TRAC) ZFNs with FokI Mutations

[0249]

[0250]

[0251] Off-target analysis was also performed to identify the top-ranked off-target cleavage sites as determined by unbiased capture analysis (PCT Patent Publication No. WO 2016 / 183298). The nuclease activities at four genomic loci (the expected target in TCRA (TRAC) and three off-targets) identified for this ZFN pair by unbiased capture assay are presented in Table 3 below.

[0252] Table 3: Cleavage Sites of TCRA (TRAC)-Specific ZFNs

[0253]

[0254] The nuclease activities at two off-target sites (designated OT11 or XVFENVRX and OT16 or XSKWTVWD, shown in Table 3 above) were analyzed at a 400 ng mRNA dose of each ZFN, where one partner of the ZFN pair had a designated mutation and the other partner retained the unmodified ELD or KKR FokI domain. The data shown (Table 4) indicate the percentages of indels (activity) observed using 400 ng of each ZFN partner. These data also show that some mutations maintain nearly equal on-target activity but show reduced off-target activity. For example, for SBS#52774_KKR_K387S, an on-target activity of 89.06% indels was shown at 400 ng (compared to 89.08 for the 52774_KKR parent) and a combined off-target activity of 6.39% indels at OT11 and OT16 (compared to 15.07% for the 52774_KKR parent). Results include FokI mutations tested in BCL11A-specific ZFNs, where changes in the positive charges at positions 416, 422, 447, 448, and 525 reduced off-target activity.

[0255] Table 4: Off-Target Cleavage of TRAC ZFN FokI Mutants

[0256]

[0257] The data were then compared with the estimated distances between the mutated amino acid residues and the DNA molecule to examine the effects on on-target and off-target cleavage ( Figure 4)。The activity of each ZFN pair is shown as a single point, and it is demonstrated that the proteins with the most desirable properties (high on-target activity and low off-target activity) are those with mutations within 10 angstroms of the DNA molecule. Figure 4 )。Data points corresponding to ZFN pairs are indicated, where one of the ZFNs carries a FokI mutation at position 416, 422, 447, 448, or 525.

[0258] These results demonstrate that mutation of one or more of residues 416, 422, 447, 448, or 525 can increase on-target activity while decreasing off-target activity.

[0259] Example 3: Design of novel engineered zinc finger backbone mutations

[0260] Previous studies have shown that there may be some interaction between positively charged amino acid residues in the 'backbone' of zinc fingers (regions of the structure that do not participate in site-specific recognition of DNA nucleotides) and the phosphate groups on the DNA molecule (Elrod-Erickson et al., supra), see Figure 5A 。The amino acids at positions -14, -9, and -5 (all relative to the conventional numbering of the alpha-helical region) are typically positively charged and can interact with the negatively charged phosphate groups in the DNA backbone (see Figure 5B )。Therefore, the 4867 zinc finger sequences were analyzed for the presence of amino acid residues at each position in the finger sequence (see Figure 6 )。At position -5, the neutral amino acids alanine, leucine, and glutamine were observed at low but non-zero frequencies, and these amino acids were therefore used for modification of the finger backbone. Positions in 6- and 5-finger ZFPs were also identified, as well as potential substitutions ( Figure 7A and Figure 7B )。

[0261] The mutations were made in the TCRA (TRAC)-specific ZFNs against SBS#52774 / SBS#52742 (see PCT publication WO2017106528). For these proteins, a total of 21 variants were generated for each monomer ((F1, F3, F5, F1+F3, F1+F5, F3+F5, F1+F3+F5) x (R->A, R->Q, R->L)). Representative data selections (Table 5) demonstrate that many pairs show reduced off-target activity against the three off-targets analyzed. In this table, variants of 52774 ZFN and 52742 ZFN are combined, which have the indicated mutations at finger 1 (F1), finger 3 (F3), and finger 5 (F5) at position -5. The type of mutation performed is also indicated, where all mutants in this dataset are arginine (R) to glutamine (Q) mutants. For example, the protein labeled 52742-F1RQ indicates a mutant where the arginine at position -5 in finger 1 has been changed to glutamine. The front part of the table shows the activity as % indel, and the second half of the table shows the activity as a fraction of the average of two replicates of the parental ZFN pair 52774 / 52742. These experiments demonstrate that these mutations may have an impact on off-target cleavage. For example, for the 52742-F1RQ; F3RQ; F5RQ mutant, while on-target cleavage is maintained at a robust level (69.96% on-target activity compared to 62.59% activity of the parental protein at a 6 μg dose), off-target cleavage decreases (OT16 shows 19.16% cleavage activity of the parental protein and 1.43% activity in the triple mutant).

[0262] Table 5: Exemplary data for zinc finger backbone mutations

[0263]

[0264] Starting with the TCRA (TRAC)-specific parental ZFNs 52742 and 52774, the arginine (R) in the first finger of each module was replaced with alanine (A), glutamine (Q), or leucine (L). The constructs were tested in CD34+ cells, where the cells were treated with 6 μg of mRNA encoding each ZFN monomer (see Figure 8A ). For these datasets, each data bar shown is the average of the data for all mutations of each type, and the error bars represent the standard error. For example, in Figure 8AFor the leftmost black bar indicating on-target activity, this value is the average of on-targeting for all 6 mutants that may have a single mutation in the ZFN pair. Referring to the diagram in Figure 7, these thus include single mutations in the N-terminal finger of module A of 52742 (F1 in the full-length protein), single mutations in the N-terminal finger of module B of 52742 (F3 in the full-length protein), single mutations in the N-terminal finger of module C of 52742 (F5 in the full-length protein), and a similar set of mutations in the 52774 protein of the partner. For the second black bar (indicating on-target activity), the on-target library is the average of all 2 mutant proteins (a set of 6 possibilities where the mutation is made in a single partner ZFN), and for the third black bar, the on-target library is the average of 2 possibilities where all three modules in the left or right ZFN are mutated. For the off-target data shown by the gray bars, similar libraries were prepared except that data from three off-target sites were combined such that the libraries of single or double mutations in the off-target data set each included 18 data points and the library of three mutations included 6 data points. The mutations resulted in up to a 4.8-fold decrease in off-target activity.

[0265] Experiments were also performed where the two partners of the ZFN pair were mutated in a similar manner ( Figure 8B right half). For example, if the N-terminal finger in module A was changed to alanine (52742-F1RA), the partner protein would also be mutated at the N-terminal finger in module A (52774-F1RA) to obtain a total of 2 mutations in the ZFN dimer. The alanine substitution was only tested simultaneously in the two ZFNs in the dimer, and the data is shown on the Figure 8B right half (indicated by 2, 4, or 6). For comparison purposes, the R->A mutation (indicated by 1, 2, or 3) made in only one ZFN in the dimer from the left third of Figure 8B is again shown in the left half of Figure 8A . Figure 8C Similar to the Figure 8A right side, except that only 2 μg (20 μg / mL) of RNA was used. These experiments demonstrated that when a total of six mutations occur in the two ZFNs in the dimer, these mutations can produce a 27-fold decrease in off-target activity.

[0266] These experiments were also performed on 51857-ELD / 51949-KKR using the above-described BCL11A-specific ZFN pair. The experimental design was similar to that described for the TCRA (TRAC)-specific ZFN pair, and the results are shown in Figure 9Among them. The presented data show the results of pairs containing the R->Q (Gln) or R->L (Leu) mutations administered at 2 μg (20 μg / mL), where the depicted off-target results are for only one off-target site (NIFMAEVG, see Table 1).

[0267] Additional amino acid variants at the -5 position are prepared by substituting R with E, N, Y, A or L. In addition to changing the amino acid at the -5 position, a series of mutations are made at positions -9 and -14 in Finger 1 of the BCL11A - specific ZFN pair 51857 / 51949. These are tested individually and in combination with the -5 changes in Fingers 2 - 6 for on - target and off - target activities as described above. The data are shown in Table 6 below. Each protein has the same DNA - specific helix region as described for the parental proteins above, but new SBS identification numbers are given to reflect the ZFP backbone mutations. Briefly, the full name of the protein reflects the variants listed in the fingers. Thus, the "cR" part of the full name refers to the changes made in the C - terminal finger of the two - finger module (see Figure 7), and "nR" refers to the changes made in the N - terminal finger of the two - finger module. The description "rQa" means the change made at the -5 position in Module A, and the finger in which it is made is defined by cR or nR. Thus, SBS#65461 (full name 51857 - NELD - cR - 5Qa) is a derivative of SBS#51857, where a change is made at the -5 position in the C - terminal finger, where Q substitutes R in Module A. This can also be seen in the table, where Q is indicated in the F2, -5 column. When a change is made at the -14 position, as for SBS#65459 (full name 51857 - NELD - nR - 14Q - 5Qabc), it is indicated. Thus, SBS#65459 is a derivative of SBS#51857, where a change is made in the N - terminal finger of the module, where -14 in Finger 1 has been changed from R to Q, and the -5 positions in the N - terminal fingers of Modules A, B and C have been changed from R to Q. In the case of SBS#65460, -14 in Finger 1 has been changed from R to S. Again, this can also be determined from the table. The names "NELD" and "CKKR" indicate the type of FokI nuclease domain ("ELD" or "KKR" FokI domain variant) and other aspects of the vector (see PCT Publication No. WO 2017 / 136049). Table 6A shows the pairing of mutations made in the SBS#51857 derivative partner with the SBS#51949 partner or an SBS#51949 partner with a change inserting Q at the -5 position in the N - terminal finger in Modules A, B and C, where this partner also contains an R416S mutation in the FokI domain, and the experiments were carried out using 2 μg of mRNA encoding each ZFN. These experiments were also carried out using the mutations in the SBS#51949 protein (see Table 6B), where mutations in the phosphate - interacting amino acids of the FokI domain were also tested in combination with the backbone mutations. These data show that the additional changes can affect the specificity of the ZFN pair.

[0268] Table 6A: Changes in the backbone positions of ZFN SBS#51857

[0269]

[0270]

[0271] Table 6B: Additional backbone positions and changes in FokI for ZFN SBS #51949

[0272]

[0273]

[0274]

[0275]

[0276] *63022-R416S is also known as SBS #65721.

[0277] **63022-K525S is also known as SBS #65722.

[0278] The experiment was repeated using CD34+ cells, where RNA was delivered to the cells using the BTX transfection system under conditions optimized according to the manufacturer's instructions. Three concentrations of RNA were used: 60, 20, and 5 μg / mL final concentration. The data are shown in Table 6C below and demonstrate that robust on-target cleavage can be detected even at very low levels of ZFN mRNA, such that off-target cleavage is significantly reduced (>100x). Mutations are indicated in the nomenclature shown. Only the parental and 3x(R->Q) / 3x(R->Q) pairs from Table 6C below were used to repeat the experiment to determine the robustness of the results. As shown in Table 6D, the results are highly reproducible. Figure 2C The results are highly reproducible. Only the parental and 3x(R->Q) / 3x(R->Q) pairs from Table 6C below were used to repeat the experiment to determine the robustness of the results. As shown in Table 6D, the results are highly reproducible.

[0279] Table 6C: Titration of on-target and off-target effects:

[0280]

[0281] Numbers in parentheses indicate values for which no evidence of nuclease cleavage was found. * indicates that the right ZFN also contains an additional R416S mutation in the FokI nuclease domain.

[0282] Table 6D: Repeated measurements of on-target activity

[0283]

[0284] * indicates that the right ZFN also contains an additional R416S mutation in the FokI nuclease domain.

[0285] Example 4: Titrating ZFN conjugates to obtain optimal on-target activity

[0286] Titrating each partner of the ZFN pair individually allows determination of the optimal concentration of each ZFN partner and thus the maximum on-target modification by the pair while minimizing off-target modification. Each individual ZFN half-domain can have its own kinetics of binding to its cognate DNA target, and thus, through individual titration of each, optimal activity can be achieved. Thus, using ZFNs introduced as mRNA, the BCL11A-specific pair SBS#51949 / SBS#51857 was used for titration studies in CD34+ cells, where high concentrations of ZFNs were used to allow detection of off-target cleavage. The experiments (Table 7 below) found that titrating the SBS#51857 partner reduced off-target cleavage (by approximately 8-fold) while maintaining robust on-target cleavage. For example, when 60 μg / mL of 51949 mRNA was combined with 6.6 μg / mL of 51857 mRNA, the on-target modification was approximately the same as when using 60 μg / mL of each ZFN (76.1% indel when using 60 μg / mL of each, 78.3% indel when 60 μg / mL of 51949 was combined with 6.6 μg / mL of 51857), while the overall off-target decreased from 32.4% indel to 4.0%. Note that reducing the mRNA input of both ZFNs also led to a gradual decline in on-target modification, and off-target modification only decreased significantly when on-target modification was reduced.

[0287] Table 7: Single ZFN titration

[0288]

[0289] Western blot analysis was also performed to demonstrate that the expression of each ZFN partner was correlated with the amount of mRNA encoding the ZFN delivered ( Figure 10 ). In this experiment, CD34+ cells were transfected with the indicated ZFNs, and after 24 hours, the expression of the ZFN protein was detected using an anti-Flag antibody (the expression construct contained an encoded Flag tag). The expression of the two proteins was also analyzed when the ZFNs were co-introduced as a single RNA separated by a 2A self-cleaving peptide sequence (see Figure 10 , lane 2).

[0290] The titration was repeated, thus independently varying the two ZFNs to observe if there was any effect on on-target or off-target cleavage activity. The results ( Figure 11 ) demonstrated that down-titration of both partners reduced off-target cleavage, but the effect was strongest for SBS#51857. Figure 11The boxes in the BCL11A targeting graph represent the maintenance of cleavage activity against the intended BCL11A target using 60 μg of SBS#51949 mRNA and 6.6 or 60 μg of SBS#51857 mRNA, while the off-target activity at the site NIFMAEVG decreased from 27% to 4% indel as the dose of SBS#51857 decreased.

[0291] Example 5: Combining ZFN Spoiler Titration and FokI-Phosphate Contact Mutation

[0292] Next, an analysis was performed to measure the activity of ZFN spoiler titration with exemplary FokI mutation combinations. BCL11A-specific ZFNs containing Fok mutations were used in the ratios shown above to maintain on-target activity while reducing off-target cleavage. The experiment was performed in CD34+ cells, where the cells were transfected with mRNA encoding the ZFNs. The data are presented in Table 8 below. The "on-target / off-target ratio" represents the ratio of on-target to all off-target activities for the indicated sample combinations. The combination of 6.6 μg / mL of SBS#51857 and 60 μg / mL of the SBS#51949-R416S FokI mutant produced similar levels of on-target activity (84.85% vs. 78.17% for the two parental ZFNs at 60 μg), while significantly reducing the activity at all five monitored off-target sites and producing a 32-fold increase in the on-target / off-target ratio (89.52 vs. 2.76 for each parental ZFN at 60 μg / mL).

[0293] Table 8: Combining Reduced Titration of FokI Mutants with ZFN Spoilers: NHEJ Activity %

[0294]

[0295] The effect of all off-target sites previously listed in Table 1 on off-target activity was examined by combining titration with the FokI mutant method (see Figure 12 ).

[0296] The data showed an overall approximately 30-fold reduction in off-target activity.

[0297] The data were also analyzed based on the number of capture events detected by the unbiased capture assay described above after cleavage. Both the parental pair (SBS51857 / SBS51949) and the variant pair (SBS63014 / SBS65721) of the above BCL11A specificities in Table 6D were used, and the number of off-target capture events was determined. In this experiment, the ZFNs were given to CD34+ cells at equal amounts (60 μg / mL final concentration) for the parental pair or at 6.6 μg and 60 μg final concentrations and for the variants at 20 μg and 60 μg final concentrations.

[0298] The results are shown in Figure 13 and demonstrate that while both the parental and variant showed robust cleavage activity (>80% indel) under both concentration conditions, off-target capture events were greatly reduced, especially for the variant pairs when delivered at unequal doses. The combination of ZFN FokI mutations with unequal concentrations of ZFN partners produced a 350-fold increase in cleavage specificity.

[0299] Example 6: Combining ZFP backbone mutations with FokI phosphate contact mutations

[0300] ZFNs were also generated that contained the zinc finger backbone mutations described in Example 3 and the FokI phosphate contact mutations described in Example 2. This combination was tested in CD34+ cells with BCL11A-specific ZFNs at two doses: 6 μg or 2 μg. The results are shown in Table 9 below and demonstrate that combining these two approaches can significantly affect the amount of off-target activity. In this table, the backbone mutations are shown as the type of mutation for each module (A, B, and / or C, see Figure 7). For example, in the sample labeled 51949LeuABC R416S, the protein contains R->L backbone substitutions in fingers 1 (module A), 3 (module B), and 5 (module C), and further carries the R416S FokI mutation. In several instances, there were ZFN pairs that had no detectable off-target activity while retaining full on-target activity. These examples are boxed in Table 9.

[0301] Table 9: Activity (% NHEJ) of ZFNs containing ZFN backbone and FokI phosphate contact mutations.

[0302]

[0303] ND: No activity detected

[0304] Example 7: Combining partner titration, ZFP backbone, and FokI phosphate contact mutations

[0305] Activity measurements were also performed where the optimal partner titration was combined with ZFNs containing ZFP backbone mutations and ZFN FokI mutations. The ZFNs were tested in CD34+ cells as described above, and the increased specificity of the ZFN pairs was demonstrated by the reduced levels of off-target activity.

[0306] Example 8: Specificity of ZFNs in CD34+ cells at the clinical scale

[0307] The specificity of the BCL11A variant for SBS63014 / SBS65722 was also tested in large-scale procedures suitable for generating cell materials for clinical trials. Briefly, according to the manufacturer's instructions, each batch of approximately 9.5 - 130 million CD34+ cells was transduced using a Maxcyte device. 80 μg / mL of SBS63014 mRNA and 20 μg / mL of SBS#65722 mRNA were used, and off-target cleavage of the cells was assayed two days later using an unbiased capture assay.

[0308] The results showed that when 47 different potential capture loci were analyzed by PCR (see Figure 14 ), no significant modifications were detected except at the on-target position, where 79.54% indels were found. This data indicates that these nucleases, as described herein, are highly specific even when used in large-scale manufacturing procedures.

[0309] Example 9: Further Specificity Studies

[0310] Specificity studies were also performed using AAVS1-targeted ZFNs with various mutations as described above. In particular, ZFNs SBS#30035 and SBS#30054 as described in US Publication No. 20150110762 were used to study various mutants, including dimerization mutants (e.g., ELD, KKR, and additional mutants), other mutations (e.g., Sharkey), and phosphate contact mutants in the activity assays as described above.

[0311] For the results in the table below, one or more of the indicated FokI mutants were introduced into the ELDFokI domain of SBS 30035 (labeled ELD_X, where X is the FokI mutation), the KKR FokI domain of SBS 30054 (labeled KKR_X, where X is the FokI mutation), or into the FokI domains of both constructs (labeled ELD_KKR_X, where X is the same FokI mutation introduced into both the ELD and KKR FokI domains). The results of the combination of the unmodified "parental constructs" 30035 and 30054 are labeled "parental", "parentals", "parental ZFN", etc. A lower dose of each parental construct is typically labeled "half dose". The nuclease-free negative control is typically labeled "GFP". The ratio of the % indels at the expected locus (usually labeled "AAVS1") in a given experiment divided by the sum of the % indels at all off-targets measured is typically labeled "ratio", "on / off ratio", etc.

[0312] The positions of the AAVS1 target and off-targets are shown below, where 'hg38' represents the assembled genomic data according to the UCSC Genome Browser database, build hg38:

[0313]

[0314] Tables 10A - 10C show the cleavage results of the on-target (AAVS1) and three off-targets (OT1, OT2, OT3) from 2 different experiments, as well as the on-target to off-target ratios of the dimerization mutants ELD, KKR, and ELD-KKR with additional substitution mutations (each amino acid of the wild type) at R416 or K525. The results using the ELD_S418D, ELD_N476D, ELD_I479T, ELD_Q481E, ELD_N527D, and ELD_Q531R mutants are also shown.

[0315] Table 10A

[0316]

[0317]

[0318]

[0319] Table 10B

[0320]

[0321]

[0322]

[0323] Table 10C

[0324]

[0325]

[0326]

[0327] Tables 11A - 11C show the cleavage results of the on-target (AAVS1) and three off-targets (OT1, OT2, OT3) from 2 different experiments, as well as the on-target to off-target ratios of the indicated mutants (including combinations of substitution mutants at 418, 422, and 525 with the dimerization mutants ELD and / or KKR).

[0328] Table 11A

[0329]

[0330]

[0331]

[0332]

[0333] Table 11B

[0334]

[0335]

[0336]

[0337]

[0338] Table 11C

[0339]

[0340]

[0341]

[0342]

[0343] Tables 12A - 12C show the cleavage results of the target (AAVS1) and three off - targets (OT1, OT2, OT3) from two different experiments, as well as the on - target to off - target ratios of the indicated mutants.

[0344] Table 12A

[0345]

[0346]

[0347]

[0348]

[0349]

[0350]

[0351] Table 12B

[0352]

[0353]

[0354]

[0355]

[0356]

[0357] Table 12C

[0358]

[0359]

[0360]

[0361]

[0362]

[0363]

[0364] Tables 13A through 13C show on-target and off-target cleavage events using ZFNs targeting AAVS1 with the indicated mutations (ELD, KKR, ELD / KKR, and other indicated mutations).

[0365] Table 13A

[0366]

[0367]

[0368]

[0369]

[0370] Table 13B

[0371]

[0372]

[0373]

[0374]

[0375] Table 13C

[0376]

[0377]

[0378]

[0379] Tables 14A - 14C show the results using the indicated mutants.

[0380] Table 14A

[0381]

[0382]

[0383] Table 14B

[0384]

[0385]

[0386]

[0387] Table 14C

[0388]

[0389]

[0390] Tables 15A through 15C show the results of repeated experiments using the indicated mutants.

[0391] Table 15A

[0392]

[0393]

[0394]

[0395] Table 15B

[0396]

[0397]

[0398] Table 15C

[0399]

[0400]

[0401]

[0402] Table 16B

[0403]

[0404]

[0405]

[0406] Tables 17A through 17C show the results using the indicated exemplary mutants (including exemplary double mutants).

[0407] Table 17A

[0408]

[0409]

[0410] Table 17B

[0411]

[0412]

[0413] Table 17C

[0414]

[0415]

[0416]

[0417] Tables 18A through 18C show the results using the indicated cleavage domain mutants.

[0418] Table 18A

[0419]

[0420]

[0421] Table 18B

[0422]

[0423]

[0424] Table 18C

[0425]

[0426] Figures 15 and 16 also show a summary of selected results using the indicated mutants.

[0427] The results demonstrate highly specific cleavage using the mutants described herein.

[0428] All patents, patent applications, and publications mentioned herein are hereby incorporated by reference in their entirety.

[0429] Although the disclosure has been provided in detail for purposes of clarity of understanding, it will be apparent to those skilled in the art that various changes and modifications can be made without departing from the spirit or scope of the disclosure. Accordingly, the foregoing description and examples should not be construed as restrictive.

Claims

1. A zinc finger protein comprising at least 3 zinc finger DNA-binding domains, wherein one or more of said zinc finger DNA-binding domains comprise mutations of amino acid residues (-5), (-9) and / or (-14) relative to the start numbering of the alpha helix region.

2. The zinc finger protein according to claim 1, wherein each zinc finger DNA-binding domain comprises two beta sheets, an alpha helix, and a recognition helix region that binds to a nucleotide sequence.

3. The zinc finger protein according to claim 1 or claim 2, wherein the mutations at (-5), (-9) and / or (-14) are mutated to alanine (A), leucine (L), serine (S), asparagine (N), glutamine (E), tyrosine (Y) and / or glutamine (Q) residues.

4. The zinc finger protein according to any one of claims 1-3, wherein the Arg (R) at (-5) is changed to Tyr (Y), Asp (N), Glu (E), Leu (L), Gln (Q) or Ala (A) residues; the Arg (R) at (-9) is replaced with Ser (S), Asp (N) or Glu (E); and / or the Arg (R) at (-14) is replaced with Ser (S) or Gln (Q) residues.

5. A zinc finger nuclease, said zinc finger nuclease comprising the zinc finger protein according to any one of claims 1-4 and a cleavage domain.

6. The zinc finger nuclease according to claim 5, wherein the cleavage domain comprises an engineered FokI cleavage domain.

7. The zinc finger nuclease according to claim 5 or claim 6, wherein the engineered cleavage half-domain comprises one or more mutations of residues 416, 421, 422, 424, 472, 478, 480, 525 or 542, or wherein the wild-type residue at position 418 is replaced with a Glu (E) or Asp (D) residue, the wild-type residue at position 446 is replaced with an Asp (D) residue, the wild-type residue at position 448 is replaced with an Ala (A) residue (K448A), the wild-type residue at position 479 is replaced with a Gln (Q) or Thr (T) residue (I479Q or I479T), the wild-type residue at position 481 is replaced with an Ala (A), Cys (C), Asp (D), Ser (S), or Glu (E) residue (Q481A, Q481C, Q481D, Q481S, Q481E), or the wild-type residue at position 476 is replaced with a Glu (E) or Glu (G) residue (N476E, N476G), wherein the amino acid residues are numbered relative to the full-length FokI wild-type cleavage domain as shown in SEQ ID NO:

1.

8. The zinc finger nuclease according to claim 7, which comprises the following mutations: mutations at residues 416 and 422; the mutation at position 416 and the K448A mutation; the K448A mutation and the I479Q mutation; the K448A and Q481A mutations; and / or the K448A mutation and the mutation at position 525.

9. The zinc finger nuclease according to claim 8, wherein: the wild-type residue at position 416 is replaced with a Glu (E), Asp (D), His (H), or Asn (N) residue (R416E, R416D, R416H, R416N); the wild-type residue at position 421 is replaced with a Ser (S) residue (D421S); the wild-type residue at position 422 is replaced with a His (H) residue (R422H); the wild-type residue at position 424 is replaced with a Phe (F) residue (L242F); the wild-type residue at position 472 is replaced with an Asp (D) residue (S472D); the wild-type residue at position 478 is replaced with an Asp (D) residue (P478D); the wild-type residue at position 480 is replaced with an Asp (D) residue (G480D); the wild-type residue at position 525 is replaced with an Ala (A), Cys (C), Glu (E), Ile (I), Ser (S), Thr (T), or Val (V) residue (K525A, K525C, K525E, K525I, K525S, K525T, K525V); and / or the wild-type residue at position 542 is replaced with an Asp (D) residue (N542D).

10. The zinc finger nuclease according to any one of claims 5-9, wherein the cleavage domain is a FokI cleavage half-domain, and the FokI cleavage half-domain further comprises additional amino acid mutations at positions 432, 441, 483, 486, 487, 490, 496, 499, 527, 537, 538, and 559.

11. A zinc finger nuclease, which comprises the first and second dimeric zinc finger nucleases and a second zinc finger nuclease according to any one of claims 5-8.

12. One or more polynucleotides, wherein the polynucleotide encodes the zinc finger protein according to any one of claims 1-4 or the zinc finger nuclease according to any one of claims 5-11.

13. An isolated cell, which comprises one or more polynucleotides according to claim 12.

14. The zinc finger nuclease according to claim 11 or one or more polynucleotides according to claim 12, for use in a method for cleaving genomic cell chromatin in a target region, the method comprising: expressing in a cell the zinc finger nuclease or the one or more polynucleotides, wherein the zinc finger nuclease is expressed in the cell and site-specifically cleaves the nucleotide sequence in the target region of the genomic cell chromatin.

15. The zinc finger nuclease or one or more polynucleotides used according to claim 14, further comprising contacting the cell with a donor polynucleotide; wherein cleavage of the cell chromatin promotes homologous recombination between the donor polypeptide and the cell chromatin.

16. An isolated cell or cell line, the isolated cell or cell line comprising a zinc finger protein according to claims 1-4, a zinc finger nuclease according to any one of claims 5-11, and / or one or more polynucleotides according to claim 12.

17. A composition, the composition comprising first and second polynucleotides encoding the first and second zinc finger nucleases of claim 11.

Citation Information

Patent Citations

  • In vitro peptide or protein expression library

    GB2338237A

  • Targeted modification of chromatin structure

    US20020115215A1

  • Modulation of gene expression using localization domains

    US20030082552A1

  • Methods and compositions for using zinc finger endonucleases to enhance homologous recombination

    US20030232410A1

  • Hybrid and single chain meganucleases and use thereof

    US20040002092A1