Homing endonuclease variants

JP2025072470A5Pending Publication Date: 2025-08-08NOVO NORDISK AS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025016020
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-10
Filing Date
2025-02-03
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Current genome editing strategies using nuclease-based tools face challenges such as reduced efficiency, specificity, stability, and delivery, which hinder their therapeutic potential for treating diseases related to genetic mutations.

Method used

Development of homing endonuclease (HE) variants and megaTALs with improved thermal stability and catalytic activity, achieved through specific amino acid substitutions, to enhance their ability to cleave target sites in the human genome.

Benefits of technology

The engineered HE variants and megaTALs demonstrate increased thermal stability and editing activity, addressing the limitations of existing genome editing tools and potentially improving their therapeutic efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000059_0000
    Figure 00000059_0000
  • Figure 00000059_0001
    Figure 00000059_0001
  • Figure 00000059_0002
    Figure 00000059_0002
Patent Text Reader

Abstract

To provide homing endonuclease variants.SOLUTION: The disclosure provides a homing endonuclease variant and MEGATAL reprogrammed to bind to a genomic polynucleotide sequence and cleave it. The homing endonuclease and MEGATAL have been engineered to increase heat stability and / or activity. The invention provides, for example, an I-Onul homing endonuclease (HE) variant comprising one or more amino acid substitutions relative to the parent I-Onul HE comprising a specific amino acid sequence, where the one or more amino acid substitutions increase the heat stability of the I-Onul HE variant as compared with the parent I-Onul HE.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 62 / 777,476, filed December 10, 2018, which is incorporated herein by reference in its entirety.

[0002] SEQUENCE LISTING STATEMENT The sequence listing associated with this application is provided in text format in lieu of a paper copy and is incorporated herein by reference. The name of the text file containing the sequence listing is BLBD_110_01WO_ST25.txt. The text file is 83KB, was created on November 26, 2019, and was submitted electronically via EFS-Web simultaneously with the filing of the specification.

[0003] background The present disclosure relates to genome editing compositions with improved stability and activity. More specifically, the present disclosure relates to nuclease variants, compositions with improved stability and / or activity, and methods of making and using the same for genome editing. [Background technology]

[0004] 2. Description of Related Art Mutations in 3000 human genes have already been associated with disease phenotypes (www.omim.org / statistics / geneMap), and disease-associated genetic variants are being uncovered at an astonishing rate, many of which are associated with monogenic diseases or cancer. Genome editing strategies based on programmable nucleases, such as meganucleases, zinc finger nucleases, transcription activator-like effector nucleases and clustered regularly interspaced short palindromic repeats (CRISPR)-associated nuclease Cas9, hold tremendous, but as yet unrealized, potential for the treatment of diseases, disorders, and conditions with a genetic component. Particular obstacles to the implementation of nuclease-based genome editing tools as therapeutic strategies include, but are not limited to, reduced genome editing efficiency, nuclease specificity, nuclease stability, and delivery challenges. The current state of the art for most genome editing strategies fails to meet some or all of these criteria. Summary of the Invention [Means for solving the problem]

[0005] This disclosure relates generally, in part, to compositions comprising homing endonuclease (HE) variants and megaTALs with improved stability and activity to cleave target sites in the human genome, and methods of using the same. In certain embodiments, the HE variants and megaTALs are engineered to improve or enhance the thermostability of the enzymes and / or improve the catalytic activity of the enzymes.

[0006] In various embodiments, the present disclosure contemplates polypeptides that include engineered homing endonucleases that have been engineered, in part, to improve stability and binding and cleavage of the target site.

[0007] In various embodiments, the I-OnuI homing endonuclease (HE) variant comprises one or more amino acid substitutions relative to a parent I-OnuI HE comprising the amino acid sequence set forth in SEQ ID NO:1, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0008] In certain embodiments, the one or more amino acid substitutions are at amino acid positions selected from the group consisting of I14, A19, V116, F168, D208, N246, and L263.

[0009] In certain embodiments, the amino acid substitutions are at amino acid positions I14, A19, F168, D208, and N246.

[0010] In some embodiments, the one or more amino acid substitutions are at an amino acid position selected from the group consisting of K108, K156, S176, E231, V261, E277, and G300.

[0011] In some embodiments, the one or more amino acid substitutions are at an amino acid position selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0012] In certain embodiments, the I-OnuI HE variant comprises three or more amino acid substitutions and is at an amino acid position selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300.

[0013] In certain embodiments, an I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0014] In a further embodiment, the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300.

[0015] In certain embodiments, an I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0016] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE. 50 At least 10°C higher than TM 50 has.

[0017] In some embodiments, the I-OnuI HE variant has the TM of the parent I-OnuI HE. 50 At least 15°C higher than TM 50 has.

[0018] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE. 50At least 20°C higher than TM 50 has.

[0019] In certain embodiments, the I-OnuI HE variant has a TM50 that is at least 25° C. higher than the TM50 of the parent I-OnuI HE.

[0020] In certain embodiments, the I-OnuI HE variant targets a site in a gene selected from the group consisting of: HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS.

[0021] In various embodiments, the I-OnuI homing endonuclease (HE) variant comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0022] In a further embodiment, the parent I-OnuI HE amino acid sequence is set forth in SEQ ID NO:1.

[0023] In further embodiments, the one or more amino acid substitutions are at an amino acid position selected from the group consisting of I14, A19, V116, F168, D208, N246, and L263.

[0024] In certain embodiments, the amino acid substitutions are at amino acid positions I14, A19, F168, D208, and N246.

[0025] In certain embodiments, the amino acid substituted for I14 is selected from the group consisting of: S, N, M, K, F, D, T, and V.

[0026] In some embodiments, the amino acid substituted for I14 is selected from the group consisting of: T and V.

[0027] In certain embodiments, the amino acid substituted for A19 is selected from the group consisting of: C, D, I, L, S, T and V.

[0028] In a further embodiment, the amino acid substituted for A19 is selected from the group consisting of: T and V.

[0029] In certain embodiments, the amino acid substituted for V116 is selected from the group consisting of: F, D, A, L and I.

[0030] In certain embodiments, the amino acid substituted for V116 is selected from the group consisting of: L and I.

[0031] In certain embodiments, the amino acid substituted for F168 is selected from the group consisting of: H, Y, I, V, P, L and S.

[0032] In a further embodiment, the amino acid substituted for F168 is selected from the group consisting of: L and S.

[0033] In some embodiments, the amino acid substituted for D208 is selected from the group consisting of: N, V, Y, and E.

[0034] In a particular embodiment, the amino acid substituted for D208 is E.

[0035] In certain embodiments, the amino acid substituted for N246 is selected from the group consisting of: H, I, D, R, S, T, V, Y, and K.

[0036] In a particular embodiment, the amino acid substituted for N246 is K.

[0037] In additional embodiments, the amino acid substituted for L263 is selected from the group consisting of: H, F, P, T, V, and R.

[0038] In a particular embodiment, the amino acid substituted for L263 is R.

[0039] In additional embodiments, the one or more amino acid substitutions are selected from the group consisting of: K108, K156, S176, E231, V261, E277, and G300.

[0040] In certain embodiments, the amino acid substituted for K108 is selected from the group consisting of: E, N, Q, R, T, V, and M.

[0041] In a particular embodiment, the amino acid substituted for K108 is M.

[0042] In some embodiments, the amino acid substituted for K156 is selected from the group consisting of N, Q, R, T, V, I, and E.

[0043] In certain embodiments, the amino acid substituted for K156 is selected from the group consisting of: I and E.

[0044] In additional embodiments, the amino acid substituted for S176 is selected from the group consisting of: P, N, and A.

[0045] In a particular embodiment, the amino acid substituted for S176 is A.

[0046] In certain embodiments, the amino acid substituted for E231 is selected from the group consisting of: D, K, V, and G.

[0047] In certain embodiments, the amino acid substituted for E231 is selected from the group consisting of: K and G.

[0048] In certain embodiments, the amino acid substituted for V261 is selected from the group consisting of: D, G, I, L, S, T, and A.

[0049] In a specific embodiment, the amino acid substituted for V261 is A.

[0050] In certain embodiments, the amino acid substituted for E277 is selected from the group consisting of: A, D, G, Q, V, and K.

[0051] In a further embodiment, the amino acid substituted for E277 is K.

[0052] In a further embodiment, the amino acid substituted for G300 is selected from the group consisting of: S, V, DC, and R.

[0053] In certain embodiments, the amino acid substituted for G300 is R.

[0054] In additional embodiments, the one or more amino acid substitutions are selected from the group consisting of: N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0055] In some embodiments, the amino acid substituted at N31 is selected from the group consisting of: D, H, I, R, K, S, T and Y.

[0056] In a further embodiment, the amino acid substituted for N31 is K.

[0057] In certain embodiments, the amino acid substituted for N33 is selected from the group consisting of: D, G, H, I, K, S, T and Y.

[0058] In certain embodiments, the amino acid substituted for N33 is K.

[0059] In certain embodiments, the amino acid substituted for K52 is selected from the group consisting of: Q, R, T, Y, N, E, and M.

[0060] In a particular embodiment, the amino acid substituted for K52 is M.

[0061] In certain embodiments, the amino acid substituted for Y97 is selected from the group consisting of: F, N and H.

[0062] In a specific embodiment, the amino acid substituted for Y97 is F.

[0063] In certain embodiments, the amino acid substituted for K124 is selected from the group consisting of: E, N, R and T.

[0064] In a particular embodiment, the amino acid substituted for K124 is N.

[0065] In certain embodiments, the amino acid substituted for K147 is selected from the group consisting of: E, I, N, R and T.

[0066] In a particular embodiment, the amino acid substituted for K147 is I.

[0067] In a further embodiment, the amino acid substituted for I153 is selected from the group consisting of: D, H, K, T, Y, S, V and N.

[0068] In another embodiment, the amino acid substituted for I153 is N.

[0069] In certain embodiments, the amino acid substituted for K209 is selected from the group consisting of: E, M, N, Q and R.

[0070] In certain embodiments, the amino acid substituted for K209 is R.

[0071] In a further embodiment, the amino acid substituted for E264 is selected from the group consisting of: A, D, G, K, Q, R and V.

[0072] In a particular embodiment, the amino acid substituted for E264 is K.

[0073] In a further embodiment, the amino acid substituted for D268 is selected from the group consisting of A, E, G, H, N, V ​​and Y.

[0074] In a particular embodiment, the amino acid substituted for D268 is N.

[0075] In certain embodiments, the I-OnuI HE variant contains three or more amino acid substitutions.

[0076] In additional embodiments, the I-OnuI HE variant comprises five or more amino acid substitutions.

[0077] In certain embodiments, the I-OnuI HE variant comprises three or more amino acid substitutions and is at an amino acid position selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300.

[0078] In a further embodiment, the I-OnuI HE variant comprises three or more amino acid substitutions and are at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions are at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0079] In some embodiments, the I-OnuI HE variant comprises five or more amino acid substitutions and is at an amino acid position selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300.

[0080] In certain embodiments, an I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0081] In additional embodiments, the I-OnuI HE variant comprises the TM of the parent I-OnuI HE. 50 At least 10°C higher than TM 50 has.

[0082] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE. 50 At least 15°C higher than TM 50 has.

[0083] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE. 50 At least 20°C higher than TM 50 has.

[0084] In some embodiments, the I-OnuI HE variant has the TM of the parent I-OnuI HE. 50 At least 25°C higher than TM 50 has.

[0085] In certain embodiments, the I-OnuI HE variant targets a site in a gene selected from the group consisting of: HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS.

[0086] In various embodiments, the I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human BCL11A gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0087] In various embodiments, an I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human PCSK9 gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0088] In certain embodiments, an I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human PDCD-1 gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0089] In some embodiments, an I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human TCR alpha gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0090] In a further embodiment, the I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human CBLB gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0091] In certain embodiments, an I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human CTLA-4 gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0092] In certain embodiments, an I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human TGFβRII gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0093] In an additional embodiment, the I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human TIM3 gene comprises one or more amino acid substitutions relative to the parent I-OnuI HE, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant.

[0094] In certain embodiments, the I-OnuI HE variant comprises three or more amino acid substitutions and is at an amino acid position selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300.

[0095] In certain embodiments, the I-OnuI HE variant comprises amino acid substitutions at amino acid positions I14, A19, F168, D208, and N246.

[0096] In certain embodiments, an I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0097] In certain embodiments, an I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300.

[0098] In a further embodiment, the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268.

[0099] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE.50 At least 10°C higher than TM 50 has.

[0100] In some embodiments, the I-OnuI HE variant has the TM of the parent I-OnuI HE. 50 At least 15°C higher than TM 50 has.

[0101] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE. 50 At least 20°C higher than TM 50 has.

[0102] In certain embodiments, the I-OnuI HE variant has a TM of the parent I-OnuI HE. 50 At least 25°C higher than TM 50 has. [Brief description of the drawings]

[0103] [Figure 1] Figure 1 shows melting curves of I-OnuI LHE and I-OnuI LHE variants reprogrammed to target the CBLB and TCRα using a yeast surface display-based assay. [Diagram 2] Figure 2 shows a stacked bar graph of the percent change at each position from eight I-OnuI LHE variants that were randomly mutated and sorted for greater stability. Positions with percent change more than two standard deviations above the mean are identified. [Diagram 3]Figure 3A shows melting curves for parental BCL11A I-OnuI LHE variant endonuclease, BCL11A I-OnuI LHE variant endonuclease with individual stabilizing point mutations (F168S, I14T, N246I, V261A), and randomly mutagenized BCL11A I-OnuI LHE variants (stable HE variants). Figure 3B shows melting curves for parental BCL11A I-OnuI LHE variant endonuclease, sorted populations of BCL11A I-OnuI LHE variants generated from a randomly mutagenized library, and sorted populations of BCL11A I-OnuI LHE variants generated from mutagenesis of amino acid positions that affect stability. FIG. 3C shows the melting curves of the parent BCL11A I-OnuI LHE variant endonuclease and the BCL11A I-OnuI LHE A5 thermostable variant. [Figure 4] Figures 4A-4C show that the A5 thermostability-enhancing mutation can stabilize I-OnuI LHE variants targeting PCDC-1 (Figure 4A), CBLB (Figure 4B), and TCRα (Figure 4C). [Figure 5-1] Figure 5A shows a diagram of the reporter construct used to assess homing endonuclease stability, and Figure 5B shows a Western blot comparing protein expression of I-OnuI, the parental BCL11A I-OnuI LHE variant, and the BCL11A I-OnuI LHE A5 variant. [Figure 5-2] FIG. 5C shows the time course of GFP reporter expression compared to I-OnuI, the parental BCL11A I-OnuI LHE variant, and the BCL11A I-OnuI LHE A5 variant. [Figure 6] FIG. 6 shows that increasing the thermostability of the PDCD-1 megaTAL significantly increases the editing activity compared to its parent PDCD-1 megaTAL. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0104] Brief description of sequence identifiers SEQ ID NO:1 is the amino acid sequence of the wild-type I-OnuI LAGLIDADG homing endonuclease (LHE). SEQ ID NO:2 is the amino acid sequence of wild-type I-OnuI LHE. SEQ ID NO:3 is the amino acid sequence of a biologically active fragment of wild-type I-OnuI LHE. SEQ ID NO:4 is the amino acid sequence of a biologically active fragment of wild-type I-OnuI LHE. SEQ ID NO:5 is the amino acid sequence of a biologically active fragment of wild-type I-OnuI LHE. SEQ ID NO:6 is the amino acid sequence of the I-OnuI LHE variant reprogrammed to bind to and cleave a target site in the human TCR alpha gene. SEQ ID NO:7 is the amino acid sequence of the I-OnuI LHE variant reprogrammed to bind to and cleave a target site in the human CBLB gene. SEQ ID NO:8 is the amino acid sequence of the I-OnuI LHE variant reprogrammed to bind to and cleave a target site in the human BCL11A gene. SEQ ID NOs:9-14 set forth the amino acid sequences of I-OnuI LHE thermostable variants reprogrammed to bind and cleave a target site in the human BCL11A gene. SEQ ID NO:15 is the amino acid sequence of the I-OnuI LHE variant reprogrammed to bind to and cleave a target site in the human PDCD-1 gene. SEQ ID NO: 16 is the amino acid sequence of the I-OnuI LHE thermostable variant reprogrammed to bind to and cleave a target site in the human PDCD-1 gene. SEQ ID NO: 17 is the amino acid sequence of the I-OnuI LHE thermostable variant reprogrammed to bind to and cleave a target site in the human TCR alpha gene. SEQ ID NO: 18 is the amino acid sequence of the I-OnuI LHE thermostable variant reprogrammed to bind to and cleave a target site in the human CBLB gene. SEQ ID NO:19 is the mRNA encoding the BCL11A I-OnuI HE variant. SEQ ID NO:20 is a codon-optimized mRNA encoding the BCL11A I-OnuI HE variant. SEQ ID NO:21 is the mRNA encoding the BCL11A I-OnuI HE thermostable variant. SEQ ID NO:22 is the amino acid sequence encoding PDCD-1 megaTAL. SEQ ID NO:23 is the amino acid sequence encoding the PDCD-1 megaTAL thermostable variant. SEQ ID NOs: 24 to 34 set forth the amino acid sequences of various linkers. SEQ ID NOs: 35 to 59 set forth the amino acid sequences of the protease cleavage site and the self-cleaving polypeptide cleavage site. In the above sequences, X, if present, refers to any amino acid or to the absence of an amino acid. (Mode for carrying out the invention)

[0105] A. Overview The present disclosure generally relates to improved genome editing compositions and methods of using the same. Genome editing enzymes hold considerable promise for treating diseases, disorders, and conditions with genetic components. Previously, genome editing enzymes engineered to bind and cleave target sites in genomes can have short half-lives and fail to cleave with high efficiency. Without intending to be bound by any particular theory, the inventors have found that homing endonuclease scaffolds can be engineered to increase thermal stability and catalytic activity, and when enzymes are engineered to have higher thermal stability, genome editing enzyme activity is unexpectedly increased. Furthermore, amino acid positions of homing endonucleases can be modified to increase thermal stability and activity, and can be used to increase the thermal stability of other homing endonucleases that are reprogrammed to bind and cleave one target site, preserved, and reprogrammed to bind and cleave other target sites.

[0106] Genome editing compositions and methods contemplated in various embodiments include nuclease variants with improved stability and activity designed to bind and cleave target sequences present in genomes.In certain embodiments, contemplated nuclease variants can be used to introduce double-strand breaks in target polynucleotide sequences, which can be repaired by non-homologous end joining (NHEJ) in polynucleotide templates, for example donor repair templates, or by homology-directed repair (HDR) in the presence of donor repair templates, i.e., non-phase recombination.In certain embodiments, contemplated nuclease variants can also be designed as nickases, which generate single-strand DNA breaks that can be repaired using the cell's base excision repair (BER) mechanism or homologous recombination in the presence of donor repair templates.NHEJ is error-prone, which frequently results in the formation of small insertions and deletions that disrupt gene function. Homologous recombination requires homologous DNA as a template for repair and can be used to generate an infinite number of modifications specified by the introduction of donor DNA containing the desired sequence at the target site, flanked on both sides by sequences that produce homology to regions flanking the target site.

[0107] In certain embodiments, the homing endonuclease comprises one or more amino acid substitutions that increase stability and / or activity, In certain embodiments, the homing endonuclease comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid substitutions that increase stability and / or activity.

[0108] In certain embodiments, the homing endonuclease comprises one or more amino acid substitutions that increase stability and / or activity and is formatted as a megaTAL, hi certain embodiments, the megaTAL comprises a homing endonuclease having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acid substitutions that increase stability and / or activity.

[0109] In certain embodiments, the genome editing compositions contemplated herein comprise a homing endonuclease variant or megaTAL modified to increase stability and / or activity, and optionally an endo-processing enzyme, such as Trex2.

[0110] In various embodiments, the cell or population of cells comprises a homing endonuclease variant or a megaTAL that has been modified to increase stability and / or activity.

[0111] Thus, the methods and compositions contemplated herein represent a quantitative improvement over existing adoptive cellular therapies.

[0112] Recombinant (i.e., engineered) DNA, peptide and oligonucleotide synthesis, immunoassays, tissue culture, transformation (e.g., electroporation, lipofection), enzymatic reactions, purification and related techniques and procedures may be generally performed as described in various general and more specific references in microbiology, molecular biology, biochemistry, molecular genetics, cell biology, virology and immunology, which are cited and discussed throughout this specification. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Current Protocols in Molecular Biology (John Wiley and Sons, updated July 2008); Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, Greene Pub. Associates and Wiley-Interscience; Glover, DNA Cloning: A Practical Approach, volumes I and II (IRL Press, Oxford Univ. Press USA, 1985); Current Protocols in Immunology (eds. John E. Coligan, Ada M. Kruisbeek, David H. Margulies, Ethan M.Shevach, Warren Strober 2001 John Wiley & Sons, NY,NY);Real-Time PCR: Current Technology and Applications, Julie Logan, Kirstin Edwards and Nick Saunders (eds.), 2009, Caister Academic Press, Norfolk, UK;Anand, Techniques for the Analysis of Complex Genomes, (Academic Press, New York, 1992);Guthrie and Fink, Guide to Yeast Genetics and Molecular Biology (Academic Press, New York, 1991);Oligonucleotide Synthesis (N. Gait, ed.), 1984;Nucleic Acid The Hybridization (B. Hames and S. Higgins, eds.), 1985;Transcription and Translation (B. Hames and S. Higgins, eds.), 1984;Animal Cell Culture (R. Freshney, ed.), 1986;Perbal, A Practical Guide to Molecular Cloning (1984); Next-Generation Genome Sequencing (Janitz, 2008 Wiley-VCH); PCR Protocols (Methods in Molecular Biology) (Park, ed., 3rd ed., 2010, Humana Press); Immobilized Cells And Enzymes (IRL Press, 1986); the treatise, Methods In Enzymology (Academic Press, Inc., NY); Gene Transfer Vectors For Mammalian Cells (JH Miller and MPCalos, ed., Cold Spring Harbor Laboratory, 1987); Harlow and Lane, Antibodies, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1998); Immunochemical Methods In Cell And Molecular Biology (Mayer and Walker, eds., Academic Press, London, 1987); Handbook Of Experimental Immunology, volumes I-IV (D. M. Weir and C. C. Blackwell, eds., 1986); Roitt, Essential Immunology, ed. 6, (Blackwell Scientific Publications, Oxford, 1988); Current Protocols in Immunology (Q. E. Coligan, A. M. Kruisbeek, D. H. Margulies, E. M. Shevach and W. Strober, eds., 1991); Annual Review of Immunology; and academic journal monographs such as Advances in Immunology.

[0113] B. Definition Before describing the present disclosure in more detail, it may be helpful to an understanding thereof to provide definitions of certain terms to be used herein.

[0114] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.Although any method and material similar or equivalent to those described herein can be used in the implementation or testing of specific embodiments, the preferred embodiments of compositions, methods and materials are described herein.For the purpose of this disclosure, the following terms are defined below.

[0115] The articles "a," "an," and "the" are used herein to refer to one or to more than one (i.e., to at least one or to one or more) of the grammatical object of the article. By way of example, "an" means one element or one or more elements.

[0116] The use of the alternative (eg, "or") should be understood to mean either one, both, or any combination thereof of the alternatives.

[0117] The term "and / or" should be understood to mean either one or both of the alternatives.

[0118] As used herein, the term "about" or "approximately" refers to an amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that varies by as much as 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1% relative to a reference amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length. In one embodiment, the term "about" or "approximately" refers to an amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length range of ±15%, ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, or ±1% of the reference amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length.

[0119] In one embodiment, ranges, for example, 1 to 5, about 1 to 5, or about 1 to about 5, refer to each of the numbers encompassed by the range. For example, in one non-limiting and merely exemplary embodiment, the range "1 to 5" is equivalent to the expressions 1, 2, 3, 4, 5, or 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, or 5.0, or 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.0.

[0120] As used herein, the term "substantially" refers to an amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that is 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more compared to a reference amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length. In one embodiment, "substantially the same" refers to an amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that produces about the same effect, e.g., a physiological effect, as the reference amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length.

[0121] Throughout this specification, unless the context requires otherwise, the words "comprise", "comprises" and "comprising" will be understood to imply the inclusion of a specified step or component or group of steps or components, but not the exclusion of any other step or component or group of steps or components. "Consisting of" means including and limited to everything that follows the phrase "consisting of". Thus, the phrase "consisting of" indicates that the recited components are required or essential, and that no other components may be present. "Consisting essentially of" means including any components recited after the phrase, and limited to other components that do not interfere with or participate in the activity or action specified in this disclosure for the recited components. Thus, the phrase "consisting essentially of" indicates that the recited components are required or essential, but that there are no other components present that would materially affect the activity or action of the recited components.

[0122] Throughout this specification, the use of "one embodiment," "an embodiment," "a particular embodiment," "a related embodiment," "a particular embodiment," "an additional embodiment," or "a further embodiment," or combinations thereof, means that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of such phrases in various places throughout this specification do not necessarily all refer to the same embodiment. Moreover, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. It is also understood that the forward recitation of a feature in an embodiment serves as a basis for excluding the feature in a particular embodiment.

[0123] The term "in vitro" generally refers to activities that occur at a location outside of an organism, such as experiments or measurements performed in an artificial environment outside of the organism, preferably with minimal changes to natural conditions, or experiments or measurements on living tissue. In certain embodiments, "in vitro" procedures include living cells or tissues taken from an organism and cultured or conditioned in a laboratory setup, usually under sterile conditions, typically for a few hours or up to about 24 hours, but up to 48 or 72 hours depending on the circumstances. In certain embodiments, such tissues or cells can be collected, frozen, and then thawed for in vitro processing. Tissue culture experiments or procedures that use living cells or tissues and last for several days or more are typically considered to be "in vitro", although in certain embodiments, the term can be used interchangeably with in vitro.

[0124] The term "in vivo" generally refers to activities that occur inside an organism. In one embodiment, the genome of a cell is manipulated, edited, or modified in vivo.

[0125] "Enhancement" or "promotion" or "increase" or "magnification" or "enhancement" generally refers to the ability of a nuclease variant to produce, induce or cause a greater response (i.e., a physiological response) compared to the response caused by either a vehicle or a control. A measurable response can include an increase in stability, e.g., the thermal stability, catalytic activity, and / or binding affinity of a homing endonuclease variant to the parent homing endonuclease from which the variant is derived. An "increased" or "enhanced" amount is typically a "statistically significant" amount and can include an increase that is 1.1, 1.2, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30-fold or more (e.g., 500, 1000-fold) (including all integers and decimal points above 1 in between, e.g., 1.5, 1.6, 1.7, 1.8, etc.) of the response caused by a vehicle or control.

[0126] "Reduction" or "decreasing" or "reducing" or "reducing" or "attenuating" or "excision" or "inhibition" or "attenuating" generally refers to the ability of a nuclease variant intended herein to produce, induce or cause a lesser response (i.e., a physiological response) compared to the response caused by either a vehicle or a control. A measurable response may include a reduction in the solubility, off-target binding affinity, or off-target cleavage specificity of a homing endonuclease variant compared to the parent homing endonuclease from which the variant is derived. The amount of "reduction" or "reduced" is typically a "statistically significant" amount and can include a reduction that is 1.1, 1.2, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30 or more fold (e.g., 500, 1000 fold) (all integers and decimals in between, 1 or greater, e.g., 1.5, 1.6, 1.7, 1.8, etc.) than the response produced by vehicle or control (reference response).

[0127] "Maintain" or "preserve" or "maintain", or "no change" or "no substantial change" or "no substantial decrease" generally refers to the ability of a nuclease variant to produce, induce, or cause a substantially similar or equivalent physiological response (i.e., a downstream effect) compared to the response caused by either a vehicle or a control. An equivalent response is not significantly or measurably different from the response of the reference.

[0128] As used herein, the terms "specific binding affinity" or "specifically binds" or "specifically bound" or "specific binding" or "specifically targeted" describe the binding of one molecule to another molecule, e.g., a DNA binding domain of a polypeptide that binds to DNA with a higher binding affinity than background binding. A binding domain may, for example, have a binding affinity of about 10 5 M -1 or higher affinity or K a(i.e., the equilibrium association constant for a particular binding interaction in units of 1 / M). In certain embodiments, a binding domain "specifically binds" to a target site if it binds to or associates with the target site with a 6 M -1 , 10 7 M -1 , 10 8 M -1 , 10 9 M -1 , 10 10 M -1 , 10 11 M -1 , 10 12 M -1 , or 10 13 M -1 More than K a A "high affinity" binding domain binds to a target site with at least 10 7 M -1 , at least 10 8 M -1 , at least 10 9 M -1 , at least 10 10 M -1 , at least 10 11 M -1 , at least 10 12 M -1 , at least 10 13 M -1 , or more K a The term refers to those binding domains that have

[0129] Alternatively, the affinity may be expressed in M ​​units (e.g., 10 -5 M~10 -13 The equilibrium dissociation constant (K d The affinity of a nuclease variant comprising one or more DNA binding domains for a DNA target site as contemplated in certain embodiments can be readily determined using conventional techniques, such as yeast cell surface display, or by binding association or displacement assays using labeled ligands.

[0130] In one embodiment, the affinity of specific binding is about 2-fold higher than background binding, about 5-fold higher than background binding, about 10-fold higher than background binding, about 20-fold higher than background binding, about 50-fold higher than background binding, about 100-fold higher than background binding, or about 1000-fold higher than background binding, or more.

[0131] The terms "selectively bind" or "selectively bound" or "selective binding" or "selectively target" describe preferential binding of one molecule to a target molecule (binding on the target) in the presence of multiple off-target molecules. In certain embodiments, the HE or megaTAL selectively binds to a DNA binding site on the target about 5-fold, 10-fold, 15-fold, 20-fold, 25-fold, 50-fold, 100-fold, or 1000-fold more frequently than the HE or megaTAL binds to an off-target DNA target binding site.

[0132] "Site on target" refers to the target site sequence.

[0133] "Off-target" refers to a sequence that is similar, but not identical, to the target site sequence.

[0134] A "target site" or "target sequence" is a chromosomal or extrachromosomal nucleic acid sequence that defines a portion of a nucleic acid to which a binding molecule binds and / or cleaves, provided that sufficient conditions for binding and / or cleavage exist. When referring to a polynucleotide sequence or SEQ ID NO: that references only one strand of a target site or target sequence, it is understood that the target site or target sequence bound and / or cleaved by a nuclease variant is double-stranded and includes the reference sequence and its complementary strand. In a preferred embodiment, the target site is a sequence of the human PDCD-1 gene.

[0135] "Protein stability" refers to the net balance of forces that determine whether a protein is in its native folded structure or in a denatured (unfolded or extended) state. Protein unfolding, either partial or complete, can result in loss of function along with degradation by cellular machinery. Polypeptide stability can be measured in response to a variety of conditions, including, but not limited to, temperature, pressure, and osmolality.

[0136] "Thermal stability" refers to the ability of a protein to fold properly in a native folded structure or resist denaturation or unfolding upon exposure to temperature fluctuations. At non-ideal temperatures, proteins may not fold efficiently into active structures or may have a tendency to unfold from active structures. Proteins with increased thermal stability fold properly and retain activity over an increased temperature range compared to proteins with low thermal stability.

[0137] "TM 50 " refers to the temperature at which 50% of the amount of protein is unfolded. In certain embodiments, TM 50 is the temperature at which that amount of protein has 50% of its maximum activity. 50 is a specific value determined by fitting multiple data points to a Boltsman sigmoid curve. In one non-limiting example, the TM 50 is measured in a yeast surface display activity assay by expressing the protein on the yeast surface at approximately 25°C, distributing the yeast into multiple wells, exposing them to a range of higher temperatures, cooling the yeast, and then measuring the cleavage activity of the enzyme. As the temperature is increased, more of the protein loses its activity confirmation and therefore fewer protein expressing cells show sufficient activity to measure cleavage by flow cytometry. The temperature at which 50% of the yeast display population is active compared to the non-heat shocked population is known as the TM. 50 It is.

[0138] "Recombination" refers to the process of exchange of genetic information between two polynucleotides, including, but not limited to, non-homologous end joining (NHEJ) and donor capture by homologous recombination. For purposes of this disclosure, "homologous recombination (HR)" refers to a specific form of such exchange that occurs, for example, during the repair of double-strand breaks in cells via the homology-directed repair (HDR) mechanism. This process is known variously as "non-crossover gene conversion" or "short tract gene conversion" because it requires nucleotide sequence homology and uses the "donor" molecule as a template to repair the "target" molecule (i.e., the one that experienced the double-strand break), leading to the transfer of genetic information from the donor to the target. Without intending to be bound by any particular theory, such transfer may involve mismatch correction of the heteroduplex DNA formed between the destroyed target and the donor, and / or "synthesis-dependent strand annealing," in which the donor is used to resynthesize the genetic information that will become part of the target, and / or related processes. Such specified HR often results in an alteration of the sequence of the target molecule such that some or all of the sequence of the donor polynucleotide is incorporated into the target polynucleotide.

[0139] "NHEJ" or "non-homologous end joining" refers to the resolution of double-stranded breaks in the absence of a donor repair template or homologous sequence. NHEJ can result in insertions and deletions at the break site. NHEJ is mediated by several sub-pathways, each of which has distinct mutational consequences. The classical NHEJ pathway (cNHEJ) requires the KU / DNA-PKcs / Lig4 / XRCC4 complex, which finally religates with minimal processing, often leading to accurate repair of the break. Alternative NHEJ pathways (altNHEJ) are also effective in resolving dsDNA breaks, but these pathways are highly mutagenic and result in imprecise repair of breaks marked by insertions and deletions. Without intending to be bound by any particular theory, it is contemplated that modification of dsDNA breaks by end-processing enzymes such as exonucleases, e.g., Trex2, can increase the likelihood of imprecise repair.

[0140] "Cleavage" refers to the destruction of the covalent backbone of a DNA molecule. Cleavage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Both single-strand and double-strand cleavage are possible. Double-strand cleavage can result from two separate single-strand cleavage events. DNA cleavage can result in the generation of either blunt ends or staggered ends. In certain embodiments, the polypeptides and nuclease variants contemplated herein, such as homing endonuclease variants, megaTALs, etc., are used for targeted double-strand DNA cleavage. The endonuclease cleavage recognition site can be on either DNA strand.

[0141] A "foreign" molecule is a molecule that is not normally present in a cell, but is introduced into a cell by one or more genetic, biochemical or other methods. Exemplary foreign molecules include, but are not limited to, small organic molecules, proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, polysaccharides, any modified derivatives of the above molecules, or any complex containing one or more of the above molecules. Methods for introducing foreign molecules into cells are known to those skilled in the art and include, but are not limited to, lipid-mediated transfer (i.e., liposomes containing neutral and cationic lipids), electroporation, direct injection, cell fusion, biolistics, biopolymer nanoparticles, calcium phosphate co-precipitation, DEAE-dextran-mediated transfer, and viral vector-mediated transfer.

[0142] An "endogenous" molecule is one that is normally present in a particular cell at a particular developmental stage under particular environmental conditions. Additional endogenous molecules can include proteins.

[0143] "Gene" refers to a DNA region that codes for a gene product, as well as all DNA regions that regulate the production of the gene product, whether or not such regulatory sequences are adjacent to the coding and / or transcribed sequence. Genes include, but are not limited to, promoter sequences, enhancers, silencers, insulators, border regions, terminators, polyadenylation sequences, post-transcriptional response elements, translation control sequences such as ribosome binding sites and internal ribosome entry sites, origins of replication, substrate binding sites, and locus control regions.

[0144] "Gene expression" refers to the conversion of information contained within a gene into a gene product. A gene product can be a direct transcription product of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, structural RNA, or any other type of RNA) or a protein produced by translation of an mRNA. Gene products also include RNA that is modified by processes such as capping, polyadenylation, methylation, and editing, as well as proteins that are modified, for example, by methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristylation, and glycosylation.

[0145] As used herein, the term "genetically engineered" or "genetically modified" refers to the chromosomal or extrachromosomal addition of extra genetic material in the form of DNA or RNA to the total genetic material in a cell. The genetic modification can be targeted to a specific site in the genome of the cell or can be non-targeted. In one embodiment, the genetic modification is site-specific. In one embodiment, the genetic modification is not site-specific.

[0146] As used herein, the term "genome editing" refers to the replacement, deletion, and / or introduction of genetic material at a target site in the genome of a cell, which restores, corrects, destroys, and / or modifies the expression and / or function of a gene or gene product. In certain embodiments, genome editing contemplated includes introducing one or more nuclease variants into a cell to generate a DNA lesion at or near a target site in the genome of the cell, optionally in the presence of a donor repair template.

[0147] As used herein, the term "gene therapy" refers to the introduction of excess genetic material into the total genetic material in a cell to restore, correct, or modify the expression of a gene or gene product, or for the purpose of expressing a therapeutic polypeptide. In certain embodiments, the introduction of genetic material into the genome of a cell by genome editing to restore, correct, disrupt, or modify the expression of a gene or gene product, or for the purpose of expressing a therapeutic polypeptide, is considered gene therapy.

[0148] C. Nuclease variants Various engineered nucleases may lack sufficient stability to be used in clinical settings. The nuclease variants contemplated herein are modified to increase thermal stability and enzymatic activity, allowing clinical use of previously unstable enzymes. The nuclease variants are suitable for genome editing of target sites and include one or more DNA binding domains and one or more DNA cleavage domains (e.g., one or more endonuclease domains and / or exonuclease domains), and optionally one or more linkers as contemplated herein. The engineered nucleases include one or more amino acid substitutions that increase thermal stability and / or activity compared to a reference or parent nuclease. The terms "reprogrammed nuclease", "engineered nuclease", or "nuclease variant" are used interchangeably and refer to a nuclease that includes one or more DNA binding domains and one or more DNA cleavage domains, where the nuclease is designed to bind and cleave double-stranded DNA target sequences and is modified to increase the nuclease and / or activity.

[0149] "Reference nuclease" or "parent nuclease" refers to a wild-type nuclease, a naturally occurring nuclease, or a nuclease or variant that has been modified to increase basal activity, affinity, specificity, selectivity, and / or stability to generate a subsequent nuclease variant.

[0150] In certain embodiments, the nuclease variant binding comprises at least one amino acid substitution that increases the stability and / or activity of the variant compared to the parent nuclease. In certain embodiments, the nuclease variant binding comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten amino acid substitutions that increase the stability and / or activity of the variant compared to the parent nuclease.

[0151] Nuclease variants may be designed and / or engineered from naturally occurring nucleases or from existing nuclease variants. In preferred embodiments, the nuclease variants comprise increased thermostability and / or enzymatic activity compared to the parent nuclease variant. In certain embodiments, contemplated nuclease variants may further comprise one or more additional functional domains, such as 5' to 3' exonuclease, 5' to 3' alkaline exonuclease, 3' to 5' exonuclease (e.g., Trex2), 5' flap endonuclease, helicase, template-dependent DNA polymerase, or an endo-processing enzyme domain of an endo-processing enzyme exhibiting template-independent DNA polymerase activity.

[0152] Examples of nuclease variants that are reprogrammed to bind and cleave target sequences and are designed to increase thermostability include, but are not limited to, homing endonuclease (meganuclease) variants and megaTALs. In certain embodiments, the nuclease variants are reprogrammed to bind to target sites or sequences of genes selected from the group consisting of: HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS.

[0153] 1. Homing endonuclease (meganuclease) variants Homing endonucleases (meganucleases) are genome editing enzymes that can be reprogrammed to bind and cleave selected target sites. However, some reprogrammed homing endonucleases are not stable enough to allow further development or clinical use. The inventors have unexpectedly discovered that certain amino acid positions of homing endonucleases, for example, affect the stability of the enzyme, and furthermore, substitution of amino acids at these positions can enhance the stability of the enzyme compared to the parent enzyme without sacrificing the affinity or activity of the enzyme.

[0154] In various embodiments, the homing endonuclease or meganuclease is reprogrammed to introduce a double-stranded break (DSB) at the target site and engineered to increase its thermostability, affinity, specificity, selectivity, and / or enzymatic activity. In a preferred embodiment, the homing endonuclease is reprogrammed to bind and cleave the target site and engineered to increase the thermostability of the enzyme relative to the thermostability of the enzyme for which it was designed.

[0155] "Homing endonucleases" and "meganucleases" are used interchangeably and refer to naturally occurring homing endonucleases that recognize cleavage sites of 12 to 45 base pairs and are commonly grouped into five families based on sequence and structural motifs: LAGLIDADG, GIY-YIG, HNH, His-Cys box, and PD-(D / E)XK.

[0156] "Reference homing endonuclease," "reference meganuclease," "parent homing endonuclease," or "parent meganuclease" refers to a wild-type homing endonuclease, a homing endonuclease occurring in nature, or a homing endonuclease that has been modified to increase basal activity, affinity, and / or stability to generate a subsequent homing endonuclease variant.

[0157] "Engineered homing endonuclease," "reprogrammed homing endonuclease," "homing endonuclease variant," "engineered meganuclease," "reprogrammed meganuclease," or "meganuclease variant" refers to a homing endonuclease that includes one or more DNA binding domains and one or more DNA cleavage domains, where the homing endonuclease has been designed and / or modified from a parent or naturally occurring homing endonuclease to bind and cleave a DNA target sequence, and optionally modified to improve one or more times affinity, selectivity, specificity and / or activity, and / or increase thermostability. A homing endonuclease variant may be designed and / or modified from a naturally occurring homing endonuclease or from another homing endonuclease variant. Homing endonuclease variants contemplated in certain embodiments may further comprise one or more additional functional domains, such as an endo-processing enzyme domain of an endo-processing enzyme that exhibits 5' to 3' exonuclease, 5' to 3' alkaline exonuclease, 3' to 5' exonuclease (e.g., Trex2), 5' flap endonuclease, helicase, template-dependent DNA polymerase or template-independent DNA polymerase activity.

[0158] Homing endonuclease variants do not exist in nature and can be obtained by recombinant DNA technology or by random mutagenesis. Homing endonuclease variants may be obtained by making one or more amino acid changes in naturally occurring homing endonucleases or homing endonuclease variants, for example, by mutating, substituting, adding, or deleting one or more amino acids. In certain embodiments, homing endonuclease variants include one or more amino acid changes to the DNA recognition interface for binding and cleaving selected target sequences, and one or more amino acid substitutions for increasing thermal stability.

[0159] Homing endonuclease variants contemplated in certain embodiments may further comprise one or more additional functional domains, for example, an endo-processing enzyme domain of an endo-processing enzyme that exhibits 5' to 3' exonuclease, 5' to 3' alkaline exonuclease, 3' to 5' exonuclease (e.g., Trex2), 5' flap endonuclease, helicase, template-dependent DNA polymerase or template-independent DNA polymerase activity. In certain embodiments, the homing endonuclease variant is introduced into a cell that has an endo-processing enzyme that exhibits 5' to 3' exonuclease, 5' to 3' alkaline exonuclease, 3' to 5' exonuclease (e.g., Trex2), 5' flap endonuclease, helicase, template-dependent DNA polymerase or template-independent DNA polymerase activity. The homing endonuclease variant and the 3' processing enzyme may be introduced separately, for example on different vectors or separate mRNAs, or together, for example together as a fusion protein, or in a polycistronic construct separated by a viral self-cleaving peptide or an IRES element.

[0160] "DNA recognition interface" refers to the homing endonuclease amino acid residues that interact with the nucleic acid target base as well as adjacent residues. For each homing endonuclease, the DNA recognition interface contains an extensive network of side chain-to-side chain and side chain-to-DNA contacts, most of which are necessarily unique to recognizing a particular nucleic acid target sequence. Thus, the amino acid sequence of the DNA recognition interface corresponding to a particular nucleic acid sequence varies widely and is a property of any natural or homing endonuclease variant. As a non-limiting example, homing endonuclease variants contemplated in certain embodiments can be derived by constructing a library of homing endonuclease variants in which one or more amino acid residues localized in the DNA recognition interface of a natural homing endonuclease (or a previously generated HE variant) are altered. The library can be screened for target cleavage activity against each target site using a cleavage assay (see, e.g., Jarjour et al., 2009, Nuc. Acids Res. 37(20):6871-6880).

[0161] LAGLIDADG homing endonucleases (LHEs) are the best studied family of homing endonucleases, are primarily encoded in archaea and in organelle DNA in green algae and fungi, and exhibit the highest overall DNA recognition specificity.

[0162] In one embodiment, the reprogrammed LHE or LHE variant engineered for enhanced thermostability is an I-OnuI HE variant (I-OnuI LHE variant), see, e.g., SEQ ID NOs:8-14 and 16-18.

[0163] In one embodiment, the reprogrammed I-OnuI HE or I-OnuI HE variant engineered to increase thermostability is generated from a naturally occurring I-OnuI, an I-OnuI HE variant, or a biologically active fragment thereof (e.g., SEQ ID NOs: 1-8 and 15). In a preferred embodiment, the reprogrammed I-OnuI HE or I-OnuI HE variant engineered to increase thermostability is generated from an existing I-OnuI HE variant. In an even more preferred embodiment, the reprogrammed I-OnuI HE or I-OnuI HE variant engineered to increase thermostability contains one, two, three, four, five, six, seven, eight, nine, or ten or more amino acid substitutions to increase the thermostability of the enzyme compared to the thermostability of the existing parent I-OnuI HE variant.

[0164] In certain embodiments, the I-OnuI HE variants comprise amino acid substitutions at at least one, at least two, at least three, at least four, at least five, at least six, or seven of the following amino acid positions that, independently and collectively, have been identified to increase homing endonuclease thermostability: I14, A19, V116, F168, D208, N246, and L263 of the representative I-OnuI amino acid sequence (SEQ ID NOs: 1-8 and 15), biologically active fragments thereof, and / or further variants thereof.

[0165] In certain embodiments, the I-OnuI HE variant comprises amino acid substitutions at the following amino acid positions that were identified as individually and collectively increasing homing endonuclease thermostability: I14, A19, V116, F168, D208, and N246 of the representative I-OnuI amino acid sequence (SEQ ID NOs: 1-8 and 15), biologically active fragments thereof, and / or further variants thereof.

[0166] In certain embodiments, the amino acid substitution at I14 is selected from the group consisting of: I14S, I14N, I14M, I14K, I14F, I14D, I14T, and I14V. In preferred embodiments, the amino acid substitution at I14 is I14T or I14V. In certain embodiments, the amino acid substitution at A19 is selected from the group consisting of: A19C, A19D, A19I, A19L, A19S, A19T, and A19V. In preferred embodiments, the amino acid substitution at A19 is A19T or A19V. In certain embodiments, the amino acid substitution at V116 is selected from the group consisting of: V116F, V116D, V116A, V116L, and V116I. In preferred embodiments, the amino acid substitution at V116 is V116L or V116I. In certain embodiments, the amino acid substitution at F168 is selected from the group consisting of: F168H, F168Y, F168I, F168V, F168P, F168L, and F168S. In preferred embodiments, the amino acid substitution at F168 is F168L and F168S. In certain embodiments, the amino acid substitution at D208 is selected from the group consisting of: D208N, D208Y, D208V, and D208E. In preferred embodiments, the amino acid substitution at D208 is D208E. In preferred embodiments, the amino acid substitution at F168 is F168L and F168S. In certain embodiments, the amino acid substitution at N246 is selected from the group consisting of: N246H, N246I, N246D, N246R, N246S, N246T, N246V, N246Y, and N246K. In preferred embodiments, the amino acid substitution at N246 is N246K. In certain embodiments, the amino acid substitution at L263 is selected from the group consisting of: L263H, L263F, L263P, L263T, L263V, and L263R. In preferred embodiments, the amino acid substitution at L263 is L263R.

[0167] In certain embodiments, the I-OnuI HE variants comprise amino acid substitutions at at least one, at least two, at least three, at least four, at least five, at least six, or seven of the following amino acid positions which, independently and collectively, have been identified to increase homing endonuclease thermostability: K108, K156, S176, E231, V261, E277, and G300 of the representative I-OnuI amino acid sequence (SEQ ID NOs: 1-8 and 15), biologically active fragments thereof, and / or further variants thereof.

[0168] In certain embodiments, the amino acid substitution at K108 is selected from the group consisting of K108E, K108N, K108Q, K108R, K108T, K108V, and K108M. In a preferred embodiment, the amino acid substitution at K108 is K108M. In certain embodiments, the amino acid substitution at K156 is selected from the group consisting of K156N, K156Q, K156R, K156T, K156V, K156I, and K156E. In a preferred embodiment, the amino acid substitution at K156 is K156I or K156E. In certain embodiments, the amino acid substitution at S176 is S176P, S176N, or S176A. In a preferred embodiment, the amino acid substitution at S176 is S176A. In certain embodiments, the amino acid substitution at E231 is selected from the group consisting of: E231D, E231V, E231K, and E231G. In preferred embodiments, the amino acid substitution at E231 is E231K or E231G. In certain embodiments, the amino acid substitution at V261 is selected from the group consisting of: V261D, V261G, V261I, V261L, V261S, V261T, and V261A. In preferred embodiments, the amino acid substitution at V261 is V261A. In certain embodiments, the amino acid substitution at E277 is selected from the group consisting of: E277A, E277D, E277G, E277Q, E277V, and E277K. In preferred embodiments, the amino acid substitution at E277 is E277K. In certain embodiments, the amino acid substitution at G300 is selected from the group consisting of: G300S, G300V, G300D, G300C, and G300R. In a preferred embodiment, the amino acid substitution at G300 is G300R.

[0169] Without intending to be bound by any particular theory, the inventors have also discovered that a homing endonuclease reprogrammed to bind and cleave a specific target sequence may have one or more additional amino acid positions that affect thermostability. In certain embodiments, the I-OnuI HE variants contain at least one, at least two, at least three, at least four, at least five, at least six, or seven, at least eight, at least nine, or ten of the following amino acid substitutions, which have been identified, independently and collectively, to increase homing endonuclease thermostability: N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268 of the representative I-OnuI amino acid sequence (SEQ ID NOs: 1-8 and 15), biologically active fragments thereof, and / or further variants thereof.

[0170] In certain embodiments, the amino acid substitution at N31 is selected from the group consisting of: N31D, N31H, N31I, N31R, N31K, N31S, N31T and N31Y. In certain embodiments, the amino acid substituted at N31 is N31K. In certain embodiments, the amino acid substitution at N33 is selected from the group consisting of: N33D, N33G, N33H, N33I, N33K, N33S, N33T and N33Y. In certain embodiments, the amino acid substituted at N33 is N33K. In certain embodiments, the amino acid substitution at K52 is selected from the group consisting of: K52Q, K52R, K52T, K52Y, K52N, K52E and K52M. In a preferred embodiment, the amino acid substitution at K52 is K52M. In certain embodiments, the amino acid substitution for Y97 is selected from the group consisting of: Y97H, Y97N, and Y97F. In certain embodiments, the amino acid substituted for Y97 is Y97F. In certain embodiments, the amino acid substitution for K124 is selected from the group consisting of: K124E, K124N, K124R, and K124T. In certain embodiments, the amino acid substituted for K124 is K124N. In certain embodiments, the amino acid substitution for K147 is selected from the group consisting of: K147E, K147I, K147N, K147R, and K147T. In certain embodiments, the amino acid substituted for K147 is K147I. In certain embodiments, the amino acid substitution at I153 is selected from the group consisting of: I153D, I153H, I153K, I153T, I153Y, I153S, I153V and I153N. In a preferred embodiment, the amino acid substitution at I153 is I153N. In certain embodiments, the amino acid substitution at K209 is selected from the group consisting of: K209E, K209M, K209N, K209Q and K209R. In certain embodiments, the amino acid substituted at K209 is K209R. In certain embodiments, the amino acid substitution at E264 is selected from the group consisting of: E264A, E264D, E264G, E264K, E264Q, E264R and E264V. In certain embodiments, the amino acid substituted at E264 is E264K.In certain embodiments, the amino acid substitution for D268 is selected from the group consisting of: D268A, D268E, D268G, D268H, D268N, D268V and D268Y. In certain embodiments, the amino acid substituted for D268 is D268N.

[0171] In certain embodiments, I-OnuI HE variants include amino acid substitutions at one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, or fourteen or more of the following amino acid positions that have been identified, individually and collectively, to increase the thermostability of a homing endonuclease: I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300 of the representative I-OnuI amino acid sequence (SEQ ID NOs: 1-8 and 15), biologically active fragments thereof, and / or further variants thereof.

[0172] In certain embodiments, the I-OnuI HE variant comprises amino acid substitutions at one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, or fourteen of the following amino acid positions, which have been individually and collectively identified to increase the thermostability of a homing endonuclease: I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, which have been identified individually and collectively to increase the thermostability of a homing endonuclease; Amino acid substitutions at one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or ten of the following amino acid positions, which in some cases have been identified collectively to increase the thermal stability of a homing endonuclease: N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268 of the representative I-OnuI amino acid sequence (SEQ ID NOs: 1-8 and 15), biologically active fragments thereof, and / or further variants thereof.

[0173] In certain embodiments, an I-OnuI HE variant that binds and cleaves a target sequence has one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten amino acid substitutions that increase thermostability and is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, or at least 98%, or at least 99% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 1-18, or comprises a biologically active fragment thereof.

[0174] In certain embodiments, the I-OnuI HE variant has increased thermostability compared to the parent I-OnuI LHE variant. In certain embodiments, the I-OnuI HE variant has increased thermostability compared to the parent I-OnuI LHE variant. 50 TM of about 5°C to about 35°C higher, about 10°C to about 35°C higher, about 10°C to about 30°C higher, about 10°C to about 25°C higher, about 15°C to about 35°C higher, about 15°C to about 30°C higher, or about 15°C to about 25°C higher. 50 has.

[0175] In certain embodiments, an I-OnuI HE variant has increased thermostability compared to a parent or reference I-OnuI LHE variant. In certain embodiments, an I-OnuI HE variant has an increased thermostability compared to a parent or reference I-OnuI LHE variant. 505°C, 6°C, 7°C, 8°C, 9°C, 10°C, 11°C, 12°C, 13°C, 14°C, 15°C, 16°C, 17°C, 18°C, 19°C, 20°C, 21°C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, 31°C, 32°C, 33°C, 34°C, 35°C, 36°C, 37°C, 38°C, 39°C, 40°C, 41°C, 42°C, 43°C, 44°C, 45°C, 46°C, 47°C, 48°C, 49°C, 50°C, 51°C, 52°C, 53°C, 54°C, 55°C, 56°C, 57°C, 58°C, 59°C, 60°C, 61°C, 62°C, 63°C, 64°C, 65°C, 66°C, 67°C, 68°C, 69°C, 70°C, 71°C, 72°C, 73°C, 74°C, 75°C, 76°C, 77°C, 78°C, 79°C, 80°C, 81°C, 82°C, 83°C, 84°C, 85°C, 86°C, 87°C, 88°C, 89°C, 90°C, 50 has.

[0176] In certain embodiments, an I-OnuI HE variant comprising one or more mutations that increase thermostability is reprogrammed to bind to a target site or sequence of a gene selected from the group consisting of: HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS.

[0177] 2. MegaTAL MegaTALs are genome editing enzymes that combine the DNA binding properties of TAL DNA binding domains with the DNA binding and cleavage activity of homing endonucleases. Without intending to be bound by any particular theory, it is believed that when a relatively unstable homing endonuclease is formatted as a megaTAL, the megaTAL is not inherently stabilized. Thus, introducing one or more stabilizing mutations into a homing endonuclease will similarly stabilize the corresponding megaTAL.

[0178] In various embodiments, the homing endonuclease or meganuclease comprises one or more TAL DNA binding domains and a homing endonuclease or meganuclease that is reprogrammed to introduce a double strand break (DSB) at a target site and engineered to increase its thermostability, affinity, specificity, selectivity, and / or enzymatic activity of the enzyme. In a preferred embodiment, the increase in thermostability of a megaTAL comprising a homing endonuclease engineered to increase the thermostability of the enzyme is relative to the thermostability of the megaTAL comprising the homing endonuclease prior to engineering to increase its thermostability.

[0179] "megaTAL" refers to a polypeptide comprising a TALE DNA binding domain and a homing endonuclease variant engineered to bind to and cleave a DNA target sequence and have increased thermal stability, and optionally comprising one or more linkers and / or additional functional domains, e.g., a 5' to 3' exonuclease, a 5' to 3' alkaline exonuclease, a 3' to 5' exonuclease (e.g., Trex2), a 5' flap endonuclease, a helicase or an endo-processing enzyme domain that exhibits template-independent DNA polymerase activity.

[0180] A "reference megaTAL" or "parent megaTAL" refers to a megaTAL that comprises a TALE DNA binding domain and a wild-type homing endonuclease, a naturally occurring homing endonuclease, or a homing endonuclease that has been modified to increase basal activity, affinity, and / or stability to generate a subsequent homing endonuclease variant.

[0181] In certain embodiments, the megaTAL can be introduced into a cell together with an endo-processing enzyme that exhibits a 5' to 3' exonuclease, a 5' to 3' alkaline exonuclease, a 3' to 5' exonuclease (e.g., Trex2), a 5' flap endonuclease, a helicase, a template-dependent DNA polymerase, or a template-independent DNA polymerase activity. The megaTAL and the 3' processing enzyme can be introduced separately, e.g., on different vectors or separate mRNAs, or together, e.g., as fusion proteins, or in a polycistronic construct separated by a viral self-cleaving peptide or an IRES element.

[0182] A "TALE DNA-binding domain" is the DNA-binding portion of a transcription activator-like effector (TALE or TAL effector) that mimics plant transcription activators and manipulates the plant transcriptome (see, e.g., Kay et al., 2007, Science 318:648-651). In certain embodiments, TALE DNA binding domains contemplated are engineered from novel or naturally occurring TALEs, such as AvrBs3 from Xanthomonas spot blight, Xanthomonas gardneri, Xanthomonas translucens, Xanthomonas axonopodis, Xanthomonas perforans, Xanthomonas alfalfa, Xanthomonas citri, Xanthomonas euvesicatoria, and Xanthomonas oryzae, and brg11 and hpx17 from Ralstonia solanacearum. Examples of TALE protein deriving and engineering DNA binding domains are disclosed in U.S. Patent No. 9,017,967, and the references cited therein, all of which are incorporated by reference in their entirety herein.

[0183] In certain embodiments, megaTALs comprise a TALE DNA binding domain that comprises one or more repeat units that are involved in the binding of the TALE DNA binding domain to its corresponding target DNA sequence. A single "repeat unit" (also called a "repeat") is typically 33-35 amino acids long. Each TALE DNA binding domain repeat unit typically contains one or two DNA binding residues that constitute a repeat variable dipeptide (RVD) at positions 12 and / or 13 of the repeat. The natural (canonical) code for DNA recognition of these TALE DNA binding domains has been determined such that the HD sequence at positions 12 and 13 directs binding to cytosine (C), NG binds to T, NI binds to A, NN binds to G or A, and NG binds to T. In certain embodiments, non-canonical (atypical) RVDs are contemplated.

[0184] Examples of non-canonical RVDs suitable for use in certain megaTALs contemplated in certain embodiments are: HH, KH, NH, NK, NQ, RH, RN, SS, NN, SN, KN for guanine (G) recognition; NI, KI, RI, HI, SI for adenine (A) recognition; NG, HG, KG, RG for thymine (T) recognition; RD, SD, HD, ND, KD, YG for cytosine (C) recognition; NV, HN for A or G recognition; and H for A or T or G or C recognition. * , H.A., K.A., N. * , N.A., N.C., N.S., R.A., S. * And ( * ) means that the amino acid at position 13 is absent. Further examples of RVDs suitable for use in certain megaTALs contemplated in certain embodiments further include those disclosed in U.S. Patent No. 8,614,092, which is incorporated herein by reference in its entirety.

[0185] In certain embodiments, megaTALs contemplated herein comprise a TALE DNA binding domain comprising 3-30 repeat units. In certain embodiments, megaTALs comprise 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 TALE DNA binding domain repeat units. In preferred embodiments, megaTALs contemplated herein comprise a TALE DNA binding domain comprising 5-15 repeat units, more preferably 7-15 repeat units, more preferably 9-15 repeat units, more preferably 9, 10, 11, 12, 13, 14, or 15 repeat units.

[0186] In certain embodiments, megaTALs contemplated herein comprise a TALE DNA binding domain comprising 3-30 repeat units and an additional single truncated TALE repeat unit comprising 20 amino acids located at the C-terminus of the set of TALE repeat units, i.e., an additional C-terminal half TALE DNA binding domain repeat unit (amino acids -20 to -1 of the C-cap disclosed above and elsewhere herein). Thus, in certain embodiments, megaTALs contemplated herein comprise a TALE DNA binding domain comprising 3.5-30.5 repeat units. In certain embodiments, the megaTAL comprises 3.5, 4.5, 5.5, 6.5, 7.5, 8.5, 9.5, 10.5, 11.5, 12.5, 13.5, 14.5, 15.5, 16.5, 17.5, 18.5, 19.5, 20.5, 21.5, 22.5, 23.5, 24.5, 25.5, 26.5, 27.5, 28.5, 29.5, or 30.5 TALE DNA binding domain repeat units. In a preferred embodiment, a megaTAL contemplated herein comprises a TALE DNA binding domain comprising 5.5 to 15.5 repeat units, more preferably 7.5 to 15.5 repeat units, more preferably 9.5 to 15.5 repeat units, more preferably 9.5, 10.5, 11.5, 12.5, 13.5, 14.5, or 15.5 repeat units.

[0187] In certain embodiments, a megaTAL comprises a TAL effector construct comprising an "N-terminal domain (NTD)" polypeptide, one or more TALE repeat domains / units, a "C-terminal domain (CTD)" polypeptide, and a homing endonuclease variant. In some embodiments, the NTD, TALE repeat, and / or CTD domains are from the same species. In other embodiments, one or more of the NTD, TALE repeat, and / or CTD domains are from different species.

[0188] As used herein, the term "N-terminal domain (NTD)" polypeptide refers to a sequence adjacent to the N-terminal portion or fragment of a naturally occurring TALE DNA binding domain. The NTD sequence, if present, may be of any length so long as the TALE DNA binding domain repeat unit retains the ability to bind DNA. In certain embodiments, the NTD polypeptide comprises at least 120 to at least 140 or more amino acids N-terminal to the TALE DNA binding domain (where 0 is amino acid 1 of the most N-terminal repeat unit). In certain embodiments, the NTD polypeptide comprises at least about 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, or at least 140 amino acids N-terminal to the TALE DNA binding domain. In one embodiment, a megaTAL contemplated herein comprises an NTD polypeptide from at least about amino acids +1 to +122 to at least about +1 to +137 of a Xanthomonas TALE protein (where 0 is amino acid 1 of the most N-terminal repeat unit). In certain embodiments, an NTD polypeptide comprises at least about 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, or 137 amino acids N-terminal to the TALE DNA binding domain of a Xanthomonas TALE protein. In one embodiment, a megaTAL contemplated herein comprises an NTD polypeptide from at least amino acids +1 to +121 of a Ralstonia TALE protein (where 0 is amino acid 1 of the most N-terminal repeat unit). In certain embodiments, the NTD polypeptide comprises at least about 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, or 137 amino acids N-terminal to the TALE DNA binding domain of a Ralstonia TALE protein.

[0189] As used herein, the term "C-terminal domain (CTD)" polypeptide refers to the sequence adjacent to the C-terminal portion or fragment of a naturally occurring TALE DNA binding domain. The CTD sequence, if present, may be of any length so long as the TALE DNA binding domain repeat unit retains the ability to bind DNA. In certain embodiments, the CTD polypeptide comprises at least 20 to at least 85 or more amino acids C-terminal to the last complete repeat of the TALE DNA binding domain (the first 20 amino acids are the C-terminal half repeat unit to the last C-terminal complete repeat unit). In certain embodiments, the CTD polypeptide comprises at least about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 443, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, or at least 85 amino acids C-terminal to the last complete repeat of the TALE DNA binding domain. In one embodiment, a megaTAL contemplated herein comprises a CTD polypeptide from at least about amino acids -20 to -1 of a Xanthomonas TALE protein (-20 being amino acid 1 of the half-repeat unit at the C-terminus of the C-terminal full repeat unit). In a particular embodiment, a CTD polypeptide comprises at least about 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acid at the C-terminus of the last full repeat of the TALE DNA binding domain of a Xanthomonas TALE protein. In one embodiment, a megaTAL contemplated herein comprises a CTD polypeptide from at least about amino acids -20 to -1 of a Ralstonia TALE protein (-20 being amino acid 1 of the half-repeat unit at the C-terminus of the last full repeat unit).In certain embodiments, the CTD polypeptide comprises at least about 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acids C-terminal to the last complete repeat of the TALE DNA binding domain of a Ralstonia TALE protein.

[0190] In certain embodiments, the megaTALs contemplated herein comprise a fusion polypeptide comprising a TALE DNA binding domain engineered to bind to a target sequence, a homing endonuclease reprogrammed to bind to and cleave the target sequence and engineered to increase enzyme stability and / or activity, and optionally an NTD and / or CTD polypeptide linked together with one or more linker polypeptides as otherwise contemplated herein. Without intending to be bound by any particular theory, it is contemplated that the megaTALs comprising the TALE DNA binding domain and, optionally, the NTD and / or CTD polypeptides are fused to a linker polypeptide that is further fused to a homing endonuclease variant. Thus, the TALE DNA binding domain binds to a DNA target sequence that is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides away from the target sequence bound by the DNA binding domain of the homing endonuclease variant. Thus, the megaTALs contemplated herein increase the specificity and efficiency of genome editing.

[0191] In one embodiment, the megaTAL comprises a TALE DNA binding domain that binds to about 2, 3, 4, 5, or 6 nucleotide sequences upstream of the binding site of the homing endonuclease variant and the reprogrammed homing endonuclease.

[0192] In certain embodiments, megaTALs contemplated herein comprise one or more TALE DNA-binding repeat units and an I-OnuI HE variant comprising increased thermal stability and / or enzymatic activity compared to the parent I-OnuI HE variant.

[0193] In certain embodiments, megaTALs contemplated herein comprise an NTD, one or more TALE DNA-binding repeat units, a CTD, and an I-OnuI HE variant that has increased thermal stability and / or enzymatic activity compared to the parent I-OnuI HE variant.

[0194] In certain embodiments, megaTALs contemplated herein comprise an NTD, about 9.5 to about 15.5 TALE DNA-binding repeat units, and an I-OnuI HE variant with increased thermal stability and / or enzymatic activity compared to the parent I-OnuI HE variant.

[0195] In certain embodiments, megaTALs contemplated herein comprise an NTD of about 122 amino acids to 137 amino acids, about 9.5, about 10.5, about 11.5, about 12.5, about 13.5, about 14.5, or about 15.5 binding repeat units, a CTD of about 20 amino acids to about 85 amino acids, and an I-OnuI HE variant comprising increased thermal stability and / or enzymatic activity compared to the parent I-OnuI HE variant. In certain embodiments, any one, two, or all of the NTD, DNA binding domain, and CTD can be designed from the same or different species, in any suitable combination.

[0196] In certain embodiments, a megaTAL comprising an I-OnuI HE variant with one or more mutations that increase thermostability is reprogrammed to bind to a target site or sequence of a gene selected from the group consisting of: HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS.

[0197] 3. Endoprocessing enzymes Genome editing compositions and methods contemplated in certain embodiments include editing a cell genome using one or more copies of an I-OnuI HE variant and an end-processing enzyme that comprises increased thermostability and / or enzymatic activity compared to the parent I-OnuI HE variant. In certain embodiments, a single polynucleotide encodes a homing endonuclease variant and an end-processing enzyme separated by a linker, a self-cleaving peptide sequence, e.g., a 2A sequence, or an IRES sequence. In certain embodiments, the genome editing composition comprises a polynucleotide encoding a nuclease variant and a separate polynucleotide encoding an end-processing enzyme. In certain embodiments, the genome editing composition comprises a polynucleotide encoding a single polypeptide fusion of a homing endonuclease variant end-processing enzyme in addition to tandem copies of the end-processing enzyme separated by a self-cleaving peptide.

[0198] The term "endo-processing enzyme" refers to an enzyme that modifies exposed ends of polynucleotide chains. Polynucleotides may be double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), RNA, double-stranded hybrids of DNA and RNA, and synthetic DNA (e.g., containing bases other than A, C, G, and T). Endo-processing enzymes may modify exposed polynucleotide chain ends by adding one or more nucleotides, removing one or more nucleotides, removing or modifying phosphate groups, and / or removing or modifying hydroxyl groups. Endo-processing enzymes may modify ends at endonuclease cleavage sites or at ends generated by shearing (e.g., by passing through a fine gauge needle, heating, sonication, mini-bead tumbling, and spraying), ionizing radiation, ultraviolet radiation, oxygen radicals, chemical hydrolysis, and other chemical or mechanical means such as chemotherapeutic agents.

[0199] In certain embodiments, genome editing compositions and methods contemplated in certain embodiments include editing a cellular genome using I-OnuI HE variants and I-OnuI HE variants that include increased thermal stability and / or enzymatic activity compared to the parent I-OnuI HE variant or megaTALs and DNA end-processing enzymes.

[0200] The term "DNA end-processing enzyme" refers to an enzyme that modifies exposed ends of DNA. DNA end-processing enzymes can modify blunt or staggered ends (ends with 5' or 3' overhangs). DNA end-processing enzymes can modify single-stranded or double-stranded DNA. DNA end-processing enzymes can modify ends at endonuclease cleavage sites or at ends generated by shearing (e.g., by passing through a fine gauge needle, heating, sonication, mini-bead tumbling, and spraying), ionizing radiation, ultraviolet radiation, oxygen radicals, chemical hydrolysis, and other chemical or mechanical means such as chemotherapeutic agents. DNA end-processing enzymes can modify exposed DNA ends by adding one or more nucleotides, removing one or more nucleotides, removing or modifying a phosphate group, and / or removing or modifying a hydroxyl group.

[0201] Examples of DNA end-processing enzymes suitable for use in certain embodiments contemplated herein include, but are not limited to, 5' to 3' exonucleases, 5' to 3' alkaline exonucleases, 3' to 5' exonucleases, 5' flap endonucleases, helicases, phosphatases, hydrolases and template-dependent DNA polymerases.

[0202] Additional examples of DNA end-processing enzymes suitable for use in certain embodiments contemplated herein include Trex2, Trex1, transmembrane domain-free Trex1, Apollo, Artemis, DNA2, Exo1, ExoT, ExoIII, Fen1, Fan1, MreII, Rad2, Rad9, TdT (terminal deoxynucleotidyl transferase), PNKP, RecE, RecJ, RecQ, Lambda exonuclease, Sox, vaccinia DNA polymerase, exonuclease I, exonuclease II, exonuclease III, exonuclease I, exonuclease III, exonuclease I, exonuclease II, exonuclease I, exonuclease III, exonuclease I ... Examples of suitable nucleases include, but are not limited to, nuclease III, exonuclease VII, NDK1, NDK5, NDK7, NDK8, WRN, T7-exonuclease gene 6, avian myeloblastosis virus integration protein (IN), Bloom, Antarctic phosphatase, alkaline phosphatase, polynucleotide kinase (PNK), ApeI, mung bean nuclease, Hex1, TTRAP (TDP2), Sgs1, Sae2, CUP, Pol Mu, Pol Lambda, MUS81, EME1, EME2, SLX1, SLX4, and UL-12.

[0203] In certain embodiments, the genome editing compositions and methods for editing a cell genome contemplated herein include a polypeptide comprising an I-OnuI HE variant or megaTAL and an exonuclease. The term "exonuclease" refers to an enzyme that cleaves a phosphodiester bond at the end of a polynucleotide chain via a hydrolysis reaction that cleaves the phosphodiester bond at either the 3' or 5' end.

[0204] Examples of exonucleases suitable for use in certain embodiments contemplated herein include, but are not limited to, hExoI, yeast ExoI, E. coli hTREX2, mouse TREX2, rat TREX2, hTREX1, mouse TREX1, rat TREX1, and rat TREX1.

[0205] In a particular embodiment, the DNA end-processing enzyme is a 3' to 5' exonuclease, preferably Trex1 or Trex2, more preferably Trex2, even more preferably human or mouse Trex2.

[0206] D. Polypeptides A variety of polypeptides are contemplated herein, including, but not limited to, homing endonuclease variants and megaTALs engineered to increase thermostability and / or enzymatic activity, as well as fusion polypeptides. In a preferred embodiment, the polypeptide comprises the amino acid sequence set forth in SEQ ID NOs: 9-14, 16-18, 22, and 23. "Polypeptide," "peptide," and "protein" are used interchangeably according to their conventional meaning, i.e., according to the sequence of amino acids, unless specified to the contrary. In one embodiment, "polypeptide" includes fusion polypeptides and other variants. Polypeptides can be prepared using any of a variety of well-known recombinant and / or synthetic techniques. Polypeptides are not limited to a particular length, for example, they may include full-length protein sequences, fragments of full-length proteins, or fusion proteins, and may include post-translational modifications of the polypeptide, such as glycosylation, acetylation, phosphorylation, and the like, as well as other modifications known in the art, both naturally occurring and non-naturally occurring.

[0207] As used herein, "isolated protein," "isolated peptide," or "isolated polypeptide," and the like, refer to the in vitro synthesis, isolation, and / or purification of a peptide or polypeptide molecule from the cellular environment and from association with other components of a cell, i.e., not significantly associated with substances in vivo. In certain embodiments, an isolated polypeptide is a synthetic polypeptide, a semi-synthetic polypeptide, or a polypeptide obtained or derived from a recombinant source.

[0208] Polypeptides include "polypeptide variants". Polypeptide variants may differ from naturally occurring polypeptides in one or more amino acid substitutions, deletions, additions and / or insertions. Such variants may be naturally occurring or synthetically generated, for example, by modifying one or more amino acids of the above polypeptide sequences. For example, in certain embodiments, it may be desirable to improve the biological properties of homing endonucleases, such as megaTALs, that bind and cleave target sites by introducing one or more substitutions, deletions, additions and / or insertions into the polypeptide. In certain embodiments, polypeptides include those having at least about 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% amino acid identity to a reference sequence, and typically the variant retains at least one biological activity of the reference sequence.

[0209] In a preferred embodiment, the polypeptide variant comprises a homing endonuclease or megaTAL engineered to increase its thermal stability and / or activity. The I-OnuI HE polypeptide or a fragment thereof can be reprogrammed to bind and cleave a target site. In certain embodiments, the reprogrammed I-OnuI HE variant has a relatively low thermal stability and / or activity compared to the parent I-OnuI HE. In a preferred embodiment, the I-OnuI homing endonuclease or a fragment thereof is engineered to bind and cleave a target site, increasing the thermal stability and / or activity of the enzyme.

[0210] Polypeptide variants include biologically active "polypeptide fragments." Examples of biologically active polypeptide fragments include DNA binding domains, nuclease domains, and the like. As used herein, the term "biologically active fragment" or "minimal biologically active fragment" refers to a polypeptide fragment that retains at least 100%, at least 90%, at least 80%, at least 70%, at least 60%, at least 50%, at least 40%, at least 30%, at least 20%, at least 10%, or at least 5% of the naturally occurring polypeptide activity. In preferred embodiments, the biological activity is binding affinity and / or cleavage activity for a target sequence. In certain embodiments, the polypeptide fragment may comprise an amino acid chain of at least 5 to about 1700 amino acids in length. In certain embodiments, the fragment comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60 , 65, 70, 75, 80, 85, 90, 95, 100, 110, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700 or more amino acids in length. In certain embodiments, the polypeptide comprises a biologically active fragment of a homing endonuclease variant. In certain embodiments, the polypeptide comprises a biologically active fragment of a homing endonuclease variant or a megaTAL. In certain embodiments, the polypeptides described herein may comprise one or more amino acids designated as "X". "X" refers to any amino acid when present in an amino acid sequence. One or more "X" residues may be present at the N-terminus and C-terminus of an amino acid sequence set forth in a particular SEQ ID NO contemplated herein.When the "X" amino acid is not present, the remaining amino acid sequence set forth in the SEQ ID NO: may be considered a biologically active fragment.

[0211] The biologically active fragment may include an N-terminal truncation and / or a C-terminal truncation. In certain embodiments, the biologically active fragment lacks or includes a deletion of 1, 2, 3, 4, 5, 6, 7, or 8 N-terminal amino acids of the homing endonuclease variant compared to the corresponding wild-type homing endonuclease sequence, and more preferably includes a deletion of 4 N-terminal amino acids of the homing endonuclease variant compared to the corresponding wild-type homing endonuclease sequence. In certain embodiments, the biologically active fragment lacks or includes a deletion of 1, 2, 3, 4, or 5 C-terminal amino acids of the homing endonuclease variant compared to the corresponding wild-type homing endonuclease sequence, and more preferably includes a deletion of 2 C-terminal amino acids of the homing endonuclease variant compared to the corresponding wild-type homing endonuclease sequence. In certain preferred embodiments, the biologically active fragment lacks or comprises a deletion of the four N-terminal amino acids and the two C-terminal amino acids of a homing endonuclease variant compared to the corresponding wild-type homing endonuclease sequence.

[0212] In certain embodiments, an I-OnuI variant comprises a deletion of 1, 2, 3, 4, 5, 6, 7, or 8 of the following N-terminal amino acids: M, A, Y, M, S, R, R, E; and / or a deletion of 1, 2, 3, 4, or 5 of the following C-terminal amino acids: R, G, S, F, V.

[0213] In certain embodiments, an I-OnuI variant comprises a deletion or substitution of 1, 2, 3, 4, 5, 6, 7, or 8 of the following N-terminal amino acids: M, A, Y, M, S, R, R, E; and / or a deletion or substitution of 1, 2, 3, 4, or 5 of the following C-terminal amino acids: R, G, S, F, V.

[0214] In certain embodiments, the I-OnuI variant comprises a deletion of 1, 2, 3, 4, 5, 6, 7, or 8 of the following N-terminal amino acids: M, A, Y, M, S, R, R, E; and / or a deletion of 1 or 2 of the following C-terminal amino acids: F, V.

[0215] In certain embodiments, an I-OnuI variant comprises a deletion or substitution of 1, 2, 3, 4, 5, 6, 7, or 8 of the following N-terminal amino acids: M, A, Y, M, S, R, R, E; and / or a deletion or substitution of 1 or 2 of the following C-terminal amino acids: F, V.

[0216] As mentioned above, polypeptides may be modified in various ways, including amino acid substitution, deletion, truncation, and insertion. Such manipulation methods are generally known in the art. For example, amino acid sequence variants of a reference polypeptide can be prepared by mutations in DNA. Methods of mutagenesis and nucleotide sequence changes are well known in the art. See, for example, Kunkel (1985, Proc. Natl. Acad. Sci. USA. 82:488-492), Kunkel et al. (1987, Methods in Enzymol, 154:367-382), U.S. Patent No. 4,873,192, Watson, JD et al. (Molecular Biology of the Gene, 4th ed., Benjamin / Cummings, Menlo Park, Calif., 1987) and references cited therein. Guidance regarding appropriate amino acid substitutions that do not affect the biological activity of the protein of interest can be found in the model of Dayhoff et al. (1978) Atlas of Protein Sequence and Structure (Natl. Biomed. Res. Found., Washington, DC).

[0217] In certain embodiments, the variant contains one or more conservative substitutions. A "conservative substitution" is one in which an amino acid is replaced with another amino acid having similar properties, such that a person skilled in the art of peptide chemistry would expect the secondary structure and hydropathic properties of the polypeptide to be substantially unchanged. Modifications may be made in the polynucleotide and polypeptide structures contemplated in certain embodiments, including polypeptides that have at least approximately the functional molecule encoding a variant or derivative polypeptide with desired characteristics, yet still obtain. When it is desired to modify the amino acid sequence of a polypeptide to generate an equivalent or improved variant polypeptide, a person skilled in the art can, for example, modify one or more of the codons of the coding DNA sequence, for example, according to Table 1. [Table 1]

[0218] Guidance for determining which amino acid residues may be substituted, inserted, or deleted in a particular embodiment without destroying biological activity can be found using computer programs well known in the art, such as DNASTAR, DNA Strider, Geneious, Mac Vector, or Vector NTI software. Conservative amino acid changes involve substituting one of a family of related amino acids in their side chains. Naturally occurring amino acids are generally divided into four families: acidic (aspartic acid, glutamic acid), basic (lysine, arginine, histidine), non-polar (alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan), and uncharged polar (glycine, asparagine, glutamine, cysteine, serine, threonine) amino acids. Phenylalanine, tryptophan, and tyrosine are sometimes classified together as aromatic amino acids. In a peptide or protein, suitable conservative substitutions of amino acids are known to those skilled in the art and can generally be made without altering the biological activity of the resulting molecule. Those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide generally do not substantially alter biological activity (see, for example, Watson et al., Molecular Biology of the Gene, 4th ed., 1987, The Benjamin / Cummings Pub.Co., p. 224).

[0219] In one embodiment, the I-OnuI variant comprises one or more non-conservative amino acid substitutions at positions that affect the thermostability of the enzyme. In one embodiment, the I-OnuI variant comprises one or more conservative and / or non-conservative amino acid substitutions at positions that affect the thermostability of the enzyme.

[0220] In certain embodiments, expression of more than one polypeptide is desired, and the polynucleotide sequences encoding them may be separated by an IRES sequence, as disclosed elsewhere herein.

[0221] In certain embodiments, the intended polypeptide comprises fusion polypeptide.In certain embodiments, fusion polypeptide and polynucleotide encoding fusion polypeptide are provided.Fusion polypeptide and fusion protein refer to a polypeptide having at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 polypeptide segments.

[0222] In another embodiment, two or more polypeptides may be expressed as a fusion protein containing one or more self-cleaving polypeptide sequences, as disclosed elsewhere herein.

[0223] In one embodiment, a fusion protein contemplated herein comprises one or more DNA binding domains and one or more nucleases, and one or more linker and / or self-cleaving polypeptides.

[0224] In one embodiment, a fusion protein contemplated herein comprises a nuclease variant; a linker or a self-cleaving peptide; and an end-processing enzyme, including, but not limited to, a 5' to 3' exonuclease, a 5' to 3' alkaline exonuclease, and a 3' to 5' exonuclease (e.g., Trex2).

[0225] Fusion polypeptides may include one or more polypeptide domains or segments, including, but not limited to, signal peptides, cell penetrating peptide domains (CPPs), DNA binding domains, nuclease domains, etc., epitope tags (e.g., moltose binding protein ("MBP"), glutathione S-transferase (GST), HIS6, MYC, FLAG, V5, VSV-G, and HA), polypeptide linkers, and polypeptide cleavage signals. Fusion polypeptides are typically linked C-terminus to N-terminus, but they may also be linked C-terminus to C-terminus, N-terminus to N-terminus, or N-terminus to C-terminus. In certain embodiments, the polypeptides of the fusion protein may be in any order. Fusion polypeptides or fusion proteins may also include conservatively modified variants, polymorphic variants, alleles, mutants, subsequences, and interspecies homologs, so long as the desired activity of the fusion polypeptide is retained. Fusion polypeptides may be produced by chemical synthesis methods or by chemical conjugation between two moieties, or generally may be prepared using other standard techniques. The ligated DNA sequence comprising the fusion polypeptide is operably linked to suitable transcriptional or translational control elements as disclosed elsewhere herein.

[0226] A fusion polypeptide may optionally include a linker that can be used to link one or more polypeptides or domains within the polypeptide. A peptide linker sequence may be utilized to separate any two or more polypeptide components by a distance sufficient to ensure that each polypeptide folds into its appropriate secondary and tertiary structure so that the polypeptide domains can perform their desired functions. Such peptide linker sequences are incorporated into the fusion polypeptide using standard techniques in the art. A suitable peptide linker sequence may be selected based on the following factors: (1) the ability to accommodate a flexible extended conformation, (2) the ability to accommodate a secondary structure that can interact with functional epitopes on the first and second polypeptides, and (3) the lack of hydrophobic or charged residues that can react with the polypeptide functional epitopes. A preferred peptide linker sequence contains Gly, Asn, and Ser residues. Other near-neutral amino acids, such as Thr and Ala, may also be used in the linker sequence. Amino acid sequences that may be usefully employed as linkers include those disclosed in Maratea et al., Gene 40:39-46, 1985; Murphy et al., Proc. Natl. Acad. Sci. USA 83:8258-8262, 1986; U.S. Patent No. 4,935,233 and U.S. Patent No. 4,751,180. Linker sequences are not required if a particular fusion polypeptide segment contains a non-essential N-terminal amino acid region that can be used to separate functional domains and prevent steric interference. Preferred linkers are typically flexible amino acid subsequences that are synthesized as part of the recombinant fusion protein. Linker polypeptides can be 1-200 amino acids long, 1-100 amino acids long, or 1-50 amino acids long, including all integer values ​​in between.

[0227] Exemplary linkers include the following amino acid sequences: glycine polymer (G) n ; glycine-serine polymer (G 1-5 S 1-5 ) n(wherein n is at least 1, 2, 3, 4, or 5); glycine-alanine polymers; alanine-serine polymers; GGG (SEQ ID NO:24); DGGGS (SEQ ID NO:25); TGEKP (SEQ ID NO:26) (e.g., Liu et al., PNAS 5525-5530 (1997)); GGRR (SEQ ID NO:27) (Pomerantz et al., 1995, supra); (GGGGS) n (wherein n=1, 2, 3, 4 or 5) (SEQ ID NO:28) (Kim et al., PNAS 93, 1156-1160 (1996); EGKSSGSGSESKVD (SEQ ID NO:29) (Chaudhary et al., 1990, Proc. Natl. Acad. Sci. USA 87:1066-1070); KESGSVSSEQLAQFRSLD (SEQ ID NO:30) (Bird et al., 1988, Science 242:423-426), GGRRGGGS (SEQ ID NO:31); LQRDGERP (SEQ ID NO:32); LRQKDGGGSERP (SEQ ID NO:33); LRQKD(GGGS)2ERP (SEQ ID NO:34). Alternatively, flexible linkers can be modeled using computer programs capable of modeling both the DNA binding site and the peptide itself (Desjarlais and Berg, PNAS 90:2256-2260 (1993), PNAS 91:11099-11103 (1994)) or rationally designed by phage display methods.

[0228] The fusion polypeptide may further comprise a polypeptide cleavage signal between each of the polypeptide domains described herein, or between the endogenous open reading frame and the polypeptide encoded by the donor repair template. In addition, a polypeptide cleavage site can be inserted into any linker peptide sequence. Exemplary polypeptide cleavage signals include polypeptide cleavage recognition sites such as protease cleavage sites, nuclease cleavage sites (e.g., rare restriction enzyme recognition sites, self-cleaving ribozyme recognition sites), and self-cleaving viral oligopeptides (see deFelipe and Ryan, 2004, Traffic, 5(8);616-26).

[0229] Suitable protease cleavage sites and self-cleaving peptides are known to those of skill in the art (see, for example, Ryan et al., 1997, J. Gener. Virol. 78, 699-722; Scymczak et al., (2004) Nature Biotech, 5, 589-594). Exemplary protease cleavage sites include, but are not limited to, cleavage sites for potyvirus NIa protease (e.g., tobacco etch virus protease), potyvirus HC protease, potyvirus P1 (P35) protease, biovirus NIa protease, biovirus RNA-2-encoded protease, aphthovirus L protease, enterovirus 2A protease, rhinovirus 2A protease, picorna 3C protease, comovirus 24K protease, nepovirus 24K protease, RTSV (Rice Tungro Spherical Virus) 3C-like protease, PYVF (Parsnip Yellow Fleck Virus) 3C-like protease, heparin, thrombin, factor Xa, and enterokinase. Due to their high cleavage stringency, TEV (Tobacco Etch Virus) protease cleavage sites are preferred in one embodiment, e.g., EXXYXQ(G / S) (SEQ ID NO: 35), e.g., ENLYFQG (SEQ ID NO: 36) and ENLYFQS (SEQ ID NO: 37), where X represents any amino acid (cleavage by TEV occurs between Q and G or Q and S).

[0230] In certain embodiments, the polypeptide cleavage signal is a viral autocleaving peptide or a ribosomal skipping sequence.

[0231] Examples of ribosomal skipping sequences include, but are not limited to, 2A or 2A-like sites, sequences, or domains (see Donnelly et al., 2001, J. Gen. Virol. 82:1027-1041). In certain embodiments, the viral 2A peptide is an aphthovirus 2A peptide, a potyvirus 2A peptide, or a cardiovirus 2A peptide.

[0232] In one embodiment, the viral 2A peptide is selected from the group consisting of a foot and mouth disease virus (FMDV) 2A peptide, an equine rhinitis A virus (ERAV) 2A peptide, a zosea signa virus (TaV) 2A peptide, a porcine teschovirus-1 (PTV-1) 2A peptide, a telirovirus 2A peptide, and an encephalopathy virus 2A peptide.

[0233] Examples of 2A sites are provided in Table 2. [Table 2-1] [Table 2-2]

[0234] E. Polynucleotides In certain embodiments, polynucleotides encoding one or more homing endonuclease variants and megaTALs engineered to increase thermostability and / or enzymatic activity as contemplated herein, as well as fusion polypeptides, are provided. As used herein, the term "polynucleotide" or "nucleic acid" refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA) and DNA / RNA hybrids. Polynucleotides may be single-stranded or double-stranded, either recombinant, synthetic, or isolated. Polynucleotides include, but are not limited to, pre-messenger RNA (pre-mRNA), messenger RNA (mRNA), RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozyme, genomic RNA (gRNA), positive strand RNA (RNA(+)), negative strand RNA (RNA(-)), tracrRNA, crRNA, single guide RNA (sgRNA), synthetic RNA, synthetic mRNA, genomic DNA (gDNA), PCR amplified DNA, complementary DNA (cDNA), synthetic DNA, or recombinant DNA. A polynucleotide refers to a polymeric form of nucleotides, either ribonucleotides or deoxyribonucleotides, of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 5000, at least 10000, or at least 15000 or more nucleotides in length, or modified forms of either type of nucleotide, as well as all intermediate lengths. It is readily understood that "intermediate length" in this context means any length between the cited values, such as 6, 7, 8, 9, etc., 101, 102, 103, etc., 151, 152, 153, etc., 201, 202, 203, etc.In certain embodiments, a polynucleotide or variant has at least or about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to a reference sequence.

[0235] In certain embodiments, the polynucleotide may be codon-optimized. As used herein, the term "codon optimization" refers to the substitution of codons in a polynucleotide encoding a polypeptide to increase the expression, stability and / or activity of the polypeptide. Factors that influence codon optimization include, but are not limited to, one or more of: (i) variation in codon bias between two or more organisms or genes, or synthetically constructed bias tables; (ii) variation in the degree of codon bias within an organism, gene, or set of genes; (iii) systematic variation of codons with context; (iv) variation of codons with their decoding tRNAs; (v) variation of codons with GC%, either overall or at any one position of a triplet; (vi) variation in the degree of similarity to a reference sequence, e.g., a naturally occurring sequence; (vii) variation of codon frequency cutoffs; (viii) structural properties of mRNA transcribed from a DNA sequence; (ix) prior knowledge of the function of the DNA sequence on which the design of the codon substitution set is to be based; (x) systematic variation of the codon set for each amino acid; and / or (xi) isolated removal of spurious translation start sites.

[0236] As used herein, the term "nucleotide" refers to a heterocyclic nitrogenous base in N-glycosidic linkage with a phosphorylated sugar. Nucleotides are understood to include natural bases and a wide variety of art-recognized modified bases. Such bases are typically located at the 1' position of the nucleotide sugar moiety. Nucleotides generally include a base, a sugar, and a phosphate group. In ribonucleic acid (RNA), the sugar is ribose, and in deoxyribonucleic acid (DNA), the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group present in ribose. Exemplary natural nitrogenous bases include the purines, adenosine (A) and guanidine (G), and the pyrimidines, cytidine (C) and thymidine (T) (or in the context of RNA, uracil (U)). The C-1 atom of the deoxyribose is linked to the N-1 of a pyrimidine or the N-9 of a purine. Nucleotides are usually monophosphates, diphosphates, or triphosphates. Nucleotides can be unmodified or modified at the sugar, phosphate and / or base moieties (also referred to interchangeably as nucleotide analogues, nucleotide derivatives, modified nucleotides, non-natural nucleotides, and non-standard nucleotides, see, for example, WO 92 / 07065 and WO 93 / 15187). Examples of modified nucleobases are reviewed by Limbach et al. (1994, Nucleic Acids Res. 22, 2183-2196).

[0237] Nucleotides may also be considered as phosphate esters of nucleosides, with esterification occurring on the hydroxyl group attached to C-5 of the sugar. As used herein, the term "nucleoside" refers to a heterocyclic nitrogenous base in N-glycosidic linkage with a sugar. Nucleosides are recognized in the art to include natural bases and also include well-known modified bases. Such bases are generally located at the 1' position of the nucleoside sugar moiety. Nucleosides generally include a base and a sugar group. Nucleosides may be unmodified or modified at the sugar and / or base moieties (also referred to interchangeably as nucleoside analogs, nucleoside derivatives, modified nucleosides, non-natural nucleosides, or non-standard nucleosides). Also as mentioned above, examples of modified nucleobases are reviewed by Limbach et al. (1994, Nucleic Acids Res. 22, 2183-2196).

[0238] Exemplary polynucleotides include, but are not limited to, polynucleotides encoding SEQ ID NOs: 9-14, 16-18, 22, and 23, and the polynucleotide sequences set forth in SEQ ID NOs: 19-21.

[0239] In various exemplary embodiments, polynucleotides contemplated herein include, but are not limited to, polynucleotides encoding homing endonuclease variants, megaTALs, endoprocessing enzymes, fusion polypeptides, and expression vectors, viral vectors, and transfer plasmids that comprise the polynucleotides contemplated herein.

[0240] As used herein, the terms "polynucleotide variant" and "variant" and the like refer to a polynucleotide that exhibits substantial sequence identity with a reference polynucleotide sequence, or hybridizes with a reference sequence under stringent conditions as defined below. These terms also encompass polynucleotides that are distinguished from a reference polynucleotide by the addition, deletion, substitution, or modification of at least one nucleotide. Thus, the terms "polynucleotide variant" and "variant" include polynucleotides in which one or more nucleotides are added or deleted, or modified, or replaced with different nucleotides. In this regard, it is well known in the art that certain changes, including mutations, additions, deletions, and substitutions, may be made to a reference polynucleotide, whereby the modified polynucleotide retains the biological function or activity of the reference polynucleotide.

[0241] In one embodiment, the polynucleotide comprises a nucleotide sequence that hybridizes to a target nucleic acid sequence under stringent conditions. To hybridize under "stringent conditions", a hybridization protocol is described in which nucleotide sequences that are at least 60% identical to each other remain hybridized. In general, stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) of a particular sequence at a defined ionic strength and pH. Tm is the temperature (under a defined ionic strength, pH and nucleic acid concentration) at which 50% of the probes that are complementary to the target sequence hybridize to the target sequence at equilibrium. The target sequence is generally present in excess, so that at Tm, 50% of the probes are occupied at equilibrium.

[0242] The "sequence identity" listed, or for example, "50% identical sequence" as used herein, refers to the degree to which sequences are identical on a nucleotide-to-nucleotide basis or an amino acid-to-amino acid basis over a comparison window.Therefore, "sequence identity percentage" can be calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions where identical nucleic acid bases (e.g., A, T, C, G, I) or identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys and Met) occur in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., window size), and multiplying the result by 100 to obtain the percentage of sequence identity. Included are nucleotides and polypeptides having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any of the reference sequences described herein, and typically the polypeptide variants retain at least one biological activity of the reference polypeptide.

[0243] Terms used to describe sequence relationships between two or more polynucleotides or polypeptides include "reference sequence," "comparison window," "sequence identity," "percentage of sequence identity," and "substantial identity." A "reference sequence" comprises nucleotides and amino acid residues that are at least 12 monomeric units in length, and often 15-18 monomeric units, and often at least 25 monomeric units in length. Two polynucleotides may each contain (1) sequences that are similar between the two polynucleotides (i.e., only a portion of the complete polynucleotide sequence), and (2) sequences that diverge between the two polynucleotides, and sequence comparison between two (or more) polynucleotides is typically performed by comparing the sequences of the two polynucleotides in a "comparison window" to identify and compare local regions of sequence similarity. A "comparison window" refers to a conceptual segment of at least 6 contiguous positions, usually about 50 to about 100, more commonly about 100 to about 150, and a sequence is compared to a reference sequence in the same number of contiguous positions after the two sequences are optimally aligned. The comparison window may contain about 20% or less additions or deletions (i.e., gaps) compared to the reference sequence (not including additions or deletions) for optimal alignment of two sequences. Optimal alignment of sequences for aligning the comparison window can be performed by computerized implementation of algorithms (GAP, BESTFIT, FASTA, and TFASTA) in Wisconsin Genetics Software Package Release 7.0, Genetics Computer Group, 575 Science Drive Madison, WI, USA, or by inspection and best alignment (i.e., resulting in the highest homology over the comparison window) generated by any of the various methods selected. For example, see the BLAST family of programs disclosed by Altschul et al., 1997, Nucl.Acids Res.25:3389.A detailed discussion of sequence analysis can be found in Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons Inc., 1994-1998, Chapter 15, Unit 19.3.

[0244] As used herein, an "isolated polynucleotide" refers to a polynucleotide that has been purified from sequences adjacent to it in its naturally occurring state, e.g., a DNA fragment that has been removed from sequences that normally flank the fragment. In certain embodiments, an "isolated polynucleotide" refers to a complementary DNA (cDNA), a recombinant polynucleotide, a synthetic polynucleotide, or other polynucleotide that does not exist in nature and is created by the hand of man. In certain embodiments, an isolated polynucleotide is a synthetic polynucleotide, a semi-synthetic polynucleotide, or a polynucleotide obtained or derived from a recombinant source.

[0245] In various embodiments, the polynucleotide comprises an mRNA encoding a polypeptide contemplated herein, including, but not limited to, a homing endonuclease variant, a megaTAL, and an endoprocessing enzyme. In certain embodiments, the mRNA comprises a cap, one or more nucleotides, and a poly(A) tail.

[0246] As used herein, the term "5' cap" or "5' cap structure" or "5' cap moiety" refers to a chemical modification incorporated at the 5' end of an mRNA. The 5' cap is involved in nuclear export, mRNA stability, and translation.

[0247] In certain embodiments, the mRNA encoding the homing endonuclease variant or megaTAL comprises a 5' cap comprising a 5'-ppp-5'-triphosphate linkage between the terminal guanosine cap residue and the 5'-terminal transcribed sense nucleotide of the mRNA molecule, which 5'-guanylate cap may then be methylated to generate an N7-methyl-guanylate residue.

[0248] Illustrative examples of 5' caps suitable for use in certain embodiments of the mRNA polynucleotides contemplated herein include unmethylated 5' cap analogs, e.g., G(5')ppp(5')G, G(5')ppp(5')C, G(5')ppp(5')A; methylated 5' cap analogs, e.g., m 7 G(5´)ppp(5´)G、m 7 G(5´)ppp(5´)C, and m 7 G(5´)ppp(5´)A; dimethylated 5´ cap analogues, e.g., m 2,7 G(5´)ppp(5´)G、m 2,7 G(5´)ppp(5´)C, and m 2,7 G(5´)ppp(5´)A; trimethylated 5´ cap analogs, e.g., m 2,2,7 G(5´)ppp(5´)G、m 2,2,7 G(5´)ppp(5´)C, and m 2,2,7 G(5´)ppp(5´)A; dimethylated symmetrical 5´ cap analogues, e.g., m 7 G(5´)pppm 7 (5´)G、m 7 G(5´)pppm 7 (5´)C, and m 7 G(5´)pppm 7 (5')A; and anti-reverse 5' cap analogs, such as anti-reverse cap analog (ARCA) caps (3'O-Me-m 7 G(5´)ppp(5´)G, 2´O-Me-m 7 G(5´)ppp(5´)G, 2´O-Me-m 7 G(5´)ppp(5´)C, 2´O-Me-m 7 G(5´)ppp(5´)A、m 7 2´d(5´)ppp(5´)G、m 7 2´d(5´)ppp(5´)C、m 7 2´d(5´)ppp(5´)A, 3´O-Me-m 7 G(5´)ppp(5´)C, 3´O-Me-m 7 G(5´)ppp(5´)A、m 7 3´d(5´)ppp(5´)G、m7 3´d(5´)ppp(5´)C、m 7 3'd(5')ppp(5')A and their tetraphosphate derivatives) (e.g., Jemielity et al., RNA, 9:1108-1122 (2003)).

[0249] In certain embodiments, the mRNA encoding the homing endonuclease variant or the megaTAL is attached to the 5' end of the first transcribed nucleotide via a triphosphate bridge, 7 G(5')ppp(5')N, where N is any nucleoside. 7 5' cap which is

[0250] In some embodiments, the mRNA encoding the homing endonuclease variant or megaTAL comprises a 5' cap, wherein the cap is a Cap0 structure (the Cap0 structure lacks a 2'-O-methyl residue of the ribose attached to bases 1 and 2), a Cap1 structure (the Cap1 structure has a 2'-O-methyl residue at base 2), or a Cap2 structure (the Cap2 structure has a 2'-O-methyl residue attached to both bases 2 and 3).

[0251] In one embodiment, the mRNA is 7 Contains a G(5')ppp(5')G cap.

[0252] In one embodiment, the mRNA includes an ARCA cap.

[0253] In certain embodiments, the mRNA encoding the homing endonuclease variant or the megaTAL comprises one or more modified nucleosides.

[0254] In one embodiment, the mRNA encoding the homing endonuclease variant or megaTAL is selected from the group consisting of pseudouridine, pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyluridine, 1-amino ... Methyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, 1-taurinomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine , 4-Methoxy-pseudouridine, 4-Methoxy-2-thio-pseudouridine, 5-Aza-cytidine, Pseudoisocytidine, 3-Methyl-cytidine, N4-Acetylcytidine, 5-Formylcytidine, N4-Methylcytidine, 5-Hydroxymethylcytidine, 1-Methyl-pseudoisocytidine, Pyrrolo-cytidine, Pyrrolo-pseudoisocytidine, 2-Thio-cytidine, 2-Thio-5-Methyl-cytidine, 4-Thio-pseudoisocytidine, 4-Thio-1-Methyl-pseudoisocytidine, 4-Thio-1-Methyl-1-Deaza-pseudoisocytidine Cytidine, 1-methyl-1-deaza-pseudoisocytidine, Zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-Diaminopurine, 7-Deaza-8-aza-2,6-diaminopurine, 1-Methyladenosine, N6-Methyladenosine, N6-Isopentenyladenosine, N6-(cis-Hydroxyisopentenyl)adenosine, 2-Methylthio-N6-(cis-Hydroxyisopentenyl)adenosine, N6-Glycinylcarbamoyladenosine, N6-Threonylcarbamoyladenosine, 2-Methylthio-N6-Threonylcarbamoyladenosine, N6,N6-Dimethyladenosine, 7-Methyladenine, 2-Methylthio-adenine, 2-Methoxy-adenine, Inosine, 1-Methyl-inosine, Wyosine, Wybutosine, 7-Deazagu The modified nucleosides include one or more modified nucleosides selected from the group consisting of inosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0255] In one embodiment, the mRNA encoding the homing endonuclease variant or megaTAL is selected from the group consisting of pseudouridine, pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio- ... The modified nucleosides include one or more modified nucleosides selected from the group consisting of nomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, and 4-methoxy-2-thio-pseudouridine.

[0256] In one embodiment, the mRNA encoding the homing endonuclease variant or megaTAL is selected from the group consisting of 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 5-methyl-cytidine, 5-hydroxymethyl- ... The modified nucleosides include one or more modified nucleosides selected from the group consisting of pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, and 4-methoxy-1-methyl-pseudoisocytidine.

[0257] In one embodiment, the mRNA encoding the homing endonuclease variant or megaTAL is selected from the group consisting of 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyl adenosine, N6 -(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, and 2-methoxy-adenine.

[0258] In one embodiment, the mRNA encoding the homing endonuclease variant or megaTAL is selected from the group consisting of inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, It comprises one or more modified nucleosides selected from the group consisting of 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, and N2,N2-dimethyl-6-thio-guanosine.

[0259] In one embodiment, the mRNA includes one or more pseudouridines, one or more 5-methyl-cytosines, and / or one or more 5-methyl-cytidines.

[0260] In one embodiment, the mRNA includes one or more pseudouridines.

[0261] In one embodiment, the mRNA includes one or more 5-methyl-cytidines.

[0262] In one embodiment, the mRNA includes one or more 5-methyl-cytosines.

[0263] In certain embodiments, the mRNA encoding the homing endonuclease variant or megaTAL comprises a poly(A) tail to help protect the mRNA from exonuclease degradation, stabilize the mRNA, and facilitate translation. In certain embodiments, the mRNA comprises a 3' poly(A) tail structure.

[0264] In certain embodiments, the length of the poly(A) tail is at least about 10, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, or at least about 500 or more adenine nucleotides, or any intervening number of adenine nucleotides. , 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 202, 203, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 2 41, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, or 275 or more adenine nucleotides.

[0265] In certain embodiments, the length of the poly(A) tail is about 10 to about 500 adenine nucleotides, about 50 to about 500 adenine nucleotides, about 100 to about 500 adenine nucleotides, about 150 to about 500 adenine nucleotides, about 200 to about 500 adenine nucleotides, about 250 to about 500 adenine nucleotides, about 300 to about 500 adenine nucleotides, about 50 to about 450 adenine nucleotides, about 50 to about 500 adenine nucleotides, about 60 to about 650 adenine nucleotides, about 70 to about 750 adenine nucleotides, about 80 to about 850 adenine nucleotides, about 90 to about 950 adenine nucleotides, about 100 to about 1000 adenine nucleotides, about 150 to about 500 adenine nucleotides, about 200 to about 500 adenine nucleotides, about 250 to about 500 adenine nucleotides, about 300 to about 500 adenine nucleotides, about 50 to about 450 adenine nucleotides, about 50 to about 500 adenine nucleotides, about 60 to about 750 adenine nucleotides, about 70 to about 850 adenine nucleotides, about 80 to about 950 adenine nucleotides, about 90 to about 1000 adenine nucleotides, about 100 to about 1000 adenine nucleotides, about 150 to about 500 adenine nucleotides, about 200 to about 500 adenine nucleotides, about 250 to about 500 adenine nucleotides, about 30 adenine nucleotides, about 50 to about 400 adenine nucleotides, about 50 to about 350 adenine nucleotides, about 100 to about 500 adenine nucleotides, about 100 to about 450 adenine nucleotides, about 100 to about 400 adenine nucleotides, about 100 to about 350 adenine nucleotides, about 100 to about 300 adenine nucleotides, about 150 to about 500 adenine nucleotides, about 150 to about 500 adenine nucleotides, About 450 adenine nucleotides, about 150 to about 400 adenine nucleotides, about 150 to about 350 adenine nucleotides, about 150 to about 300 adenine nucleotides, about 150 to about 250 adenine nucleotides, about 150 to about 200 adenine nucleotides, about 200 to about 500 adenine nucleotides, about 200 to about 450 adenine nucleotides, about 200 to about 400 adenine nucleotides leotide, about 200 to about 350 adenine nucleotides, about 200 to about 300 adenine nucleotides, about 250 to about 500 adenine nucleotides, about 250 to about 450 adenine nucleotides, about 250 to about 400 adenine nucleotides, about 250 to about 350 adenine nucleotides, or about 250 to about 300 adenine nucleotides, or any intervening range of adenine nucleotides.

[0266] Terms describing the orientation of a polynucleotide include 5' (usually the end of a polynucleotide having a free phosphate group) and 3' (usually the end of a polynucleotide having a free hydroxyl (OH) group). A polynucleotide sequence may be annotated in a 5' to 3' or 3' to 5' direction. For DNA and mRNA, the 5' to 3' strand is designated the "sense", "plus", or "coding" strand because its sequence is identical to that of the pre-messenger (pre-mRNA) [except for uracil (U) in RNA instead of thymine (T) in DNA]. For DNA and mRNA, the complementary 3' to 5' strand, which is the strand transcribed by RNA polymerase, is designated as the "template", "antisense", "minus", or "non-coding" strand. As used herein, the term "reverse" refers to a 5' to 3' sequence written in a 3' to 5' direction or a 3' to 5' sequence written in a 5' to 3' direction.

[0267] The terms "complementary" and "complementarity" refer to polynucleotides (i.e., a sequence of nucleotides) related by the base-pairing rules. For example, the complementary strand of the DNA sequence 5'AGTCTATTG 3' is 3'TCAGTAC 5'. The latter sequence is often written as the reverse complement, 5'ATGGACT 3', with the 5' end on the left and the 3' end on the right. A sequence equivalent to its reverse complement is said to be a palindromic sequence. Complementarity can be "partial," where only a portion of the nucleic acid bases match according to the base-pairing rules. Alternatively, there is "complete" or "total" complementarity between the nucleic acids.

[0268] In certain embodiments, the intended polynucleotides, regardless of the length of the coding sequence itself, may be combined with other DNA sequences, and may include promoters and / or enhancers, untranslated regions (UTRs), Kozak sequences, polyadenylation signals, restriction enzyme sites, multiple cloning sites, internal ribosome entry sites (IRES), recombinase recognition sites (e.g., LoxP, FRT, and Att sites), stop codons, transcription termination signals, post-transcriptional response elements, and polynucleotides encoding self-cleaving polypeptides, epitope tags, as disclosed elsewhere herein or known in the art, and thus may vary greatly in overall length. Thus, in certain embodiments, polynucleotide fragments of almost any length may be used, with the overall length preferably being limited by the ease of preparation and use in the intended recombinant DNA protocol.

[0269] Polynucleotides can be prepared, manipulated, expressed, and / or delivered using any of a variety of established techniques known and available in the art.To express a desired polypeptide, the nucleotide sequence encoding the polypeptide can be inserted into an appropriate vector.A desired polypeptide can also be expressed by delivering the mRNA encoding the polypeptide into cells.

[0270] Examples of vectors include, but are not limited to, plasmids, autonomously replicating sequences, and transposable elements, e.g., Sleeping Beauty, PiggyBac.

[0271] Additional examples of vectors include, but are not limited to, plasmids, phagemids, cosmids, artificial chromosomes such as yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), or P1-derived artificial chromosomes (PACs), bacteriophages such as lambda or M13 phages, and animal viruses.

[0272] Examples of viruses useful as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpes viruses (e.g., herpes simplex viruses), pox viruses, baculoviruses, papilloma viruses, and papova viruses (e.g., SV40).

[0273] Exemplary expression vectors include, but are not limited to, pClneo vector (Promega) for expression in mammalian cells, pLenti4 / V5-DEST™, pLenti6 / V5-DEST™, and pLenti6.2 / V5-GW / lacZ (Invitrogen) for lentivirus-mediated gene transfer and expression in mammalian cells. In certain embodiments, the coding sequence of the polypeptide disclosed herein can be ligated into such an expression vector for expression of the polypeptide in mammalian cells.

[0274] In certain embodiments, the vector is an episomal vector or a vector that is maintained extrachromosomally. As used herein, the term "episomal" refers to a vector that can replicate without integration into the host's chromosomal DNA and without gradual loss from dividing host cells, meaning that the vector replicates extrachromosomally or episomally.

[0275] An "expression control sequence," "control element," or "regulatory sequence" present in an expression vector is an untranslated region of the vector, including, but not limited to, origins of replication, selection cassettes, promoters, enhancers, translation initiation signals (Shine Dalgarno or Kozak sequences), introns, post-transcriptional regulatory elements, polyadenylation sequences, 5' and 3' untranslated regions, that interact with host cell proteins to effect transcription and translation. Such elements may vary in their strength and specificity. Depending on the vector system and host utilized, any number of suitable transcription and translation elements may be used, including ubiquitous and inducible promoters.

[0276] The term "operably linked" refers to a juxtaposition where the components described are in a relationship permitting them to function in their intended manner. In one embodiment, the term refers to a functional linkage between a nucleic acid expression control sequence (such as a promoter and / or enhancer) and a second polynucleotide sequence, e.g., a polynucleotide of interest, where the expression control sequence directs transcription of the nucleic acid corresponding to the second sequence.

[0277] Elements that direct efficient termination and polyadenylation of heterologous nucleic acid transcripts increase heterologous gene expression. Transcription termination signals are commonly found downstream of polyadenylation signals. In certain embodiments, the vector includes a polyadenylation sequence 3' of the polynucleotide encoding the polypeptide to be expressed. As used herein, the term "polyA site" or "polyA sequence" refers to a DNA sequence that directs both the termination and polyadenylation of the nascent RNA transcript by RNA polymerase II. Polyadenylation sequences can promote mRNA stability by adding a polyA tail to the 3' end of the coding sequence, thus contributing to improved translation efficiency. Cleavage and polyadenylation are directed by poly(A) sequences in the RNA. The core poly(A) sequence of mammalian pre-mRNA has two recognition elements adjacent to the cleavage polyadenylation site. Typically, a nearly invariant AAUAAA hexamer is present 20-50 nucleotides upstream of a more variable element that is rich in U or GU residues. Cleavage of the nascent transcript occurs between these two elements, resulting in the attachment of up to 250 adenosines to the 5' cleavage product. In certain embodiments, the core poly(A) sequence is an ideal polyA sequence (e.g., AATAAA, ATTAAA, AGTAAA). In certain embodiments, the poly(A) sequence is an SV40 polyA sequence, a bovine growth hormone polyA sequence (BGHpA), a rabbit β-globin polyA sequence (rβgpA), a variant thereof, or another suitable heterologous or endogenous polyA sequence known in the art. In certain embodiments, the poly(A) sequence is synthetic.

[0278] In certain embodiments, polynucleotides encoding one or more nuclease variants, megaTALs, endo-processing enzymes, or fusion polypeptides may be introduced into cells by both non-viral and viral methods.

[0279] The term "vector" is used herein to refer to a nucleic acid molecule that can transmit or transport another nucleic acid molecule. The transmitted nucleic acid is generally inserted, for example, into a vector nucleic acid molecule. The vector may contain sequences that direct self-replication in a cell, or may contain sequences sufficient to allow integration into host cell DNA. In certain embodiments, non-viral vectors are used to deliver one or more polynucleotides contemplated herein to T cells.

[0280] Examples of non-viral vectors include, but are not limited to, plasmids (eg, DNA or RNA plasmids), transposons, cosmids, and bacterial artificial chromosomes.

[0281] Exemplary methods of non-viral delivery of polynucleotides contemplated in certain embodiments include, but are not limited to, electroporation, sonoporation, lipofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, nanoparticles, polycation or lipid:nucleic acid conjugates, naked DNA, artificial virions, DEAE-dextran mediated delivery, gene guns, and heat shock.

[0282] Examples of viral vector systems suitable for use in certain embodiments contemplated herein include, but are not limited to, adeno-associated virus (AAV), retroviruses, e.g., lentiviruses, herpes simplex viruses, adenoviruses, and vaccinia virus vectors.

[0283] F. Compositions and Formulations Compositions contemplated in certain embodiments may include one or more homing endonuclease variants engineered to increase thermostability and / or enzymatic activity and megaTALs contemplated herein, polynucleotides, vectors comprising same, and genome editing compositions and genome editing cell compositions. Genome editing compositions and methods contemplated in certain embodiments are useful for editing target sites in the human genome in a cell or cell population.

[0284] "Isolated cells" refer to cells obtained from an in vivo tissue or organ and substantially free of extracellular matrix that do not occur in nature, e.g., non-naturally occurring cells, modified cells, engineered cells, recombinant cells, etc.

[0285] As used herein, the term "population of cells" refers to a plurality of cells that may consist of any number and / or combination of homogenous or heterogeneous cell types.

[0286] In certain embodiments, the genome editing compositions are used to edit target sites in embryonic stem cells, or adult stem or progenitor cells.

[0287] In certain embodiments, the genome editing composition is used to edit the target site of a stem or progenitor cell selected from the group consisting of mesodermal stem or progenitor cells, endodermal stem or progenitor cells, and ectodermal stem or progenitor cells. Illustrative examples of mesodermal stem or progenitor cells include, but are not limited to, bone marrow stem or progenitor cells, umbilical cord blood stem or progenitor cells, adipose tissue-derived stem or progenitor cells, hematopoietic stem or progenitor cells (HSPCs), mesenchymal stem or progenitor cells, muscle stem or progenitor cells, kidney stem or progenitor cells, osteoblastic stem or progenitor cells, chondroblastic or endodermal progenitor cells, and the like. Illustrative examples of ectodermal stem or progenitor cells include, but are not limited to, neural stem or progenitor cells, retinal stem or progenitor cells, skin stem or progenitor cells, and the like. Illustrative examples of endodermal stem or progenitor cells include, but are not limited to, liver stem or progenitor cells, pancreatic stem or progenitor cells, epithelial stem or progenitor cells, and the like.

[0288] In certain embodiments, the genome editing composition is used to edit a target site in a bone cell, a bone cell, an osteoblast, an adipocyte, a chondrocyte, a chondroblast, a muscle cell, a skeletal muscle cell, a myoblast, a muscle cell, a smooth muscle cell, a bladder cell, a bone marrow cell, a central nervous system (CNS) cell, a peripheral nervous system (PNS) cell, a glial cell, an astrocyte, a neuron, a pigment cell, an epithelial cell, a skin cell, an endothelial cell, a vascular endothelial cell, a breast cell, a colon cell, an esophageal cell, a gastrointestinal cell, a stomach cell, a colon cell, a head cell, a neck cell, a gingival cell, a tongue cell, a kidney cell, a liver cell, a lung cell, a nasopharyngeal cell, an ovarian cell, a follicular cell, a cervical cell, a vaginal cell, a uterine cell, a pancreatic cell, a pancreatic parenchymal cell, a pancreatic duct cell, a pancreatic islet cell, a prostate cell, a penile cell, a gonadal cell, a testicular cell, a hematopoietic cell, a lymphatic cell, or a bone marrow cell.

[0289] In a preferred embodiment, the genome editing composition is used to edit a target site in a hematopoietic cell, such as a hematopoietic stem cell, a hematopoietic progenitor cell, a CD34+ cell, an immune effector cell, a T cell, a NKT cell, a NK cell, or the like.

[0290] In various embodiments, the compositions contemplated herein comprise an I-OnuI HE variant engineered to increase thermostability and / or enzymatic activity, and optionally an endo-processing enzyme, e.g., a 3'-5' exonuclease (Trex2). The I-OnuI HE variant may be in the form of an mRNA that is introduced into a cell via the polynucleotide delivery methods disclosed above, e.g., electroporation, lipid nanoparticles, etc. In one embodiment, a composition comprising an I-OnuI HE variant or a megaTAL encoding an mRNA, and optionally a 3'-5' exonuclease, is introduced into a cell via the polynucleotide delivery methods disclosed above. The composition may be used to generate a genome-edited cell or a population of genome-edited cells by error-prone NHEJ.

[0291] In various embodiments, the compositions contemplated herein include a donor repair template. The compositions may be delivered to cells that express or will express an I-OnuI HE variant, and optionally an end-processing enzyme. In one embodiment, the compositions may be delivered to cells that express or will express an I-OnuI HE variant or a megaTAL, and optionally a 3'-5' exonuclease. Expression of gene editing enzymes in the presence of a donor repair template can be used to generate genome-edited cells or genome-edited cell populations by HDR.

[0292] In certain embodiments, the composition comprises cells containing one or more homing endonuclease variants engineered to increase thermostability and / or enzymatic activity and megaTALs, polynucleotides, vectors comprising same. In certain embodiments, the cells may be autologous / autologous (autologous) or non-autologous (non-autologous, e.g., allogeneic, syngeneic, or xenogeneic). As used herein, "autologous" refers to cells from the same subject. As used herein, "allogeneic" refers to cells of the same species that are genetically different from the compared cells. As used herein, "syngeneic" refers to cells of a different species than the compared cells, but are genetically identical to the compared cells. As used herein, "xenogeneic" refers to cells of a different species than the compared cells. In preferred embodiments, the cells are obtained from a mammalian subject. In more preferred embodiments, the cells are obtained from a primate subject, optionally a non-human primate. In the most preferred embodiments, the cells are obtained from a human subject.

[0293] In certain embodiments, the compositions contemplated herein comprise a population of cells, an I-OnuI HE variant, and optionally a donor repair template. In certain embodiments, the compositions contemplated herein comprise a population of cells, an I-OnuI HE variant, an end-processing enzyme, and optionally a donor repair template. The I-OnuI HE and / or the end-processing enzyme may be in the form of mRNA that is introduced into the cell via the polynucleotide delivery method disclosed above.

[0294] In certain embodiments, the compositions contemplated herein comprise a population of cells, an I-OnuI HE variant or a megaTAL engineered to increase the thermostability and / or activity of the enzyme, and optionally a donor repair template. In certain embodiments, the compositions contemplated herein comprise a population of cells, an I-OnuI HE variant or a megaTAL, a 3'-5' exonuclease, and optionally a donor repair template. The I-OnuI HE variant, the megaTAL, and / or the 3'-5' exonuclease may be in the form of mRNA that is introduced into the cells via the polynucleotide delivery methods disclosed above.

[0295]

[0296] All publications, patent applications, and issued patents cited in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or issued patent was specifically and individually indicated to be incorporated by reference.

[0297] Although the foregoing embodiments have been described in some detail by way of explanation and example for purposes of clarity of understanding, it will be readily apparent to those skilled in the art that, in light of the teachings contemplated herein, certain changes and modifications may be made without departing from the spirit or scope of the appended claims. The following examples are provided for illustrative purposes only, and not for purposes of limitation. Those skilled in the art will readily recognize various non-critical parameters that can be changed or modified to produce essentially similar results. EXAMPLES

[0298] Example 1 Identification of amino acid positions in LADLIDADG homing endonucleases that increase thermostability A yeast surface display assay was used to identify mutations that increase the thermostability of LAGLIDADG homing endonucleases. First, the stability of I-OnuI (e.g., SEQ ID NO: 1) and engineered nucleases (e.g., SEQ ID NOs: 6 and 7) was measured. After nuclease surface expression was induced in yeast, each yeast population was subjected to heat shock at multiple temperatures for 15 minutes, and the percentage of nuclease expressing cells that were still able to cleave their DNA target was measured by flow cytometry. This assay determined that for each endonuclease, the associated TM 50 Generating a standard protein melting curve with values. Figure 1.

[0299] To identify mutations that confer improved stability, multiple I-OnuI-derived homing endonucleases were subjected to random mutagenesis via PCR across the entire open reading frame. These mutant libraries were expressed in yeast and the TMs of the libraries were analyzed. 50 The above were selected for active nuclease activity after heat shock. After two rounds of selection, the HE variants were sequenced using either PacBio or Sanger sequencing to determine the identity and frequency of mutations at each position. The cumulative mutation frequencies are shown in Figure 2 as stacked bar graphs, and the top mutations at each position are shown in Table 2. Figure 3A shows the BCL11A HE variants with single amino acid substitutions and their effect on thermostability (SEQ ID NO: 8), as well as an overall TM that was slightly higher than the individual variants. 50 Figure 3A shows the thermostability of a representative I-OnuI HE variant (SEQ ID NO: 13) targeting parent I-OnuI HE variant generated from random mutagenesis with the following sequence:

[0300] Example 2 Combinatorial I-OnuI HE stabilizing mutations increase thermostability To further increase the thermostability of I-OnuI HE variants, the most frequently mutated amino acid positions were combined in a single library. Starting with the BCL11A I-OnuI HE variant (SEQ ID NO: 8), residues 14, 153, 156, 168, 178, 208, 261, and 300 were mutated using degenerate codons and PCR (subset 1 mutation library). Screening this library at a relatively permissive temperature of 46°C resulted in a population of variants that were 10°C more stable than the products from the random mutation library (Figure 3B). One representative BCL11A I-OnuI HE variant from the subset 1 mutant library (A5 SEQ ID NO: 14) showed an unexpected increase in thermostability of 22°C compared to the parent I-OnuI HE variant (Figure 3C). Combinatorial mutations derived from either random or directed mutagenesis increase the thermostability of I-Onu HE variants.

[0301] Example 3 Mutations that increase thermostability can be transferred between I-OnuI HE variants To determine whether the stabilizing mutations were unique to each reprogrammed I-OnuI HE variant or could be transferred between enzymes, mutations from the BCL11A A5 I-OnuI HE mutant (SEQ ID NO: 14) were transferred to I-OnuI HE mutants targeting PDCD-1 (SEQ ID NO: 6), TCRα (SEQ ID NO: 7), or CBLB (SEQ ID NO: 15). These mutations increased the TM of I-OnuI HE variants targeting the human PDCD-1 gene. 50 (SEQ ID NO: 16) and the TM of the I-OnuI HE variant targeting the human TCRα gene. 50 (SEQ ID NO: 17), which increased the TM of the I-OnuI HE variant targeting the human CBLB gene by approximately 14°C. 50 (SEQ ID NO: 18) increased the thermal stability by 19° C. (FIGS. 4A-4C). Mutations increasing thermostability were transferable between different I-OnuI HE variants.

[0302] Example 4 Increased thermostability extends the duration of I-OnuI HE variant expression The thermostability of I-OnuI HE variants was also assessed by measuring the expression of the enzyme in 293T cells. Briefly, each I-OnuI HE variant was formatted as an mRNA with a c-terminal HA tag followed by tracking mRNA transfection efficiency by T2A GFP (Figure 5A). mRNA was prepared by in vitro transcription, co-transcriptionally capped with an anti-reverse cap analog, and enzymatically polyadenylated with poly(A) polymerase. mRNA was purified and equal amounts were electroporated into 293T cells (protein SEQ ID NOs: 2, 8, 14: mRNA SEQ ID NOs: 19-21). At each time point, cells were run on a cytometer to measure GFP expression, and then lysed and frozen for western blot analysis.

[0303] The dynamics of GFP protein expression were similar for each polycistronic mRNA, whereas the amount of HA-tagged HE protein varied (Figure 5B and 5C). Four hours after electroporation, the amount of stabilized BCL11A A5 HE protein was significantly higher compared to the amount of parental BCL11A HE protein when normalized to the actin loading control. Furthermore, at later time points, for example, at 21 hours, the amount of parental BCL11A HE protein was undetectable, in contrast to the BCL11A A5 HE variant, which was still close to its peak expression level. Overall, stabilized BCL11A A5 HE protein was significantly higher than the amount of parental BCL11A HE protein when normalized to the actin loading control. The BCL11A A5 HE variant persisted in cells for nearly twice as long as the parental HE.

[0304] Example 5 I-OnuI HE variants engineered to increase thermostability show increased catalytic activity The effect of the stabilizing mutation on PDCD-1 editing was measured by comparing the editing rate of the parental megaTAL lacking the stabilizing mutation (SEQ ID NO: 22) with the megaTAL containing the stabilizing mutation (SEQ ID NO: 23). The megaTAL mRNA was prepared by in vitro transcription, co-transcribed with anti-inverted cap analog (ARCA), and enzymatically polyadenylated with poly(A) polymerase. The purified mRNA was used to measure the PDCD-1 editing efficiency in primary human T cells.

[0305] Primary human peripheral blood mononuclear cells (PBMCs) from two donors were activated with anti-CD3 and anti-CD28 antibodies and cultured in the presence of 250 U / mL IL-2. Three days after activation, cells were electroporated with megaTAL mRNA. Transfected T cells were expanded for an additional 7–10 days and editing efficiency was measured using sequencing and Tracking of Indels by Decomposition (TIDE, see Brinkman et al., 2014) across the PDCD-1 target site (Figure 6). Without the stabilizing mutation, the PDCD-1megaTAL showed low levels of editing (<20%), while the stabilized PDCD-1megaTAL increased editing activity to nearly 80%.

[0306] Generally, in the following claims, the terms used should not be construed to limit the claims to the specific embodiments disclosed in the specification and the claims, but should be construed to include all possible embodiments, along with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by this disclosure. The present invention provides, for example, the following items. (Item 1) An I-OnuI homing endonuclease (HE) variant comprising one or more amino acid substitutions relative to a parent I-OnuI HE comprising the amino acid sequence set forth in SEQ ID NO:1, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 2) 2. The I-OnuI HE variant of item 1, wherein the one or more amino acid substitutions are at amino acid positions selected from the group consisting of I14, A19, V116, F168, D208, N246, and L263. (Item 3) 2. The I-OnuI HE variant of item 1, wherein the I-OnuI HE variant contains amino acid substitutions at amino acid positions I14, A19, F168, D208, and N246. (Item 4) 2. The I-OnuI HE variant of item 1, wherein the one or more amino acid substitutions are at amino acid positions selected from the group consisting of K108, K156, S176, E231, V261, E277, and G300. (Item 5) 2. The I-OnuI HE variant of item 1, wherein the one or more amino acid substitutions are at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 6) 2. The I-OnuI HE variant of item 1, wherein the I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300. (Item 7) 2. The I-OnuI HE variant of item 1, wherein the I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 8) 2. The I-OnuI HE variant of item 1, wherein the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300. (Item 9) 2. The I-OnuI HE variant of item 1, wherein the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 10) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 10°C higher than TM 50 10. The I-OnuI HE variant according to any one of items 1 to 9, having the following structure: (Item 11) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 15°C higher than TM 50 10. The I-OnuI HE variant according to any one of items 1 to 9, having the following structure: (Item 12) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 20°C higher than TM 50 10. The I-OnuI HE variant according to any one of items 1 to 9, having the following structure: (Item 13) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 25°C higher than TM 50 10. The I-OnuI HE variant according to any one of items 1 to 9, having the following structure: (Item 14) 14. The I-OnuI HE variant of any one of items 1 to 13, wherein said I-OnuI HE variant targets a site of a gene selected from the group consisting of HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS. (Item 15) An I-OnuI homing endonuclease (HE) variant comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 16) 16. The I-OnuI HE variant of item 15, wherein the parent I-OnuI HE amino acid sequence is set forth in SEQ ID NO:1. (Item 17) 17. The I-OnuI HE variant of item 15 or 16, wherein the one or more amino acid substitutions are at amino acid positions selected from the group consisting of I14, A19, V116, F168, D208, N246, and L263. (Item 18) 18. The I-OnuI HE variant according to any one of items 15 to 17, wherein the I-OnuI HE variant comprises amino acid substitutions at amino acid positions I14, A19, F168, D208, and N246. (Item 19) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted with I14 is selected from the group consisting of S, N, M, K, F, D, T, and V. (Item 20) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted with I14 is selected from the group consisting of T and V. (Item 21) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted with A19 is selected from the group consisting of C, D, I, L, S, T and V. (Item 22) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted for A19 is selected from the group consisting of T and V. (Item 23) 18. The I-OnuI HE variant of item 17, wherein the amino acid substituted for V116 is selected from the group consisting of F, D, A, L, and I. (Item 24) 18. The I-OnuI HE variant of item 17, wherein the amino acid substituted for V116 is selected from the group consisting of L and I. (Item 25) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted for F168 is selected from the group consisting of H, Y, I, V, P, L and S. (Item 26) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted for F168 is selected from the group consisting of L and S. (Item 27) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted for D208 is selected from the group consisting of N, V, Y, and E. (Item 28) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted for D208 is E. (Item 29) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted at N246 is selected from the group consisting of H, I, D, R, S, T, V, Y, and K. (Item 30) 19. The I-OnuI HE variant according to item 17 or 18, wherein the amino acid substituted at N246 is K. (Item 31) 18. The I-OnuI HE variant of item 17, wherein the amino acid substituted for L263 is selected from the group consisting of H, F, P, T, V, and R. (Item 32) 18. The I-OnuI HE variant according to item 17, wherein the amino acid substituted for L263 is R. (Item 33) 17. The I-OnuI HE variant of item 15 or 16, wherein the one or more amino acid substitutions are at amino acid positions selected from the group consisting of K108, K156, S176, E231, V261, E277, and G300. (Item 34) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for K108 is selected from the group consisting of E, N, Q, R, T, V, and M. (Item 35) 34. The I-OnuI HE variant according to item 33, wherein the amino acid substituted for K108 is M. (Item 36) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for K156 is selected from the group consisting of N, Q, R, T, V, I, and E. (Item 37) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted with K156 is selected from the group consisting of I and E. (Item 38) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for S176 is selected from the group consisting of P, N and A. (Item 39) 34. The I-OnuI HE variant according to item 33, wherein the amino acid substituted for S176 is A. (Item 40) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for E231 is selected from the group consisting of D, K, V, and G. (Item 41) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for E231 is selected from the group consisting of K and G. (Item 42) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for V261 is selected from the group consisting of D, G, I, L, S, T, and A. (Item 43) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted at V261 is A. (Item 44) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for E277 is selected from the group consisting of A, D, G, Q, V, and K. (Item 45) 34. The I-OnuI HE variant according to item 33, wherein the amino acid substituted for E277 is K. (Item 46) 34. The I-OnuI HE variant of item 33, wherein the amino acid substituted for G300 is selected from the group consisting of S, V, D, C, and R. (Item 47) 34. The I-OnuI HE variant according to item 33, wherein the amino acid substituted with G300 is R. (Item 48) 17. The I-OnuI HE variant of item 15 or 16, wherein the one or more amino acid substitutions are at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 49) 49. The I-OnuI HE variant of item 48, wherein the amino acid substituted at N31 is selected from the group consisting of D, H, I, R, K, S, T and Y. (Item 50) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted at N31 is K. (Item 51) 49. The I-OnuI HE variant of item 48, wherein the amino acid substituted at N33 is selected from the group consisting of D, G, H, I, K, S, T and Y. (Item 52) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted at N33 is K. (Item 53) 49. The I-OnuI HE variant of item 48, wherein the amino acid substituted for K52 is selected from the group consisting of Q, R, T, Y, N, E, and M. (Item 54) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted with K52 is M. (Item 55) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted with Y97 is selected from the group consisting of F, N and H. (Item 56) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for Y97 is F. (Item 57) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for K124 is selected from the group consisting of E, N, R and T. (Item 58) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for K124 is N. (Item 59) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for K147 is selected from the group consisting of E, I, N, R and T. (Item 60) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for K147 is I. (Item 61) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for I153 is selected from the group consisting of D, H, K, T, Y, S, V and N. (Item 62) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for I153 is N. (Item 63) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for K209 is selected from the group consisting of E, M, N, Q and R. (Item 64) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for K209 is R. (Item 65) 49. The I-OnuI HE variant of item 48, wherein the amino acid substituted for E264 is selected from the group consisting of A, D, G, K, Q, R and V. (Item 66) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for E264 is K. (Item 67) 49. The I-OnuI HE variant of item 48, wherein the amino acid substituted for D268 is selected from the group consisting of A, E, G, H, N, V ​​and Y. (Item 68) 49. The I-OnuI HE variant according to item 48, wherein the amino acid substituted for D268 is N. (Item 69) 69. The I-OnuI HE variant according to any one of items 15 to 68, wherein the I-OnuI HE variant comprises three or more amino acid substitutions. (Item 70) 69. The I-OnuI HE variant according to any one of items 15 to 68, wherein the I-OnuI HE variant comprises five or more amino acid substitutions. (Item 71) 17. The I-OnuI HE variant of item 15 or 16, wherein the I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300. (Item 72) 17. The I-OnuI HE variant of item 15 or 16, wherein the I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 73) 17. The I-OnuI HE variant of item 15 or 16, wherein the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300. (Item 74) 17. The I-OnuI HE variant of item 15 or item 16, wherein the I-OnuI HE variant contains five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid positions are at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 75) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 10°C higher than TM 5075. The I-OnuI HE variant according to any one of items 15 to 74, having the following structure: (Item 76) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 15°C higher than TM 50 75. The I-OnuI HE variant according to any one of items 15 to 74, having the following structure: (Item 77) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 20°C higher than TM 50 75. The I-OnuI HE variant according to any one of items 15 to 74, having the following structure: (Item 78) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 25°C higher than TM 50 75. The I-OnuI HE variant according to any one of items 15 to 74, having the following structure: (Item 79) 79. The I-OnuI HE variant of any one of items 15 to 78, wherein the I-OnuI HE variant targets a site of a gene selected from the group consisting of HBA, HBB, HBG1, HBG2, BCL11A, PCSK9, TCRA, TCRB, B2M, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, CIITA, AHR, PD-1, CTLA4, TIGIT, TGFBR2, LAG-3, TIM-3, BTLA, IL4R, IL6R, CXCR1, CXCR2, IL10R, IL13Rα2, TRAILR1, RCAS1R, and FAS. (Item 80) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human BCL11A gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 81) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human PCSK9 gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 82) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human PDCD-1 gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 83) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human TCR alpha gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 84) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human CBLB gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 85) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human CTLA-4 gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 86) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human TGFβRII gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 87) An I-OnuI homing endonuclease (HE) variant that cleaves a target site in the human TIM3 gene, comprising one or more amino acid substitutions relative to a parent I-OnuI HE sequence, wherein the one or more amino acid substitutions increase the thermal stability of the I-OnuI HE variant compared to the parent I-OnuI HE variant. (Item 88) 88. The I-OnuI HE variant of any one of items 80 to 87, wherein the I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300. (Item 89) 88. The I-OnuI HE variant of any one of items 80 to 87, wherein the I-OnuI HE variant comprises amino acid substitutions at amino acid positions I14, A19, F168, D208, and N246. (Item 90) 88. The I-OnuI HE variant of any one of items 80 to 87, wherein the I-OnuI HE variant comprises three or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 91) 88. The I-OnuI HE variant of any one of items 80 to 87, wherein the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300. (Item 92) 88. The I-OnuI HE variant of any one of items 80 to 87, wherein the I-OnuI HE variant comprises five or more amino acid substitutions at amino acid positions selected from the group consisting of I14, A19, K108, V116, K156, F168, S176, D208, E231, N246, V261, L263, E277, and G300, and one or more amino acid substitutions at amino acid positions selected from the group consisting of N31, N33, K52, Y97, K124, K147, I153, K209, E264, and D268. (Item 93) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 10°C higher than TM 50 93. The I-OnuI HE variant according to any one of items 80 to 92, having the following structure: (Item 94) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 15°C higher than TM 50 93. The I-OnuI HE variant according to any one of items 80 to 92, having the following structure: (Item 95) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 20°C higher than TM 50 93. The I-OnuI HE variant according to any one of items 80 to 92, having the following structure: (Item 96) The I-OnuI HE variant is a TM of the parent I-OnuI HE 50 At least 25°C higher than TM 5093. The I-OnuI HE variant according to any one of items 80 to 92, having the following structure:

Claims

[Claim 1] The invention described in the specification.