Tale protein scaffolds involving fusions of monopartite and bipartite nls

By integrating monopartite and bipartite NLS into TALE fusion proteins, the specificity and activity of these proteins are enhanced, addressing safety and efficacy concerns in gene therapy.

WO2026046724A1PCT designated stage Publication Date: 2026-03-05CELLECTIS SA
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/073157
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-30
Filing Date
2025-08-12
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing TALE fusion proteins used in gene therapy do not meet the required levels of specificity and efficiency, often resulting in off-target binding and reduced catalytic activity, which poses safety concerns for therapeutic applications.

Method used

Incorporation of specific combinations of monopartite and bipartite nuclear localization signals (NLS), such as SV40 T antigen, C-myc, and Nucleoplasmin NLS, into the N-terminus of TALE fusion proteins to enhance their specificity and activity.

Benefits of technology

The modified TALE fusion proteins demonstrate improved on-target activity and reduced off-target binding, making them safer and more effective for genetic modifications in mammalian cells, particularly for gene therapy applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025073157_05032026_PF_FP_ABST
    Figure EP2025073157_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the design of improved TALE protein fusions useful as sequence-specific genomic reagents, such as TALE-nucleases and TALE base editors, comprising series of nucleus localization signals (NLS), especially a fusion of at least two monopartite NLS, such as from C-myc (C-myc NLS) and / or from SV40 T antigen (SV40 T NLS), and at least one bi-partite NLS, such as from Nucleoplasmin (Nucleoplasmin NLS). The goal of these fusions is to produce safer TALE reagents to genetically modify genomes and / or its organisation in different types of cells for their potential use in gene therapy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] IMPROVED TALE PROTEIN SCAFFOLDS INVOLVING FUSIONS OF MONOPARTITE AND BIPARTITE NLS

[0002] Field of the invention

[0003] The present invention relates to improved TALE fusion proteins involving a combination of monopartite and bipartite nuclear localization signals (NLS) to display better activity and specificity when they are expressed into eucaryotic cells. The resulting reagents are more convenient and safer to engineer the genomes of different types of cells, especially mammalian cells, and can be used in gene therapy.

[0004] Background of the invention

[0005] Artificial transcription-activator-like effectors (TALE) form a special class of proteins that can bind DNA originally derived from the phytopathogenic bacterial genus Xanthomonas [Kay S. et al. (2007) A bacterial effector acts as a plant transcription factor and induces a cell size regulator. Science 318: 648-651], Artificial TALE proteins have emerged to be versatile and sequence specific gene tools offering flexible applications upon elucidation of a DNA recognition ‘code’, linking the amino-acid sequence of the TALE with its bound genomic DNA sequence [Moscou J.M. et al. (2009) A Simple Cipher Governs DNA Recognition by TAL Effectors. Science. 326:1501],

[0006] TALE binding is driven by a series of 33 to 35 amino-acid-long repeats that differ at essentially two positions, the so-called repeat variable dipeptide (RVD). Each base of one strand in the DNA target is contacted by a single repeat, with predictable specificity resulting from the linear arrangement of RVDs. The biochemical structure-function studies suggest that the amino acid present at position 13 uniquely identifies a nucleotide on the DNA target major groove [Deng D., et al. (2012) Structural basis for sequence-specific recognition of DNA by TAL effectors. Science 335:720-723; Stella S., et al. (2013) Structure of the AvrBs3-DNA complex provides new insights into the initial thymine-recognition mechanism. Acta Crystallogr Sect D Biol Crystallogr 69(9):1707-1716], This DNA-protein interaction unit is stabilized by the amino acid at position 12. For the creation of TALEs with variable precision and binding affinity, six conventional RVDs are generally used (NG, HD, Nl, NK, NH, and NN). HD and NG are associated with cytosine (C) and thymine (T) respectively. NN is a degenerate RVD showing binding affinity for both guanine (G) and adenine (A), but its specificity for guanine is reported to be stronger. RVD Nl binds with A and NK binds with G. It is worth noting that the binding affinity of TALE is influenced by the methylation status of the target DNA sequence [Streubel J, et al. (2012) TAL effector RVD specificities and efficiencies. Nat Biotechnol 30(7):593-595.]. Methylated cytosine is not efficiently bound by the canonical RVDs. However, they can be accommodated by a certain degree of degeneracy in TALEs as described by Valton J, et al. [Overcoming transcription activator- 1 ike effector (TALE) DNA binding domain sensitivity to cytosine methylation (2012) J. Biol. Chem. 287(46): 38427-38432], This code was adopted to effectively engineer TALE DNA-binding scaffold specificity via modular assembly in order to form different associations of TALE proteins with various enzymatic domains, such as transcriptional activators, repressors, base editors or nucleases with potential ability to act on genomic sequences [Voytas et al. (2011) TAL effectors: Customizable proteins for DNA targeting. Science 333(6051): 1843-6], In comparison to Zine- Finger protein fusions, TALE-proteins have significantly emerged as critical DNA-binding scaffolds governed by a simple cipher without significant restrictions. Their compatibility with a broad range of epigenetic modifiers is commendable [Laufer B.I., et al. (2015) Strategies for precision modulation of gene expression by epigenome editing: an overview. Epigenetics Chromatin 8(1):34.] and it is considered that, with these DNA-binding proteins, it is possible to target an epigenetic effector domain to any locus in the genome [Cano-Rodriguez D., Rots M.G. (2016) Epigenetic editing: on the verge of reprogramming gene expression at will. Curr Genet Med Rep 4(4):170-179.].

[0007] Natural TAL effectors produced by bacteria are injected into plant cells via their type III secretion system to act as transcriptional regulators in the plant nucleus. They originally comprise their own nuclear localization domains into their extended C-terminal DNA binding region at around positions 750-800. However, artificial TALE fusions proteins used for gene editing have been made more compact by truncation of most of the C-terminal region and the original nuclear localization domain has been replaced by a unique SV40 monopartite NLS moved to the N- terminus of the TALE fusion proteins.

[0008] For instance, TALEN monomer constructs, which are generally based on truncated version of the TALE binding domain from the AvrBs3 protein fused to the catalytic domain of Fok1 , such as initially described by Voytas et al. in WO2011072246, referred to herein as “canonical”, are typically encoded by polynucleotides comprising from 5’ to 3’: (1) SV40 NLS, (2) truncated N- terminal region from AvrBs3 comprising at least the 150 amino acids that are proximal to the binding domain; (3) an engineered central DNA-binding domain, which generally comprises between 12 to 28 repeats that are assembled to target a genomic nucleotide sequence; (4) a wild type half repeat of about 20 amino acids from AvrBs3 designed to bind the 3'-end of the targeted DNA sequence; (5) an optional linker sequence of preferably from about 10 to about 80 amino acids, or preferably of at least 40 amino acids from the C-terminal wild type region of AvrBs3, fused to (6) the wild type Fok1 nuclease catalytic domain.

[0009] Enhancements to the core TALE domain via various truncations have been proposed in several studies [Miller, J. C. et al. (2011) A TALE nuclease architecture for efficient genome editing. Nat. Biotechnol. 29, 143] along with the use of additional or alternative RVDs, which have shown to improve specificity and efficacy of these programmable TALE DNA-binding domain [Juillerat A, et al. (2015) Optimized tuning of TALEN specificity using non-conventional RVDs. Sci Rep 5(1)] and also in WO2023094435 the use of particular RVDs and C-terminal mutated truncations have been shown to lower off-target genetic modifications.

[0010] As reviewed by Becker, S. and Boch, J. [TALE and TALEN genome editing technologies, Gene and Genome editing (2021) 2:100007] TALE protein fusions may result in extensive variety of genome engineering tools, such as TALE-nucleases, TALE-base editors, TALE transposases, TALE transcription regulators, TALE recombinases, TALE transcriptome modifiers, TALE-based epigenomic modifiers.

[0011] For instance, TALE nucleases can be generated by the fusion of TALE with various nuclease catalytic domains. The popularly used TALEN® system, which provides specific nucleases as a fusion of TALE scaffolds with the catalytic domain of the Fok1 restriction enzyme has proven to be very specific through many studies, as it combines two TALE dimers that bind together at the selected locus. The TALEN heterodimers (right and left) generally bind on opposite strands at about 10-20 pb away from each other (spacer) to allow the nuclease Fok1 to dimerize and induce double strands cleavage between the binding sites within the spacer. This heterodimeric setting allows an increased sequence specificity based on the extended target sequence encompassed by the two TALE binding sites that can span up to 40 base pairs. Such TALE-nucleases are currently developed as therapeutic grade nuclease reagents in gene therapy, especially to produce allogeneic CAR-T cells [Poirot et al. (2015) Multiplex Genome- Edited T-cell Manufacturing Platform for “Off-the-Shelf’ Adoptive T-cell Immunotherapies Cancer Res 75(18):3853-3864; Quasim W. et al. (2017) Molecular remission of infant B-ALL after infusion of universal TALEN gene-edited CAR T cells. Science translational medicine (9)374],

[0012] TALE Artificial transcription factors, which have been generated by the fusion of TALE with a 16 amino acid peptide (VP16) from herpes simplex virus as a transactivation domain [Zhang, F. et al. Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. Nature Biotechnol. 29:149-153], By contrast to zinc-fingers binding domains, which have encountered many off-target effects, TALE transcriptional activators are efficient transcription modulators with only 10.5 repeats with an effector module fused to the carboxyl terminal [Miller, J., et al. (2011) A TALE nuclease architecture for efficient genome editing. Nat Biotechnol. 29, 143-148], TALEs in the form of activators can also be used to control the gene expression in case of external stimuli like a chemical change, or optical stimulus in various organisms including plants and animals.

[0013] TALE repressors can be generated by the fusion of TALE with either Kruppel-associated box (KRAB), Sid4, or EAR-repression domain (SRDX) repressors [Cong L, et al. (2012) Comprehensive interrogation of natural TALE DNA-binding modules and transcriptional repressor domains. Nat Commun 3(1):968],

[0014] TALE base editors can be generated by the fusion of TALE with deaminase, and sometimes, to other DNA repair proteins. Base editor catalytic domains can introduce singlenucleotide variants at desired loci in DNA (nuclear or organellar) or RNA of both dividing and nondividing cells. Broadly, there are two types of DNA base editors that directly induce targeted point mutations in DNA, and RNA base editors that convert one ribonucleotide to another in RNA. Currently available DNA base editors can be further categorized into cytosine base editors (CBEs), adenine base editors (ABEs), C-to-G base editors (CGBEs), dual-base editors and organellar base editors. For instance, Mok et al. [A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing (2020) Nature. 583:631-637] recently developed a base editing approach using the bacterial cytidine deaminase toxin, DddAtox, to demonstrate efficient C-to-T base conversions in vitro. In this approach, split DddAtox nontoxic halves fused to transcription activator-like effector (TALE) proteins, which can be custom-designed to recognize predetermined target DNA sequences, form a functional cytosine deaminase within the editing window to induce C-to-T base editing at the target site in genomic DNA. Such DddA-TALE fusion deaminase constructs have since achieved mitochondrial DNA editing in mice [Lee, H., et al. (2021) Mitochondrial DNA editing in mice with DddA-TALE fusion deaminases. Nat Commun 12: 1190],

[0015] Such bespoke TALE proteins have proven to be robust reagents for targeting genomic DNA sequences of interest in almost every cell types [Weeks D.P; et al. Use of designer nucleases for targeted gene and genome editing in plants (2016) Plant Biotechnology Journal.14:483-495; Mussolino C. et al. (2014) TALENs facilitate targeted genome editing in human cells with high specificity and low cytotoxicity. Nucleic Acids Res 42(10):6762-6773],

[0016] Nevertheless, with the development of TALE-nucleases and TALE-base editors for human gene therapy, standard TALE constructs do not always meet the specificity and efficiency levels required for therapeutic safety. Depending on the sequences to be targeted in the genome and their intrinsic variability in human populations, TALE scaffolds sometimes need further refinements to reduce potential off-target binding and increase their catalytic activity. In fact, specificity and catalytic activity are often in balance and it may be difficult to find a good compromise that preserves safety and efficiency.

[0017] Although NLS had been considered so far as a minor aspect of TALE constructs, the present invention is based on the unexpected finding that certain combinations of monopartite and bipartite nuclear localization signals (NLS) have a significant impact on the specificity and / or catalytic activity of the TALE fusion proteins.

[0018] Based on these considerations, the inventors have designed new TALE scaffolds that are advantageous for genetic modifications in mammalian cells, while retaining most of their catalytic activities, and remain adaptable to any target sequence and RVD adjustment.

[0019] Their invention can be combined with the previous approaches prevailing in the art for rational designing TALE catalytic proteins of higher therapeutic grade.

[0020] Summary of the invention

[0021] The present invention aims at improving the specificity and / or activity of existing or new TALE fusion proteins by inclusion of combinations of monopartite and bipartite nuclear localization signals (NLS) with the effect of improving the overall activity and specificity of said proteins.

[0022] The present invention thus relates to the design of improved TALE protein fusions useful as sequence-specific genomic reagents, such as TALE-nucleases and TALE base editors, comprising series of nucleus localization signals (NLS), especially a fusion of at least two monopartite NLS, such as from C-myc (C-myc NLS) (SEQ ID NO:1) and / or from SV40 T antigen (SV40 T NLS) (SEQ ID NO:2), and at least one bi-partite NLS, such as from Nucleoplasmin (Nucleoplasmin NLS) (SEQ ID NO:3). The goal of these fusions is to produce safer TALE reagents to genetically modify the genomes and / or its organisation in different types of cells for their potential use in gene therapy.

[0023] These monopartite and bi-partite NLS are preferably fused together and more preferably located at the N-terminus region of the TALE fusion protein.

[0024] In some embodiments, the TALE fusion protein comprises three NLS, two monopartite ones and at least one monopartite. In some embodiments, the TALE fusion protein comprise a combination of three NLS in the order monopartite / bi-partite / monopartite, with a preference for the combination SV40 T NLS / Nucleoplasmin NLS / C-myc NLS referred to in the experimental section as S-N-C or N-S-C, which has preferably at least 80%, 90%, 95% or 99% identity with SEQ ID NO:11 or SEQ ID NO: 12.

[0025] In some embodiments, the TALE fusion protein is a TALE repressor including a functional domain comprises a recruiting domain, such as a VP64 or Krabb.

[0026] In some embodiments, the TALE fusion protein of the present invention includes a catalytic domain and in some cases a combination of catalytic domain(s) and / or recruiting domain(s). In general, the catalytic domain aims at modifying in-situ the genome and / or its organisation, such as by mutating, cleaving, nicking, or introducing epigenetic changes. Such catalytic domain may come from an enzyme selected from a methyltransferase, demethylase, exonuclease, endonuclease, nickase, deaminase, polymerase, transposase, integrase, recombinase, histone methyltransferase / acetyl transferase, reverse transcriptase or helicase.

[0027] The resulting TALE fusion proteins according to the invention are useful to engineer the genome and / or its organisation, such as by mutating, cleaving, nicking, or introducing epigenetic changes.

[0028] For instance, the TALE fusion protein can be a TALE-nuclease, a TALE-base editor, a TALE-transcriptional modulator, a TALE-integrase, a TALE-transposase or a TALE-recombinase.

[0029] The C-terminal region of the TALE fusion protein generally comprises from about 10 to 80 amino acids. According to preferred embodiments, the C-terminal region comprises about 10 to 40 amino acids, and preferably about 10 amino acids, such as that comprising SEQ ID NO:282 (C11-AA).

[0030] In some embodiments, the C-terminal region of a TALE fusion protein according to the invention preferably consists of a polypeptide sequence from 40 to 80 residues comprising a sequence having at least 85% identity with one of the following polypeptide sequences:

[0031] SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GL

[0032] (SEQ ID NO:13)

[0033] SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVXIX2GLPHAPALIX3RT

[0034] (SEQ ID NO:14), or

[0035] SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GLPHAPALIX3RTNRRIPERTSH (SEQ ID NO:15), wherein Xi, X2, and X3, are any amino acid, preferably H (histidine) or a R (arginine) residue.

[0036] In some embodiments, the C-terminal region is not necessary, and the TALE binding domain can be directly fused with the functional domain, which is generally a catalytic domain from a known enzyme.

[0037] In some embodiments, the TALE fusion protein comprises at least one or a majority of TALE repeat(s) comprising D (aspartic acid) residues at positions 4 and 32 with respect to any of the canonical sequence of AvrBs3 of SEQ ID NO: 16 to SEQ ID NO:22.

[0038] In some embodiments, the TALE fusion protein according to the invention comprises at least one, or a majority of TALE repeat(s) represented by one the following sequences:

[0039] LTPQQVVAIASX4X5GGKQALETVQRLLPVLCQAHG (SEQ ID NO: 16),

[0040] LTPEQVVAIASX4X5GGKQALETVQALLPVLCQAHG (SEQ ID NO: 17),

[0041] LTPDQWAIASX4X5GGKQALETVQQLLPVLCQDHG (SEQ ID NO: 18),

[0042] LTPDQLVAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NO:19),

[0043] LTPDQMVAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NQ:20),

[0044] LTPDQWAIASX4X5GGKQALETVQRLLPVLCQDQG (SEQ ID NO:21), or

[0045] LTLDQWAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NO:22), wherein X4X5 is an amino acid forming a variable di-residue.

[0046] In preferred embodiments, the TALE fusion protein combines such TALE repeats and at least one C-terminal region of SEQ ID NO:282, and preferably SEQ ID NO:13, SEQ ID NO:14 or SEQ ID NO:15.

[0047] The present application provides a large number of representative TALE fusion proteins according to the invention including in the experimental section, in particular in Table 4 as well as their equivalents that have display at least 80%, preferably at least 85%, more preferably at least 90%, even more preferably at least 95% or 99% identity with one of their polypeptide sequence or polynucleotide sequence encoding thereof.

[0048] The TALE fusion protein according to the present invention are useful over a large array of genetic applications ranging from gene therapy to the production of engineered plant cells. In preferred embodiments the TALE fusion protein is used to target a nuclear genomic sequence, preferably expressed in human NK or T-cells, selected from one encoding TCR, B2M, CIITA, CD52, GR, CCR5, CS1 / SLAMF7, CD38, TGFBRII, GMSCF, HLA-A, HLA-B, HLA-C, PD1 , CTLA4, LAG3, TIGIT, CISH, IgH, GADPH, HVACR2, CD3E, CDKN2A, CDKN2B, MTAP, which are referred to in the literature as being relevant to the production of CAR immune cells, the invention being not limited to these examples. In preferred embodiments the TALE fusion protein is used to target a nuclear genomic sequence, preferably expressed in human hepatocytes such as ALB or APOC3.

[0049] In most preferred embodiments the TALE fusion protein according to the invention are TALE-nucleases used to inactivate one or several genes involved into resistance of cells to senescence, such as CDKN2A, CDKN2B or MTAP, especially in primary T-cells or NK cells for their subsequence use in cell therapy.

[0050] In further preferred embodiments the TALE fusion protein according to the invention are TALE-nucleases or TALE-base editors used to inactivate genes such as B2M, CIITA or TRAC to improve the persistence of allogeneic immune cells in patients, such as T-cells or NK cells for their use in cell therapy.

[0051] In further preferred embodiments the TALE fusion protein according to the invention are TALE-nucleases or TALE-base editors used to inactivate genes such as ApoC3 and ALB (Albumin) in patients or ex vivo in patients’ hepatocytes, and optionally to insert an exogenous nucleic acid sequence at these loci, such as a transgene of therapeutic interest, for treating said patient. The invention encompasses vectors comprising the polynucleotide sequences as well as the polypeptide sequences or reagents obtainable by the present invention, as well as their use for cell transformation and gene modification.

[0052] Description of figures and tables

[0053] Figure 1 : Fold increase of indels induced by TALEN targeting B2M, CD3E or CS1 loci and comprising different 3xNLS combinations (S-N-C, S-C-N, N-S-C, N-C-S, C-S-N, C-N-S) relative to classical TALEN (comprising a single SV40 NLS).

[0054] Figure 2: Fold increase of indels induced by 10 TALEN targeting 10 different loci in the HVACR2 gene and comprising S-N-C or N-S-C combinations relative to classical TALEN (comprising a single SV40 NLS). Figure 3: Indels frequencies induced upon no (Mock) or transfection with TALEN fusion proteins comprising the S-N-C combination at the 4 indicated loci (CDKN2A isoforms, MTAP and CDKN2B) making them useful to make immune cells resistant to senescence.

[0055] Figure 4: Structure of an illustrative TALE-nuclease protein fusion as per the present invention.

[0056] Figure 5: Percentage of indels induced after transfection of primary T-cells, or not (Mock), with mRNA encoding TALEN heterodimer comprising 3 NLS (S-N-C combination) on one arm, targeting TRAC locus (TALEN treated), as illustrated in Example 5.

[0057] Figure 6: Oligo Capture Assay (OCA) score calculated at the on-site as well as at the twenty first potential off-target sites of the TALEN targeting TRAC locus upon transfection with mRNA encoding a TALEN heterodimer comprising 3 NLS (SNC combination) on one arm (TALEN treated), as illustrated in Example 5.

[0058] Figure 7: Percentage of indels induced after transfection of primary T-cells, or not (Mock), with mRNA encoding TALEN heterodimer comprising a SNC combination on both arms, targeting B2M locus, as illustrated in Example 5.

[0059] Figure 8: Oligo Capture Assay (OCA) score calculated at the on-site as well as at the twenty first potential off-target sites of the TALEN targeting B2M locus upon transfection with mRNA encoding a TALEN heterodimer comprising 3 NLS (S-N-C combination) on both arms, as illustrated in Example 5.

[0060] Figure 9: Percentage of indels induced after transfection, or not (Mock), of HEP-G2 cells with mRNA encoding a TALEN heterodimer comprising a 3xNLS (S-N-C combination) on one arm, targeting ALB locus (TALEN treated) as illustrated in Example 6.

[0061] Figure 10: Oligo Capture Assay (OCA) score calculated at the on-site as well as at the twenty first potential off-target sites of the TALEN targeting ALB locus after transfection with mRNA encoding a TALEN heterodimer comprising 3 NLS (S-N-C combination) on one arm, as illustrated in Example 6.

[0062] Figure 11 : Percentage of indels at the on-site, and potential off-sites OT-2, OT-4, OT-6 loci (normalized to the on-site) upon transfection with mRNA encoding TALEN heterodimer comprising either a single SV40 NLS on both arms or a 3xNLS (SNC combination) on one arm, targeting the ALB locus as illustrated in Example 6. Figure 12: Onsite activity of 3NLS CIITA TALEN in T-cells as illustrated in Example 7:

[0063] A. Schematic outlining transfection of T cells with CIITA TALEN and subsequent analysis for genomic knockout and functional downregulation of activity. B Percentage of indels induced after transfection, or not (Mock) with TALEN targeting CIITA locus comprising an S-N-C combination on one arm (TALEN treated).

[0064] Figure 13: Representative flow cytometry plots showing MHCII expression on T-cells post transfection with CIITA TALEN as illustrated in Example 7, accompanied by graphical representation of MHCII (-) T cell population in mock or different CIITA TALEN transfected groups.

[0065] Figure 14: Loss of MHCII from T-cell surface prevents their depletion by alloreactive T cells shown in Example 7: A. Schematic illustrating the generation of alloresponsive T-cells and subsequent Mix lymphocyte reaction (MLR) performed with CIITA knock out T-cells.

[0066] B. Representative flow cytometry plots showing the evolution of TCRaP(+); MHCII(+ / -) T-cell targets when co-cultured in the presence of allogeneic CD4+ T-cells for 24 hours. Cells are pregated on single, viable Cell Trace violet positive T cells. C. Box plot representing MHCII(-) enrichment when engineered TCRaP(+); MHCII(+ / -) T-cell targets were co-cultured with alloreactive T-cells at different ratios (3 technical replicates, with 3 donors for T-cell effector and 3 donors for engineered T-cell targets).

[0067] Table 1 : Example of monopartite and bipartite NLS polypeptide sequences.

[0068] Table 2: Example of linkers that may be included in the TALE fusion proteins.

[0069] Table 3: Example of catalytic domains fused to the TALE proteins as per the present invention.

[0070] Table 4: Examples of TALE fusion proteins according to the present invention useful in gene therapy or adoptive immune cells therapy as illustrated in the examples.

[0071] Table 5: Polynucleotide sequences used in the Examples.

[0072] Table 6: Results obtained in connection with the experimental results detailed in Example 5 regarding on-site and the five first off-target sites sequencing after transfection of mRNA in T- cells encoding TRAC TALEN heterodimer comprising 3xNLS on one arm.

[0073] Table 7: Results obtained in connection with the experimental results detailed in Example 5 regarding on-site and the first five off-target sites sequencing after transfection of mRNA in T-cells encoding B2M TALEN heterodimer comprising 3xNLS on both arms. Detailed description of the invention:

[0074] Unless specifically defined herein, all technical and scientific terms used have the same meaning as commonly understood by a skilled artisan in the fields of gene therapy, biochemistry, genetics, and molecular biology.

[0075] All methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, with suitable methods and materials being described herein. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will prevail. Further, the materials, methods, and examples are illustrative only and are not intended to be limiting, unless otherwise specified.

[0076] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and immunology, which are within the skill of the art. Such techniques are explained fully in the literature. See, for example, Current Protocols in Molecular Biology [Frederick M. AUSUBEL, 2000, Wiley and son Inc, Library of Congress, USA); Molecular Cloning: A Laboratory Manual, Third Edition, (Sambrook et al, 2001 , Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press; Oligonucleotide Synthesis (M. J. Gait ed., 1984); Mullis et al. U.S. Pat. No. 4,683,195; Nucleic Acid Hybridization (B. D. Harries & S. J. Higgins eds. 1984); Transcription And Translation (B. D. Hames & S. J. Higgins eds. 1984); Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987); Immobilized Cells And Enzymes (IRL Press, 1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); the series, Methods In ENZYMOLOGY (J. Abelson and M. Simon, eds. -in-chief, Academic Press, Inc., New York), specifically, Vols.154 and 155 (Wu et al. eds.) and Vol. 185, "Gene Expression Technology" (D. Goeddel, ed.); Gene Transfer Vectors For Mammalian Cells (J. H. Miller and M. P. Calos eds., 1987, Cold Spring Harbor Laboratory); Immunochemical Methods In Cell And Molecular Biology (Mayer and Walker, eds., Academic Press, London, 1987); Handbook Of Experimental Immunology, Volumes l-IV (D. M. Weir and C. C. Blackwell, eds., 1986); and Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1986],

[0077] The present invention has thus for object methods to design and produce TALE proteins that display reduced off-target DNA binding, which can be fused to various catalytic domains in view of forming highly specific and active TALE fusion proteins, in particular TALE-nucleases and TALE-base editors.

[0078] Several websites and tools assist scientists with the proper positioning and the specific design of TALE scaffolds. Many of these tools offer a variety of different functions, but their main purpose is the suggestion of possible TALE binding sites in a particular DNA sequence, the selection of suitable RVDs and, in some cases, the prediction of unwanted off-targets. To name a few of them: TALE-NT, idTALE, PROGNOS, TALENoffer, E-TALEN, Mojo Hand and ChopChop are well known by one skilled in the art and can be used in conjunction with the present invention.

[0079] Meanwhile, it is easily possible to design TALE fusion proteins without the need for bioinformatic tools - even for skilled in the art who are new to the field, by following the following basic guidelines [Becker, S. and Boch, J. (2021) TALE and TALEN genome editing technologies, Gene and Genome editing 2:100007], although they might not allow to utilize their full flexibility:

[0080] 1) A TALE monomer is generally composed of 11.5 to 17.5 repeats, preferably

[0081] 2) The nucleotide preceding the region specified by the RVDs is generally a T (TO).

[0082] 3) If a TO is not available, other nucleotides at position zero can be used when, for instance: a) a modified N-terminal region is used, generally by deleting the first 152 amino acids, or b) NS is usually the first repeat.

[0083] 4) Use preferentially HD, NN, NG and Nl for the binding of C, G, T and A, respectively.

[0084] 5) A TALE core binding domain generally contains at least 3 strong RVDs (HD or NN).

[0085] 6) better avoid long stretches of A or T in the beginning or end of a target sequence.

[0086] In this context, a monopartite Nuclear Localization Signal (NLS) has been generally added to the TALE monomers at the N-terminus of the modified N-terminal region of the TALE monomer and mostly or from SV40 T antigen (SV40 T NLS).

[0087] Nuclear Localization Signals (NLS) are generally short peptides that act as a signal fragment that mediates the transport of proteins from the cytoplasm into the nucleus. This NLS- dependent protein recognition, a process necessary for cargo proteins to pass the nuclear envelope through the nuclear pore complex, is facilitated by members of the importin superfamily.

[0088] The NLS encompass two categories, termed “monopartite” (MP) and “bipartite” (BP) as defined by Bradley, K. et al. [Parafibromin is a nuclear protein with a functional monopartite nuclear localization signal (2007) Oncogene. 26, 1213-1221],

[0089] MP NLS are a single cluster composed of 4-8 basic amino acids, which generally contains 4 or more positively charged residues, that is, arginine (R) or lysine (K). The characteristic motif of MP NLS is usually defined as K (K / R) X (K / R), where X can be any residue. For example, the NLS of SV40 large T-antigen is126PKKKRKV132, with five consecutive positively charged amino acids (KKKRK). By contrast, BP NLS are characterized by two clusters of 2-3 positively charged amino acids that are separated by a 9-12 amino-acid linker region, which contains several proline (P) residues. The consensus sequence can be expressed as R / K(X)10-12KRXK. Notably, in BP NLS, the upstream and downstream clusters of amino acids are interdependent and indispensable, and jointly determine the localization of the protein in the cell.

[0090] For instance, the BP NLS at the C-terminus of nucleoplasmin, whose sequence is generally155KRPAATKKAGQAKKKK170, can guide the protein into the nucleus. In addition to nucleoplasmin, 53BP1 (TP53-binding protein 1) also has a classical BP NLS with the sequence1666GKRKLITSEEERSPAKRGRKS1686. Its upstream (1667KRK1669) and downstream (1681 KRGRK1685) clusters are generally used for proper localization of 53BP1 and maintenance of genomic integrity. ING4 contains the potential BP NLS128KGKKGRTQKEKKAARARSKGKN149, among which142RARSK146mainly binds to p53 and mediates the nuclear localization of ING4 and p53. The extracellular signal regulated kinase 5 (ERK5) is known to contain a classical BP NLS. Examples of preferred NLS according to the present invention are reported in Table 1. Table 1 : Examples of monopartite and bipartite NLS sequences The present invention has thus for object transcriptional Activator-like Effector (TALE) fusion proteins that generally comprise (1) a core binding domain comprising TALE repeats, (2) N-terminal and (3) C-terminal regions, wherein said N-terminal region comprises series of nucleus localization signals (NLS). This series is generally a fusion of at least two monopartite NLS, such as from C-myc (C-myc NLS) and / or from SV40 T antigen (SV40 T NLS), and at least one bipartite NLS, such as from Nucleoplasmin (Nucleoplasmin NLS).

[0091] These monopartite and bi-partite NLS are preferably fused together and more preferably located at the N-terminus region of the TALE fusion protein.

[0092] In some embodiments, the TALE fusion protein comprises three NLS, two monopartite ones and at least one monopartite.

[0093] In some embodiments, the TALE fusion protein comprises a combination of three NLS in the order monopartite / bi-partite / monopartite, with a preference for the combination fusion SV40 T NLS / Nucleoplasmin NLS / C-myc NLS referred to in the experimental section as S-N-C, which has preferably at least 80 %, preferably 90%, 95% or 99% identity with SEQ ID NO:11. or N-S-C which has preferably at least 80 %, 90 %, 95% or 99% identity with SEQ ID NO: 12.

[0094] By “TALE fusion protein” is meant an engineered polypeptide comprising a TALE binding domain resulting from the assembly of polypeptides from different proteins.

[0095] In general, TALE fusions protein results from the heterologous expression of a polynucleotide, such as DNA or RNA transfected into living cells.

[0096] Such TALE fusion proteins of the present invention have their N-terminal, or more commonly their C-terminal region fused to (4) a functional domain that confers a particular function to the TALE fusion, in particular by modifying the genome and / or its organization, such as by mutating, cleaving, nicking, or introducing epigenetic changes.

[0097] By “TALE binding domain” is meant a core DNA binding domain comprising repeats, each of said repeats has the ability to match a given nucleic acid base into a target genetic sequence. In general, such core DNA binding domain has at least 50%, preferably at least 60%, 70%, 80% or 90% identity with the DNA binding domain of wild-type AvrBs3 [also called TalC Uniprot - G7TLQ9], which represents the archetype of the family of transcription activator-like (TAL) effectors from phytopathogenic Xanthomonas campestris. Such DNA binding domain is characterized by repeated sequences of about 30 and 34 amino acids comprising variable diresidues usually found in positions 12 and 13. A consensus sequence for these repeats, also called RVDs, has been established for each targeted base A, C, G and T, which are respectively:

[0098] LTPQQVVAIASNIGGKQALETVQRLLPVLCQQHG (SEQ ID NO:26) for targeting A; LTPQQVVAIASHDGGKQALETVQRLLPVLCQQHG (SEQ ID NO:27) for targeting C;

[0099] LTPQQVVAIASNNGGKQALETVQRLLPVLCQQHG (SEQ ID NO:28) for targeting G;

[0100] LTPQQVVAIASNGGGKQALETVQRLLPVLCQQHG (SEQ ID NO:29) for targeting T.

[0101] By “AvrBs3-like repeats” are meant artificial arrays of about 30 to 33 amino acids, which typically comprise variable di-residues in positions 12 and 13 interacting with A, C, G orT, similarly as the above consensus AvrBs3 repeats. In other words, AvrBs3-like repeats are similar and can be combined with AvrBs3 repeats, but are generally not identical to the consensus or to the wildtype AvrBs3 repeats. It shall be noted that, in some instances, di-residues in positions 12 or 13 may be absent - so-called * (star) - to accommodate methylated bases in genomic DNA as described by [Valton et al. (2012) Overcoming Transcription Activator-like Effector (TALE) DNA Binding Domain Sensitivity to Cytosine Methylation. DNA and Chromosomes. 287(46): 38427],

[0102] In some embodiments the TALE binding domain is fused to a functional domain that is act as a recruiting domain, such as a VP64 or Krabb, to activate transcriptional activity.

[0103] In some embodiments the TALE binding domain is fused to a functional domain that is a catalytic domain that has the ability to promote enzymatic reactions. Examples of catalytic domain are detailed in Table 3. They can confer the TALE fusion proteins a variety of enzymatic activities such as methyltransferase, demethylase, exonuclease, endonuclease, nickase, deaminase, polymerase, transposase, integrase, recombinase, histone methyltransferase / acetyl transferase, reverse transcriptase or helicase.

[0104] In some embodiments the TALE binding domain is fused to one or several functional domain(s) to form a combination of catalytic domain(s) and / or recruiting domain(s).

[0105] In some preferred embodiments, the (TALE) fusion protein according to the present invention modifies the genome and / or its organisation, such as by mutating, cleaving, nicking, or introducing epigenetic changes. Accordingly, Such TALE fusion proteins can be a TALE- nuclease, TALE-base editor, TALE-transcriptional modulator, TALE-integrase, TALE- transposase or a TALE-recombinase as described herein.

[0106] The AvrBs3-like repeats of the present invention generally display at least 60%, preferably at least 70%, 75%, 80%, 90% or 95% identity with either of the above AvrBs3 consensus repeats sequences of SEQ ID NO:26 to 29. They generally comprise D4 and D32 substitutions, such as in the following repeat sequences SEQ ID NO: 16 to 22 of the present invention:

[0107] LTPQQVVAIASX4X5GGKQALETVQRLLPVLCQAHG (SEQ ID NO: 16), LTPEQVVAIASX4X5GGKQALETVQALLPVLCQAHG (SEQ ID NO: 17), LTPDQWAIASX4X5GGKQALETVQQLLPVLCQDHG (SEQ ID NO: 18), LTPDQLVAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NO:19), LTPDQMVAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NQ:20), LTPDQWAIASX4X5GGKQALETVQRLLPVLCQDQG (SEQ ID NO:21), or LTLDQWAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NO:22), wherein X4X5 are the di-residues interacting with a given nucleotide base pair in the targeted sequence. X4 and X5 can be any amino acid or null (referred to as * (star) to designate a missing residue in the RVD). X4 and X5 can be identical or different.

[0108] The AvrBs3-like repeats are generally represented by polypeptide sequences, in which X4 and X5 are respectively Nl (to preferably target A), HD (to preferably target C), (to preferably target G) NN and NG (to preferably target T), such as in SEQ ID NO:26, 27, 28 and 29 referred to before.

[0109] "Identity" throughout the present specification refers to sequence identity between two nucleic acid molecules or polypeptides. Identity can be determined by comparing a position in each sequence which may be aligned for purposes of comparison. When a position in the compared sequence is occupied by the same base, then the molecules are identical at that position. A degree of similarity or identity between nucleic acid or amino acid sequences is a function of the number of identical or matching nucleotides at positions shared by the nucleic acid sequences. Various alignment algorithms and / or programs may be used to calculate the identity between two sequences, including FASTA, or BLAST which are available as a part of the GCG sequence analysis package (University of Wisconsin, Madison, Wis.), and can be used with, e.g., default setting. The present specification generally encompasses polypeptides and polynucleotides having at least 70%, 85%, 90%, 95%, 98% or 99% identity with the specific polypeptides and polynucleotides sequences described herein, exhibiting substantially the same functions or that can be considered as equivalents.

[0110] In some embodiments, the invention also provides a recombinant transcriptional activatorlike Effector (TALE) protein comprising one or several AvrBs3-like repeats comprising D (aspartic acid) residues at positions 4 and 32, such as in the above polynucleotide sequences SEQ ID NO: 16 to 22. Such AvrBs3-like repeats can be further mutated into 1 to 5 amino acid positions, including or in addition to the D4 and D32 positions. Such recombinant transcriptional activatorlike Effector (TALE) proteins can comprise one or several of such repeats, to form polypeptides comprising generally from 8 to 20 repeats, preferably from 8 to 18, more preferably from 10 to 16, and alternatively from 5 to 12 repeats in situations where smaller genomes are considered, such as for instance mitochondrial genomes.

[0111] The variable di-residues (X4X5) present in the AvrBs3-like repeats and associated with recognition of the different nucleotides are generally HD for recognizing C, NG for recognizing T, Nl for recognizing A, NN for recognizing G or A, NS for recognizing A, C, G or T, HG for recognizing T, IG for recognizing T, NK for recognizing G, HA for recognizing C, ND for recognizing C, HI for recognizing C, HN for recognizing G, NA for recognizing G, SN for recognizing G or A and YG for recognizing T, TL for recognizing A, VT for recognizing A or G and

[0112] SW for recognizing A. More preferably, RVDs associated with recognition of the nucleotides C, T, A, G / A and G respectively are selected from the group consisting of NN or NK for recognizing G, HD for recognizing C, NG for recognizing T and Nl for recognizing A, TL for recognizing A, VT for recognizing A or G and SW for recognizing A. More generally, RVDs associated with recognition of nucleotide C are selected from the group consisting of N*, RVDs associated with recognition of the nucleotide T are selected from the group consisting of N* and H*, where * may denote a gap in the repeat sequence that corresponds to a lack of amino acid residue at the second position of the RVD. In some embodiments, X4Xscan represent unusual or unconventional amino acid residues in order to modulate their specificity towards nucleotides A, T, C and G as described in J uillerat et al. [Optimized tuning of TALEN specificity using non-conventional RVDs (2015) Sci Rep 5:8150],

[0113] Although not mandatory, the core DNA binding domain generally comprises a half RVD made of 20 amino acids located at the C-terminus. Said core DNA binding domain thus comprises between 8.5 and 30.5 RVDs, more preferably between 8.5 and 20.5 RVDs, and even more preferably, between 10,5 and 15.5 RVDs.

[0114] As per the present invention, the core DNA binding domain as previously described, preferably comprising RVDs bearing D4 and / or D32 substitutions, is flanked by N-terminal and C- terminal sequences, said N-terminal and C-terminal sequences having preferably one of the following features detailed below.

[0115] In some embodiments, the N-terminal sequence is derived from the N-terminal domain of a naturally occurring TAL effector such as AvrBs3. In another embodiment, said additional N- terminus domain is the full-length N-terminus domain of a naturally occurring TAL effector N- terminus domain. In a further embodiment, said additional N-terminus domain is a variant which allows overcoming sequence constraints associated with the so-called “RVD0” (i.e. first cryptic repeat), such as for instance the necessity to have a T required as the first base on the binding nucleic acid sequence.

[0116] In another embodiment, said N-terminal sequence is derived from a naturally occurring TAL effector or a variant thereof. In another embodiment, said N-terminal sequence is a truncated N-terminus of such naturally occurring TAL effector or variant. In another embodiment, said additional domain is a truncated version of AvrBs3 TAL effector. In another embodiment, said truncated version lacks its N-terminal segment distal from the core TALE binding domain, such as the first 152 N-terminal amino acids residues of the wild type AvrBs3, or at least the 152 amino acids residues.

[0117] In some preferred embodiments, said N-terminal sequence comprises a polypeptide sequence showing at least 85%, preferably at least 90%, more preferably at least 95% identity with SEQ ID NO:30.

[0118] In some embodiments, the C-terminal sequence corresponds to a full or preferably truncated C-terminal region of a naturally occurring TAL effector such as AvrBs3. In general, said C-terminal sequence is a truncated version of AvrBs3 TAL effector C-terminal region, proximal to the core TALE binding domain, such as SEQ ID NO:13 (40 amino acids), SEQ ID NO: 14 (50 amino acids) or SEQ ID NO:15 (60 amino acids) or a natural variant thereof. Accordingly, said C-terminal sequence generally comprises or consists of a polypeptide sequence from 40 to 80 residues comprising a sequence having at least 85% identity with the below SEQ ID NO:13, SEQ ID NO:14 or SEQ ID NO:15:

[0119] - SEQ ID NO:13 (C-40 AA):

[0120] SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GL

[0121] - SEQ ID NO:14 (C-50 AA):

[0122] SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GLPHAPALIX3RT

[0123] - SEQ ID NO:15 (C-60 AA):

[0124] SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GLPHAPALIX3RTNRRIPERTSH In the above sequences, X1 , X2 and X3 represent an amino acid substitution introduced into the wild type AvrBs3 C-terminal polypeptide sequence, which is preferably R (arginine) or H (histidine) residue, most preferably R, instead of originally K. X1 , X2 and X3 can be identical or different.

[0125] In some preferred embodiments, the C-terminal sequence is a shorter truncated version of AvrBs3 TAL effector C-terminal region, proximal to the core TALE binding domain, such as (C- 11 AA) represented below, which does not comprise the above X1 , X2 and X3, such as that represented by SEQ ID NO:282: SIVAQLSRPDP (C-11 AA).

[0126] By “TALE fusion protein” is meant a TALE-protein which is linked to a polypeptide domain that confers a catalytic activity to said TALE protein. A TALE fusion protein can be for instance a sequence-specific reagent that processes DNA at the locus specified by the TALE binding domain. The fusion with the TALE protein can be made with the catalytic domain from an existing protein, such as a DNA processing enzyme, especially one having an activity selected from the group consisting of nuclease activity, polymerase activity, deaminase activity, kinase activity, phosphatase activity, methylase activity, topoisomerase activity, integrase activity, transposase activity, ligase activity, helicase activity, reverse transcriptase and recombinase activity. In some embodiments, the TALE fusion protein according to the present invention can comprise a peptide linker to fuse the catalytic domain to said previously described core scaffold, or more preferably to link the C-terminal or N-terminal of said TALE protein to said catalytic domain. Such linker is generally flexible. Such as one linker sequence selected from the group consisting of NFS1 , NFS2, CFS1 , RM2, BQY, QGPSG, LGPDGRKA, 1a8h_1, 1dnpA_1, 1d8cA_2, 1ckqA_3, 1sbp_1, 1ev7A_1, 1alo_3, 1amf_1, 1adjA_3, 1fcdC_1, 1al3_2, 1g3p_1, 1acc_3, 1ahjB_1, 1acc_1 , 1af7_1 , 1heiA_1 , 1bia_2, 1igtB_1 , 1nfkA_1, 1au7A_1, 1bpoB_1 , 1b0pA_2, 1cO5A_2, 1gcb_1 , 1bt3A_1, 1b3oB_2, 16vpA_6, 1dhx_1, 1b8aA_1, 1qu6A_1 optionally comprising SGGSGS stretches at either or both N and C-terminal ends that surround a variable region of 3 to 28 amino acids as exemplified in Table 2 below (SEQ ID NO:31 to 104).

[0127] Table 2: Examples of peptide linkers.

[0128] In some embodiments, said peptide linker can comprise a calmodulin domain that changes TALE fusion protein conformation under calcium stimulation. Other protein domains inducing conformational changes under a specific metabolite interaction can also be used. Such linker can comprise, for instance, a light sensitive domain that allows a change from a folded inactive state toward an unfolded active state under light stimulation, or reverse. Other examples of “switch linkers can be reactive to small molecules such as Chemical Inducers of Dimerization (CID).

[0129] In preferred embodiments, as illustrated herein with the preferred C-terminal sequences previously described, a linker may not be necessary to fuse the TALE core binding domain with the catalytic domain, as the C-terminal sequences can have enough flexibility to achieve an optimal conformation of the TALE fusion protein.

[0130] The present invention encompasses TALE fusion proteins comprising a variety of functional domains, such as catalytic domains obtainable from different enzymes. Such catalytic domains can be unspecific endonucleases such as for instance Fok-1 , clo51 or I-Tev1 , or specific endonuclease, such as engineered meganucleases (e.g. derived from I-Cre1 , I-Onu1 , I-Bmo1 , HmuL..), exonucleases such as human Trex2, transcription repressors (e.g. KRAB) or transcription activators such as VP64, or VP16, deaminases such as for example cytosine deaminase 1 (pCDM), adenosine deaminase, such as TadA ou TadA7.10, Apolipoprotein B mRNA editing enzyme catalytic polypeptide-like (APOBEC), Activation-induced cytidine deaminase (AICDA), DddA (double strand DNA cytidine deaminase) that is generally fused to Uracil Glycosylase Inhibitors (UGI), nickases derived from Cas9 or Cpfl , transposase, integrase, topoisomerase and reverse transcriptase (e.g. Moloney murine leukemia virus RT enzyme), their functional mutants, variants or derivatives thereof.

[0131] Exemplary polypeptides sequences that can be included in the TALE fusion proteins of the present invention are listed in Table 3 (SEQ ID NO: 105 to 133).

[0132] Table 3: exemplary catalytic domains of the TALE proteins of the present invention

[0133] In another embodiment, the TALE fusion protein according to the present invention comprises a catalytic domain that is a polypeptide comprising an amino acid sequence having at least 80%, preferably at least 90%, more preferably at least 95% identity with any of SEQ ID NO: 105 to 133.

[0134] Since gene editing reagents can cause unintended interruptions in the genome, gene editing is crucial and as multiplex methods become more widely used, the likelihood of off-targets and the downstream consequences of such off-target activity grow. Minimizing such undesired cleavage (off-targets) is a matter of utmost importance for any genome-engineering applications, especially in the therapeutic domain. Undesired double-stranded breaks in the genome may lead to chromosome translocation, and cellular toxicity [Cantoni O., et al. (1996) Cytotoxic impact of DNA single vs double strand breaks in oxidatively injured cells. Arch Toxicol Suppl 18:223-235], There are currently a variety of techniques available to predict and quantify off-target by analysing secondary target locations and establish on-target / off target ratios, such as those described by Tsai S., et al. [CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets (2017) Nat Methods 14(6):607-614], Hockemeyer D, et al. [Genetic engineering of human pluripotent cells using TALE nucleases (2011) Nat. Biotechnol. 29(8):731- 734] and Wienert B, et al. [Unbiased detection of CRISPR off-targets in vivo using DISCOVER- Seq (2019) Science 364(6437): 286-289], As previously mentioned, TALE proteins have a well-defined DNA base-pair choice, offering a basic strategy for scientific researchers and engineers to design and construct TALE fusion proteins for genome alteration. A TALE repeat tandem is responsible for recognizing individual DNA base pairs. Such tandem is made up of a pair of alpha helices linked by a loop of three-residue of RVDs in the shape of a solenoid. For the creation of TALE proteins with variable precision and binding affinity, the six conventional RVDs (NG, HD, Nl, NK, NH, and NN) are frequently used. HD and NG are associated with cytosine (C) and thymine (T) respectively. These associations are strong and exclusive [Streubel J, et al. (2012) TAL effector RVD specificities and efficiencies. Nat Biotechnol 30(7): 593-595], NN is a degenerate RVD usually showing binding affinity for both guanine (G) and adenine (A), but its specificity for guanine is reported to be stronger. RVD Nl binds with A and NK binds with G. These associations are exclusive but the binding affinity between these pairs is less due to which they are considered weak. Therefore, it is recommended to use RVD NH which binds with G with medium affinity. It is also worth noting that the binding affinity of TALE is influenced by the methylation status of the target DNA sequence.

[0135] The TALEN code is degenerate, which means that certain RVDs can bind to multiple nucleotides with a diverse spectrum of efficiency. The binding ability of the NN (for A and G) and NS (A, C, and G) repeat variable di-residue empowers the TALE proteins to encode degeneracy for the target DNA. This degeneracy may although be useful in targeting hyper variable sites. TALE proteins technology is the only known genome editing tool which can be engineered in a way that can be easily used for the escape mutations in a genome. This unique feature make them a more flexible and reliable tool in the field of genome editing specifically in clinical applications to tolerate predicted mutations [Strong C.L., et al. (2015) Damaging the integrated HIV proviral DNA with TALENs. PLoS One 10(5):e0125652],

[0136] A typical TALE protein usually consists of 18 repeats of 34 amino acids. A TALEN pair must bind to the target site on opposite sides, separated by a “spacer” of 14-20 nucleotides as an offset since Fokl requires dimerization for operation. As a whole, such a long (approximately 36 bp) DNA binding site is predicted to appear in genomes as being very rare.

[0137] According to some embodiments, the invention provides methods for designing a TALE protein for introducing a genetic modification into a polynucleotide sequence, said method comprising one or several of the following steps: a) selecting a polynucleotide target sequence on which the genetic modification is intended; b) assembling polynucleotide sequences encoding AvrBs3-like repeat(s) to form a polynucleotide encoding a TALE-binding domain to bind said selected polynucleotide target sequence; c) fusing to said polynucleotide encoding the TALE-binding domain at least:

[0138] (1) a polynucleotide sequence encoding a N-terminal domain including a series of NLS as per the present invention comprising a sequence having at least 85% identity with SEQ ID NO: 11 or SEQ ID NO:12 and

[0139] (2) a polynucleotide sequence encoding a C-terminal domain consisting of a polypeptide sequence from 40 to 80 residues, said polypeptide sequence comprising preferably a sequence having at least 85%, preferably 90%, more preferably 95% and even more preferably 99% identity with SEQ ID NO: 13, SEQ ID NO: 14 or SEQ ID NO: 15; Xi, X2, X3 in these sequences representing any amino acids and preferably R (arginine) or H (histidine) or said C-terminal domain consists of a polypeptide sequence of less than 40, preferably less than 30, more preferably less than 20, such as about 11 residues, such as C11-AA (SEQ ID NO:282);

[0140] In general, the above steps can be performed in-silico and the final polynucleotide sequence synthetised or cloned according to methods well known in the art, such as explained for instance in WQ2013017950.

[0141] By « genetic modification » is intended any enzymatic reaction voluntarily induced at a given locus, such as a mutation, methylation, transcriptional modulation, in view of obtaining an effect on gene expression.

[0142] According to some embodiments, the methods of the invention comprise one or several of the steps consisting of: a) selecting a cleavage site in a target polynucleotide sequence, such as into a genome, where cleavage is intended; b) selecting a polynucleotide sequence located between 5 and 25 bp upstream and / or downstream of said cleavage site; c) assembling polynucleotide sequences encoding AvrBs3-like repeat(s) to encode a TALE-binding domain to bind said selected polynucleotide sequence, wherein at least one AvrBs3-like repeat(s) comprises D substitutions at positions 4 (D4) and 32 (D32) in its polypeptide sequence, such as one sequence selected from SEQ ID NO: 16 to 22; d) fusing said TALE-binding domain to at least (1) a polynucleotide sequence encoding a N-terminal domain, preferably comprising a sequence having at least 85%, preferably at least 90%, more preferably at least 95% identity with SEQ ID NQ:30 and (2) a polynucleotide sequence encoding a C-terminal domain preferably of a polypeptide sequence from 40 to 80 residues comprising a sequence having at least 85% identity with SEQ I D NO: 13, SEQ I D NO: 14 or SEQ I D NO: 15, (Xi , X2, X3 in these sequences representing any amino acids, preferably R (arginine) or H (histidine)) or less than 40 residues, such as about 11 residues like in SEQ ID NO:282 as indicated before; e) fusing the polynucleotide sequence obtained in d) with another polynucleotide sequence encoding a functional domain. f) fusing the preferred series of NLS as per the present invention referred to herein.

[0143] Optionally, said N-terminal sequence or C-terminal sequence can further comprise a localization sequence toward a given organelle within an organism, a tissue or a cell. Non-limiting examples of such localization signals are chloroplastic localization signals to target plant organelles, or mitochondrial localization signals, such as the superoxide dismutase 2 mitochondrial addressing signals LSRAVCGTSRQLAPVLGYLGSRQKHSLPD (SEQ ID NO: 134) or the cytochrome c oxidase subunit 8A mitochondrial addressing signals SVLTPLLLRGLTGSARRLPVPRAKIHSL (SEQ ID NO:135). In another embodiment, said additional sequence can comprise a nuclear export signal having the opposite effect of a nuclear localization signal to help targeting organelles such as chloroplasts or mitochondria.

[0144] The present method can also comprise optional steps, wherein, for instance, the polynucleotide sequence that is fused to the TALE protein and encode the catalytic domain can be mutated to introduce amino acid substitutions into said catalytic domain. ;ific TALE-nucleases

[0145] By following the above teachings, highly specific TALE-nucleases can be produced according to the present invention allowing high degree of cleavage specificity and low cytotoxicity in diverse cell types, especially plant or mammalian cells.

[0146] According to some embodiments, the TALE-fusion protein of the present invention is a TALE-nuclease obtained by fusion of a TALE protein as described herein with the nuclease catalytic domain of a non-specific nuclease, such as Fok-1 (SEQ ID NQ:109) or Tev-1 (SEQ ID NO:114) as described with classical TALE scaffolds for instance in Beurdeley, M. et al. [Compact designer TALENs for efficient genome engineering (2013) Nat Commun 4:1762], In preferred embodiments, said nuclease catalytic domain is Fok1 , i.e. comprises a polypeptide showing at least 80% identity with SEQ ID NO.1, and more preferably comprising at least one of the amino acid substitutions: 13, 52, 57, 59, 61 , 65, 84, 85, 88, 91 , 92, 95, 98, 103, 109, 110, 111 , 113, 119, 143, 148, 152, 158, 159, 160, 167, 169, 170 and 194 into SEQ ID NQ:105.. Preferred substitutions are introduced at positions 84, 85, 88, 95, 98, 91 , 103, 109, 148, 152 and 158, and most preferred ones are in positions 84, 88, 91 , 103 and 152. According to some embodiments, the TALE-fusion protein of the present invention is a TALE-nuclease obtained by fusion of a TALE protein as described herein with a nickase, in particular a Cas9 nickase. Such Cas9 nickase are generally Cas9 proteins which are mutated in their RuvC or HNH domains, for instance by introducing mutations D10A in RuvC and H840A in HNH. In general, TALE-Cas9 nickase fusions are used by pairs as formerly described with classical TALE scaffolds by Guilinger, J., et al. [Fusion of catalytically inactive Cas9 to Fokl nuclease improves the specificity of genome modification (2014) Nat. Biotechnol. 32, 577-582],

[0147] In some other embodiments, the TALE-fusion protein of the present invention is a TALE- nuclease obtained by fusion of a TALE protein as described herein with a specific nuclease, preferably a customized rare-cutting endonuclease, such as a meganuclease variant. In preferred embodiments, said rare-cutting endonuclease can be a variant of LADLIDADG, such as l-crel or l-Onul, as previously described for instance in EP3320910 and EP3004338.

[0148] On another hand, a TALE-nuclease according to the present invention has also the ability to efficiently manipulate mtDNA (mitochondrial DNA) as a treatment for treating human mitochondrial diseases triggered by mitochondrial pathogenic mutations. So called “Mito-TALEN” (mitochondrial-targeted TALENs) have been proven to be effectively treating human mitochondrial disorders affected by mtDNA mutations, such as Leber’s hereditary optic neuropathy, ataxia, neurogenic muscle fatigue, and retinal pigmentosa [Gammage, P.A., et al. (2018) Mitochondrial Genome Engineering: The Revolution May Not Be CRISPR-lzed. Trends in Genetics, 34(2):101-110], Plastid engineering has also demonstrated competent results in varieties of plants for crop improvements [Piatek AA, Lenaghan SC, Neal Stewart C. (2018) Advanced editing of the nuclear and plastid genomes in plants. Plant Sci 273:42-49],

[0149] Many examples of TALE-nuclease as per the present invention are herein described to be used as therapeutic reagent to induce highly specific cleavage in a selection of genes in human cells, especially blood cells. More particularly, improved TALE nuclease reagents have been synthetized and tested pursuant to the present teachings in order to cleave gene targets in primary cells, especially in T-cells or NK cells, such as TCR, B2M, CD52, GR, CCR5, CS1 / SLAMF7, CD38, TGFBRII, GMSCF, HLA-A, HLA-B, HLA-C, PD1 , CTLA4, LAG3, TIGIT, CISH, IgH, GADPH HVACR2, CD3E, CDKN2A, CDKN2B or MTAP. The interest of targeting these genomic sequences in NK and T-cells to produce therapeutic cells useful in immunotherapy have been well established as reviewed for instance by Moradi V, et al. [Progress and pitfalls of gene editing technology in CAR-T cell therapy: a state-of-the-art review (2024) Front Oncol.14: 1388475],

[0150] In some particular embodiments, the fusion proteins of the present invention have been specifically designed to target CDKN2A, CDKN2B and MTAP genes, separately or concomitantly, in immune cells , especially the TALE-nucleases of any of SEQ ID NO:243 to 250 referred to herein, is of particular interest to inactivate these genes in T-cells and generate immune cells resistant to senescence as initially described in WO2023025862.

[0151] In some particular embodiments, the fusion proteins of the present invention have been specifically designed to target CIITA and are TALE-nuclease heterodimers comprising SEQ ID NO:278 and SEQ ID NO:279 and / or SEQ ID NQ:280 and SEQ ID NO:281.

[0152] Other examples of TALE-nucleases as per the present invention can be used as therapeutic agent to induce highly specific cleavage in a selection of genes in human cells which are not blood cells, like hepatocytes to cleave, for instance, within the Albumin locus or the APOCIII locus. In particular, improved TALE nuclease reagents have been synthetized and tested pursuant to the present teachings to cleave the Albumin locus (see, for instance, Example 6). In some particular embodiments, the fusion proteins of the present invention have been specifically designed to target the ALB locus and are TALE-nuclease heterodimers comprising SEQ ID NO:275 and SEQ ID NO:277.

[0153] The polypeptide sequences of some exemplary TALE fusion proteins obtained as per the present invention, as well as their target sequences (polynucleotide sequence spanning the two left and right heterodimeric binding sites) are listed in Table 4. The present invention encompasses any homologue variants or equivalents which amino acid sequence would be at least 90%, 95%, 98 or 99% identical to same.

[0154] In some preferred embodiments, the TALE-proteins of the present invention can be used by pairs, each member of this pair binding DNA close to each other, side-by-side or on opposite DNA strands, in such a way they are co-localized in the genome with the effect of directing the catalytic activity induced by the catalytic domain at a specified locus. For instance, a pair of TALE- proteins fused to the homodimerizing Fok1 nuclease domain, also referred to as “left-” and “right- ” TALE-Nuclease monomers, form heterodimers that induce DNA double strand break cleavage. In such instances, the invention provides that one monomer as per the present invention can be used with another monomer that is based on a conventional TALE-Nuclease scaffold using canonical AvrBs3 sequences. Indeed, as shown in the experimental section herein, one TALE- nuclease monomer of the present invention is sufficient to have an overall effect on the heterodimeric specificity.

[0155] The present invention thus provides a number of new TALE fusion monomers based on the TALE-fusion proteins described herein, comprising such proteins fused with a nuclease or deaminase domain, for their use in genetic therapeutic modifications, in-vivo or in-vitro, as well as for the ex-vivo preparation of therapeutic cells. According to some particular aspects, the invention provides with fusion of TALE-protein monomers to introduce a genetic modification, preferably a mutation, into one of the two isoforms of the CDKN2A gene locus, preferably into a target sequence comprising SEQ ID NO:264 or SEQ ID NO:265 as shown in Example 4, wherein said TALE fusion protein comprises (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO:16 to 22, (2) an optional C-terminal polypeptide sequence from 10 to 80 residues preferably comprising a sequence having at least 85% identity with SEQ ID NO:13, 14, or 15, (3’) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS of the present as described herein, preferably comprising SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C), preferably at the N-terminal end of said TALE fusion protein. Said TALE-protein is preferably a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:243, 244, 245 or 246. In particular, the invention provides with TALE-nuclease heterodimers comprising at least one monomer polypeptides having at least 90%, preferably 95% or 99% identity with SEQ ID NO:243, and SEQ ID NO:244, cleaving CDKN2A target sequence which at least 90% identical to the polynucleotide sequence SEQ ID NO:264 and also TALE-nuclease heterodimers comprising at least one monomer polypeptides having at least 90%, preferably 95% or 99% identity with SEQ ID NO:245, and SEQ ID NO:246, cleaving CDKN2A target sequence which at least 90% identical to the polynucleotide sequence SEQ ID NO:265.

[0156] According to some particular aspects, the invention provides with fusion of TALE-protein monomers to introduce a genetic modification, preferably a mutation, into the MTAP gene locus, preferably into a target sequence comprising SEQ ID NO:266, as shown in Example 4, wherein said TALE fusion protein comprises (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO:16 to 22, (2) an optional C-terminal polypeptide sequence from about 10 to 80 residues preferably comprising a sequence having at least 85% identity with SEQ ID NO: 13, 14, or 15, (3) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS of the present as described herein, preferably comprising SEQ ID NO:11 (S-N-C) or SEQ ID NO: 12 (N-S-C) preferably at the N-terminal end of said TALE fusion protein. Said TALE-protein is preferably a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:247 or 248. In particular, the invention provides with TALE-nuclease heterodimers comprising at least one monomer polypeptides having at least 90%, preferably 95% or 99% identity with SEQ ID NO:247 and / or SEQ ID NO:248, cleaving MTAP target sequence which is at least 90% identical to the polynucleotide sequence SEQ ID NO:266. According to some particular aspects, the invention provides with fusion of TALE-protein monomers to introduce a genetic modification, preferably a mutation, into the CDKN2B gene locus, preferably into a target sequence comprising SEQ ID NO:267, as shown in Example 4, wherein said TALE fusion protein comprises (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO: 16 to 22, (2) an optional C-terminal polypeptide sequence from 10 to 80 residues preferably comprising a sequence having at least 85% identity with SEQ ID NO: 13, 14, 15 or 282, (3) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS of the present as described herein, preferably comprising SEQ ID NO: 11 (S-N-C) or SEQ ID NO: 12 (N-S-C), preferably at the N-terminal end of said TALE fusion protein. Said TALE-protein is preferably a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:249 or 250. In particular, the invention provides with TALE-nuclease heterodimers comprising at least one monomer polypeptides having at least 90%, preferably 95% or 99% identity with SEQ ID NO:249 and / or SEQ ID NQ:250, cleaving CDKN2B target sequence which is at least 90% identical to the polynucleotide sequence SEQ ID NO:267.

[0157] The above TALE-nuclease targeting respectively CDKN2A isoforms, MTAP and CDKN2B which are illustrated in Example 4, are useful to inactivate those genes and induce resistance to senescence in immune cells, in particular when they are expressed in primary T-cells, especially human primary T-cells.

[0158] According to further aspects, the present invention provides TALE fusion proteins particularly active and specific to introduce a genetic modification into the TRAC gene, such as more specifically illustrated in Example 5 of the present application. As a preferred embodiment, is a TALE nuclease monomer comprising a 3-NLS sequence such as SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C). Such fusion proteins typically comprise (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO: 16 to 22, (2) a C-terminal polypeptide sequence from about 10 to 80 residues comprising a sequence having at least 85% identity with SEQ ID NO: 13, 14, 15 or 282 (3’) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS of the present as described herein, preferably comprising SEQ ID NO:11 (S-N-C) or SEQ ID NO: 12 (N-S-C) preferably at the N-terminal end of said TALE fusion protein. Said TALE-protein is preferably a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:271. In particular, the invention provides TALE-nuclease heterodimers comprising at least one such monomer polypeptide, which preferably cleave a target sequence, which is at least 90%, 95% or 99% identical to the polynucleotide sequence SEQ ID NO:268 (TRAC_OS1). According to further aspects, the present invention provides TALE fusion proteins particularly active and specific to introduce a genetic modification into the B2M gene, such as more specifically illustrated in Example 5 of the present application. As a preferred embodiment, is a TALE nuclease monomer comprising a 3-NLS sequence such as SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C). Such fusion proteins typically comprise (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO:16 to 22, (2) an optional C-terminal polypeptide sequence from about 10 to 80 residues comprising a sequence having at least 85% identity with SEQ ID NO:13, 14, 15 or 282, and (3’) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS as described herein, preferably comprising SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C), preferably at the N-terminal end of said TALE fusion protein. Said TALE-protein is preferably a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:273 or SEQ ID NO:274. In particular, the invention provides TALE-nuclease heterodimers comprising at least one such monomer polypeptide, which preferably cleave a target sequence, which is at least 90%, 95% or 99% identical to the polynucleotide sequence SEQ ID NO:269 (hsB2M_TO2.1).

[0159] According to further aspects, the above preferred TALE-nucleases comprising a 3-NLS sequence as per the present invention targeting respectively TRAC and B2M genes are used separately or concomitantly in the same population of cells to limit off-site cleavage, especially to produce engineered cells, such as primary T-cells [TRAC]negative[B2M]negative, which can be useful in allogeneic therapeutic settings.

[0160] According to further aspects, the present invention provides TALE fusion proteins particularly active and specific to introduce a genetic modification into the CIITA gene, such as more specifically illustrated in Example 7 of the present application. As a preferred embodiment, is a TALE nuclease monomer comprising a 3-NLS sequence such as SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C). Such fusion proteins typically comprise (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO:16 to 22, (2) an optional C-terminal polypeptide sequence from about 10 to 80 residues comprising a sequence having at least 85% identity with SEQ ID NO:282, 13, 14, or 15, and (3’) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS as described herein, preferably comprising SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C), preferably at the N-terminal end of said TALE fusion protein. Said TALE-protein is preferably a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:278, SEQ ID NO:279, SEQ ID NQ:280 or SEQ ID NO:281. In particular, the invention provides TALE-nuclease heterodimers comprising at least one such monomer polypeptide, which preferably cleaves a target sequence, which is at least 90%, 95% or 99% identical to the polynucleotide sequence SEQ ID NO:286 (T008446) or SEQ ID NO:287 (T008447) in the CIITA human gene.

[0161] According to further aspects, the above preferred TALE-nucleases comprising a 3-NLS sequence as per the present invention targeting respectively TRAC and CIITA genes are used separately or concomitantly in the same population of cells to limit off-site cleavage, especially to produce engineered cells, such as primary T-cells [TRAC]negative[CIITA]negative, which can be useful in allogeneic therapeutic settings.

[0162] An aspect of the present invention thus concerns the use of the above TALE fusion proteins, independently or altogether, to manufacture therapeutic immune cells, for example CAR T-cells, CAR NK-cells, TILL or immune cells expressing recombinant TCRs to make them resistant to senescence and / or to improve the persistence of allogeneic immune cells in patients, as well as kits comprising such TALE fusion proteins in view of manufacturing engineered cells or any polynucleotide or vectors encoding thereof for their expression in those cells.

[0163] According to further aspects, the present invention provides TALE fusion proteins particularly active and specific to introduce a genetic modification into the Albumin (ALB) gene, such as more particularly illustrated in Example 6 of the present application.

[0164] A preferred embodiment is a TALE nuclease monomer comprising a 3-NLS sequence such as SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S-C). Such TALE nuclease monomer typically comprises (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14 or 15 TALE repeats, preferably comprising at least one repeat of SEQ ID NO:16 to 22, (2) an optional C-terminal polypeptide sequence from about 10 to 80 residues comprising a sequence having at least 85% identity with SEQ ID NO:282, 13, 14, or 15, (3’) a functional domain, such as nuclease or a base editor, and (4) series of bipartite and monopartite NLS as described herein, preferably comprising SEQ ID NO:11 (S-N-C) or SEQ ID NO:12 (N-S- C) preferably at the N-terminal end of said TALE fusion protein. Preferred is a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:275 or SEQ ID NO:277.

[0165] In particular, the invention provides TALE-nuclease heterodimers comprising at least a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO:277 and, optionally, a TALE-nuclease monomer comprising a polypeptide sequence showing at least 90% identity, preferably 95% or 99% identity with SEQ ID NO: 275, wherein said TALE-nuclease heterodimer preferably cleaves within a target sequence preferably at least 90%, 95% or 99% identical to the polynucleotide sequence SEQ ID NO:270. In a preferred embodiment, a TALEN-heterodimer targeting the ALB locus comprises a TALEN-monomer comprising a binding sequence of SEQ ID NO: 275 and a TALEN-monomer comprising a binding sequence of SEQ ID NO. 277.

[0166] Such TALE fusion proteins are especially useful to perform in vivo gene therapy, such as to insert a transgene at the ALB locus. The present invention thus encompasses methods wherein a TALE fusion protein (such as a TALEN), or a mRNA encoding thereof, such as described above, is loaded into a lipidic vector, in particular lipid nanoparticles (LNP), preferably along with a DNA template comprising an exogenous coding sequence for its insertion in the locus targeted by the TALE fusion protein, more particularly into the ALB locus. factors

[0167] By following the previous teachings, the TALE proteins according to the invention can also be fused to desired transcriptional activator and repressor protein domains to create specific trans-activator or repressor reagents in view of controlling endogenous gene expression.

[0168] As an example, artificial transcription factors can be obtained by fusion of a TALE protein of the present invention with VP64 or the 16 amino acid peptide VP16 (SEQ ID NQ:120) from herpes simplex virus as described by Miller J. C., et al. [A TALE nuclease architecture for efficient genome editing (2011) Nat Biotechnol 29(2): 143-148],

[0169] To accomplish repression of a gene, the TALE proteins of the present invention can be fused for example with Kruppel-associated box (KRAB), Sid4, or EAR-repression domain (SRDX), which have been previously reported as being strong pleiotropic repressors [Cong L, et al. (2012) Comprehensive interrogation of natural TALE DNA-binding modules and transcriptional repressor domains. Nat Commun 3( 1 ) : 968].

[0170] Development of TALE- base editors

[0171] By following the previous teachings, the TALE proteins according to the invention can also be fused to desired base editors.

[0172] The term “base editor” as used herein, refers to a catalytic domain capable of making a modification to a base ( e.g ., A, T, C, G, or U) within a nucleic acid sequence that converts one base to another (e.g., A to G, A to C, A to T, C to T C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). Adenine and cytosine base editors catalytic domains are described, for instance, in Rees & Liu [Base editing: precision chemistry on the genome and transcriptome of living cells (2018) Nat. Rev. Genet. 19(12):770-788], Catalytic base editors can include cytidine deaminase that convert target C / G to T / A and adenine base editors that convert target A / T to G / C. Preferred cytosine deaminase can be cytosine deaminase 1 (pCDM) or Activation-induced cytidine deaminase (AICDA). Preferred adenosine deaminase can be TadA (SEQ ID NO: 121) or its variant TadA7.10 as described by Jeong, Y.K., et al. [Adenine base editor engineering reduces editing of bystander cytosines (2021) Nat. Biotechnol. https: / / doi.org / 10.1038 / s41587-021-00943]. Different members of Apolipoprotein B mRNA editing enzyme (APOBEC) family can be used convert cytidines to thymidines, such as the murine rAPOBECI and the human APOBEC3G (SEQ ID NO:130) as developed by Lee et al. [Single C-to-T substitution using engineered APOBEC3G-nCas9 base editors with minimum genome- and transcriptome-wide off-target effects (2020) Science Advances. 6(29)].

[0173] In preferred embodiments, base editor catalytic domain converts a C to T (cytidine deaminase) that catalyzes the chemical reaction “cytosine + H2O -> uracil + NH3” or “5-methyl- cytosine + H2O -> thymine + NH3.” As it may be apparent from the reaction formula, such chemical reactions result in a C to ll / T nucleobase change. In the context of a gene, such a nucleotide change, or mutation, may in turn lead to an amino acid change in the protein, which may affect the protein’s function, e.g., loss-of-function or gain-of-function.

[0174] In some embodiments, the TALE-base editors according to the present invention can comprise a domain that inhibits uracil glycosylase referred to as “UGI”, and / or a nuclear localization signal. The term “uracil glycosylase inhibitor” or “UGI,” as used herein, refers to a protein that is capable of inhibiting a uracil-DNA glycosylase base-excision repair enzyme. In some embodiments, a UGI domain comprises a wild-type UGI or a canonical UGI as set forth in SEQ ID NO:132. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to a UGI or a UGI fragment comprising an amino acid sequence that comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, of the amino acid sequence as set forth in SEQ ID NO:132. TALE base editors according to the present invention comprising UGI are useful to improve the specificity of base editing performed at a predetermined locus.

[0175] In some embodiments, the base editor catalytic domain is a double- stranded DNA deaminase (“DddA”) to precisely install nucleotide changes and / or correct pathogenic mutations, rather than destroying DNA with double-strand breaks (DSBs). In preferred embodiments, DddAtox is generally split into inactive fragments which can be separately delivered to a target deamination site on separate TALE-base editor constructs that will co-localize each fragment of the DddA on site, such as on either side of a target edit site, where they reform a functional DddA that is capable deaminating a target site on the double-stranded DNA molecule. In certain embodiments, the programmable DNA binding proteins can be engineered to comprise one or more mitochondrial localization signals (MLS), in such a way that the DddA domains become translocated into the mitochondria, thereby providing a means by which to conduct base editing directly on the mitochondrial genome.

[0176] An example of a double stranded DNA cytosine deaminase is the DddAtox protein from Burkholderia cenocepacia, the amino acid sequence of which is available under UniProt accession number P0DLIH5 and is represented by SEQ ID NO:283.

[0177] For instance, in the case of the DddA that is the cytosine deaminase from Burkholderia cenocepacia (“DddAtox”) of SEQ ID NO:283, the truncation can occur at any position within the catalytic domain spanning from amino acid position 1290 and amino acid position 1427 in SEQ ID NO:283, such as any position between amino acid at position 1333 and amino acid at position 1397 of SEQ ID NO:283, for instance just after amino acid G1322, G1333, A1343, N1357, G1371 , N1387, or G1397, relative to SEQ ID NO:283. In preferred embodiments, the truncation of DddA occurs just after the amino acid residue G1333, A1343, N1357, or G1397, said positions being relative to the amino acid sequence of SEQ ID NO:283. More preferably, the truncation occurs just after amino acid G at position 1333 or just after amino acid G at position 1397, relative to SEQ ID NO:283.

[0178] Accordingly, truncation of the DddA catalytic domain can occur between amino acid G at position 33 and amino acid G at position 34, between amino acid G at position 44 and amino acid P at position 45 of SEQ ID NO:284, between amino acid A at position 54 and amino acid G at position 55 of SEQ ID NO:284, between amino acid N at position 68 and amino acid G at position 69 of SEQ ID NO:284, between amino acid G at position 82 and amino acid T at position 83 of SEQ ID NO:284, between amino acid N at position 98 and amino acid A at position 99 of SEQ ID NO:284, or between amino acid G at position 108 and amino acid A at position 109 of SEQ ID NO:284. More preferably, truncation of the DddA catalytic domain takes place between amino acid G at position 44 and amino acid P at position 45 of SEQ ID NO:284, or between amino acid G at position 108 and amino acid A at position 109 of SEQ ID NO:284.

[0179] In another instance, a mutated DddA comprising the mutated catalytic domain of SEQ ID NO:285 can be used. In this embodiment, the truncation can occur at any position within said catalytic domain, such as between amino acid G at position 44 and amino acid P at position 45 of SEQ ID NO:285, or between amino acid G at position 108 and amino acid A at position 109 of SEQ ID NO:285.

[0180] In certain embodiments, when the DddA is separated into two fragments by dividing the DddA at one of these split sites these polypeptides may be referred to as “DddA-N half’ and “DddA-C half’. According to preferred embodiments said “DddA-N half’ and “DddA-C half’ comprise an amino acid sequence that respectively share at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, with the amino acid sequence SEQ ID NO.130 and SEQ ID NO: 131. Also two TALE proteins acting by pairs respectively comprising N and C- DddA halves can be used to co-localize and induce on-site nucleobase change.

[0181] TALE-base editors of the present invention can also be used by pairs, each member comprising different but complementary catalytic domains in view of obtaining a given base editing reaction at one precise locus.

[0182] Development of TALE- transposase or integrase

[0183] By following the previous teachings, the TALE proteins according to the invention can also be fused to a transposase or an integrase in order to perform site-directed integration of transgenes into the genome.

[0184] As an example, the TALE protein according to the invention can be fused to the PiggyBac transposase as described for instance by Owens, J.B. et al. [Transcription activator like effector (TALE)-directed piggyBac transposition in human cells (2013) N.A.R. 41 (19):9197-9207], The PiggyBac transposase is autonomously functional in such system so that a co-transfected transposon is able to integrate into any genomic location specified by the TALE protein. This system can permanently introduce large cassettes (>100 kb) encoding numerous components such as multiple transgenes, insulators and inducible or endogenous promoters and allows to potentially target integrations to nearly any genomic region. This system is especially worth in situations where safe single- targeted insertions need to be verified ex vivo, and cells be amplified and re-infused into patients. Targeted transposition could be used to intentionally disrupt endogenous coding regions or to direct insertions to user-defined genomic safe harbours to protect the cargo from unknown chromosomal position effects and to circumvent accidental mutation of target cells.

[0185] Still following the previous teachings, TALE-protein fusions can be made by fusion with catalytic domains that can modulate the expression of a gene without altering the DNA sequence, especially by remodelling chromatin.

[0186] In this regard, TALE proteins as per the present invention can be fused to methyltransferase obtain histone methylation and / or with a p300 effector domain that enhances histone acetyltransferase. Conversely, TALE protein can be fused to the catalytic domain thymidine DNA glycosylase (TDG) to abolish the DNA methylation and induce gene expression. Unwanted DNA methylations are associated with many neurodegenerative diseases. TALE protein could be fused to TET domain (ten-eleven translocation methylcytosine dioxygenase 2) as an example, for targeting epigenetically silenced cancer gene (ICAM-1) and induce its expression in cancerous cells. TET1 can also be used used in the treatment of many diseases like diabetes (inducing p cell replication) and cancer (inhibiting cell proliferation) [Ou K., et al. (2019) Targeted demethylation at the CDKN1C / p57 locus induces human cell replication. J Clin Invest 129(1):209-214],

[0187] The present invention encompasses the polynucleotides, in particular DNA or RNA encoding the polypeptides and proteins previously described, as well as any intermediary products involved in any aspects and steps of the methods described herein. These polynucleotides may be included in vectors, more particularly plasmids or virus, in view of being expressed in prokaryotic or eukaryotic cells.

[0188] The terms "vector" or “vectors” refer to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. A “vector” in the present invention includes, but is not limited to, a viral vector, a plasmid, a RNA vector or a linear or circular DNA or RNA molecule which may consists of a chromosomal, non-chromosomal, semi-synthetic or synthetic nucleic acids. Preferred vectors are those capable of autonomous replication (episomal vector) and / or expression of nucleic acids to which they are linked (expression vectors). Large numbers of suitable vectors are known to those of skill in the art and commercially available. Viral vectors include retrovirus, adenovirus, especially AAV6 vectors, parvovirus (e. g. adenoassociated viruses), coronavirus, negative strand RNA viruses such as orthomyxovirus (e. g., influenza virus), rhabdovirus (e. g., rabies and vesicular stomatitis virus), paramyxovirus (e. g. measles and Sendai), positive strand RNA viruses such as picornavirus and alphavirus, and double-stranded DNA viruses including adenovirus, herpesvirus (e. g., Herpes Simplex virus types 1 and 2, Epstein-Barr virus, cytomegalovirus), and poxvirus (e. g., vaccinia, fowlpox and canarypox). Other viruses include Norwalk virus, togavirus, flavivirus, reoviruses, papovavirus, hepadnavirus, and hepatitis virus, for example. Examples of retroviruses include: avian leukosis-sarcoma, mammalian C-type, B-type viruses, D type viruses, HTLV-BLV group, lentivirus, spumavirus (Coffin, J. M., Retroviridae: The viruses and their replication, In Fundamental Virology, Third Edition, B. N. Fields, et al., Eds., Lippincott- Raven Publishers, Philadelphia, 1996).

[0189] As per the present invention, the TALE proteins or polynucleotide encoding thereof, especially mRNA, can also be loaded into nanoparticles for their effective delivery into cells. A variety of nanoparticles are described in the art to target particular tissues of cell types [Friedman A.D. et al. (2013) The Smart Targeting of Nanoparticles Curr Pharm Des. 19(35): 6315-6329], Preferred nanoparticles are positively charged nanoparticles, such as silica based nanoparticles or LNP (Lipid nanomolar nanoparticles) as described in the art with other types of nucleases [Conway, A. et al. (2019) Non-viral Delivery of Zinc Finger Nuclease mRNA Enables Highly Efficient In Vivo Genome Editing of Multiple Therapeutic Gene Targets, Molecular Therapy 27(4):866-877],

[0190] Alternatively, the polynucleotides encoding the present TALE proteins of the present invention, especially under mRNA form can be electroporated directly into blood cells by electroporation, by using for instance the steps described in WO2013176915 on pages 29 and 30 incorporated herein by reference.

[0191] The present invention also relates to methods for use of said polypeptides polynucleotides and proteins previously described for various applications ranging from targeted nucleic acid cleavage to targeted gene regulation. In genome engineering experiments, the efficiency of the nuclease fusion proteins as referred to in the present patent application, e.g. their ability to induce a desired event (Homologous gene targeting, targeted mutagenesis, sequence removal or excision, base editing) at a locus, depends on several parameters, including the specific activity of the nuclease, probably the accessibility of the target, and the efficacy and outcome of the repair pathway(s) resulting in the desired event (homologous repair for gene targeting, NHEJ pathways for targeted mutagenesis) , which can be assessed by standard techniques known in the art. The present invention more particularly relates to a method for modifying the genetic material of a cell within or adjacent to a nucleic acid target sequence by using one TALE fusion protein of the present invention. The double strand breaks caused by a TALE-nuclease, for instance, are commonly repaired through non-homologous end joining (NHEJ). NHEJ comprises at least two different processes. Mechanisms involve rejoining of what remains of the two DNA ends through direct re-ligation or via the so-called microhomology-mediated end joining. Repair via non- homologous end joining (NHEJ) often results in small insertions or deletions and can be used for the creation of specific gene knockouts.

[0192] Other aspects of the present disclosure relate to pharmaceutical compositions comprising any of the various components of the TALE proteins obtainable by the methods of the present invention (e.g., TALE-nuclease, TALE-deaminase, TALE-transcriptase, TALE-methylase, TALE- transposase...).

[0193] The term “pharmaceutical composition”, as used herein, refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the pharmaceutical composition comprises additional agents (e.g. for specific delivery, increasing half-life, or other therapeutic compounds). In some embodiments, the pharmaceutical composition are provided as reagents to correct genetic deficiencies, which can be used in vivo or ex-vivo, especially in gene therapy.

[0194] In preferred embodiments, the TALE proteins of the present invention are used to genetically modify blood cells ex-vivo, especially immune cells such as T-cells and NK cells, preferably primary cells to produce therapeutic cells for immunotherapy.

[0195] In some embodiments, the pharmaceutical composition is formulated in accordance with routine procedures as a composition adapted for intravenous or subcutaneous administration to a subject (e.g., a human). In some embodiments, pharmaceutical composition for administration by injection are solutions in sterile isotonic aqueous buffer. Where necessary, the pharmaceutical can also include a solubilizing agent and a local anesthetic such as lidocaine to ease pain at the site of the injection. Generally, the ingredients are supplied either separately or mixed together in unit dosage form, for example, as a dry lyophilized powder or water free concentrate in a hermetically sealed container such as an ampoule or sachette indicating the quantity of active agent. Where the pharmaceutical is to be administered by infusion, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline.

[0196] The pharmaceutical composition can be contained within a lipid particle or vesicle, such as a liposome or microcrystal, which is also suitable for parenteral administration. The particles can be of any suitable structure, such as unilamellar or plurilamellar, so long as compositions are contained therein. Compounds can be entrapped in “stabilized plasmid-lipid particles” (SPLP) containing the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipid, and stabilized by a polyethyleneglycol (PEG) coating (Zhang Y. P. et ah, Gene Ther. 1999, 6:1438-47). Positively charged lipids such as N-[l-(2,3-dioleoyloxi)propyl]-N,N,N- trimethyl-amoniummethylsulfate, or “DOTAP,” are particularly preferred for such particles and vesicles. The preparation of such lipid particles is well known. See, e.g., U.S. Patent Nos. 4,880,635; 4,906,477; 4,911 ,928; 4,917,951 ; 4,920,016; and 4,921 ,757; each of which is incorporated herein by reference.

[0197] The pharmaceutical composition described herein may be administered or packaged as a unit dose, for example. The term “unit dose” when used in reference to a pharmaceutical composition of the present disclosure refers to physically discrete units suitable as unitary dosage for the subject, each unit containing a predetermined quantity of active material calculated to produce the desired therapeutic effect in association with the required diluent; i.e., carrier, or vehicle.

[0198] Further, the pharmaceutical composition can be provided as a pharmaceutical kit comprising for example: (a) a container containing a compound of the invention in lyophilized form; and (b) a second container containing a pharmaceutically acceptable diluent (e.g., sterile water) for injection. The pharmaceutically acceptable diluent can be used for reconstitution or dilution of the lyophilized compound of the invention. Optionally associated with such container(s) can be a notice in the form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceuticals or biological products, which notice reflects approval by the agency of manufacture, use or sale for human administration.

[0199] EXAMPLES

[0200] Example 1 : MATERIALS AND METHODS

[0201] T cell culture

[0202] Cryopreserved human PBMCs were acquired from Allcells (Alameda, California, USA), PBMCs were cultured in X-vivo-15 media (Lonza Group), containing 20 ng / ml human IL-2 (Miltenyi Biotec), and 5% human serum AB (Seralab). Human T cell activator TransAct (Miltenyi Biotec) was used to activate T cells at 25pl TransAct per million CD3+ cells the day after thawing the PBMCs. TransAct was kept in the culture media for 72 hours.

[0203] TALEN production

[0204] TALEN (fusion TALE Nter (delta152)-repeats15,5-Cter(40)-Fok1 nuclease domain) were assembled using standard molecular biology and / or microbiology technics such as enzymatic restriction digestion, ligation, bacterial transformation and plasmid DNA extraction (NEB 10-beta competent E.coli for ccdB selection or NEB stable competent E.coli for blue / white screening) and plasmid DNA extraction. TALE DNA targeting array were assembled and cloned in respective TALEN backbones containing a single SV40 NLS or in the six combinations of three distinct NLS derived from SV40, Nucleoplasmin, and c-Myc e.g. S-N-C, S-C-N, N-S-C, N-C-S, C-S-N and C- N-S (SEQ ID NO:11, 12, 137 to 140 displayed in Table 4).

[0205] Small scale mRNA production

[0206] Plasmids of the different TALEN used in this study containing a T7 promoter and a polyA sequence were produced according to standard procedures. The plasmids were linearized with Sapl (NEB) and mRNA was produced by in vitro transcription (NEB HiScribe ARCA, NEB).

[0207] Small scale TALEN testing

[0208] Activated T cells were transferred into fresh complete media containing 20ng / ml human IL-2 (Miltenyi Biotec), and 5% human serum AB (Seralab) 10-12hrs before transfection. Cells were then harvested and washed once with warm PBS. One million PBS washed cells were pelleted and resuspended in 25pl BTXpress High Performance Buffer (BTX). 1 jig of mRNA encoding left or right TALEN monomers per million cells was mixed with the cells, the total volume brought up to 100pl with BTXpress High Performance Buffer, and then the cell mixture was electroporated using the BTX ECM 830 Square Wave Electroporation System (single pulse, 720V amplitude, 0.6ms duration). After electroporation, 80 pl warm complete media was added to the cuvette to dilute the electroporation buffer, the mixture was then carefully transferred to 400ml pre-warmed complete media in 48-well plates. TALEN transfected cells were incubated at 30°C for an overnight culture and then harvested post transfection for gDNA extraction and NGS analysis.

[0209] Large scale TALEN testing

[0210] T cells activated with TransACT for 3 days were transferred into fresh complete media containing 20ng / ml human IL-2 (Miltenyi Biotec), and 5% human serum AB (Seralab) 10-12hrs before transfection.

[0211] The harvested cells were washed twice with Cytoporation Media T (BTXpress, 47-0002). 5E6 washed cells were pelleted and resuspended in 180pl Cytoporation Media T. 5 .g / arm / million cells of TALEN mRNA was mixed with the cells to a final volume of 200 il and then the cell / mRNA mixture was electroporated using the BTX Pulse Agile in 0.4 cm gap cuvettes. After electroporation, 180 pl warm complete media was added to the cuvette to dilute the electroporation buffer, and the mixture was then carefully transferred to 2 ml pre-warmed complete media in 12-well plates. TALEN transfected cells were incubated at 30°C for an overnight culture, passaged into fresh complete media containing 20ng / ml human IL-2 (Miltenyi Biotec), and 5% human serum AB (Seralab), returned to 37°C, and then harvested at 72 hours post transfection for gDNA extraction and NGS analysis.

[0212] Genomic DNA extraction

[0213] Cells were harvested and washed once with PBS. Genomic DNA extraction was performed using Mag-Bind Blood & Tissue DNA HDQ kits (Omega Bio-Tek) following the manufacturer’s instructions.

[0214] Targeted PCR and NGS

[0215] 100pg genomic DNA was used per reaction in a 50pl reaction with Phusion High-Fidelity PCR Master Mix (NEB). The PCR condition was set to 1 cycle of 30s at 98°C; 30 cycles of 10s at 98°C, 30s at 60°C, 30s at 72°C; 1 cycle of 5 min at 72°C; hold at 4°C. The PCR product was then purified with Omega NGS beads (1:1.2 ratio) and eluted into 30pl of 10mM Tris buffer pH7.4. The second PCR which incorporates NGS indices was then performed on the purified product from the first PCR. 15 ul of the first PCR product were set in a 50pl reaction with Phusion High-Fidelity PCR Master Mix (NEB). The PCR condition was set to 1 cycle of 30s at 98°C; 8 cycles of 10s at 98°C, 30s at 62°C, 30s at 72°C; 1 cycle of 5 min at 72°C; hold at 4°C. Purified PCR products were sequenced on MiSeq (Illumina) on a 2x250 nano V2 cartridge. Oligo Capture Assay (OCA) sequencing and analysis

[0216] The OCA assays were performed as previously described, as for instance in Sachdeva, M. et al. [Repurposing endogenous immune pathways to tailor and control chimeric antigen receptor T cell functionality (2019) Nat Commun 10, 5100] in order to detect potential genome wide off-site cleavage due to TALEN heterodimers.

[0217] Example 2: 3xNLS combinations

[0218] A first set of experiments was performed to assess the effect of three NLS on TALEN activity compared to classical TALEN comprising a single SV40 derived NLS. Such test was performed in different genomic contexts with different TALEN constructs targeting either the B2M locus (SEQ target sequence ID NO:251), the CD3E locus (target sequence SEQ ID NO:252), or the CS1 / SLAMF7 locus (target sequence SEQ ID NO:253). For each locus, the various TALEN comprising different NLS combinations (S-N-C, S-C-N, N-S-C, N-C-S, C-S-N and C-N-S) were tested (except S-C-N combination for B2M targeting TALEN). Sequences are shown in Table 4: SEQ ID NO: 143 to 154 (TALEN B2M + 3NLS (S-N-C; S-C-N; N-S-C; N-C-S; C-S-N or C-N-S), SEQ ID NO: 157 to 168 (TALEN CD3E + 3NLS (S-N-C; S-C-N; N-S-C; N-C-S; C-S-N or C-N-S)), and SEQ ID NO: 171 to 182 (TALEN CS1 + 3NLS (S-N-C; S-C-N; N-S-C; N-C-S; C-S-N or C-N- S)). These heterodimers were compared to the classical TALEN heterodimers comprising a single SV40 NLS (SEQ ID NO: 141 + 142, SEQ ID NO: 155 + 156 and SEQ ID NO: 169 + 170 respectively) also displayed in Table 4.

[0219] The results in Figurel showed that S-N-C and N-S-C combinations were able to stimulate CS1 TALEN activity and importantly were the most active on B2M TALEN activity stimulation. In addition, these combinations were the only ones not having a negative impact on TALEN activity at the CD3E locus. The S-N-C and N-S-C combinations have therefore resulted into an improvement of TALEN activity, or at least did not impact TALEN activity. Thus, these combinations were selected to be studied further on a larger number of TALEN.

[0220] Example 3: 3NLS S-N-C and N-S-C comparison

[0221] A set of 10 targets within the HAVCR2 gene (SEQ ID NO:254 to 263) was studied. For each target, classical TALEN comprising SV40 NLS, were tested and compared to 3-NLS TALENs with S-N-C or N-S-C combinations at the N-terminus of each (SEQ ID NO: 183 to 242 as detailed in Table 4). The results shown in Figure 2 demonstrated that the TALEN displaying the S-N-C combination was showing the most robust increase in TALEN activity relative to the classical single NLS control. Indeed, such S-N-C combination could increase the TALEN activity up to 8 times and, at worse, would have no impact on TALEN activity.

[0222] Example 4: SNC combination results in highly active TALEN

[0223] A set of 4 heterodimers TALEN fusion proteins were designed under the S-N-C NLS combination scaffold tested before, cloned and transcribed under mRNA for expression in primary T-cells as set forth in Example 1 , to target respectively the CDKN2A gene and cleave the respective target sequences (SEQ ID NO:264 and 265) of the two isoforms of this gene, the MTAP gene and cleave the target sequence SEQ ID:266, and the CDKN2B gene to cleave the target sequence SEQ ID NO:267. The polypeptide sequences of the TALEN fusion monomers comprising the SNC NLS at the N-terminus designed to form the heterodimers cleaving each of the CDKN2A isoforms, MTAP and CDKN2B are listed in Table 4 (SEQ ID: 243 to 250).

[0224] Transfections were carried out in primary T-cells (PBMC) populations originating from three different donors for each tested TALEN heterodimers and mock sample.

[0225] The results in Figure 3 demonstrate that each TALEN heterodimers comprising at least a monomer with the SNC combination showed robust and high onsite activity. All these TALEN fusion proteins were able to induce above 90% of indels on average into their respective CDKN2A isoforms, MTAP and CDKN2B target sequences. Thus, such TALEN fusion proteins are useful to inactivate those respective genes, for instance in T-cells or NK cells to make those cells resistant to senescence as described in WQ2023025862.

[0226] Example 5: 3NLS (S-N-C) combination results in highly active and specific TALEN targeting TRAC and B2M genomic loci

[0227] Experiments were performed on T-cells to assess the effect of the three NLS (S-N-C combination) on TALEN activity and specificity at the TRAC locus (target sequence SEQ ID NO:268). The SNC polypeptide sequence was included on a single TALE arm (SEQ ID NO: 271) in combination with a second TALE arm encoded with a single SV40 NLS (SEQ ID NO: 272). The results in Figure 5 showed that such TRAC TALEN heterodimer with a single arm containing 3NLS and the second arm with containing a single SV40 NLS could induce more than 91% Indels at the target site.

[0228] OCA was performed in T-cells as referred to in Example 1 using same TALEN heterodimer, which led to the identification of 20 potential off-target site with very low OCA score (Figure 6). The first five potential off target sites were analyzed at the molecular level. As shown in Table 6, none of these potential off-target sites presented Indels formation (Table 6).

[0229] Table 6: On-Site and first five potential Off-Target sites sequencing obtained with TRAC TALEN displaying 3xNLS on one arm

[0230] Therefore, this 3-NLS TALEN targeting TRAC locus demonstrated a high specificity.

[0231] In previous experiments, when OCA was performed upon use of TALEN harboring a single SV40 NLS on both arms targeting the same sequence (SEQ ID NO:268), one real off-target site could be detected among the first five potential off-target sites, meaning that the addition of 3-NLS (S- N-C combination) on at least one arm contributed to increasing the specificity of the TALEN heterodimer.

[0232] Additional similar experiments were performed on T-cells to assess the effect of three NLS (SNC combination) on TALEN heterodimer activity and specificity when targeting the B2M locus (SEQ ID NO: 269). In these experiments, both arms included the 3-NLS structure (S-N-C) (SEQ ID 273 and 274). The results in Figure 7 showed that such B2M TALEN heterodimer containing three NLS (S-N-C combination) could induce more than 97% of Indels formation at the target site. In addition, OCA was performed using such TALEN heterodimers and led to the identification of 20 potential off- target sites with very low OCA score (Figure 8). The first five potential off target sites were analyzed at the molecular level. None of these potential off-target sites presented Indels formation (Table 7).

[0233] Table 7: On-Site and first five potential Off-Target sites sequencing obtained with B2M TALEN displaying 3xNLS on both arms

[0234] Therefore, this TALEN heterodimer with 3 NLS on both arms demonstrated not only a high activity but also a high specificity.

[0235] Example 6: 3NLS (S-N-C) combination results in highly active and specific TALEN targeting ALB genomic locus

[0236] An experiment was performed on HEP-G2 cell line to assess the effect of three NLS (SNC combination) inclusion on TALEN heterodimers activity and specificity. This was included on one TALE arm (SEQ ID NO: 277) in combination with a second TALE arm including a single SV40 NLS (SEQ ID NO: 276). The activity and specificity of such TALEN heterodimer were compared with those of a classical TALEN heterodimer comprising a single SV40 NLS on both arms (SEQ ID NO: 275, SEQ ID NO: 276), targeting the same polynucleotide sequence at the ALB locus (SEQ ID NO: 270). The results in Figure 9 showed that both ALB TALEN heterodimers, either with single SV40 NLS or with the combination of 3xNLS (SNC combination) and a single SV40 NLS, could induce more than 80 % and up to about 85% of indels on the target site with respect to the 3-NLS heterodimer. In addition, when both TALEN heterodimers were tested for their specificity by OCA, it was observed that the off-target sites (OT-2, OT-4 and OT-6) detected at low levels with the SV40 heterodimer were not detected after transfection with the TALEN heterodimer having the 3xNLS. None of the 30 potential off-target sites were matching the OT-2, OT-4 or OT-6 previously identified (Figure 10). In addition, and most importantly, when quantifying at the molecular level, OT-2, OT-4 and OT-6, no mutation could be detected at the molecular level when treated with the TALEN heterodimer comprising an arm with the 3xNLS combination and an arm with the SV40 NLS combination (Figure 11).

[0237] Altogether these experiments demonstrate that the S-N-C combination is not only able to improve (or maintain a high) activity of the TALEN but can also improve their specificity.

[0238] Example 7: 3NLS (S-N-C) combination results in highly active and specific TALEN targeting CIITA genomic locus with consequent loss of MHCII expression on T cells and protection against alloreactive CD4+T cells.

[0239] CIITA is a master transcriptional regulator of MHC II (MHC-DR, DP, DQ) gene expression in immune cells. TALEN-mediated knockout of CIITA in T-cells has been investigated to downregulate their surface expression of MHCII, thereby making them invisible to alloreactive CD4+ T cells. These experiments were done with a view to increasing the persistence of CIITA KO CAR T-cells in an allogeneic setting, thereby providing patients with engineered CAR-T cells with a longer window of activity.

[0240] As outlined in Figure 12A, transfections were carried out in primary T-cells (PBMC) populations originating from different donors. Each population was transfected with mRNA encoding TALEN targeting into the CIITA locus (target sequences T008446 or T008447, respectively referred to as SEQ ID NO:286 and 287 in Table 5) or without mRNA (mock sample). The results in Figure 12B demonstrated that both CIITA TALEN heterodimers comprising 3NLS (SEQ ID NO:278-279 and SEQ ID NQ:280-281 respectively) showed robust and high on-site activity (more than 50% indels) at the CIITA genomic locus, especially the TALEN heterodimer targeting T008446 (SEQ ID NO:278-279). Furthermore, CIITA knockout resulted in the downregulation of surface MHC-II expression on transfected T cells, as determined by flow cytometry (Figure 13). The functional effect of CIITA knockout in T-cells on their susceptibility to be alloreactive with respect to other CD4+T cells was further assessed using Mixed Lymphoycte Reaction. As outlined in Figure 14A, to generate alloreactive CD4+T cells, PBMC from donor A was thawed and mixed at a 1 :1 ratio with irradiated TRACKOB2MKOT cells from donor B at day 0. Donor B cells were knocked out for B2M to remove stimulation and subsequent expansion of Donor A CD8+cells. Cells were co-cultivated for 7 days before being re-stimulated with irradiated TRACKOB2MKOT cells from Donor B at day 7. The same procedure was reiterated one more time before harvesting Donor A alloresponsive T-cells at day 21. In parallel, Donor B CIITA knockout T-cells (target) generated by transfection with CIITA TALEN T008446 were labeled with CTV dye and eventually mixed with alloresponsive Donor A T-cells (effector) at an effector to target ratio of 0:1 , 1 :1 , 2:1 and 4:1. Cells were recovered after 24 hours of co-culture and analyzed by flow cytometry to determine the frequency of viable CTV(+) MHCII(+) and viable CTV(+) MHCII(-) T- cell targets which were remaining (Figure 14B). Resistance of CIITA KO T-cells to alloreactive CD4+ effector T cells is graphically represented as MHCII(-) enrichment (Figure 14C). Taken together, the data validates the genomic knockout at the CIITA locus using 3NLS CIITA TALEN, the downregulation of surface MHCII expression and resistance of CIITAKOT cells to alloreactive CD4+effector T cells.

[0241] Table 4: Polypeptide sequences used in the Examples

[0242] Table 5: Polynucleotide sequences used in the Examples

Claims

93CLAIMS1. A transcriptional Activator-like Effector (TALE) fusion protein comprising a core binding domain comprising TALE repeats, a N-terminal region and a C-terminal region, said N- terminal region comprising nucleus localization signals (NLS) and said C-terminal region being fused to a functional domain, wherein said N-terminal region comprises a fusion of at least two monopartite NLS, such as from C-myc (C-myc NLS) and / or from SV40 T antigen (SV40 T NLS), and at least one bipartite NLS, such as from Nucleoplasmin (Nucleoplasmin NLS).

2. A T ranscriptional Activator-like Effector (TALE) fusion protein according to claim 1 , wherein said N-terminal region comprises a fusion comprising SV40 T antigen NLS (SEQ ID NO:1), Nucleoplasmin NLS (SEQ ID NO:2) and C-myc NLS (SEQ ID NO:3).

3. A Transcriptional Activator- 1 ike Effector (TALE) fusion protein according to claim 1 or 2, wherein said fusion of three NLS has at least 80 %, preferably 90 %, 95%, or 99% identity with SEQ ID NO: 11 (in the order S-N-C) or SEQ ID NO: 12 (in the order N-S-C).

4. A T ranscriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 3, wherein said functional domain comprises a recruiting domain, such as a VP64 or Krabb.

5. A T ranscriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 3, wherein said functional domain comprises a catalytic domain.

6. A T ranscriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 5, wherein said functional domain comprises a combination of catalytic domain(s) and / or recruiting domain(s).

7. A T ranscriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 6, wherein said catalytic domain is from an enzyme selected from a methyltransferase, demethylase, exonuclease, endonuclease, nickase, deaminase, polymerase, transposase, integrase, recombinase, histone methyltransferase / acetyl transferase, reverse transcriptase or helicase.

948. A T ranscriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 7, wherein said catalytic domain modifies the genome and / or its organisation, such as by mutating, cleaving, nicking, or introducing epigenetic changes.

9. A T ranscriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 8, wherein said TALE fusion protein is a TALE-nuclease, TALE-base editor, TALE- transcriptional modulator, TALE-integrase, TALE-transposase or a TALE-recombinase.

10. A Transcriptional Activator- 1 ike Effector (TALE) fusion protein according to any one of claims 1 to 9, wherein said N-terminal region comprises a polypeptide sequence showing at least 85% sequence identity with SEQ ID NO:30 (classical N-ter N152).11 . A Transcriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 10, wherein said C-terminal region consists of a polypeptide sequence from about 10 to 40 residues, preferably comprising SEQ ID NO:282.

12. A Transcriptional Activator-like Effector (TALE) fusion protein according to any one of claims 1 to 10, wherein said C-terminal region consists of a polypeptide sequence from 40 to 80 residues comprising a sequence having at least 85% identity with:SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GL (SEQ ID NO:13)SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GLPHAPALIX3RT(SEQ ID NO: 14), orSIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX1X2GLPHAPALIX3RTNRRIPERTSH (SEQ ID NO:15), wherein Xi, X2, and X3, are any amino acid, preferably H (histidine) or a R (arginine) residue.

13. The transcriptional activator-like Effector (TALE) fusion protein according to any one of claims 1 to 12, wherein at least one of said TALE repeats comprises D (aspartic acid) residues at positions 4 and 32 with respect to any of the canonical sequence of AvrBs3 of SEQ ID NO:26 to 29.

14. The transcriptional activator-like Effector (TALE) fusion protein according to any one of claims 1 to 13, wherein at least one of said TALE repeats comprises one of the sequences:LTPQQVVAIASX4X5GGKQALETVQRLLPVLCQAHG (SEQ ID NO: 16), LTPEQVVAIASX4X5GGKQALETVQALLPVLCQAHG (SEQ ID NO: 17), LTPDQWAIASX4X5GGKQALETVQQLLPVLCQDHG (SEQ ID NO:18), LTPDQLVAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NO:19),95LTPDQMVAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NQ:20),LTPDQWAIASX4X5GGKQALETVQRLLPVLCQDQG (SEQ ID NO:21), LTLDQWAIASX4X5GGKQALETVQRLLPVLCQDHG (SEQ ID NO:22), wherein X4X5 is an amino acid forming a variable di-residue.

15. The activator-like Effector (TALE) fusion protein according to any one of claims 1 to 14, which has at least 90% identity with one polypeptide sequence selected from Table 4.

16. The activator-like Effector (TALE) fusion protein according to any one of claims 1 to 14, wherein said TALE is fused to a nuclease domain to form a TALE-nuclease.

17. The TALE-nuclease according to claim 16, wherein said nuclease domain comprises a catalytic domain from Fok-1.

18. The TALE-nuclease according to claim 17, wherein said nuclease domain comprises a polypeptide sequence that shows at least 85% identity, preferably at least 90%, more preferably at least 95%, even more preferably 99% identity with SEQ ID NO: 105 (Fok1 catalytic domain).

19. The TALE-nuclease according to any one of claims 16 to 18, wherein said nuclease domain has at least one amino acid substitution at positions corresponding to 13, 52, 57, 59, 61 , 65, 84, 85, 88, 91 , 92, 95, 98, 103, 109, 110, 111 , 113, 119, 143, 148, 152, 158, 159, 160, 167, 169, 170, and 194 into SEQ ID NQ:105.

20. The TALE-nuclease according to any one of claims 16 to 19, which has at least 90% identity with a polypeptide sequence selected from SEQ ID NO:243 to 250.

21. The TALE-nuclease according to claim 20 targeting CDKN2A, which has at least 90% identity with a polypeptide sequence selected from SEQ ID NO:243 to 246.

22. The TALE-nuclease according to claim 20 targeting MTAP, which has at least 90% identity with a sequence selected from SEQ ID NO:247 and 248.

23. The TALE-nuclease according to claim 20 targeting CDKN2B, which has at least 90% identity with a polypeptide sequence selected from SEQ ID NO:249 and 250.

24. The TALE-nuclease according to any one of claims 16 to 19 targeting B2M, which has at least 90% identity with a polypeptide sequence selected from SEQ ID NO:273 and 274.9625. The TALE-nuclease according to any one of claims 16 to 19 targeting TRAC, which has at least 90% identity with the polypeptide sequence SEQ ID NO: 271.

26. The TALE-nuclease according to any one of claims 16 to 19 targeting ALB, which has at least 90% identity with the polypeptide sequence SEQ ID NO:277.

27. The TALE-nuclease according to any one of claims 16 to 19 targeting CIITA, which has at least 90% identity with a sequence selected from SEQ ID NO:278, 279, 280 and 281.

28. The transcriptional activator-like Effector (TALE) protein according to any one of claims 1 to 15, wherein said TALE protein is fused to a deaminase domain to form a TALE-base editor.

29. The transcriptional activator-like Effector (TALE) protein according to any one of claims 1 to 15, wherein said TALE protein is fused to a transcriptional modulator domain to form a TALE-transcriptional modulator, such as a TALE-transcriptional activator or a TALE- transcriptional repressor.

30. The TALE fusion protein according to any one of claims 1 to 29, for use in the treatment of a genetic disease.

31. The TALE fusion protein according to any one of claims 1 to 29, for use in gene therapy.

32. The TALE fusion protein according to any one of claims 1 to 29, for use in cell therapy.

33. The TALE fusion protein according to any one of claims 1 to 29, for use in the manufacture of gene edited cells.

34. The TALE fusion protein according to any one of claims 1 to 29, for use in the production of plant engineered cells.

35. A polynucleotide encoding the TALE fusion protein according to any one of claims1 to 29.

36. A vector comprising the polynucleotide according to claim 35.9737. A cell comprising the polynucleotide according to claim 35, a vector according to claim 36 or TALE fusion protein according to any one of claims 1 to 19.

38. A method for modifying the genome of a cell and / or its organisation, said method comprising the step of expressing or introducing in said cell the polynucleotide according to claim 35, vector according to claim 36, or TALE fusion protein according to any one of claims 1 to 19.

39. The method according to claim 38, wherein said cell is an immune cell, such as a T-cell or a NK cell.

40. The method according to claim 38, wherein said cell is a hematopoietic stem cell (HSC).

41. The method according to claim 38, wherein said cell is a hepatocyte.

42. The method according to any one of claims 38 to 41 , wherein the cells are primary cells.

43. The method according to any one of claims 38 to 42, wherein the cell is an iPS cell.

44. The method according to any one of claims 38 to 39, 40, 42 or 43, wherein said method further comprises expressing a chimeric antigen receptor or a recombinant TCR.

45. The method according to any one of claims 38 to 44, further comprising the steps of expanding the modified cells to produce a therapeutic composition of cells.

46. The method according to any one of claims 38 to 45, further comprising the steps of conditioning the cells for therapeutic use, such as freezing injectable doses of the cells.

Citation Information

Patent Citations

  • A laglidadg homing endonuclease cleaving the t cell receptor alpha gene and uses thereof

    EP3004338A1

  • Method for the generation of compact tale-nucleases and uses thereof

    EP3320910A1

  • Process for amplifying, detecting, and / or-cloning nucleic acid sequences

    US4683195A

  • Dehydrated liposomes

    US4880635A

  • Antineoplastic agent-entrapping liposomes

    US4906477A