A novel TALE protein scaffold with improved on-target / off-target activity ratio
Patent Information
- Application Number
- JP2024530470
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-03-15
- Filing Date
- 2022-11-23
- Publication Date
- 2025-12-01
AI Technical Summary
Current TALE protein constructs for gene therapy do not meet the required specificity and efficiency standards, often leading to off-target binding and reduced therapeutic safety.
A novel TALE scaffold design incorporating specific mutations in the AvrBs3 repeats, combined with a C-terminal and N-terminal sequence, enhances the specificity and activity of TALE fusion proteins by improving their interaction with target sequences while maintaining catalytic function.
The novel TALE scaffold achieves a higher on-target/off-target activity ratio, ensuring safer and more effective genetic modifications in mammalian cells.
Smart Images

Figure 00000000_0001_ABST 
Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] FIELD OF THEINVENTION The present invention relates to the design of improved TALE protein fusions useful as sequence-specific genomic reagents that exhibit higher on-target / off-target activity ratios, with the goal of generating said reagents for genetically modifying the genomes of different types of cells, particularly mammalian cells, in order to use safer reagents, especially in gene therapy. [Background technology]
[0002] 2. Background of the Invention Artificial transcription activator-like effectors (TALEs) form a special class of proteins capable of binding to DNA, originally derived from the plant pathogenic bacterium Xanthomonas (Kay S. et al. (2007) A bacterial effector acts as a plant transcription factor and induces a cell size regulator. Science 318: 648-651). Artificial TALE proteins have emerged as versatile and sequence-specific tools offering flexible applications, based on the elucidation of the DNA recognition "code" that links the amino acid sequence of the TALE with its binding genomic DNA sequence (Moscou JM et al. (2009) A Simple Cipher Governs DNA Recognition by TAL Effectors. Science. 326:1501).
[0003] TALE binding is essentially driven by a series of 33-35 amino acid long repeats that differ at two positions, the so-called repeat variable dipeptides (RVDs). Each base of one strand in the DNA target is contacted by a single repeat with predictable specificity due to the linear arrangement of the RVDs. Biochemical structure-function studies suggest that the amino acid present at position 13 uniquely identifies the nucleotide in the major groove of the DNA target [Deng D., et al. (2012) Structural basis for sequence-specific recognition of DNA by TAL effectors. Science 335:720-723 (Non-Patent Document 3); Stella S., et al. (2013) Structure of the AvrBs3-DNA complex provides new insights into the initial thymine-recognition mechanism. Acta Crystallogr Sect D Biol Crystallogr 69(9):1707-1716 (Non-Patent Document 4)]. This DNA-protein interaction unit is stabilized by an amino acid at position 12. For the generation of TALEs with variable accuracy and binding affinity, six conventional RVDs are commonly used (NG, HD, NI, NK, NH, and NN). HD and NG bind cytosine (C) and thymine (T), respectively. NN is a degenerate RVD that shows binding affinity for both guanine (G) and adenine (A), but its specificity for guanine has been reported to be stronger. RVD NI binds A and NK binds G. It is noteworthy that the binding affinity of TALEs is affected by the methylation state of the target DNA sequence [Streubel J, et al. (2012) TAL effector RVD specificities and efficiencies. Nat Biotechnol 30(7):593-595]. Methylated cytosines are not efficiently bound by canonical RVDs.However, this can be accommodated by some degree of degeneracy in TALEs, as described by Valton J, et al. [Overcoming transcription activator-like effector (TALE) DNA binding domain sensitivity to cytosine methylation (2012) J. Biol. Chem. 287(46):38427-38432 (Non-Patent Document 6)]. This code has been adopted to effectively engineer the specificity of TALE DNA binding scaffolds through modular assembly to form different combinations of TALE proteins with various enzymatic domains, such as transcription activators, repressors, base editors, or nucleases with the potential to act on genomic sequences (Voytas et al. (2011) TAL effectors: Customizable proteins for DNA targeting. Science 333(6051):1843-6 (Non-Patent Document 7)]. Compared to zinc finger protein fusions, TALE proteins have remarkably emerged as important DNA binding scaffolds governed by a simple code without significant restrictions. Their compatibility with a wide range of epigenetic modifiers is remarkable [Laufer BI, et al. (2015) Strategies for precision modulation of gene expression by epigenome editing: an overview. Epigenetics Chromatin 8(1):34 (Non-Patent Document 8)], and it is believed that these DNA-binding proteins can be used to target epigenetic effector domains to any locus in the genome [Cano-Rodriguez D., Rots MG (2016) Epigenetic editing: on the verge of reprogramming gene expression at will. Curr Genet Med Rep 4(4):170-179 (Non-Patent Document 9)].
[0004] Such TALE protein fusions can result in TALE artificial transcription factors, which have been generated by fusing TALEs with a 16-amino acid peptide from herpes simplex virus (VP16) as a transactivation domain [Zhang, F. et al. Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. Nature Biotechnol. 29:149-153 (Non-Patent Document 10)]. In contrast to zinc finger binding domains, which have encountered many off-target effects, TALE transcription activators are efficient transcription modulators with only 10.5 repeats with effector modules fused to the carboxyl terminus [Miller, J., et al. (2011) A TALE nuclease architecture for efficient genome editing. Nat Biotechnol. 29, 143-148 (Non-Patent Document 11)]. TALEs in the form of activators can also be used to control gene expression upon external stimuli, such as chemical changes or light stimuli, in various organisms, including plants and animals.
[0005] TALE repressors can be generated by fusing a TALE to either Krüppel-binding box (KRAB), Sid4, or EAR repression domain (SRDX) repressors [Cong L, et al. (2012) Comprehensive interrogation of natural TALE DNA-binding modules and transcriptional repressor domains. Nat Commun 3(1):968 (Non-Patent Document 12)].
[0006] TALE base editors can be generated by fusion of TALEs with deaminases and occasionally with other DNA repair proteins. Base editor catalytic domains can introduce single nucleotide variants at desired loci in DNA (nuclear or organelle) or RNA in both dividing and non-dividing cells. Broadly, there are two types: DNA base editors that directly induce targeted point mutations in DNA and RNA base editors that convert one ribonucleotide to another in RNA. Currently available DNA base editors can be further classified into cytosine base editors (CBEs), adenine base editors (ABEs), C to G base editors (CGBEs), double base editors, and organelle base editors. For example, Mok et al. [A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing (2020) Nature. 583:631-637] recently developed a base editing approach using a bacterial cytidine deaminase toxin, i.e., DddAtox, to demonstrate efficient C to T base conversion in vitro. In this approach, the non-toxic half of the truncated DddAtox is fused to a transcription activator-like effector (TALE) protein, which can be custom designed to recognize a given target DNA sequence and form a functional cytosine deaminase within the editing window to induce C to T base editing at the target site in the genomic DNA. Such DddA-TALE fusion deaminase constructs were subsequently used to achieve mitochondrial DNA editing in mice [Lee, H., et al. (2021) Mitochondrial DNA editing in mice with DddA-TALE fusion deaminases. Nat Commun 12: 1190 (Non-Patent Document 14)].
[0007] TALE nucleases can be generated by fusion of TALEs with various nuclease catalytic domains. The commonly used TALEN® system, which provides specific nucleases as fusions of the TALE scaffold with the catalytic domain of the Fok1 restriction enzyme, has proven through many studies to be highly specific since it combines two TALE dimers that bind together at a selected locus. TALEN heterodimers (right and left) generally bind to opposite strands about 10-20 pb apart from each other (spacer), allowing the nuclease Fok1 to dimerize and induce double-stranded cleavage between the binding sites in the spacer. This heterodimer setup allows for increased sequence specificity based on an extended target sequence of up to 40 base pairs surrounded by the binding sites of the two TALEs. Such TALE nucleases are currently being developed as therapeutic grade nuclease reagents in gene therapy, especially for generating allogeneic CAR-T cells [Poirot et al. (2015) Multiplex Genome-Edited T-cell Manufacturing Platform for “Off-the-Shelf” Adoptive T-cell Immunotherapies Cancer Res 75(18):3853-3864 (Non-Patent Document 15); Quasim W. et al. (2017) Molecular remission of infant B-ALL after infusion of universal TALEN gene-edited CAR T cells. Science translational medicine (9)374 (Non-Patent Document 16)]. Classical TALEN monomer constructs are generally based on truncated versions of the TALE binding domain from AvrBs3 protein fused to the catalytic domain of Fok1, such as those first described by Voytas et al. in WO2011072246 (Patent Document 1).Such TALE-nuclease fusion proteins, referred to herein as "canonical", typically contain from 5' to 3': (1) a truncated N-terminal region from AvrBs3, including at least 150 amino acids proximal to the binding domain; (2) an engineered central DNA-binding domain, generally including 12-28 repeats that are assembled to target a genomic nucleotide sequence; these selected repeats are followed by wild-type half-repeats of only 20 amino acids from AvrBs3, designed to bind to the 3' end of the target DNA sequence; and (4) a linker sequence of at least 40 amino acids from the C-terminal wild-type region of AvrBs3, fused to a wild-type Fok1 nuclease catalytic domain. Generally, the fusion protein further contains a nuclear localization signal (NLS) of AvrBs3 fused to the truncated N-terminal region. These programmable TALE DNA-binding domains have been shown to improve specificity and efficacy [Juillerat A, et al. (2015) Optimized tuning of TALEN specificity using non-conventional RVDs. Sci Rep 5(1)(Non-Patent Document 17)], and several studies have proposed enhancing the core TALE domain by various truncations along with the use of additional or alternative RVDs [Miller, JC et al. (2011) A TALE nuclease architecture for efficient genome editing. Nat. Biotechnol. 29, 143(Non-Patent Document 11)].
[0008] Such custom TALE proteins have proven to be powerful reagents for targeting genomic DNA sequences of interest in almost all cell types [Weeks DP,. et al. Use of designer nucleases for targeted gene and genome editing in plants (2016) Plant Biotechnology Journal.14:483-495(Non-Patent Document 18); Mussolino C. et al. (2014) TALENs facilitate targeted genome editing in human cells with high specificity and low cytotoxicity. Nucleic Acids Res 42(10):6762-6773(Non-Patent Document 19)]. Moreover, TALE proteins engineered according to this standard scheme are highly similar to each other in terms of structure and sequence identity. In fact, only amino acids at positions 12 and 13 of each repeat in the central DNA binding domain need to differ in order to adapt the scaffold to a novel target sequence.
[0009] Nevertheless, with the development of TALE nucleases for human gene therapy, standard TALE constructs do not always meet the specificity and efficiency levels required for therapeutic safety. Depending on the sequence targeted in the genome and its inherent variability in human populations, TALE scaffolds sometimes require further improvement to reduce potential off-target binding and increase their catalytic activity. Previous methods that consist in including additional or non-conventional RVDs may not be sufficient in all situations. In fact, specificity and catalytic activity are often in balance, and it can be difficult to find a good compromise that preserves safety and efficiency.
[0010] To go beyond the current high standards of engineered TALE proteins, we designed novel TALE scaffolds that combine different sets of mutations. The resulting TALE fusion proteins based on these novel scaffolds exhibit better specificity and remain adaptable to any target sequence and RVD tuning while retaining most of their catalytic activity. Thus, our invention provides a platform for the rational design of higher therapeutic grade TALE catalytic proteins. [Prior art documents] [Patent documents]
[0011] [Patent Document 1] WO2011072246 [Non-patent literature]
[0012] [Non-Patent Document 1] Kay S. et al. (2007) A bacterial effector acts as a plant transcription factor and induces a cell size regulator. Science 318: 648-651 [Non-Patent Document 2] Moscou JM et al. (2009) A Simple Cipher Governs DNA Recognition by TAL Effectors. Science. 326:1501 [Non-Patent Document 3] Deng D., et al. (2012) Structural basis for sequence-specific recognition of DNA by TAL effectors. Science 335:720-723 [Non-Patent Document 4] Stella S., et al. (2013) Structure of the AvrBs3-DNA complex provides new insights into the initial thymine-recognition mechanism. Acta Crystallogr Sect D Biol Crystallogr 69(9):1707-1716 [Non-Patent Document 5] Streubel J, et al. (2012) TAL effector RVD specificities and efficiencies. Nat Biotechnol 30(7):593-595 [Non-Patent Document 6] Valton J, et al. [Overcoming transcription activator-like effector (TALE) DNA binding domain sensitivity to cytosine methylation (2012) J. Biol. Chem. 287(46):38427-38432 [Non-Patent Document 7] Voytas et al. (2011) TAL effectors: Customizable proteins for DNA targeting. Science 333(6051):1843-6 [Non-Patent Document 8] Laufer BI, et al. (2015) Strategies for precision modulation of gene expression by epigenome editing: an overview. Epigenetics Chromatin 8(1):34 [Non-Patent Document 9] Cano-Rodriguez D., Rots MG (2016) Epigenetic editing: on the verge of reprogramming gene expression at will. Curr Genet Med Rep 4(4):170-179 [Non-Patent Document 10] Zhang, F. et al. Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. Nature Biotechnol. 29:149-153 [Non-Patent Document 11] Miller, J., et al. (2011) A TALE nuclease architecture for efficient genome editing. Nat Biotechnol. 29, 143-148 [Non-Patent Document 12] Cong L, et al. (2012) Comprehensive interrogation of natural TALE DNA-binding modules and transcriptional repressor domains. Nat Commun 3(1):968 [Non-Patent Document 13] Mok et al. [A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing (2020) Nature. 583:631-637 [Non-Patent Document 14] Lee, H., et al. (2021) Mitochondrial DNA editing in mice with DddA-TALE fusion deaminases. Nat Commun 12: 1190 [Non-Patent Document 15] Poirot et al. (2015) Multiplex Genome-Edited T-cell Manufacturing Platform for “Off-the-Shelf” Adoptive T-cell Immunotherapies Cancer Res 75(18):3853-3864 [Non-Patent Document 16] Quasim W. et al. (2017) Molecular remission of infant B-ALL after infusion of universal TALEN gene-edited CAR T cells. Science translational medicine (9)374 [Non-Patent Document 17] Juillerat A, et al. (2015) Optimized tuning of TALEN specificity using non-conventional RVDs. Sci Rep 5(1) [Non-Patent Document 18] Weeks DP,. et al. Use of designer nucleases for targeted gene and genome editing in plants (2016) Plant Biotechnology Journal.14:483-495 [Non-Patent Document 19] Mussolino C. et al. (2014) TALENs facilitate targeted genome editing in human cells with high specificity and low cytotoxicity. Nucleic Acids Res 42(10):6762-6773 Summary of the Invention
[0013] The present invention aims to improve the specificity and / or activity of TALE fusion proteins whose binding domains are generally based on assemblies of AvrBs3 repeats derived from the original Xanthomonas genomic sequence.
[0014] According to the present invention, the original AvrBs3 repeat of the TALE core binding domain has the following SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:4, in which X1, X2, and X3 represent H (histidine) or R (arginine), preferably R: TIFF2024540639000002.tif35153. X1, X2, and X3 may be the same or different.
[0015] Generally, the TALE core binding domain is fused to an N-terminal region, which preferably comprises or consists of a polypeptide sequence exhibiting at least 85%, preferably at least 90%, more preferably at least 95% identity to SEQ ID NO:1.
[0016] According to a preferred embodiment, the TALE core binding domain comprises an AvrBs3-like repeat, e.g., one that comprises a D (aspartic acid) at position 4 (D4) and / or a D (aspartic acid) at position 32 (D32) amino acid substitution at position 32 of the polypeptide sequence of the AvrBs3-like repeat.
[0017] In some embodiments, the AvrBs3-like repeat has the following polypeptide sequence: TIFF2024540639000003.tif44134, where X4X5 are two residues that interact with a given nucleotide base pair in the target sequence. X4 and X5 may be any amino acid or may be null (to specify missing residues in the RVD). *(indicated as an asterisk). X4 and X5 may be the same or different.
[0018] These selected sequences, and especially their combinations, were found by the inventors to improve the overall TALE protein structure, leading to closer interactions with its target sequence reflecting higher specificity, while the structure remains flexible enough to maintain the activity of the catalytic domain fused to said binding domain and efficiently process DNA upstream or downstream of the binding site.
[0019] The present invention also encompasses methods of producing or expressing TALE fusion proteins, such as TALE nucleases, TALE base editors, or TALE transcriptional modulators, in cells to target genomic sequences.
[0020] In particular, the present invention provides a method for designing a TALE protein for introducing a genetic modification into a polynucleotide sequence, comprising the steps of: (a) selecting a polynucleotide target sequence for which genetic modification is intended; (b) assembling polynucleotide sequences encoding AvrBs3-like repeats to form a polynucleotide encoding a TALE binding domain for binding to said selected polynucleotide target sequence; (c) a polynucleotide encoding the TALE binding domain, comprising at least (1) a polynucleotide sequence encoding an N-terminal domain comprising a sequence having at least 85% identity to SEQ ID NO:1; and (2) A polynucleotide sequence encoding a C-terminal domain consisting of a polypeptide sequence of 40 to 80 residues containing a sequence having at least 85%, preferably 90%, more preferably 95%, and even more preferably 99% identity to SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:4, in which X1, X2, and X3 represent R (arginine) or H (histidine). and optionally, (d) fusing a polynucleotide sequence encoding a catalytic domain, e.g., a nuclease or deaminase, to a polynucleotide sequence encoding the C-terminal domain; (e) fusing a polynucleotide encoding an NLS (nuclear localization signal), such as those listed in Table 1, to the polynucleotide sequence encoding the N-terminal domain.
[0021] The methods of the present invention are directed to producing polynucleotides encoding TALE fusion proteins and the polypeptides resulting from their expression.
[0022] The TALE proteins according to the present invention generally exhibit improved on-target / off-target activity ratios against target genomic sequences compared to prior art TALE fusion proteins.
[0023] The method of the present invention may further include a step of expressing the novel polynucleotide sequence in a cell to obtain, for example, cleavage, base substitution, or transcriptional activation at the target genomic locus, and comparing its efficiency with other TALE proteins to select one with a higher on-target / off-target activity ratio.
[0024] The method of the invention can also include a step in which, in addition to the D4 and D32 substitutions, at least one of said AvrBs3-like repeats is further mutated at 1, 2, 3, up to 5 amino acid positions.
[0025] The method of the present invention may also include a step in which the C-terminal domain of the TALE protein is mutated to introduce 1 to 5 positively charged amino acids, such as lysine (K), arginine (R), or histidine (H), in addition to the X1, X2, and X3 positions mentioned above.
[0026] The method of the present invention may also comprise the further step of introducing amino acid substitutions into the catalytic domain of the TALE protein to enhance its catalytic activity.
[0027] In a further aspect, the present invention relates to recombinant Transcription Activator-Like Effector (TALE) proteins comprising one or several AvrBs3-like repeats, typically comprising 8-20 repeats, preferably 8-18, more preferably 10-16, or alternatively 5-12 repeats, in the context of smaller genomes such as mitochondrial genomes being considered.
[0028] In some embodiments, the TALE proteins according to the present invention combine RVD repeats, preferably AvrBs3-like repeats containing the above amino acid substitutions, with a C-terminal sequence such as SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:4, and an N-terminal sequence comprising SEQ ID NO:1.
[0029] The recombinant core TALE proteins of the present invention are intended to be fused to various catalytic domains already described in the prior art (see WO2012138939), in particular catalytic domains from nucleases, such as Fok1 or Tev1, deaminases, such as cytidine deaminase toxins, and transcriptional modulators, such as the transactivator VP16.
[0030] In some cases, the TALE protein of the present invention is a TALE nuclease comprising a polypeptide sequence exhibiting at least 85% identity, preferably at least 90%, more preferably at least 95%, even more preferably 99% identity to SEQ ID NO:109, said polypeptide sequence corresponding to the catalytic domain of Fok-1 in which amino acid substitutions have been introduced to enhance the cleavage activity and improve the specificity of the TALE nuclease.
[0031] The present application discloses numerous examples of TALE proteins, in particular TALE base editors and TALE nucleases generated according to the principles of the present invention, also referred to as "TALE V2", directed to gene loci selected from TCRα, B2m, PD1, CTLA4, CISH, LAG3, TGFBRII, TIGIT, CD38, IgH, GADPH S100A9, PIK3CD, AAVS1, and CCR5, such as those listed in Tables 4 and 5.
[0032] The present invention encompasses vectors comprising the polynucleotide sequences and polypeptide sequences or reagents obtainable by the present invention, and their use for cell transformation and genetic modification. [Brief description of the drawings]
[0033] [Figure 1] Structure of an exemplary TALE-nuclease protein fusion according to the present invention. [Diagram 2] A chart comparing the % indels (cleavage activity) obtained by the VO, V0.1, and VO.2 TALE protein structures, as detailed in the Examples. [Diagram 3] Diagram comparing the overall off-site cleavage derived from oligo capture analysis (OCA) obtained with the VO and V0.1 TALE protein structures. [Figure 4] Diagram comparing indel formation of V1 and V1.2 TALE proteins according to the invention with the canonical TALE structure VO. A: % indels compared to VO (which maintains cleavage activity at the CS1 target site), B: % indels observed at off-site locus OS1, C: % indels observed at off-site locus OS2 (V1 and V1.2 TALE structures abolish off-site cleavage). [Diagram 5] Graph showing the reduction in overall off-site cleavage using the V1 and V1.2 TALE protein structures according to the present invention (oligo capture assay), as detailed in the Examples. [Figure 6-1]FIG. 6: Diagram showing the % of indels obtained on-site (CS1 target sequence) and off-site (OS1 and OS2 loci) when alanine substitutions were introduced into the amino acid sequence of Fok1 at the positions indicated on the x-axis (relative to wild-type Fok1). [Figure 6-2] See description of Figure 6-1. [Figure 6-3] See description of Figure 6-1. [Figure 6-4] See description of Figure 6-1. [Figure 6-5] See description of Figure 6-1. [Figure 7] Diagram showing the fold reduction in on-site indels compared to WT Fok1 (black bars) and off-site indels observed in OS1 compared to WT (white bars) when using TALE nucleases with the best substitution positions introduced into the Fok1 catalytic domain. [Figure 8] Schematic diagram of a TALE base editor scaffold according to the invention for inactivating the CD52 gene as described in Example 5. [Figure 9] Histogram comparing % indels (cleavage activity) obtained with TALE nucleases targeting TGFBRII with either VO-VO, V1.2-V0, or V1.2-V1.2 heterodimer structures at on-target (on-site) or off-target sites (OT#). V1.2 contains a TALE structure according to the present invention, as detailed in Example 6. [Figure 10] 1 is a graph showing the results of an oligo capture assay (OCA) performed on cells transfected with TALE nuclease V2 designed according to the present invention to target TIGIT. [Figure 11] This is a graph showing the results of an oligo capture assay (OCA) performed on cells transfected with TALE nuclease V2 designed according to the present invention to target CISH (against three different target sequences 1, 2, and 3). [Figure 12]This is a graph showing the results of an oligo capture assay (OCA) performed on cells transfected with TALE nuclease V2 designed according to the present invention to target CD38 (for two different target sequences 1 and 2). [Figure 13] This is a graph showing the results of an oligo capture assay (OCA) performed on cells transfected with TALE nuclease V2 designed according to the present invention to target IgH (against two different target sequences 1 and 2). [Figure 14] This is a graph showing the results of an oligo capture assay (OCA) performed on cells transfected with TALE nuclease V2 designed according to the present invention to target GAPDH (against two different target sequences 1 and 2). [Figure 15] Percentage of indels measured in cells transfected with each TALE nuclease V2 according to the present invention presented in Example 7. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0034] Table Description Table 1: Examples of NLS polypeptide sequences Table 2: Examples of linkers that can be included in TALE fusion proteins Table 3: Examples of catalytic domains Table 4: Examples of TALE proteins according to the present invention useful in gene therapy or adoptive cell therapy Table 5: Polypeptide sequences used in the examples Table 6: Polynucleotide sequences used in the examples
[0035] Detailed Description of the Invention Unless otherwise defined herein, all technical and scientific terms used have the same meaning as commonly understood by one of ordinary skill in the art of gene therapy, biochemistry, genetics, and molecular biology.
[0036] All methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, and suitable methods and materials are described herein.All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety.In case of discrepancy, the present specification, including definitions, takes precedence.Furthermore, materials, methods, and examples are illustrative only and are not intended to be limiting unless otherwise specified.
[0037] The practice of the present invention employs, unless otherwise indicated, conventional techniques of cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and immunology, which are within the skill of the art, and such techniques are fully explained in the literature. For example, Current Protocols in Molecular Biology [Frederick M. AUSUBEL, 2000, Wiley and son Inc, Library of Congress, USA); Molecular Cloning: A Laboratory Manual, Third Edition, (Sambrook et al, 2001, Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press; Oligonucleotide Synthesis (MJ Gait ed., 1984); Mullis et al. al. U.S. Patent No. 4,683,195; Nucleic Acid Hybridization (BD Harries & SJ Higgins eds. 1984); Transcription And Translation (BD Hames & SJ Higgins eds. 1984); Culture Of Animal Cells (RI Freshney, Alan R. Liss, Inc., 1987); Immobilized Cells And Enzymes (IRL Press, 1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); Methods In ENZYMOLOGY (J. Abelson and M. Simon, eds.-in-chief, Academic Press, Inc., New York), in particular, volumes 154 and 155 (Wu et al. eds.) and volume 185, "Gene Expression Technology" (D. Goeddel, ed.); Gene Transfer Vectors For Mammalian Cells (JH Miller and MP Calos eds., 1987, Cold Spring Harbor Laboratory); Immunochemical Methods In Cell And Molecular Biology (Mayer and Walker, eds., Academic Press, London, 1987); Handbook Of Experimental Immunology, Volumes I-IV (DM Weir and CC Blackwell, eds., 1986); and Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1986].
[0038] Thus, the present invention provides methods for purposefully designing and generating TALE proteins that exhibit reduced off-target DNA binding, which can be fused to various catalytic domains to form highly specific and active TALE fusion proteins, in particular TALE nucleases and TALE base editors.
[0039] According to some embodiments, the present invention provides a method for designing a TALE protein for introducing genetic modifications into a polynucleotide sequence, comprising one or several of the following steps: (a) selecting a polynucleotide target sequence for which genetic modification is intended; (b) assembling polynucleotide sequences encoding AvrBs3-like repeats to form a polynucleotide encoding a TALE binding domain for binding to said selected polynucleotide target sequence; (c) a polynucleotide encoding the TALE binding domain, comprising at least (1) a polynucleotide sequence encoding an N-terminal domain comprising a sequence having at least 85% identity to SEQ ID NO:1; and (2) A polynucleotide sequence encoding a C-terminal domain consisting of a polypeptide sequence of 40 to 80 residues containing a sequence having at least 85%, preferably 90%, more preferably 95%, and even more preferably 99% identity to SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:4, in which X1, X2, and X3 represent R (arginine) or H (histidine). and optionally, fusing the
[0040] Generally, the above steps can be performed in silico and the final polynucleotide sequence can be synthesized or cloned according to methods well known in the art, for example those described in WO2013017950.
[0041] By "genetic modification" is intended any enzymatic reaction that is spontaneously induced at a given locus, e.g. mutation, methylation, transcriptional modulation, with a view to obtaining an effect on gene expression.
[0042] According to some embodiments, the method of the present invention comprises one or several of the steps consisting of: (a) selecting a cleavage site in a target polynucleotide sequence, e.g., in a genome, where cleavage is intended; (b) selecting a polynucleotide sequence located 5 to 25 bp upstream and / or downstream of the cleavage site; (c) assembling a polynucleotide sequence encoding AvrBs3-like repeats so as to encode a TALE binding domain for binding to the selected polynucleotide sequence, wherein at least one AvrBs3-like repeat comprises a D substitution at position 4 (D4) and a D substitution at position 32 (D32) of a polypeptide sequence of the AvrBs3-like repeat, for example, a sequence selected from SEQ ID NOs:5-11; (d) fusing said TALE binding domain to at least (1) a polynucleotide sequence encoding an N-terminal domain, preferably comprising a sequence having at least 85%, preferably at least 90%, more preferably at least 95% identity to SEQ ID NO:1, and (2) a polynucleotide sequence encoding a C-terminal domain of a 40-80 residue polypeptide sequence, preferably comprising a sequence having at least 85% identity to SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:4 (X1, X2, X3 in these sequences represent R (arginine) or H (histidine)); (e) fusing the polynucleotide sequence obtained in (d) with another polynucleotide sequence encoding a nuclease, such as a type II endonuclease, in particular Fok1.
[0043] The method may also include an optional step in which, for example, the polynucleotide sequence fused to the TALE protein and encoding the catalytic domain may be mutated to introduce amino acid substitutions in said catalytic domain. This approach is exemplified in the experimental part of the present application, where amino acids in the Fok1 catalytic domain (SEQ ID NO:109) have been substituted by alanine residues with the effect of obtaining optimal nuclease activity of the TALE nuclease according to the present invention. Such individual substitutions in the Fok1 catalytic domain that were found to reduce off-site activity were in particular those at positions 13, 52, 57, 59, 61, 65, 84, 85, 88, 91, 92, 95, 98, 103, 109, 110, 111, 113, 119, 143, 148, 152, 158, 159, 160, 167, 169, 170, and 194 in SEQ ID NO:109. Preferred substitutions are at positions 84, 85, 88, 95, 98, 91, 103, 109, 148, 152, and 158 in SEQ ID NO:109, and most preferred substitutions are at positions 84, 88, 91, 103, and 152.
[0044] By "TALE protein" is meant herein a polypeptide that typically comprises a core DNA-binding domain having at least 50%, preferably at least 60%, 70%, 80%, or 90% identity to the DNA-binding domain of wild-type AvrBs3 [also called TalC Uniprot-G7TLQ9], which represents the prototype of a family of transcription activator-like (TAL) effectors from the plant pathogen Xanthomonas campestris. Such DNA-binding domains are characterized by a repeat sequence of approximately 30 and 34 amino acids that contains two variable residues usually found at positions 12 and 13. Consensus sequences of these repeats, also called RVDs, have been established for each target base A, C, G, and T, which correspond to, respectively, To target A TIFF2024540639000004.tif4128, for targeting C TIFF2024540639000005.tif4128, for targeting G TIFF2024540639000006.tif4128, for targeting T The file is TIFF2024540639000007.tif4128.
[0045] "AvrBs3-like repeat" refers to an artificial array of about 30-33 amino acids that typically contains two variable residues at positions 12 and 13 that interact with A, C, G, or T, similar to the consensus AvrBs3 repeat described above. In other words, AvrBs3-like repeats are similar and can combine with AvrBs3 repeats, but are generally not identical to the consensus or wild-type AvrBs3 repeats. In some cases, as described by [Valton et al. (2012) Overcoming Transcription Activator-like Effector (TALE) DNA Binding Domain Sensitivity to Cytosine Methylation. DNA and Chromosomes. 287(46):38427], the two residues at positions 12 or 13 are absent to accommodate methylated bases in genomic DNA - the so-called * (asterisk) - Note that this is a possibility.
[0046] The AvrBs3-like repeats of the present invention generally exhibit at least 60%, preferably at least 70%, 75%, 80%, 90% or 95% identity to any of the above AvrBs3 consensus repeat sequences of SEQ ID NOs: 31 to 34. They generally exhibit at least 60%, preferably at least 70%, 75%, 80%, 90% or 95% identity to any of the following repeat sequences of the present invention, SEQ ID NOs: 5 to 11: TIFF2024540639000008.tif44133, where X4X5 are the two residues that interact with a given nucleotide base pair in the target sequence. X4 and X5 can be any amino acid or can be null (to specify the missing residue in the RVD). * (indicated as an asterisk). X4 and X5 may be the same or different.
[0047] AvrBs3-like repeats are generally represented by polypeptide sequences in which X4 and X5 are NI (preferably for target A), HD (preferably for target C), NN (preferably for target G) and NG (preferably for target T), respectively, such as in SEQ ID NOs:24, 25, 26, and 27.
[0048] "Identity" refers throughout the present specification to sequence identity between two nucleic acid molecules or polypeptides. Identity can be determined by comparing positions in each sequence, which may be aligned for comparison purposes. If a position in the compared sequences is occupied by the same base, the molecules are identical at that position. The degree of similarity or identity between nucleic acid or amino acid sequences is a function of the number of identical or matching nucleotides at positions common to the nucleic acid sequences. A variety of alignment algorithms and / or programs can be used to calculate the identity between two sequences, including FASTA, or BLAST, which are available as part of the GCG sequence analysis package (University of Wisconsin, Madison, Wis.), and can be used, for example, with default settings. The present specification encompasses polypeptides and polynucleotides that generally have at least 70%, 85%, 90%, 95%, 98%, or 99% identity with the specific polypeptide and polynucleotide sequences described herein, and exhibit substantially the same function or can be considered equivalent.
[0049] In some embodiments, the present invention also provides recombinant Transcription Activator-Like Effector (TALE) proteins comprising one or several AvrBs3-like repeats comprising D (aspartic acid) residues at positions 4 and 32, such as in the above polynucleotide sequences SEQ ID NO:NO:5-11. Such AvrBs3-like repeats may be further mutated at 1-5 amino acid positions, including or in addition to positions D4 and D32. Such recombinant Transcription Activator-Like Effector (TALE) proteins may comprise one or several of such repeats to form a polypeptide comprising generally 8-20 repeats, preferably 8-18, more preferably 10-16, or alternatively 5-12 repeats, in the context of smaller genomes considered, such as mitochondrial genomes.
[0050] The variable two residues (X4X5) present in the AvrBs3-like repeats and associated with the recognition of different nucleotides are generally HD for recognizing C, NG for recognizing T, NI for recognizing A, NN for recognizing G or A, NS for recognizing A, C, G or T, HG for recognizing T, IG for recognizing T, NK for recognizing G, HA for recognizing C, ND for recognizing C, HI for recognizing C, HN for recognizing G, NA for recognizing G, SN for recognizing G or A, and YG for recognizing T, TL for recognizing A, VT for recognizing A or G, and SW for recognizing A. More preferably, the RVDs associated with the recognition of the nucleotides C, T, A, G / A, and G are selected from the group consisting of NN or NK for recognizing G, HD for recognizing C, T for recognizing NG, and NI for recognizing A, TL for recognizing A, VT for recognizing A or G, and SW for recognizing A. More commonly, the RVDs associated with the recognition of the nucleotide C are generally ... * and the RVD associated with recognition of the nucleotide T is selected from the group consisting of N * and H * is selected from the group consisting of *can represent a gap in the repeat sequence corresponding to the absence of an amino acid residue at the second position of the RVD. In some embodiments, X4X5 can represent unusual or atypical amino acid residues to modulate its specificity for the nucleotides A, T, C, and G, as described in Juillerat et al. [Optimized tuning of TALEN specificity using non-conventional RVDs (2015) Sci Rep 5:8150].
[0051] Although not essential, the core DNA binding domain generally comprises a half RVD of 20 amino acids located at the C-terminus, and therefore comprises 8.5 to 30.5 RVDs, more preferably 8.5 to 20.5 RVDs, even more preferably 10.5 to 15.5 RVDs.
[0052] According to the invention, said core DNA binding domain, preferably comprising an RVD with a D4 and / or D32 substitution, is flanked by N- and C-terminal sequences, said N- and C-terminal sequences preferably having one of the following features as detailed below:
[0053] In some embodiments, the N-terminal sequence is derived from the N-terminal domain of a naturally occurring TAL effector, such as AvrBs3. In another embodiment, said further N-terminal domain is the full-length N-terminal domain of a naturally occurring TAL effector N-terminal domain. In a further embodiment, said further N-terminal domain is a variant that allows overcoming sequence constraints associated with the so-called "RVD0" (i.e., the first cryptic repeat), such as the need to have a required T as the first base on the binding nucleic acid sequence.
[0054] In another embodiment, the N-terminal sequence is derived from a naturally occurring TAL effector or a variant thereof. In another embodiment, the N-terminal sequence is a truncated N-terminus of such a naturally occurring TAL effector or variant. In another embodiment, the additional domain is a truncated version of the AvrBs3 TAL effector. In another embodiment, the truncated version lacks its N-terminal segment distal to the core TALE binding domain, for example, the first 152 N-terminal amino acid residues, or at least 152 amino acid residues, of wild-type AvrBs3.
[0055] In some preferred embodiments, the N-terminal sequence comprises a polypeptide sequence that exhibits at least 85%, preferably at least 90%, and more preferably at least 95% identity to SEQ ID NO:1.
[0056] In some embodiments, the C-terminal sequence corresponds to the complete or preferably truncated C-terminal region of a naturally occurring TAL effector such as AvrBs3. Generally, the C-terminal sequence is a truncated version of the AvrBs3 TAL effector proximal to the core TALE binding domain, such as SEQ ID NO:28 (40 amino acids), SEQ ID NO:29 (50 amino acids), or SEQ ID NO:30 (60 amino acids), or naturally occurring variants thereof. Thus, the C-terminal sequence generally corresponds to the following SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:4: It comprises or consists of a polypeptide sequence of 40 to 80 residues including a sequence having at least 85% identity to TIFF2024540639000009.tif35153.
[0057] In the above sequence, X1, X2, and X3 represent amino acid substitutions introduced into the C-terminal polypeptide sequence of wild-type AvrBs3, preferably an R (arginine) or H (histidine) residue, most preferably R, in place of the original K. X1, X2, and X3 may be the same or different.
[0058] Said N- or C-terminal sequence may comprise a localization sequence (or signal) that allows targeting said chimeric protein to a given organelle in an organism, tissue or cell. Non-limiting examples of such localization signals are nuclear localization signals, chloroplast localization signals, or mitochondrial localization signals. In another embodiment, said further N-terminal domain may comprise a nuclear export signal that has the opposite effect of a nuclear localization signal to assist in targeting organelles such as chloroplasts or mitochondria. Further C- or N-terminal sequences with a combination of several localization signals are also encompassed within the scope of the present invention. Such combinations may be, as non-limiting examples, nuclear localization signals (NLS) and / or tissue-specific signals to assist in addressing said fusion protein of the present invention to the nucleus of tissue-specific cells. In a preferred embodiment, the NLS is generally comprised in the N-terminal region of the TALE protein. Preferred NLS sequences include SEQ ID NO:12 derived from SV40, SEQ ID NO:13 derived from C-Myc, or SEQ ID NO:14 polypeptide sequence derived from nucleoplasmin.
[0059] Table 1. Examples of NLS sequences TIFF2024540639000010.tif121158
[0060] "TALE fusion protein" refers to a TALE protein linked to a polypeptide domain that confers catalytic activity to the TALE protein. A TALE fusion protein can be, for example, a sequence-specific reagent that processes DNA at a locus specified by the TALE binding domain. Fusion with a TALE protein can be with a catalytic domain from an existing protein, for example, a DNA processing enzyme, particularly one with an activity selected from the group consisting of nuclease activity, polymerase activity, deaminase activity, kinase activity, phosphatase activity, methylase activity, topoisomerase activity, integrase activity, transposase activity, ligase activity, helicase activity, reverse transcriptase, and recombinase activity.
[0061] In some embodiments, the TALE fusion proteins according to the present invention may comprise a peptide linker to fuse the catalytic domain to the aforementioned core scaffold, or more preferably to link the C-terminus or N-terminus of said TALE protein to said catalytic domain. Such linkers are generally flexible. For example, NFS1, NFS2, CFS1, RM2, BQY, QGPSG, LGPDGRKA, 1a8h_1, 1dnpA_1, 1d8cA_2, 1ckqA_3, 1sbp_1, 1ev7A_1, 1alo_3, 1amf_1, 1adjA_3, 1fcd, optionally including a SGGSGS stretch at either or both the N-terminus and C-terminus surrounding a variable region of 3 to 28 amino acids, as exemplified in Table 2 below (SEQ ID NOs: 35-108). The linker sequence is one selected from the group consisting of: C_1, 1al3_2, 1g3p_1, 1acc_3, 1ahjB_1, 1acc_1, 1af7_1, 1heiA_1, 1bia_2, 1igtB_1, 1nfkA_1, 1au7A_1, 1bpoB_1, 1b0pA_2, 1c05A_2, 1gcb_1, 1bt3A_1, 1b3oB_2, 16vpA_6, 1dhx_1, 1b8aA_1, and 1qu6A_1.
[0062] Table 2: Examples of peptide linkers TIFF2024540639000011.tif157160TIFF2024540639000012.tif244160TIFF2024540639000013.tif204160
[0063] In some embodiments, the peptide linker can include a calmodulin domain that changes the conformation of the TALE fusion protein under calcium stimulation. Other protein domains that induce conformational changes under specific metabolite interactions can also be used. Such linkers can include, for example, a light-sensitive domain that allows for a change from a folded, inactive state to an unfolded, active state, or vice versa, under light stimulation. Other examples of "switch" linkers can respond to small molecules such as Chemical Inducers of Dimerization (CID).
[0064] In a preferred embodiment, a linker may not be necessary to fuse the TALE core binding domain to the catalytic domain, as the C-terminal sequence may have sufficient flexibility to achieve optimal conformation of the TALE fusion protein, as exemplified herein with the preferred C-terminal sequences discussed above.
[0065] The present invention encompasses TALE fusion proteins that contain various functional domains, e.g., catalytic domains obtainable from different enzymes, such as non-specific endonucleases, e.g., Fok-1, clo51, or I-Tev1, or specific endonucleases, e.g., engineered meganucleases (e.g., from I-Cre1, I-Onu1, I-Bmo1, HmuI, etc.), exonucleases, e.g., human Trex2, transcriptional repressors (e.g., KRAB), or transcriptional activators, e.g., VP64 or VP16, deaminases, e.g., cytosine deaminase 1 (pCDM), adenosine deaminases, e.g., TadA ou TadA7.10, apolipoprotein B, The enzyme may be an mRNA editing enzyme catalytic polypeptide-like (APOBEC), activation-induced cytidine deaminase (AICDA), DddA (double-stranded DNA cytidine deaminase) potentially associated with uracil glycosylase inhibitor (UGI), a nickase derived from Cas9 or Cpf1, a transposase, an integrase, a topoisomerase, and a reverse transcriptase (e.g., Moloney murine leukemia virus RT enzyme), or a functional mutant, variant, or derivative thereof.
[0066] Exemplary polypeptide sequences that can be included in the TALE fusion proteins of the present invention are listed in Table 3 (SEQ ID NOs:109-137).
[0067] Table 3: Exemplary catalytic domains of the TALE proteins of the present invention TIFF2024540639000014.tif92170TIFF2024540639000015.tif234170TIFF2024540639000016.tif237170 TIFF2024540639000017.tif229170TIFF2024540639000018.tif236170TIFF2024540639000019.tif81170
[0068] In another aspect, a TALE fusion protein according to the present invention comprises a catalytic domain that is a polypeptide comprising an amino acid sequence having at least 80%, preferably at least 90%, more preferably at least 95% identity to any of SEQ ID NOs:109-137.
[0069] Gene editing is crucial because gene editing reagents can cause unintended disturbances in genomes, and as multiplexing methods become more widely used, the possibility of off-targets and the downstream impact of such off-target activity increases. Minimizing such undesired breaks (off-targets) is a crucial issue for any genome engineering application, especially in the therapeutic field. Undesired double-strand breaks in genomes can lead to chromosomal translocations and cytotoxicity [Cantoni O., et al. (1996) Cytotoxic impact of DNA single vs double strand breaks in oxidatively injured cells. Arch Toxicol Suppl 18:223-235]. Currently, there are various techniques available to predict and quantify off-targets by analyzing secondary target locations and determining on-target / off-target ratios, such as those described by Tsai S., et al. [CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets (2017) Nat Methods 14(6):607-614], Hockemeyer D, et al. [Genetic engineering of human pluripotent cells using TALE nucleases (2011) Nat. Biotechnol. 29(8):731-734] and Wienert B, et al. [Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq (2019) Science 364(6437):286-289].
[0070] As mentioned above, TALE proteins have well-defined DNA base pair preferences, providing a basic strategy for scientific researchers and engineers to design and construct TALE fusion proteins for genome modification. TALE repeat tandems are involved in the recognition of individual DNA base pairs. Such tandems consist of a pair of α-helices connected by a three-residue loop of a solenoid-shaped RVD. To generate TALE proteins with variable precision and binding affinity, six conventional RVDs (NG, HD, NI, NK, NH, and NN) are frequently used. HD and NG bind to cytosine (C) and thymine (T), respectively. These bindings are strong and exclusive [Streubel J, et al. (2012) TAL effector RVD specificities and efficiencies. Nat Biotechnol 30(7):593-595]. NN is a degenerate RVD and usually shows binding affinity to both guanine (G) and adenine (A), but its specificity for guanine has been reported to be stronger. RVD NI binds A and NK binds G. Although these bindings are exclusive, the binding affinity between these pairs is low and therefore the binding is considered weak. Therefore, it is recommended to use RVD NH, which binds G with moderate affinity. It is also worth noting that the binding affinity of TALEs is affected by the methylation state of the target DNA sequence.
[0071] The code of TALEN is degenerate, which means that a certain RVD can bind to multiple nucleotides with a diverse range of efficiencies. The binding ability of NN (for A and G) and NS (A, C, and G) repeat variable dipeptide allows TALE proteins to code degenerately for target DNA. This degeneracy can be useful for targeting hypervariable sites. TALE protein technology is the only known genome editing tool that can be engineered in a way that can be easily used for escape mutations in the genome. This unique feature makes it a more flexible and reliable tool for tolerating predicted mutations, especially in the field of genome editing in clinical applications [Strong CL, et al. (2015) Damaging the integrated HIV proviral DNA with TALENs. PLoS One 10(5):e0125652.].
[0072] A typical TALE protein usually consists of 18 repeats of 34 amino acids. A pair of TALENs must bind to opposite target sites, separated by a 14-20 nucleotide "spacer" as an offset, since FokI requires dimerization to act. Overall, such long (approximately 36 bp) DNA binding sites are expected to occur very rarely in genomes.
[0073] Development of specific TALE nucleases By following the above teachings, highly specific TALE nucleases can be generated in accordance with the present invention, allowing high cleavage specificity and low cytotoxicity in a variety of cell types, particularly plant or mammalian cells.
[0074] According to some embodiments, the TALE fusion proteins of the present invention are TALE nucleases obtained by fusion of a TALE protein as described herein with the nuclease catalytic domain of a non-specific nuclease, such as, for example, Fok-1 (SEQ ID NO:109) or Tev-1 (SEQ ID NO:114) as described in classical TALE scaffolds in Beurdeley, M. et al. [Compact designer TALENs for efficient genome engineering (2013) Nat Commun 4:1762]. In a preferred embodiment, as exemplified herein in the Examples, said nuclease catalytic domain is Fok1, i.e. exhibits at least 80% identity with SEQ ID NO.1, and more preferably comprises a polypeptide comprising at least one of the following amino acid substitutions in SEQ ID NO: 109: 13, 52, 57, 59, 61, 65, 84, 85, 88, 91, 92, 95, 98, 103, 109, 110, 111, 113, 119, 143, 148, 152, 158, 159, 160, 167, 169, 170, and 194. Preferred substitutions are introduced at positions 84, 85, 88, 95, 98, 91, 103, 109, 148, 152, and 158, and most preferred substitutions are present at positions 84, 88, 91, 103, and 152.
[0075] According to some embodiments, the TALE fusion protein of the present invention is a TALE nuclease obtained by fusion of a TALE protein described herein with a nickase, in particular a Cas9 nickase. Such a Cas9 nickase is generally a Cas9 protein that is mutated in its RuvC or HNH domain, for example by introducing the mutation D10A in RuvC and the mutation H840A in HNH. Generally, TALE-Cas9 nickase fusions are used in pairs, as previously described by Guilinger, J., et al. [Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification (2014) Nat. Biotechnol. 32, 577-582] in classical TALE scaffolds.
[0076] In some other embodiments, the TALE fusion protein of the present invention is a TALE nuclease obtained by fusion of a TALE protein described herein with a specific nuclease, preferably a customized rare-cutting endonuclease, such as a meganuclease variant. In a preferred embodiment, said rare-cutting endonuclease may be a variant of LADLIDADG, such as I-creI or I-OnuI, as previously described, for example, in EP3320910 and EP3004338.
[0077] On the other hand, the TALE nuclease according to the present invention also has the ability to efficiently manipulate mtDNA (mitochondrial DNA) as a treatment for treating human mitochondrial diseases induced by mitochondrial pathogenic mutations. The so-called "Mito-TALEN" (mitochondrial-targeted TALEN) has been proven to effectively treat human mitochondrial disorders affected by mtDNA mutations, such as Leber's hereditary optic neuropathy, ataxia, neurogenic muscle fatigue, and retinal dysplasia [Gammage, PA, et al. (2018) Mitochondrial Genome Engineering: The Revolution May Not Be CRISPR-Ized. Trends in Genetics, 34(2):101-110]. Plastid engineering has also demonstrated satisfactory results in various plants for crop improvement [Piatek AA, Lenaghan SC, Neal Stewart C. (2018) Advanced editing of the nuclear and plastid genomes in plants. Plant Sci 273:42-49].
[0078] Many examples of TALE nucleases according to the present invention are described herein for use as therapeutic reagents to induce highly specific cleavage in a selection of genes in human cells, particularly blood cells. More specifically, improved TALE nuclease reagents have been synthesized and tested according to the present teachings to cleave gene targets such as, for example, TCRα, B2m, PD1, CTLA4, CISH, LAG3, TGFBRII, TIGIT, CD38, IgH, GADPH, and CCR5 in primary cells, particularly in T cells or NK cells.
[0079] The polypeptide sequences of these TALE proteins obtained by the present invention, as well as their target sequences (polynucleotide sequences spanning the two left and right heterodimer binding sites), are listed in Tables 4 and 5 below and in Tables 5 and 6 in the Examples section.
[0080] Table 4. Examples of TALE proteins useful in therapy TIFF2024540639000020.tif226166TIFF2024540639000021.tif245166TIFF2024540639000022.tif245166TIFF2024540639000023.tif245166TIFF202 4540639000024.tif245166TIFF2024540639000025.tif245166TIFF2024540639000026.tif245166TIFF2024540639000027.tif245166TIFF2024540639 000028.tif245166TIFF2024540639000029.tif245166TIFF2024540639000030.tif245166TIFF2024540639000031.tif245166TIFF2024540639000032. tif245166TIFF2024540639000033.tif244166TIFF2024540639000034.tif236166TIFF2024540639000035.tif242166TIFF2024540639000036.tif81166
[0081] Table 4 (continued) Genomic polynucleotide sequences targeted by the above TALE proteins TIFF2024540639000037.tif245166
[0082] In some preferred embodiments, the TALE proteins of the present invention can be used in pairs, where each member of the pair binds to DNA near each other, side by side, or on opposite DNA strands in such a way that they co-localize in the genome with the effect of directing the catalytic activity induced by the catalytic domain at a specific locus. For example, a pair of TALE proteins fused to homodimerized Fok1 nuclease domains, also referred to as "left" and "right" TALE nuclease monomers, form a heterodimer that induces DNA double-strand break cleavage. In such cases, the present invention provides that one monomer according to the present invention can be used together with another monomer based on a conventional TALE nuclease scaffold using the canonical AvrBs3 sequence. Indeed, as shown in the experimental section herein, one TALE nuclease monomer of the present invention is sufficient to have an overall effect on the specificity of the heterodimer.
[0083] Thus, the present invention provides several novel TALE fusion monomers based on the TALE proteins listed in Table X, including such proteins fused with nuclease or deaminase domains, for their use in in vivo or in vitro gene therapeutic modification and for ex vivo preparation of therapeutic cells.
[0084] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the CTLA4 locus, preferably comprising SEQ ID NO:231, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:138 or SEQ ID NO:139. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:174 and SEQ ID NO:175.
[0085] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in a CISH locus, preferably comprising SEQ ID NO:232, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:140 or SEQ ID NO:141. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:176 and SEQ ID NO:177.
[0086] According to a particular aspect, the present invention provides a TALE protein monomer for introducing a genetic modification, preferably a mutation, into a target sequence in the LAG3 locus, preferably comprising SEQ ID NO:233, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:142 or SEQ ID NO:143. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:178 and SEQ ID NO:179.
[0087] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the TGFBRII locus, preferably comprising SEQ ID NO:234, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. Said TALE protein preferably comprises SEQ ID NO:144 or SEQ ID NO:145. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:180 and SEQ ID NO:181.
[0088] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the CCR5 locus, preferably comprising SEQ ID NO:235, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:146 or SEQ ID NO:147. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:182 and SEQ ID NO:183.
[0089] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the B2m locus, preferably comprising SEQ ID NO:236 or SEQ ID NO:237, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NOs:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:148, SEQ ID NO:149, SEQ ID NO:150, or SEQ ID NO:151. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:184, SEQ ID NO:185, SEQ ID NO:186, and SEQ ID NO:187.
[0090] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into the TCR alpha locus, preferably into a target sequence comprising SEQ ID NO:238, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. Said TALE protein preferably comprises SEQ ID NO:152 or SEQ ID NO:153. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:188 and SEQ ID NO:189.
[0091] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the PD1 locus, preferably comprising SEQ ID NO:239, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:154 or SEQ ID NO:155. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:190 and SEQ ID NO:191.
[0092] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the PIK3CDex8 locus, preferably comprising SEQ ID NO:240, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:156 or SEQ ID NO:157. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:192 and SEQ ID NO:193.
[0093] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the PIK3CDex17 locus, preferably comprising SEQ ID NO:241, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. Said TALE protein preferably comprises SEQ ID NO:158 or SEQ ID NO:159. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:194 and SEQ ID NO:195.
[0094] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the S100A9 locus, preferably comprising SEQ ID NO:242, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:160 or SEQ ID NO:161. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:196 and SEQ ID NO:197.
[0095] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the AAVS1 locus, preferably comprising SEQ ID NO:243, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:162 or SEQ ID NO:163. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:198 and SEQ ID NO:199.
[0096] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the CD52 locus, preferably comprising SEQ ID NO:244, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:164 or SEQ ID NO:165. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:200 and SEQ ID NO:201.
[0097] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the TCR alpha locus, preferably comprising SEQ ID NO:245, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. Said TALE protein preferably comprises SEQ ID NO:166 or SEQ ID NO:167. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:202 and SEQ ID NO:203.
[0098] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the TGFBRII locus, preferably comprising SEQ ID NO:246, 247, or 248, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3, or 4. Said TALE protein preferably comprises SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, or SEQ ID NO:173. In particular, the present invention provides TALE nuclease monomers consisting of or comprising a polypeptide sequence at least 90%, preferably 95% or 99% identical to a sequence selected from SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:208, and SEQ ID NO:209, respectively.
[0099] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the TIGIT locus, preferably comprising or consisting of SEQ ID NO:289, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence having at least 90%, preferably 95% or 99% identity to SEQ ID NO:269 and / or SEQ ID NO:270.
[0100] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in a CISH locus, preferably comprising or consisting of SEQ ID NOs:290, 291, and / or 292, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 repeats comprising SEQ ID NOs:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NOs:2, 3, or 4. In particular, the present invention provides TALE nuclease monomers consisting of or comprising a polypeptide sequence having at least 90%, preferably 95% or 99% identity to SEQ ID NO:271, SEQ ID NO:272, SEQ ID NO:273, SEQ ID NO:274, SEQ ID NO:275 and / or SEQ ID NO:276.
[0101] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the CD38 locus, preferably comprising or consisting of SEQ ID NO:293 and / or SEQ ID NO:294, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence having at least 90%, preferably 95% or 99% identity to SEQ ID NO:277, SEQ ID NO:278, SEQ ID NO:279 and / or SEQ ID NO:280.
[0102] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in an IgH locus, preferably comprising or consisting of SEQ ID NO:295 and / or SEQ ID NO:296, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence having at least 90%, preferably 95% or 99% identity to SEQ ID NO:281, SEQ ID NO:282, SEQ ID NO:283 and / or SEQ ID NO:284.
[0103] According to a particular aspect, the present invention provides a TALE protein monomer for introducing genetic modifications, preferably mutations, into a target sequence in the GADPH locus, preferably comprising or consisting of SEQ ID NO:297 and / or SEQ ID NO:298, said TALE protein comprising (1) a TALE binding domain comprising at least 3, preferably at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 repeats comprising SEQ ID NO:5-11, and (2) a C-terminal polypeptide sequence of 40-80 residues comprising a sequence having at least 85% identity to SEQ ID NO:2, 3 or 4. In particular, the present invention provides a TALE nuclease monomer consisting of or comprising a polypeptide sequence having at least 90%, preferably 95% or 99% identity to SEQ ID NO:285, SEQ ID NO:286, SEQ ID NO:287 and / or SEQ ID NO:288. "Mutation" generally refers herein to any change in one or more nucleotides in a characterized polynucleotide sequence (wild type) in the genomic sequence of a cell, including deletion or substitution of said nucleotides (or base pairs), deletion insertion, integration, or translocation of a polynucleotide fragment, oligonucleotide, or foreign sequence, e.g., a transgene. Such mutations generally lead to a correction, loss, or gain of function by the cell whose genome is modified.
[0104] Development of TALE transcription factors Following the previous teachings, in consideration of the control of endogenous gene expression, the TALE proteins according to the present invention can also be fused to desired transcription activator and repressor protein domains to generate specific transactivator or repressor reagents.
[0105] As an example, artificial transcription factors can be obtained by fusing the TALE protein of the present invention with VP64 or the 16 amino acid peptide VP16 (SEQ ID NO:120) from herpes simplex virus, as described by Miller JC, et al. [A TALE nuclease architecture for efficient genome editing (2011) Nat Biotechnol 29(2):143-148].
[0106] To achieve gene repression, the TALE proteins of the present invention can be fused, for example, to Krüppel-binding box (KRAB), Sid4, or EAR repression domain (SRDX), which have previously been reported to be strong pleiotropic repressors [Cong L, et al. (2012) Comprehensive interrogation of natural TALE DNA-binding modules and transcriptional repressor domains. Nat Commun 3(1):968].
[0107] Development of TALE base editor Following the previous teachings, the TALE proteins according to the present invention can also be fused to a desired base editor.
[0108] As used herein, the term "base editor" refers to a catalytic domain that can make modifications to bases (e.g., A, T, C, G, or U) within a nucleic acid sequence, converting one base to another (e.g., A to G, A to C, A to T, C to T, C to G, C to A, G to A, G to C, G to T, T to A, T to C, T to G). Adenine and cytosine base editor catalytic domains are described, for example, in Rees & Liu [Base editing: precision chemistry on the genome and transcriptome of living cells (2018) Nat. Rev. Genet. 19(12):770-788].
[0109] Catalytic base editors can include cytidine deaminases that convert targeted C / G to T / A, and adenine base editors that convert targeted A / T to G / C. A preferred cytosine deaminase can be cytosine deaminase 1 (pCDM) or activation-induced cytidine deaminase (AICDA). A preferred adenosine deaminase can be TadA (SEQ ID NO:121) or its variant TadA7.10 described by Jeong, YK, et al. [Adenine base editor engineering reduces editing of bystander cytosines (2021) Nat. Biotechnol. https: / / doi.org / 10.1038 / s41587-021-00943]. Different members of the apolipoprotein B mRNA editing enzyme (APOBEC) family, such as mouse rAPOBEC1 and human APOBEC3G (SEQ ID NO:130) developed by Lee et al. [Single C-to-T substitution using engineered APOBEC3G-nCas9 base editors with minimum genome- and transcriptome-wide off-target effects (2020) Science Advances. 6(29)], can be used to convert cytidine to thymidine.
[0110] In a preferred embodiment, the base editor catalytic domain converts C to T (cytidine deaminase), which catalyzes the chemical reaction "Cytosine + HO->Uracil + NH3" or "5-methyl-cytosine + HO->Thymine + NH3". As may be evident from the reaction equation, such a chemical reaction results in a nucleobase change from C to U / T. In the context of a gene, such a nucleotide change or mutation may in turn lead to an amino acid change in a protein, which may affect the function of the protein, such as a loss or gain of function.
[0111] In some embodiments, the TALE base editor according to the present invention can include a domain that inhibits uracil glycosylase, referred to as "UGI", and / or a nuclear localization signal. As used herein, the term "uracil glycosylase inhibitor" or "UGI" refers to a protein that can inhibit the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain includes wild-type UGI or the canonical UGI set forth in SEQ ID NO:136. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments, including amino acid sequences that include at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% of the amino acid sequence set forth in SEQ ID NO:136. TALE base editors according to the present invention, including UGIs, are useful for improving the specificity of base editing performed at a given locus.
[0112] In some embodiments, the base editor catalytic domain is a double-stranded DNA deaminase ("DddA") to precisely incorporate nucleotide changes and / or correct pathogenic mutations rather than breaking DNA at double-strand breaks (DSBs). In preferred embodiments, DddAtox is generally split into inactive fragments that can be delivered separately to the target deamination site on separate TALE base editor constructs that co-localize each fragment of DddA at a site, such as either side of the target editing site, where they reform a functional DddA that can deaminate the target site on a double-stranded DNA molecule. In certain embodiments, programmable DNA binding proteins can be engineered to contain one or more mitochondrial localization signals (MLS) in such a way that the DddA domain is transferred into mitochondria, thereby providing a means to directly base edit the mitochondrial genome.
[0113] Fragments of DddA can be formed by truncating DddAtox (i.e., splitting or splitting the DddA protein) at specific amino acid residues, for example, amino acid residues selected from the group including 62, 71, 73, 84, 94, 108, 110, 122, 135, 138, 148, and 155. In preferred embodiments, truncation of DddA occurs at residue 148. In certain embodiments, splitting DddA at one of these split sites can separate DddA into two fragments to form the N-terminal and C-terminal portions of DddA, which may be referred to as the "DddA-N half" and "DddA-C half." According to a preferred embodiment, said "DddA-N half" and "DddA-C half" comprise an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% common to the amino acid sequences SEQ ID NO.134 and SEQ ID NO:135, respectively. As shown in Figure 8, two pairwise acting TALE proteins comprising N-DddA and C-DddA halves, respectively, can be co-localized and used to induce on-site nucleobase changes.
[0114] The TALE base editors of the present invention can also be used in pairs, with each member containing a different but complementary catalytic domain, allowing for a given base editing reaction to be obtained at one precise locus.
[0115] Development of TALE-transposases or integrases Following previous teachings, the TALE protein according to the present invention can also be fused to a transposase or integrase to perform site-specific integration of the transgene into the genome.
[0116] As an example, the TALE protein according to the invention can be fused to the PiggyBac transposase described, for example, by Owens, JB et al. [Transcription activator like effector (TALE)-directed piggyBac transposition in human cells (2013) NAR 41(19):9197-9207]. The PiggyBac transposase is autonomously functional in such a system, so that the co-transfected transposon can be integrated at any genomic location specified by the TALE protein. This system can permanently introduce large cassettes (>100 kb) encoding multiple components, for example, multiple transgenes, insulators, and inducible or endogenous promoters, allowing for potentially targeting integration into almost any genomic region. This system is particularly valuable in situations where a safe single-targeted insertion needs to be verified ex vivo, the cells expanded, and re-infused into the patient. Targeted transposition can be used to deliberately disrupt endogenous coding regions or to direct insertion into user-defined genomic safe harbors to protect the cargo from unknown chromosomal position effects and avoid accidental mutations in target cells.
[0117] Development of TALE proteins for epigenome editing Further following the above teachings, TALE protein fusions can be generated, particularly by fusion with catalytic domains that can modulate gene expression without altering the DNA sequence by chromatin remodeling.
[0118] In this regard, the TALE proteins according to the present invention can be fused to methyltransferases to obtain histone methylation, and / or p300 effector domains to enhance histone acetyltransferases.
[0119] Conversely, TALE proteins can be fused to the catalytic domain thymidine DNA glycosylase (TDG) to abolish DNA methylation and induce gene expression. Unwanted DNA methylation is associated with many neurodegenerative diseases. As an example, TALE proteins can be fused to the TET domain (ten-eleven translocation methylcytosine dioxygenase 2) to target epigenetically silenced cancer genes (ICAM-1) and induce their expression in cancerous cells. TET1 can also be used in the treatment of many diseases such as diabetes (induction of β cell replication) and cancer (inhibition of cell proliferation) [Ou K., et al. (2019) Targeted demethylation at the CDKN1C / p57 locus induces human β cell replication. J Clin Invest 129(1):209-214].
[0120] The present invention encompasses polynucleotides, particularly DNA or RNA encoding the aforementioned polypeptides and proteins, as well as any intermediates involved in any aspect and step of the methods described herein. These polynucleotides can be included in vectors, more particularly plasmids or viruses, allowing for their expression in prokaryotic or eukaryotic cells.
[0121] The term "vector" or "vectors" refers to a nucleic acid molecule capable of transporting another nucleic acid linked to it. In the present invention, "vector" includes, but is not limited to, viral vectors, plasmids, RNA vectors, or linear or circular DNA or RNA that may consist of chromosomal, non-chromosomal, semisynthetic, or synthetic nucleic acid. Preferred vectors are those capable of autonomous replication (episomal vectors) and / or expression of the nucleic acid linked to it (expression vectors). Many suitable vectors are known to those of skill in the art and are commercially available. Viral vectors include retroviruses, adenoviruses, particularly AAV6 vectors, parvoviruses (e.g., adeno-associated viruses), coronaviruses, negative-stranded RNA viruses, such as orthomyxoviruses (e.g., influenza viruses), rhabdoviruses (e.g., rabies virus and vesicular stomatitis virus), paramyxoviruses (e.g., measles and Sendai), positive-stranded RNA viruses, such as picornaviruses and alphaviruses, and double-stranded DNA viruses, such as adenoviruses, herpesviruses (e.g., herpes simplex virus types 1 and 2, Epstein-Barr virus, cytomegalovirus), and poxviruses (e.g., vaccinia, fowlpox, and canarypox).Other viruses include, for example, Norwalk virus, togavirus, flavivirus, reovirus, papovavirus, hepadnavirus, and hepatitis virus. Examples of retroviruses include avian leukosis sarcoma, mammalian type C, B, and D viruses, the HTLV-BLV complex, lentiviruses, and spumaviruses (Coffin, JM, Retroviridae: The viruses and their replication, In Fundamental Virology, Third Edition, BN Fields, et al., Eds., Lippincott-Raven Publishers, Philadelphia, 1996).
[0122] According to the present invention, the TALE protein or a polynucleotide encoding it, particularly mRNA, can also be loaded into nanoparticles for effective delivery into cells. Various nanoparticles for targeting cell types of specific tissues have been described in the art [Friedman AD et al. (2013) The Smart Targeting of Nanoparticles Curr Pharm Des. 19(35): 6315-6329]. Preferred nanoparticles are positively charged nanoparticles, such as silica-based nanoparticles, or LNPs (Lipid nanomolar nanoparticles) that have been described in the art with other types of nucleases [Conway, A. et al. (2019) Non-viral Delivery of Zinc Finger Nuclease mRNA Enables Highly Efficient In Vivo Genome Editing of Multiple Therapeutic Gene Targets, Molecular Therapy 27(4):866-877].
[0123] Alternatively, a polynucleotide encoding the TALE protein of the present invention, particularly in the form of mRNA, can be directly electroporated into blood cells by electroporation, for example by using the steps described on pages 29 and 30 of WO2013176915, which is incorporated herein by reference.
[0124] The present invention also relates to methods for the use of the aforementioned polypeptides, polynucleotides, and proteins for various applications ranging from targeted nucleic acid cleavage to targeted gene regulation. In genome engineering experiments, the efficiency of the nuclease fusion proteins referred to in this patent application, for example, the ability to induce a desired event (homologous gene targeting, targeted mutagenesis, sequence removal or excision, base editing) at a locus, depends on several parameters, including the specific activity of the nuclease, the likely accessibility of the target, and the effectiveness and success of the repair pathway that leads to the desired event (homologous repair for gene targeting, NHEJ pathway for targeted mutagenesis), which can be evaluated by standard techniques known in the art. The present invention more particularly relates to a method for modifying the genetic material of a cell within or adjacent to a nucleic acid target sequence using one of the TALE fusion proteins of the present invention. For example, double-strand breaks caused by TALE nucleases are generally repaired through non-homologous end joining (NHEJ). NHEJ includes at least two different processes. The mechanism involves rejoining what remains of the two DNA ends, either through direct religation or via so-called microhomology-mediated end joining. Repair via non-homologous end joining (NHEJ) often results in small insertions or deletions and can be used to create specific gene knockouts.
[0125] Another aspect of the present disclosure relates to pharmaceutical compositions comprising any of the various components of the TALE proteins obtainable by the methods of the present invention (e.g., TALE nuclease, TALE deaminase, TALE transcriptase, TALE methylase, TALE transposase, etc.).
[0126] As used herein, the term "pharmaceutical composition" refers to a composition formulated for pharmaceutical use. In some embodiments, the pharmaceutical composition further comprises a pharma- ceutical acceptable carrier. In some embodiments, the pharmaceutical composition comprises an additional agent (e.g., for specific delivery, increasing half-life, or other therapeutic compounds).
[0127] In some embodiments, pharmaceutical compositions are provided as reagents to correct genetic defects that can be used in vivo or ex vivo, particularly in gene therapy.
[0128] In a preferred embodiment, the TALE proteins of the present invention are used to genetically modify blood cells, particularly immune cells, such as T cells and NK cells, ex vivo, preferably primary cells, to generate therapeutic cells for immunotherapy.
[0129] In some embodiments, pharmaceutical compositions are formulated according to conventional procedures as compositions suitable for intravenous or subcutaneous administration to subjects (e.g., humans). In some embodiments, pharmaceutical compositions for administration by injection are solutions in sterile isotonic aqueous buffer. If necessary, said pharmaceutical compositions can also contain a solubilizing agent and a local anesthetic, such as lidocaine, to ease pain at the injection site. In general, ingredients are supplied either separately or mixed together in unit dosage form, for example, as lyophilized powder or water-free concentrate in a sealed container, such as an ampoule or sachet indicating the amount of active agent. When said pharmaceutical composition is administered by injection, it can be dispensed with an infusion bottle containing sterile pharmaceutical grade water or saline.
[0130] The pharmaceutical composition can be included in lipid particles or vesicles, such as liposomes or microcrystals, which are also suitable for parenteral administration. The particles can be of any suitable structure, such as unilamellar or multilamellar, so long as the composition is contained therein. The compound can be encapsulated in "stabilized plasmid lipid particles" (SPLPs) (Zhang YP et al., Gene Ther. 1999, 6:1438-47), which contain the fusogenic lipid dioleoylphosphatidylethanolamine (DOPE), low levels (5-10 mol%) of cationic lipids, and are stabilized by a polyethylene glycol (PEG) coating. Positively charged lipids, such as N-[1-(2,3-dioleoyloxy)propyl]-N,N,N-trimethyl-ammonium methylsulfate, or "DOTAP", are particularly preferred for such particles and vesicles. Preparation of such lipid particles is well known. See, e.g., U.S. Patent Nos. 4,880,635; 4,906,477; 4,911,928; 4,917,951; 4,920,016; and 4,921,757, which are incorporated herein by reference.
[0131] The pharmaceutical composition described herein can be administered or packaged as, for example, unit dose.When used in relation to the pharmaceutical composition of the present disclosure, the term "unit dose" refers to a physically separate unit suitable as a unitary dosage for subject, each unit containing a predetermined amount of active material calculated to produce desired therapeutic effect, together with necessary diluent; that is, carrier or vehicle.
[0132] Furthermore, the pharmaceutical composition can be provided as a pharmaceutical kit, for example, comprising (a) a container containing the compound of the present invention in lyophilized form; and (b) a second container containing a pharma- ceutically acceptable diluent for injection (e.g., sterile water). The pharma-ceutically acceptable diluent can be used for reconstituting or diluting the lyophilized compound of the present invention. Such container can optionally be accompanied by a notice in the form prescribed by a governmental agency regulating the manufacture, use, or sale of pharmaceutical or biological products, which notice reflects the approval by the agency of manufacture, use, or sale for human administration. EXAMPLES
[0133] Example 1: Methods Construction of TALE nuclease heterodimer Mutations were introduced into the DNA targeting module, linker domain, or FokI domain using de novo synthesis (Integrated DNA Technologies or Genescript), and TALE nuclease monomers were assembled using standard molecular biology techniques, such as enzymatic restriction digestion, ligation, and bacterial transformation. The integrity of all sequences was assessed by Sanger sequencing.
[0134] Generation of TALE nuclease fusion mRNA XL1 Blue competent bacteria were transformed with the plasmid encoding the TALE nuclease heterodimer according to standard molecular biology procedures. At least two colonies were picked from the agarose plate as miniprep cultures and DNA was extracted by QIAprep 96 plus Miniprep kit according to the manufacturer's protocol (Qiagen). Sequence verified plasmids were linearized using standard molecular biology techniques and purified using Nucleospin Gel and PCR Clean-up kit (Macherey-Nagel). mRNA was generated using HiScribe T7 ARCA mRNA kit according to the manufacturer's protocol (NEB) and purified using Mag-Bind Total Pure NGS magnetic beads (Omega) on a KingFisher Flex System (Thermo Fisher Scientific) according to the manufacturer's instructions.
[0135] cell Cryopreserved human PBMCs were cultured in X-vivo-15 medium (Lonza Group) containing IL-2 (Miltenyi Biotech,) and human serum AB (Seralab). T cells were activated for 3 days using Dynabeads Human T-Activator CD3 / CD28 for T Cell Expansion and Activation (Thermo Fisher Scientific) according to the supplier's protocol and then passaged into fresh medium.
[0136] Electroporation of TALE nuclease Two different protocols were alternatively used in different sets of experiments. (A) Four days after activation, human T lymphocytes were transfected by electroporation using the AgilePulse MAX system (Harvard Apparatus). Cells were pelleted and plated at >28 × 10 6A total of 10 μg of the indicated TALE nuclease mRNA (5 μg each of left and right monomers) and 5 × 10 6 Cells were mixed in a 0.4 cm cuvette. In parallel, mock transfections (without mRNA) were performed. Electroporation consisted of two 0.1 ms pulses at 800V followed by four 0.2 ms pulses at 130V. Following electroporation, cells were split in half and diluted into 1.2 mL of fresh warm culture medium in separate plates and incubated overnight at 30°C / 5% CO2. Cells were passaged into complete medium and maintained at 37°C / 5% CO2 for 2 days. (B) Alternatively, 4 days after activation, human T lymphocytes were transfected by electroporation (program code EO 115) using the Lonza 4D Nucleofector (Lonza). Cells were pelleted, washed with PBS, and transfected at >50 × 10 6 The cells were resuspended at 1 × 10 cells / ml in a 96-well Shuttle add-on for the 4D Nucleofector system (Lonza). 6 Cells were mixed with 1-3 μg of total mRNA (0.5-1.5 ug each of left and right monomer). In parallel, mock transfections (without mRNA) were performed. Following electroporation, cells were transferred into 96-well or 48-well culture plates containing warm fresh culture medium, which was incubated overnight at 30 °C / 5% CO2. Cells were passaged into complete medium and maintained at 37 °C / 5% CO2 for 2 days.
[0137] Cells were pelleted by centrifugation, and genomic DNA was extracted using the Mag-Bind Blood & Tissue DNA HDQ 96 Kit (Omega) on a KingFisher Flex System (Thermo Fisher Scientific) according to the manufacturer's instructions.
[0138] Targeted PCR of the endogenous locus was performed using Phusion High Fidelity PCR Master Mix with HF Buffer (NEB) to amplify a region of approximately 300 bp surrounding the TALE nuclease cut on. PCR products were purified using Mag-Bind Total Pure NGS magnetic beads (Omega) on a KingFisher Flex System (Thermo Fisher Scientific) according to the manufacturer's instructions. Amplification products were further analyzed by deep sequencing (Illumina).
[0139] Evaluation of the cleavage specificity of TALE nucleases The oligo capture assay was adapted from (Tsai et al., GUIDE-seq paper) and performed on a Fluent Automation Workstation liquid handler robot (Tecan).
[0140] TALE nucleases were co-electroporated with PCR amplifiable non-specific oligonucleotides and cells were transferred into 96- or 48-well culture plates containing warm fresh culture medium, which was incubated overnight at 30°C / 5% CO2. Cells were passaged into complete medium and maintained at 37°C / 5% CO2 for 2 days. Cells were pelleted by centrifugation and genomic DNA was extracted using the Mag-Bind Blood & Tissue DNA HDQ 96 Kit (Omega) on a KingFisher Flex System (Thermo Fisher Scientific) according to the manufacturer's instructions.
[0141] The final library was further analyzed by deep sequencing (Illumina).
[0142] Example 2: Effect of mutations in the C-terminal domain Starting with the canonical TALE-nuclease fusion (SEQ ID NO:210 and SEQ ID NO:211) containing the heterodimeric TALE-FokI nuclease (V0) described by Christian et al. [Targeting DNA Double-Strand Breaks with TAL Effector Nucleases (2010) Genetics 186:757-761], by targeting a 49 base pair sequence into the human CS1 gene (SEQ ID NO:228), two sets of substitutions were made in the C-terminal sequence between the DNA-binding core and the FokI catalytic head at positions K37 and K38 (relative to the canonical AvrBs3 C40 SEQ ID NO:109), (i) two histidines (HH, V0.1; SEQ ID NO:212 and SEQ ID NO:213) and (ii) two arginines (RR, V0.2; SEQ ID NO:214 and SEQ ID NO:215). NO:215) was introduced.
[0143] The activity of the resulting TALE nucleases, including either or both monomers with mutations, was evaluated in primary T cells as described in Example 1. The presence of single heterodimers with the above substitutions HH and RR, respectively, led to higher activity, as demonstrated by the indel frequency (Figure 2). TALE nuclease activity was also improved in the presence of both RR mutated TALE nuclease heterodimers.
[0144] Importantly, single mutant TALE monomers bearing HH had enhanced activity and simultaneously improved genome-wide specificity profiles as assessed by oligo-capture assays (Figure 3).
[0145] Example 3: Effects of Amino Acid Changes in the DNA Targeting Repeat and C-Terminal Domains Starting with the same canonical TALE nuclease heterodimers (SEQ ID NO:210 and SEQ ID NO:211) targeting the 49 base pair target sequence in CS1 (SEQ ID NO:228), a series of substitutions were introduced in the DNA binding repeats (SEQ ID NO:24-27) to obtain the V1 heterodimeric TALE nucleases (SEQ ID NO:216 and SEQ ID:217).
[0146] Starting with V1, further arginine (R) mutations were introduced at positions K37 and K38 in the C-terminal sequence to obtain V1.2 (SEQ ID NO:218 and SEQ ID NO:219).
[0147] The activity of the resulting TALE nucleases V1 and V1.2 and the original TALEN (V0) was evaluated in primary T cells as described in Example 1. Matching activity to the V0 TALEN was restored by using the V1.2 TALE nuclease, as demonstrated by the indel frequency (Figure 4). The indel frequency was further evaluated for two off-site targets, OS1 and OS2 (SEQ ID NO: 229 and SEQ ID NO: 230). Figure 5 shows that the indel frequency for both targets was reduced to background by using both the V1 and V1.2 TALE nucleases.
[0148] Finally, the genome-wide specificity profile, assessed by oligo-capture assay, was improved by using the V1 and V1.2 heterodimer construct when compared to VO (Figure 6), with activity only detected at the specific original CS1 target sequence.
[0149] Example 4: Effect of Amino Acid Changes in the FokI Catalytic Head A library of monomers of the VO structure (SEQ ID NO:210) was generated by substituting each amino acid in the wild-type FokI catalytic domain (SEQ ID NO:109) one by one with alanine.
[0150] The TALE nuclease activity generated by heterodimers formed by the resulting substituted V0 monomer of SEQ ID NO:210 and each of the other intact monomers was assessed by indel formation in the "on-site" target (SEQ ID NO:228) and two "off-site" targets, OS1 and OS2 (SEQ ID NO:229 and SEQ ID NO:230).
[0151] The detection of "on-site" and "off-site" indels for each variant in the library was normalized to the indels obtained with wild-type Fok1 (pCLS32855 and pCLS31911) (SEQ ID NO: 210 and SEQ ID NO: 211) (Figure 6).
[0152] As shown in Figure 7, several substitutions into the Fok1 catalytic domain were found to correlate with reduced indel formation in the predicted off-target OS1 while maintaining substantial nuclease activity of over 70% compared to the wild-type Fok1 sequence. These alanine substitutions into SEQ ID NO:109 were for amino acid positions 13, 52, 57, 59, 61, 65, 84, 85, 88, 91, 92, 95, 98, 103, 109, 110, 111, 113, 119, 143, 148, 152, 158, 159, 160, 167, 169, 170, and 194. Several substitutions were found to reduce indel formation while maintaining full nuclease activity, such as substitutions introduced at positions 84, 85, 88, 95, 98, 91, 103, 109, 148, 152, and 158, and further resulting in increased nuclease activity (greater than 100% activity) at positions 84, 88, and 91.
[0153] Example 5: TALE base editor for introducing nonsense mutations into the CD52 gene Construction of TALE base editor heterodimer Polynucleotide sequences were designed to target CD52 target sequences SEQ ID NOs:249-252, also referred to in Table 6, and convert one or more nucleobases C to T in these target sequences, disrupting splice sites or introducing mutations into those target sequences, with the aim of inactivating the surface presentation of CD52 in primary T cells, taking into account the expression of the heterodimeric structure depicted in FIG. 8.
[0154] One polynucleotide sequence encodes a first monomer comprising a TALE protein fused at its N-terminus to an NLS and at its C-terminus to an N-split DddA deaminase plus UGI (SEQ ID NO:220, SEQ ID NO:222, SEQ ID NO:224, and SEQ ID NO:226, respectively).
[0155] Other polynucleotide sequences encode a second monomer comprising a TALE protein fused at its N-terminus to an NLS and at its C-terminus to a C-split DddA deaminase plus UGI (SEQ ID NO:221, SEQ ID NO:223, SEQ ID NO:225 and SEQ ID NO:227, respectively).
[0156] Standard molecular biology techniques using enzymatic restriction digestion, ligation, and bacterial transformation were used to assemble the polynucleotide sequences of the above TALE proteins. The integrity of all polynucleotide sequences was assessed by Sanger sequencing.
[0157] Polynucleotide sequences encoding the above monomers were cloned into plasmids for production in suitable bacteria such as XL1-Blue.
[0158] Generation of TALE nuclease fusion mRNA XL1 Blue competent bacteria were transformed with the plasmid encoding the TALE nuclease heterodimer according to standard molecular biology procedures. At least two colonies were picked from the agarose plate as miniprep cultures and DNA was extracted by QIAprep 96 plus Miniprep kit according to the manufacturer's protocol (Qiagen). Sequence verified plasmids were linearized using standard molecular biology techniques and purified using Nucleospin Gel and PCR Clean-up kit (Macherey-Nagel). mRNA was generated using HiScribe T7 ARCA mRNA kit according to the manufacturer's protocol (NEB) and purified using Mag-Bind Total Pure NGS magnetic beads (Omega) on a KingFisher Flex System (Thermo Fisher Scientific) according to the manufacturer's instructions.
[0159] cell Cryopreserved human PBMCs were cultured in X-vivo-15 medium (Lonza Group) containing IL-2 (Miltenyi Biotech,) and human serum AB (Seralab). T cells were activated for 3 days using Dynabeads Human T-Activator CD3 / CD28 for T Cell Expansion and Activation (Thermo Fisher Scientific) according to the supplier's protocol and then passaged into fresh medium.
[0160] Electroporation of TALE base editor nucleases Four days after activation, human T lymphocytes were transfected by electroporation using the AgilePulse MAX system (Harvard Apparatus). Cells were pelleted and >28 × 10 6 A total of 10 μg of the indicated TALE nuclease mRNA (5 μg each of left and right monomers) and 5 × 106 Cells were mixed in a 0.4 cm cuvette. In parallel, mock transfections (without mRNA) were performed. Electroporation consisted of two 0.1 ms pulses at 800V followed by four 0.2 ms pulses at 130V. Following electroporation, cells were split in half and diluted into 1.2 mL of fresh warm culture medium in separate plates and incubated overnight at 30°C / 5% CO2. Cells were passaged into complete medium and maintained at 37°C / 5% CO2 for 2 days.
[0161] Cells were pelleted by centrifugation, and genomic DNA was extracted using the Mag-Bind Blood & Tissue DNA HDQ 96 Kit (Omega) on a KingFisher Flex System (Thermo Fisher Scientific) according to the manufacturer's instructions.
[0162] Targeted PCR of the endogenous locus was performed using Phusion High Fidelity PCR Master Mix with HF Buffer (NEB) according to the manufacturer's instructions to amplify a ∼300 bp region spanning the CD52 target sequence (SEQ ID NOs: 249, 250, 251, and 252). Amplification products were further analyzed by deep sequencing (Illumina) for detection of mutational events (nucleobase conversions).
[0163] Example 6: Improving the specificity of TALE nucleases targeting TGFBRII gene sequences Off-target analysis of TALE nucleases targeting TGFBRII The "classical" version (V0) of the TALEN monomer targeting the TGFBRII gene sequence (SEQ ID NO:234) was compared with the improved TALEN monomer version V1.2 according to the present invention containing tandem DD-RR mutations, and its specificity was tested by oligo capture assay.
[0164] mRNA encoding the "classical" TALE nuclease (V0) and DD-RR (V1.2) monomers targeting the TGFBRII gene sequence SEQ ID NO:234 was prepared using the mMessage mMachine T7 Ultra kit (Life Technologies) as described by Poirot et al. [Cancer Res (2015) 75 (18): 3853-3864], purified on RNeasy columns (Qiagen), and eluted in water or cytoporation medium T (Harvard Apparatus).
[0165] To perform oligo-capture assay analysis at the predicted off-site genomic locations, the heterodimer pairs V0-V0, V0-V1.2, and V1.2-V1.2 were each co-electroporated with a non-specific oligonucleotide that could be amplified by PCR. These predicted off-site locations were previously identified for the V0-V0 TALEN monomer.
[0166] The polypeptide sequences of the left and right monomers are provided in Table 5 below.
[0167] Cryopreserved human PBMCs were cultured in X-vivo-15 medium (Lonza Group) containing IL-2 (Miltenyi Biotech,) and human serum AB (Seralab). T cells were activated using Dynabeads Human T-Activator CD3 / CD28 for T Cell Expansion and Activation (Thermo Fisher Scientific) according to the provider's protocol. Six days after activation, T lymphocytes were electroporated using the AgilePulse MAX system (Harvard Apparatus) with different TALE nuclease versions targeting the same TGFBRII target sequence (SEQ ID NO:234). The TALE nucleases used either contained no mutations corresponding to SEQ ID NO:267 and SEQ ID NO:268 (VO-VO), or contained half-TALE nucleases containing DD-RR mutations corresponding to SEQ ID NO:181 and SEQ ID NO:268 (V1.2-V0), or finally contained both half-TALE nucleases containing DD-RR mutations corresponding to SEQ ID NO:181 and SEQ ID NO:180 (V1.2-V1.2). T cells were pelleted, resuspended in cytoporation medium T and incubated for 10 min with 0.5 μg of each half-TALE nuclease indicated. 6Cells were electroporated. Electroporation consisted of two 0.1 ms pulses at 800V followed by four 0.2 ms pulses at 130V. Following electroporation, cells were incubated at 30°C / 5% CO2 for 18 hours. Cells were passaged into complete medium and maintained at 37°C / 5% CO2 for 1 day and grown for 18 days. Genomic DNA (gDNA) was extracted using the Qiagen DNeasy blood & tissue kit according to the manufacturer's protocol. 200 ng of gDNA was used for high fidelity PCR amplification of on- and off-site loci using the primers listed in Table 6. Amplification products were further analyzed by deep sequencing (Illumina) to identify potential insertions at the given off-site loci.
[0168] As shown in the graphical representation in Figure 9, the percentage of indels induced by each TALE nuclease onsite was comparable, while indels induced at the different off-target sites (OT#) analyzed were no longer detectable in T cells transfected with at least one V1.2 TALE nuclease monomer containing tandem DD-RR mutations, thereby demonstrating the improved specificity of the TALEN monomers according to the present invention.
[0169] Example 7: TALE nucleases designed under V1.2 targeting TIGIT, CISH, CD38, IgH, and GADPH gene sequences TALE nucleases were designed and tested for their specificity as described in Example 1 to target the genomic sequences of each TIGIT, CISH, CD38, IgH, and GADPH human gene. The polynucleotide sequences targeted in these genes are shown in Table 6. The polypeptide sequences of the left and right TALE nuclease heterodimers are provided in Table 5. The results of the oligo capture assay for each TALEN V2 / target sequence couple are shown in Figures 10-14, which demonstrate the high specificity of the TALE scaffolds of the present invention, and the consistently high activity (activity % higher than 50%, often over 70% as shown in Figure 15).
[0170] Table 5: Polypeptide sequences used in the examples TIFF2024540639000038.tif236170TIFF2024540639000039.tif243170TIFF2024540639000040.tif243170TIFF2024540639 000041.tif247170TIFF2024540639000042.tif243170TIFF2024540639000043.tif238170TIFF2024540639000044.tif24317 0TIFF2024540639000045.tif243170TIFF2024540639000046.tif243170TIFF2024540639000047.tif243170TIFF2024540639 000048.tif243170TIFF2024540639000049.tif243170TIFF2024540639000050.tif243170TIFF2024540639000051.tif96170
[0171] Table 6: Polynucleotide sequences used in the examples TIFF2024540639000052.tif189170
Claims
1. 1. A transcription activator-like effector (TALE) protein comprising a core binding domain comprising AvrBs3-like repeats, the core-binding domain is located between the N-terminal region and the C-terminal region; wherein the N-terminal region comprises a polypeptide sequence that exhibits at least 85% sequence identity to SEQ ID NO:1; and the C-terminal region is a polypeptide sequence of 40 to 80 residues comprising a sequence having at least 85% identity to where X 1 , X 2 , and X 3 is an H (histidine) or R (arginine) residue; The transcription activator-like effector (TALE) protein.
2. The transcription activator-like effector (TALE) protein of claim 1, wherein the C-terminal region comprises SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:
4.
3. 3. The transcription activator-like effector (TALE) protein of claim 1 or 2, wherein at least one of the AvrBs3-like repeats, preferably at least two, more preferably at least three, and even more preferably at least five of the AvrBs3-like repeats, comprises a D (aspartic acid) residue at positions 4 and 32 relative to any of the AvrBs3 canonical sequences of SEQ ID NOs:31-34.
4. At least one of the AvrBs3-like repeats has the sequence: including one of where X 4 X 5 are two variable residues, 3. The transcription activator-like effector (TALE) protein of claim 1 or 2.
5. The activator-like effector (TALE) protein of claim 1 or 2, having at least 90% identity to one sequence selected from Table 4 or Table 5.
6. The activator-like effector (TALE) protein of claim 1 or 2, wherein the TALE is fused to a catalytic domain to form a TALE fusion protein.
7. The activator-like effector (TALE) protein of claim 6, wherein the catalytic domain is a nuclease domain, a deaminase domain, or a transcription modulator domain.
8. 8. The activator-like effector (TALE) protein of claim 7 for use in gene therapy.
9. 8. The activator-like effector (TALE) protein of claim 7 for use in cell therapy.
10. 10. Use of the activator-like effector (TALE) protein of claim 7 in the production of gene-edited cells.
11. 10. Use of the Activator-Like Effector (TALE) protein of claim 7 for generating engineered plant cells.
12. A polynucleotide encoding the TALE protein of claim 1 or 2.
13. A vector comprising the polynucleotide of claim 12.
14. A cell comprising the TALE protein of claim 1 or 2.
15. A cell containing the polynucleotide described in claim 12.
16. A cell comprising the vector described in claim 13.