TALE Base Editors for Gene and Cell Therapy

JP2025518278A5Pending Publication Date: 2026-05-19SELECTIS SOCIETY ANONYM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SELECTIS SOCIETY ANONYM
Filing Date
2023-06-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Current TALE base editors face challenges in efficiently targeting specific cytosine residues in human cells, particularly in primary hematopoietic stem cells and immune cells, due to off-target mutations and difficulty in defining the editing window.

Method used

The development of a novel TALE base editor scaffold, referred to as 'TALEB', which is designed to specifically target genomic sequences defined by specific rules, including the use of a heterodimeric TALE base editor generated by fusing a transcription activator-like effector array protein (TALE) with split DddA deaminase and a uracil glycosylase inhibitor (UGI).

Benefits of technology

TALEB achieves high specificity and efficiency in knocking out genes such as CD52 and β2m in primary T cells, with up to 87% phenotypic and 86% genomic level editing, while minimizing off-target mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023233003000001
    Figure 2023233003000001
  • Figure 2023233003000002
    Figure 2023233003000002
  • Figure 2023233003000003
    Figure 2023233003000003
Patent Text Reader

Abstract

The present invention relates to methods of using base editors for the efficient genetic engineering of cells, particularly primary hematopoietic stem cells (HSCs) and primary immune cells. In particular, the present invention is directed to rules for designing highly active and specific TALE base editors that exhibit an improved on-target / off-target activity ratio, useful for producing therapeutic-grade complex gene-edited cells or for performing in vivo gene therapy. The resulting TALE base editors can be used alone or in combination with rare-cut endonucleases in various gene therapy approaches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods of using base editors to efficiently genetically engineer cells, particularly primary hematopoietic stem cells (HSCs) and primary immune cells. In particular, the present invention is directed to rules for designing highly active and specific TALE base editors that exhibit an improved on-target / off-target activity ratio, useful for producing therapeutic-grade complex gene-edited cells or for performing in vivo gene therapy. The resulting TALE base editors can be used alone or in combination with rare-cutting endonucleases in various gene therapy approaches.

Background Art

[0002] Artificial transcription activator-like effectors (TALEs) form a special class of proteins that can bind to DNA originally derived from the plant pathogenic bacterium Xanthomonas [Kay S. et al. (2007) A bacterial effector acts as a plant transcription factor and induces a cell size regulator. Science 318: 648-651]. Artificial TALE proteins emerged as versatile, sequence-specific gene tools that provide flexible applications based on the elucidation of a DNA recognition "code" that links the amino acid sequence of the TALE to the genomic DNA sequence it binds [Moscou J.M. et al. (2009) A Simple Cipher Governs DNA Recognition by TAL Effectors. Science. 326:1501].

[0003] TALE binding is essentially driven by a series of 33 - 35 amino acid long repeats that differ at two positions, the so-called repeat variable dipeptide (RVD). Each base of one strand in the DNA target is contacted by a single repeat with predictable specificity due to the linear arrangement of the RVDs. Biochemical structure - function studies have suggested that the amino acid present at position 13 uniquely identifies the nucleotide in the major groove of the DNA target [Deng D., et al. (2012) Structural basis for sequence-specific recognition of DNA by TAL effectors. Science 335:720 - 723; Stella S., et al. (2013) Structure of the AvrBs3-DNA complex provides new insights into the initial thymine-recognition mechanism. Acta Crystallogr Sect. D. Biol. Crystallogr. 69(9):1707 - 1716]. This DNA - protein interaction unit is stabilized by the amino acid at position 12. For the generation of TALEs with variable precision and binding affinity, six conventional RVDs are commonly used (NG, HD, NI, NK, NH, and NN). HD and NG are associated with cytosine (C) and thymine (T), respectively. NN is a degenerate RVD that shows binding affinity for both guanine (G) and adenine (A), but its specificity for guanine is reported to be stronger. RVD NI binds to A and NK binds to G. It is noteworthy that the binding affinity of TALEs is affected by the methylation state of the target DNA sequence [Streubel J, et al. (2012) TAL effector RVD specificities and efficiencies. Nat Biotechnol 30(7):593 - 595]. Canonical RVDs do not bind efficiently to methylated cytosine.However, as described by Valton J, et al. [Overcoming transcription activator-like effector (TALE) DNA binding domain sensitivity to cytosine methylation (2012) J. Biol. Chem. 287(46):38427-38432], they can be adapted by a certain degree of degeneracy in TALEs. This code was adopted to effectively manipulate TALE DNA-binding scaffold specificity through modular assembly to form different linkages of TALE proteins with various enzyme domains, e.g., transcription activators, repressors, base editors, or nucleases with the potential ability to act on genomic sequences [Voytas et al. (2011) TAL effectors: Customizable proteins for DNA targeting. Science. 333(6051):1843-6].

[0004] TALE base editors (BEs) have more recently emerged as fusions of TALEs with deaminases and sometimes other DNA repair proteins. Base editor catalytic domains can introduce single-base mutations at desired loci in DNA (nuclear or organellar) or RNA in both dividing and non-dividing cells. Broadly, DNA base editors can be classified into cytosine base editors (CBEs), adenine base editors (ABEs), C-to-G base editors (CGBEs), dual base editors, and organellar base editors.

[0005] Mok et al. [A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing (2020) Nature. 583:631-637] recently developed a base editing approach by fusing a TALE binding domain to the bacterial cytidine deaminase toxin DddAtox and demonstrating efficient C-to-T base conversion in the mitochondrial genome in vitro. In this approach, DddAtox was split into non-toxic halves, each fused to the C-terminus of the paired (left and right) TALE binding domains to form a heterodimeric TALE base editor.

[0006] In such a setup, the deaminase DddAtox is activated by forming a functional heterodimeric cytosine deaminase that converts the C base located between the two binding sites to T when its two halves are brought sufficiently close by the TALE binding domains that recognize a given target DNA sequence in the genome. Such DddA-TALE fusion deaminase constructs have achieved mitochondrial DNA editing to some extent in mice [Lee, H., et al. (2021) Mitochondrial DNA editing in mice with DddA-TALE fusion deaminases. Nat Commun 12:1190].

[0007] However, the mitochondrial genome is considerably smaller than the nuclear genome of human cells.

[0008] It has been found that the use of such base editors is extremely difficult in human cells, especially cells for immunotherapy. Particularly in human gene therapy, the definition of the editing window for inducing C-to-T base editing at the target site is critically important to avoid unwanted substitutions of any C bases located elsewhere in the proximal genomic region.

[0009] Depending on the targeted sequences in the genome and their inherent variability in the human population, TALE base editors need further improvement to utilize their activity and reduce the risk of potential off-target substitutions.

[0010] As shown in the experimental section herein, the inventors conducted extensive investigations to define rules that enable the determination of the best target genomic sequences in the context of the design of efficient TALE base editors. They combined the screening of dozens of TALE base editors targeting various endogenous loci with the development of medium / high-throughput cell-based assays that utilize biases due to confounding effects such as epigenetic factors or modifications. This approach relied on creating a pool of cells containing artificial targets for the base editors. The cells were generated by inserting a carefully designed collection of BE target sequences (30 - 191 members) into predefined genomic loci. The pool of cells was then treated with various TALE base editors to perform gene editing on different collections of target sequences. Next-generation sequencing (NGS) analysis of the editing frequencies at the BE targets enabled better characterization of TALE base editor activity and substrate specificity within the editing window. The accumulated findings were then used to create a novel TALE base editor scaffold, herein referred to as "TALEB", which efficiently knocked out several genes in primary T cells, particularly the CD52 gene (up to 87% phenotypic and 86% genomic level editing) and the β2m gene, potential target genes for allogeneic CAR T cell adoptive therapy. The findings obtained from this study revealed editing guidelines and rules useful for the development of the TALE base editors of the present invention and their application to therapeutic immune cells. The present invention provides a platform for the rational design of higher therapeutic-grade TALE base editors based on the selection of appropriate endogenous genomic targets, beyond the novel scaffold TALEB.

Summary of the Invention

[0011] The TALE-recombinant cytosine base editor derived from DddA is a heterodimer generated by the fusion of a transcription activator-like effector array protein (TALE), half of the split DddA deaminase, and a uracil glycosylase inhibitor (UGI). It is a recent improvement of available base editor tools that can directly edit double-stranded DNA and convert cytosine (C) to thymine (T). Such TALE base editors have been used to perform editing within mitochondria, resulting in heritable modifications. However, the editing rules of this particular base editor have not been fully elucidated. To further investigate the editing rules of TALE base editors, the inventors utilized nuclease-based targeted knock-in technology to generate a pool of cells with BE target sequences unique to the same genomic locus. These cells were then treated with the TALE base editor, followed by NGS analysis of the mutation patterns at the target sequences. As shown in the experimental section herein, such an approach made it possible to generate a large and diverse pool of TALE base editor targets while excluding confounding factors, such as epigenetic and microenvironmental differences, between different genomic loci, and to gain deep insights into the editing rules in cells. Armed with the findings from this innovative approach, the inventors designed a novel scaffold called "TALEB" for a range of endogenous genes, such as those encoding CD52, TCR, B2M, and PD1, that are useful for knocking out in therapeutic immune cells.

[0012] In one aspect, the present invention relates to the identification of target sequences in the genome that specifically enable a TALE base editor to specifically focus on a desired cytosine (C) that is converted to thymine (T) while restricting off-target mutations. Such "sharper" target sequences are defined as follows: [Chemical formula] [wherein, N can be A, T, C, or G, R can be G or A, preferably G, Y can be C or T, preferably C, N left can be a polynucleotide sequence containing nucleotides between 9 and 20, and each individual nucleotide can be A, T, C, or G, N right can be a polynucleotide sequence containing nucleotides between 9 and 20, and each individual nucleotide can be A, T, C, or G, G is the complementary base of C, x = 2 to 6 y = 6 to 10 Preferably x + y ≥ 11, more preferably x + y = 12.

[0013] As shown in Figure 3, according to the above general formula, the surface of the double-stranded DNA accessible to deaminase, which represents the best target window in the genomic sequence for targeting a desired C using a TALE base editor with an approximate length of 7 nucleotides (L = 0.34 nm × 6 = 2.4 nm), is decoded. This surface has an arc f = 4 × 34.3° = 136.8° = 2.38 radians (the angle spanning 5 nucleotide bases). Assuming the radius of the double DNA helix is about 1 nm, the surface targeting C is L × R × f = 2.4 × 1 × 2.38 = 5.71 nm 2 corresponds to.

[0014] Therefore, the target surface surrounded by the diagonal line connecting the bases at positions N11, N-13 (reverse strand) and N9, N-9 (reverse strand) is such that when N left and N right are 15 base intervals apart, it is about 4.87 nm 2 in size.

[0015] According to the experiments shown in the examples, a more specific TALE base editor, for example, the "TALEB" of the present invention exemplified, can be designed to more specifically target a genomic sequence defined as follows

Chemical formula

[0016] Binding site N left and N right The spacer defined as the number of base pairs between is preferably 13 or 15 bp.

[0017] As shown by experiments, TALE C-terminals containing less than 40 amino acids, for example, the TALE base editor monomers of the present invention containing C40 and C11 exemplified herein, show high specificity for target sequences containing a 15 bp spacer. TALE C-terminals containing less than 12 amino acids, for example, the TALEB monomers containing C11 exemplified herein, showed the highest specificity not only with a 15 pb spacer but also with a 13 bp spacer. Therefore, such TALE base editor monomers are particularly suitable for target sequences containing a spacer of about 10 to 20 pb, more preferably 13 to 16 pb, and even more preferably 12 to 15 bp.

[0018] TALE C-terminal of less than 12 amino acids, in particular, the TALE base editor monomer containing C11 exemplified herein, also seems to be more discriminative when a stretch of C, for example, at least 2 CCs, 3 CCCs or 4 CCCCs is present in the target sequence. This stretch of C may or may not be preceded by T, but preferably is preceded by T, and the first C is generally the one that is converted to T (C>T) by the TALE base editor.

[0019] As an advantage of such an embodiment, there is a possibility provided by the TALE base editor of the present invention that targets genomic sequences where there will be no "T" immediately preceding the C to be edited, or genomic sequences where there is a stretch of CCCs after such a "T". Thus, the present invention broadens the number of sequences that can be edited with the TALE base editor.

[0020] In view of these findings, the present invention provides a method for designing a TALE base editor that sharply targets the C position in a gene sequence, said method comprising the following steps: i) Identifying a target sequence as defined above in the genome; ii) Synthesizing a polynucleotide sequence encoding left and right TALE binding polypeptides that bind to the polynucleotide sequences of N left and N right respectively; iii) Fusing the polynucleotide sequence encoding the left TALE binding polypeptide to a polynucleotide encoding an N-terminal split DddAtox; iv) Fusing the polynucleotide sequence encoding the right TALE binding polypeptide to a polynucleotide encoding a C-terminal split DddAtox; v) Fusing a polynucleotide sequence encoding a polypeptide that prevents uracil glycosylation, for example, UGI (uracil glycosylase inhibitor), to at least one polynucleotide sequence encoding the polynucleotide sequence obtained from ii) and iii). (vi) Optionally, simultaneously expressing the two obtained polynucleotide sequences to obtain a TALE base editor heterodimer.

[0021] According to some embodiments, the left and right TALE binding polypeptides comprise a C-terminus of 1 to 50 amino acids, preferably 8 to 40, more preferably 10 to 30, even more preferably about 11 amino acids or about 40 amino acids.

[0022] According to some embodiments, the left and right TALE binding polypeptides comprise a C-terminus of about 13 or 40 amino acids from the original AvrBs3 TALE protein that is generally at least 90%, 95% or 99% identical to SEQ ID NO: 4.

[0023] According to some embodiments, the Nter and / or Cter members of the split DddAtox comprise at least one mutation that reduces the affinity of the two split DddAtox members for each other to avoid non-specific binding of DddAtox in the genome that is independent of the TALE, thereby increasing TALE base editor specificity.

[0024] According to some embodiments, the present invention is a method for introducing a mutation into the genome of a cell, comprising introducing or expressing intracellularly a TALE base editor consisting of heterodimeric fusions of left and right TALE binding polypeptides having a C-terminal domain of about 1 to 50 amino acids with C-terminal and N-terminal split DddATox, respectively, wherein the heterodimeric TALE base editor can be considered to bind to a genomic sequence as previously defined.

[0025] As a preferred embodiment, in particular, there is a method of gene editing using the TALE base editor of the present invention in gene therapy for manipulating and producing primary cells ex vivo, more specifically, HSCs and immune cells for cell therapy, such as T cells and NK cells. The step of manufacturing the therapeutic cells more specifically includes the step in which a base editor is used to disrupt the TCR or B2M gene to make them allogeneic and / or unrecognizable to the patient's immune system, and other steps in which a rare-cut endonuclease is used for gene targeting insertion or replacement, for example, at the locus of an immune checkpoint gene.

[0026] Such manufacturing strategies are particularly effective when they combine TCR inactivation by using a base editor and the insertion / replacement of chimeric antigen receptors or recombinant TCRs at different loci such as B2M or PD1. As another example, there is the reverse strategy of B2M inactivation by using a base editor and insertion at the TCR locus.

[0027] One preferred method includes the step of creating cells resistant to immunosuppressive agents by inactivating a gene such as CD52 by using a base editor and incorporating an exogenous polynucleotide sequence into another locus by using a rare-cut endonuclease. Such steps can be carried out simultaneously by co-electroporating immune cells or their precursors with a base editing reagent and at least one nuclease reagent.

[0028] In this regard, the present invention provides special reagents and target sequences for successfully achieving the manufacture of such therapeutic immune cells, as well as various examples of TALE base editor proteins designed according to the principles and rules of the present invention.

[0029] The TALE base editor according to the present invention can also be used for in vivo gene therapy to correct or inactivate a hereditary defective gene, for example, a mutation in ApoC3, in liver cells.

[0030] The present invention encompasses vectors containing polynucleotide sequences obtainable by the present invention, as well as polypeptide sequences or reagents and their use for cell transformation and gene modification.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3

Figure 4-1

Figure 4-2

Figure 5-1

Figure 5-2

Figure 5-3

Figure 6-1

Figure 6-2

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24-1

Figure 24-2

Figure 24-3

Figure 24-4

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32-1

Figure 32-2

Mode for Carrying Out the Invention

[0032] [Explanation of Tables] [Table 1] 37 genomic target sequences used in Example 2. [Table 2] Sequences of 2×15 individual ssODNs used to identify editing windows with 15bp spacers in Example 3. [Table 3] Sequences of 191 individual ssODNs used to evaluate the effect of spacer length on editing in Example 3. [Table 4] Sequences of individual ssODNs used to evaluate the TC context in the TALE base editor target sequences in Example 4. [Table 5] Examples of KO CD52 TALEB polypeptides and target polynucleotides according to the present invention [Table 6] Predicted potential off-target sites of four TALEBs targeting CD52 evaluated in Example 4. [Table 7] List of TALEB target sequence windows according to the rules of the present invention for introducing mutations into the TRAC gene. [Table 8] List of TALEB target sequence windows according to the rules of the present invention for introducing mutations into the CD52 gene. [Table 9] List of TALEB target sequence windows according to the rules of the present invention for introducing mutations into the PD1 gene. [Table 10] List of TALEB target sequence windows according to the rules of the present invention for introducing mutations into the B2m gene. [Table 11] List of TALEB target sequence windows according to the rules of the present invention for introducing mutations into the ApoC3 gene. [Table 12] Base editor target sites in exon 1, 2, or 3 of the PK13 gene by the method of the combined gene therapy (nuclease + base editor) of the present invention exemplified in Example 5 herein. [Table 13] Polypeptide sequences of different TALE C-terminal lengths used in TALEB called C40, C11, and C0 skeletons. [Table 14] TALEB heterodimers tested in Example 6. [Table 15] Library of ssODNs containing 5'TC at position 11, adjacent to the optimal spacer length (either 13 or 15 bp spacer length) incorporated at the TCR locus targeted by the STAT3 TALEB target. [Table 16 Library of ssODNs for evaluating the effect of the context around TC in the 15 bp spacer length in Example 6. [Table 17 Library of ssODNs for evaluating the effect of the context around TC in the 13 bp spacer length in Example 6. [Table 18] Polynucleotide and polypeptide sequences used in Example 7. [Table 19] Exemplary diseases and alleles that can be cured by a gene therapy approach as exemplified in FIGS. 12 - 17, which may consist of combining a site-specific nuclease for targeted insertion of a modified rewritten gene sequence and a sequence-specific base editor that inactivates the remaining endogenous harmful allele sequence.

[0033] [Detailed Description of the Invention] Unless specifically defined herein, all technical and scientific terms used have the same meaning as commonly understood by one of ordinary skill in the fields of gene therapy, biochemistry, genetics, and molecular biology.

[0034] All methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, but suitable methods and materials are described herein. All publications, patent applications, patents, and other references cited herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Further, unless otherwise noted, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0035] In the practice of the present invention, unless otherwise specified, conventional techniques in cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and immunology within the scope of those skilled in the art are used. Such techniques are well described in the literature. For example, Current Protocols in Molecular Biology [Frederick M. AUSUBEL, 2000, Wiley and son Inc, Library of Congress, USA); Molecular Cloning: A Laboratory Manual, Third Edition, (Sambrook et al, 2001, Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press, Oligonucetide Synthesis (M. J. Gait ed., 1984), Mullis et al. U.S. Patent No. 4,683,195, Nucleic Acid Hybridization (B. D. Harries & S. J. Higgins eds. 1984), Transcription And Translation (B. D. Hames & S. J. Higgins eds. 1984), Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987), Immobilized Cells And Enzymes (IRL Press, 1986), B. Perbal, A Practical Guide To Molecular Cloning (1984), the series, Methods In ENZYMOLOGY (J. Abelson and M. Simon, eds.-in-chief, Academic Press, Inc., New York), specifically, Vols. 154 and 155 (Wu et al. eds.) and Vol. 185, "Gene Expression Technology" (D. Goeddel, ed.); See Gene Transfer Vectors For Mammalian Cells (J. H. Miller and M.P. Calos eds., 1987, Cold Spring Harbor Laboratory), Immunochemical Methods In Cell And Molecular Biology (Mayer and Walker, eds., Academic Press, London, 1987), Handbook Of Experimental Immunology, Volumes I-IV (D. M. Weir and C. C. Blackwell, eds., 1986), and Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1986).

[0036] The present invention thus has as an object a method of designing and producing a TALE protein that converts a specific C or its complementary G position in a double-stranded nucleic acid sequence to A / T. Although not always explicitly stated throughout this document, this teaching targeting a desired C position can be easily replaced with a G on the reverse DNA strand.

[0037] According to some embodiments, the method of the present invention includes the step of identifying a target sequence in a polynucleotide sequence, for example, a genomic sequence, which has the following characteristics: 5’-T 0 -N left -N y -RTC-N X -N right -A 0 -3’, or 5’-T 0 -N left -N x -GAY-N y -N right -A 0 -3’ [wherein, N can be A, T, C or G, R can be G or A, preferably G, Y can be C or T, N left can be a polynucleotide sequence containing nucleotides between 9 and 20, and each individual nucleotide can be A, T, C, or G, N right can be a polynucleotide sequence containing nucleotides between 9 and 20, and each individual nucleotide can be A, T, C, or G, G is the complementary base of C, x = 2 to 6 y = 6 to 10 preferably x + y ≥ 11, more preferably x + y = 12.

[0038] Also, it is preferable that x is included between 2 and 5, more preferably between 3 and 5.

[0039] The inventors also found that the TALE base editor, especially the TALEB of the present invention, is more specific for the polynucleotide target sequence represented by formula i) or ii): i) 5’-T 0 -N left -N y -RTCC-N X -N right -A 0 -3’, or ii) 5’-T 0 -N left -N x -GGAY-N y -N right -A 0 -3’, It is shown to be even more specific for the target sequences represented by formulas iii) and iv): iii) 5’-T 0 -N left -N y -GTCC-N X -N right -A 0 -3’, or iv) 5’-T 0 -N left -Nx -GGAC-N y -N right -A 0 -3’ [wherein, N can be A, T, C or G, R can be G or A, preferably G, Y can be C or T, N left can be a polynucleotide sequence containing 9 to 20 nucleotides, and each individual nucleotide can be A, T, C or G, N right can be a polynucleotide sequence containing 9 to 20 nucleotides, and each individual nucleotide can be A, T, C or G, G is the complementary base of C, x and y are preferably defined as follows x = 2 to 4 y = 6 to 8 11 ≧ x + y ≧ 9, more preferably, x + y = 9].

[0040] Such improved target sequences according to the present invention are particularly useful for designing and expressing corresponding appropriate specific base editor tools by synthesizing polynucleotide sequences encoding left and right TALE binding polypeptides that bind to the polynucleotide sequences of N left and N right as defined above. A polynucleotide sequence encoding such left and right TALE binding polypeptides can be fused to a polynucleotide sequence encoding a member of split DddAtox to form a TALE-DddATox heterodimer, which is generally carried out by fusing the member of split DddaTox to the C-terminus of the TALE binding polypeptide. The method of the present invention generally further includes the step of fusing a polynucleotide sequence encoding UGI (uracil glycosylase inhibitor) to one monomer of the TALE-DddATox heterodimer, as exemplified in FIG. 1.

[0041] According to some embodiments, the left and right TALE binding polypeptides are linked to split DddAtox by a TALE C-terminus of 1 to 50 amino acids, preferably 8 to 40, more preferably 10 to 30, and even more preferably about 40 amino acids. The present invention provides optimal scaffolds containing a C-terminal linker of about 11 amino acids or about 40 amino acids, which are generally derived from the AvrBs3-like TALE proteins of Xanthomonas [Christian, M. et al. TAL effector nucleases create targeted DNA double-strand breaks (2010) Genetics 186: 757-761].

[0042] As used herein, a "TALE protein" generally refers to a polypeptide comprising a core DNA-binding domain having at least 50%, preferably at least 60%, 70%, 80% or 90% identity to the DNA-binding domain of wild-type AvrBs3 [also called TalC Uniprot-G7TLQ9], which represents the prototype of the family of transcription activator-like (TAL) effectors of the plant pathogen Xanthomonas campestris. Such DNA-binding domains are characterized by repeated sequences of about 30 and 34 amino acids, including variable 2 residues usually found at positions 12 and 13.

[0043] The "AvrBs3-like repeat" refers to an artificial array of approximately 30-33 amino acids that, like the above-described consensus AvrBs3 repeat, typically contains two variable residues that interact with A, C, G, or T at positions 12 and 13. In other words, the AvrBs3-like repeat is similar to the AvrBs3 repeat and can be combined with the AvrBs3 repeat, but is generally not identical to the consensus or wild-type AvrBs3 repeat. Sometimes, as described by [Valton et al. (2012) Overcoming Transcription Activator-like Effector (TALE) DNA Binding Domain Sensitivity to Cytosine Methylation. DNA and Chromosomes. 287(46):38427], so-called * it must be noted that there may be cases where it is a (star).

[0044] The AvrBs3-like repeats of the present invention generally exhibit at least 60%, preferably at least 70%, 75%, 80%, 90%, or 95% identity to any of the above-described AvrBs3 consensus repeat sequences of SEQ ID NOs: 12-15. They generally contain D4 and D32 substitutions and are, for example, among the following repeat sequences of SEQ ID NOs: 5-11 of the present invention: LTP D QVVAIASX 12 X 13 GGKQALETVQRLLPVLCQ D HG (SEQ ID NO: 5), LTP D QVVAIASX 12 X 13 GGKQALETVQALLPVLCQDHG (SEQ ID NO: 6) LTP D QVVAIASX 12 X 13 GGKQALETVQQLLPVLCQDHG (SEQ ID NO: 7), LTP D QLVAIASX 12X 13 GGKQALETVQRLLPVLCQDHG (SEQ ID NO: 8), LTP D QMVAIASX 12 X 13 GGKQALETVQRLLPVLCQDHG (SEQ ID NO: 9), LTP D QVVAIASX 12 X 13 GGKQALETVQRLLPVLCQDQG (SEQ ID NO: 10), or LTL D QVVAIASX 12 X 13 GGKQALETVQRLLPVLCQDHG (SEQ ID NO: 11), where X 12 X 13 are two residues that interact with a given nucleotide base pair in the target sequence.

[0045] The variable two residues (X 12 X 13 ) present in the AvrBs3-like repeat and associated with the recognition of different nucleotides are generally HD for recognizing C, NG for recognizing T, NI for recognizing A, NN for recognizing G or A, NS for recognizing A, C, G or T, HG for recognizing T, IG for recognizing T, NK for recognizing G, HA for recognizing C, ND for recognizing C, HI for recognizing C, HN for recognizing G, NA for recognizing G, SN for recognizing G or A, YG for recognizing T, TL for recognizing A, VT for recognizing A or G, and SW for recognizing A. More preferably, the RVDs associated with the recognition of nucleotides C, T, A, G / A, and G are selected from the group consisting of NN or NK for recognizing G, HD for recognizing C, NG for recognizing T, NI for recognizing A, TL for recognizing A, VT for recognizing A or G, and SW for recognizing A. More generally, the RVD associated with the recognition of nucleotide C is selected from the group consisting of N * and the RVD associated with the recognition of nucleotide T is N* and H * selected from the group consisting of, wherein * may represent a gap in the repeat sequence corresponding to the absence of an amino acid residue at the second position of the RVD. In some embodiments, as described in Juillerat et al. [Optimized tuning of TALEN specificity using non-conventional RVDs (2015) Sci Rep 5:8150], X 12 X 13 may represent an amino acid residue that is unusual or non-conventional to modulate its specificity for nucleotides A, T, C, and G.

[0046] The AvrBs3-like repeats are generally, as in SEQ ID NOs: 12, 13, 14, and 15, X 12 and X 13 are represented by polypeptide sequences that are NI (preferably targeting A), HD (preferably targeting C), (preferably targeting G) NN, and NG (preferably targeting T), respectively.

[0047] In some embodiments, the invention also provides a recombinant transcriptional activator-like effector (TALE) base editor comprising one or several AvrBs3-like repeats containing D (aspartic acid) residues at positions 4 and 32, as in the polynucleotide sequence SEQ ID NOs: 5-11 above. Such AvrBs3-like repeats can be further mutated at positions 1-5 in addition to or including positions D4 and D32. Such a recombinant transcriptional activator-like effector (TALE) base editor may contain one or several of such repeats that bind to Nleft and Nright and generally forms a polypeptide containing 9-20 repeats, preferably 10-18, more preferably 11-15, or 5-12 repeats in situations where a smaller genome, such as the mitochondrial genome, is considered.

[0048] Although not required, the core DNA binding domain generally includes a half RVD consisting of 20 amino acids located at the C-terminus. The core DNA binding domain thus includes an RVD between 9.5 and 20.5, more preferably between 10.5 and 18.5, even more preferably between 11.5 and 15.5.

[0049] According to the present invention, preferably, in the core DNA binding domain as described above that includes an RVD having D4 and / or D32 substitution, the N-terminal and C-terminal sequences are adjacent, and the N-terminal and C-terminal sequences preferably have one of the following characteristics detailed below.

[0050] In some embodiments, the N-terminal sequence is derived from a naturally occurring TAL effector, such as the N-terminal domain of AvrBs3. In another embodiment, the additional N-terminal domain is the full-length N-terminal domain of a naturally occurring TAL effector N-terminal domain. In a further embodiment, the additional N-terminal domain is a variant that enables overcoming sequence constraints associated with so-called "RVD0" (i.e., the first hidden repeat), such as the requirement for a T as the first base of the binding nucleic acid sequence.

[0051] In another embodiment, the N-terminal sequence is derived from a naturally occurring TAL effector or a variant thereof. In another embodiment, the N-terminal sequence is a truncated N-terminal of such a naturally occurring TAL effector or variant. In another embodiment, the additional domain is a truncated version of the AvrBs3 TAL effector. In another embodiment, the truncated version lacks its N-terminal segment distal to the core TALE binding domain, such as the first 152 N-terminal amino acid residues or at least 152 amino acid residues of wild-type AvrBs3.

[0052] In some embodiments, the C-terminal sequence corresponds to the full or preferably truncated C-terminal region of a naturally occurring TAL effector, such as AvrBs3. Generally, said C-terminal sequence is a truncated version of the AvrBs3 TAL effector proximal to the core TALE binding domain, e.g., SEQ ID NO: 2 (11 amino acids), SEQ ID NO: 3 (40 amino acids) or SEQ ID NO: 4 (50 amino acids) or a natural variant thereof. Thus, said C-terminal sequence generally comprises or consists of a polypeptide sequence having at least 85%, 90%, 95% or 99% identity with the following SEQ ID NO: 2, SEQ ID NO: 3 or SEQ ID NO: 4: · SEQ ID NO: 2 (C-11 AA) SIVAQLSRPDP · SEQ ID NO: 3 (C-40 AA): SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX 1 X 2 GL · SEQ ID NO: 4 (C-50 AA): SIVAQLSRPDPALAALTNDHLVALACLGGRPALDAVX 1 X 2 GLPHAPALIX 3 RT

[0053] In the above sequences, X 1 , X 2 and X 3 represent amino acid substitutions introduced into the wild-type AvrBs3 C-terminal polypeptide sequence, which are K, or preferably, R (arginine) or H (histidine) residues, most preferably R. X 1 , X 2 and X 3 may be the same or different.

[0054] The N-terminal sequence or C-terminal sequence may include a localization sequence (or signal) that enables targeting of the chimeric protein to a given organelle, tissue, or cell within an organism. Non-limiting examples of such localization signals include a nuclear localization signal, a chloroplast localization signal, or a mitochondrial localization signal. In another embodiment, the additional N-terminal domain may include a nuclear export signal having the opposite effect of a nuclear localization signal that serves to target an organelle such as a chloroplast or mitochondrion. The scope of the present invention also encompasses additional C-terminal or N-terminal sequences having combinations of several localization signals. Such combinations can be, by way of non-limiting example, a nuclear localization signal (NLS) and / or a tissue-specific signal that serves to address the fusion protein of the present invention in the nucleus of tissue-specific cells. In a preferred embodiment, the NLS is generally included in the N-terminal region of the TALE-protein.

[0055] As used throughout this specification, "identity" refers to sequence identity between two nucleic acid molecules or polypeptides. Identity can be determined by comparing the positions in each sequence that can be aligned for comparison purposes. A molecule is identical at a position if the position in the sequences being compared is occupied by the same base. The degree of similarity or identity between nucleic acid or amino acid sequences is a function of the number of identical or matching nucleotides at positions shared by the nucleic acid sequences. Various alignment algorithms and / or programs, including FASTA or BLAST, which are available as part of the GCG sequence analysis package (University of Wisconsin, Madison, Wis.) and can be used, for example, with default settings, can be used to calculate the identity between two sequences. This specification generally encompasses polypeptides and polynucleotides having at least 70%, 85%, 90%, 95%, 98% or 99% identity with the specific polypeptide and polynucleotide sequences described herein, exhibiting substantially the same function, or considered equivalents.

[0056] In the present invention, DddAtox refers to the wild-type cytidine deaminase of SEQ ID NO: 1 (Uniprot number P0DUH5) from the microorganism Burkholderia cenocepacia, as described by Mok et al. [A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing (2020) Nature. 583:631-637], which can be split into two inactive halves called DddAtoxspNter (SEQ ID NO: 28) and DddAtoxspCter (SEQ ID NO: 29) at residue 1333 or 1397. These halves reconstitute deaminase activity when assembled adjacent to each other on target DNA driven by the TALE binding domain. In a preferred embodiment, DddAtox is split at residue 1397.

[0057] According to a preferred embodiment that can be considered an invention in itself, TALE base editor specificity can be further enhanced by introducing mutations into DddAtoxspNter (SEQ ID NO: 28) and DddAtoxCter (SEQ ID NO: 29) to reduce the stability of the interaction between the two splits. In such a method, only the stronger interaction induced by TALE-mediated binding between the mutated split monomers can predominate. As a result, deamination will occur more specifically at the appropriate targeted C position. The mutations can be introduced at any position in SEQ ID NO: 28 (DddAtoxNter split) and / or SEQ ID NO: 29 (DddAtoxCter split), preferably at any position in the DddAtoxCter split of SEQ ID NO: 29. Also, in the method according to the present invention, the TALE base editor monomer preferably comprises at least one mutation or modification that preferably reduces the affinity of the two split DddAtox members for each other, in the Nter and / or Cter members of the split DddAtox.

[0058] As another method for enhancing TALE base editor specificity, which can be regarded as a further invention, there is a method for reducing off-target genome mutations, which involves mutating the polypeptide sequence of the TALE base editor heterodimer to reduce the interaction with auxiliary proteins such as CTCF (CCCTC-binding factor). CTCF is a well-known transcription factor in the organization of the 3D genome structure and forms loop domains in processes involving the cohesin complex [Merkenschlager, M. & Nora, E. P. (2016) CTCF and Cohesin in Genome Folding and Transcriptional Gene Regulation. Annu Rev Genomics Hum Genet 17:17-43]. Recently, Lei, Z. et al. [Mitochondrial base editor induces substantial nuclear off-target mutations. Nature. (2022) doi.org / 10.1038 / s41586-022-04836-5] discovered that CTCF recognition sites can bias specific TALE base editor binding to its target sites, resulting in the potential for significant off-targets across the genome. Therefore, a method that includes the step of selecting an appropriate target sequence according to the present invention in combination with the step of reducing the interaction of the TALE base editor with CTCF is predicted to not only significantly increase the frequency of the desired mutation but also reduce off-target mutations within the nuclear genome.

[0059] The method of the present invention includes the step of expressing the polynucleotide construct (as DNA or mRNA) described herein in a cell, obtaining its transcription and / or translation, and obtaining a polypeptide that introduces a mutation into the genome of the cell.

[0060] The present invention also relates, inter alia, to any polypeptide or polypeptide sequence involved in the methods described herein, in particular, those encoding a TALE base editor that is active on the genomic target sequences defined herein, and also to cells transformed or engineered with these sequences or comprising said genomic target sequences.

[0061] Indeed, the present invention also relates, in particular, to a method for introducing mutations into the genome of a cell by converting C to A or G to T, comprising introducing or expressing intracellularly a polynucleotide encoding a TALE base editor as described heretofore, for example, a fusion of a left and / or right TALE binding polypeptide having a C-terminal domain of about 1 to 50 amino acids with a C-terminal and / or N-terminal split DddAtox, respectively. Such a method preferably comprises targeting a genomic sequence selected from: 5’-T 0 -N left -N y -RTC-N X -N right -A 0 5’-T 0 -N left -N x -GAY-N y -N right -A 0 [wherein, N can be A, T, C or G, R can be G or A, Y can be C or T, N left can be a polynucleotide sequence containing A, T, C or G between 9 and 20, N right can be a polynucleotide sequence containing A, T, C or G between 9 and 20, G is the complementary base of C, x = 2 to 6 y = 6 to 10 preferably, x + y ≧ 11, more preferably, x + y = 12, The heterodimeric TALE base editor binds to an N left and an N right polynucleotide sequence.

[0062] According to a preferred embodiment, the left and right TALE binding polypeptides of the TALE base editor are linked to the split deaminase via C-termini of 1 to 50 amino acids, preferably 8 to 4, more preferably 10 to 30, even more preferably about 11 amino acids or about 40 amino acids.

[0063] According to a preferred embodiment, x that determines the number of nucleotide bases in the spacer is included between 2 and 5, preferably 3 and 5, to obtain optimal specificity.

[0064] According to a preferred embodiment, the TALE base editor of the present invention has a structure comprising a TALE C-terminus containing about 11 amino acids, for example, SEQ ID NO: 2 or SEQ ID NO: 551, the latter of which contains an additional GGS linker. Such a TALE base editor structure is particularly suitable for target sequences represented by formulas i), ii), iii) and iv) as defined above, more specifically, 11 ≧ x + y ≧ 9, more preferably, x + y = 9, iii) and iv). It may be advantageous that the present invention can be implemented to introduce specific mutations into living cells ex vivo or in vivo to produce therapeutic cells, particularly therapeutic immune cells.

[0065] "Immune cells" refer to hematopoietic origin cells that are functionally involved in the initiation and / or execution of natural and / or adaptive immune responses, such as, usually, CD3 or CD4 positive cells. The immune cells according to the present invention can be T cells selected from the group consisting of dendritic cells, killer dendritic cells, mast cells, NK cells, B cells or inflammatory T lymphocytes, cytotoxic T lymphocytes, regulatory T lymphocytes or helper T lymphocytes. The cells can be obtained from several non-limiting sources including peripheral blood mononuclear cells, bone marrow, lymph node tissue, cord blood, thymus tissue, tissue from the site of infection, ascites, pleural effusion, spleen tissue and tumors, such as tumor infiltrating lymphocytes. In some embodiments, the immune cells can be derived from a healthy donor, a patient diagnosed with cancer, or a patient diagnosed with an infectious disease. In another embodiment, the cells are part of a mixed population of immune cells that exhibit different phenotypic characteristics, such as CD4, CD8 and CD56 positive cells.

[0066] In a preferred embodiment, the immune cells are tumor infiltrating lymphocytes (TIL), which include, but are not limited to, CD8+ cytotoxic T cells (lymphocytes), Th1 and Th17 CD4+ T cells, natural killer cells, dendritic cells and M1 macrophages. TIL can generally be defined biochemically using cell surface markers or functionally by its ability to infiltrate tumors and achieve therapy. TIL can generally be classified by expressing one or more of the following biomarkers: CD4, CD8, TCR αβ, CD27, CD28, CD56, CCR7, CD45Ra, CD95, PD-1 and CD25. Further, and / or, TIL can be functionally defined by its ability to infiltrate solid tumors when reintroduced into a patient.

[0067] In a preferred embodiment, the therapeutic cells are primary cells obtained from a healthy donor. "Primary cells" are intended to mean cells directly harvested from a biological tissue (e.g., a biopsy specimen) and established for in vitro growth for a limited time, which means that they can undergo a limited number of population doublings. Primary cells are in contrast to continuously tumorigenic or artificially immortalized cell lines. Non-limiting examples of such cell lines include CHO-K1 cells, HEK293 cells, Caco2 cells, U2-OS cells, NIH 3T3 cells, NSO cells, SP2 cells, CHO-S cells, DG44 cells, K-562 cells, U-937 cells, MRC5 cells, IMR90 cells, Jurkat cells, HepG2 cells, HeLa cells, HT-1080 cells, HCT-116 cells, Hu-h7 cells, Huvec cells, Molt 4 cells. Primary cells are generally used in cell therapy because they are thought to be more functional and less tumorigenic.

[0068] Generally, primary immune cells are provided from donors or patients by various methods known in the art, such as by leukapheresis techniques as reviewed by Schwartz J.et al. (Guidelines on the use of therapeutic apheresis inclinical practice-evidence-based approach from the Writing Committee of theAmerican Society for Apheresis: the sixth special issue (2013) J Clin Apher.28(3):145-284). The primary immune cells according to the present invention can also be differentiated from stem cells, such as umbilical cord blood stem cells, progenitor cells, bone marrow stem cells, hematopoietic stem cells (HSCs) and induced pluripotent stem cells (iPS).

[0069] In a preferred embodiment, the therapeutic cells of the method are T cells or NK cells that may be endowed with a chimeric antigen receptor (CAR) or a recombinant TCR as described in the prior art, such as in International Publication No. WO 2013 / 176915.

[0070] By following the teachings of the present invention, preferentially safer TALE base editor target sequences have been identified in various genes for generating engineered therapeutic immune cells.

[0071] In a preferred embodiment, the method can be used to suppress or inactivate genes encoding components of the TCR, such as those encoding TCR alpha or TCR beta, in T cells to produce T cells that are less alloreactive and can be used in allogeneic therapeutic settings. More specifically, the present invention provides a list of target window sequences (Table 7) for the TCRalpha (TRAC) gene to which TALE base editors are particularly accessible for introducing specific mutations while reducing the risk of off-target mutations across the human genome.

[0072] In a preferred embodiment, the method can be used to suppress or inactivate genes encoding targets of immunosuppressive drugs, such as alemtuzumab, e.g., CD52. By inactivating such genes, therapeutic cells can become resistant to drugs used in standard-of-care cancer treatments. In other preferred embodiments, the GR or DCK genes are inactivated by mutations, respectively, rendering the cells resistant to glucocorticoids and purine analogs.

[0073] In a preferred embodiment, the method of the present invention comprises introducing into immune cells a TALE base editor that binds to a genomic sequence contained in a gene encoding a target of an immunosuppressive drug, such as CD52. More specifically, the present invention provides a list of target window sequences for the splice acceptor site of exon 2 and into the signal peptide, to which TALE base editors are particularly accessible for introducing specific mutations while reducing the risk of off-target mutations across the human genome, into the CD52 gene (Table 8).

[0074] In a further embodiment, the method of the invention comprises introducing into immune cells a TALE base editor that binds to a genomic sequence contained in a gene encoding an immune checkpoint protein, such as PD1, CISH, CTLA4, TIM3 or LAG3. More specifically, the invention provides a list of target window sequences (Table 9) for the PD1 gene to which a TALE base editor is particularly accessible for introducing specific mutations while reducing the risk of off-target mutations throughout the human genome.

[0075] In a further embodiment, the method of the invention comprises introducing into immune cells a TALE base editor that binds to a genomic sequence contained in a gene encoding beta2-microglobulin (B2M) or human leukocyte antigen (HLA). More specifically, the invention provides a list of target window sequences (Table 10) for the B2M gene to which a TALE base editor is particularly accessible for introducing specific mutations while reducing the risk of off-target mutations throughout the human genome.

[0076] The target window sequence means a genomic sequence covered by the following general formula defined above: 5’-T 0 -N left -N y -RTC-N X -N right -A 0 -3’, or 5’-T 0 -N left -N x -GAY-N y -N right -A 0 -3’, which can be spanned by one or several TALE base editor heterodimers according to the invention, taking into account the variations of x and y and the number of nucleotides included in the N left and N right sequences.

[0077] Further examples of mutations to immune checkpoint genes that give rise to various properties of engineered immune cells for therapy are provided in the literature, in particular in WO 2019 / 016360.

[0078] According to a preferred embodiment, the method combines the use of a TALE base editor and a rare-cut endonuclease, in particular a TALE nuclease, for multiplexing gene editing in immune cells.

[0079] In some embodiments, the TALE base editor and the rare-cut endonuclease can be co-expressed, co-transfected, or sequentially introduced by minimizing the risk of chromosomal defects.

[0080] For example, as shown in Example 4, certain combinations resulted in very low levels of translocations, off-targets, and / or chromosomal rearrangements: · Inactivation of TCR using a rare-cut endonuclease and introduction of one or several point mutations into the CD52 gene using a TALE base editor, · Inactivation of TCR using a rare-cut endonuclease and introduction of one or several point mutations into the TGFBRII gene using a TALE base editor, · Inactivation of an immune checkpoint gene, such as PD1, CISH, CTLA4, TIM3 or LAG3, using a rare-cut endonuclease and introduction of one or several point mutations into the TCR by using a TALE base editor, · Inactivation of TCR using a rare-cut endonuclease and introduction of one or several point mutations into an immune checkpoint gene, such as PD1, CISH, CTLA4, TIM3 or LAG3, by using a TALE base editor, · Inactivation of TCR using a rare-cut endonuclease and introduction of one or several point mutations into a gene component of MHC, such as HLA-A, HLA-B, HLA-C or B2M, by using a TALE base editor, · Inactivation of gene component(s) of MHC, such as HLA-A, HLA-B, HLA-C or B2M, using a rare-cut endonuclease and introduction of one or several point mutations into TCR by using a TALE base editor.

[0081] This combinatorial approach, which is an important part of the present invention, is particularly useful for multiplexing knock-in (e.g., targeted gene insertion) and / or knockout (e.g., gene inactivation) in immune cells. In particular, a rare-cut endonuclease can be used to introduce an exogenous polynucleotide sequence into the genome at a first locus by site-specific gene integration, and a TALE base editor can be used simultaneously to introduce one or several point mutations at another locus, particularly at a locus that needs to be inactivated.

[0082] For example: · Using a rare-cut endonuclease to inactivate B2M expression and introducing an exogenous polynucleotide sequence encoding HLA-E at this locus to make the cell invisible to NK cells, while in the meantime, using a TALE base editor to introduce one or several point mutations into TCR and / or CD52 as previously proposed. · Using a rare-cut endonuclease to inactivate an immune checkpoint gene, such as PD1, CISH, CTLA4, TIM3 or LAG3, and introducing an exogenous polynucleotide sequence encoding a chimeric antigen receptor (CAR) at such a locus, while in the meantime, using a TALE base editor to introduce one or several point mutations into TCR and / or CD52 as previously proposed. ·Using a rare-cut endonuclease to inactivate immune checkpoint genes, such as PD1, CISH, CTLA4, TIM3, or LAG3, and introducing an exogenous polynucleotide sequence encoding a cytokine, such as IL-2, IL-12, IL-18... at such a locus, while in the meantime, using a TALE base editor to introduce one or several point mutations into TCR and / or CD52 as previously proposed. ·Using a rare-cut endonuclease to inactivate the expression of a component of TCR, such as TRAC, and introducing an exogenous polynucleotide sequence encoding a CAR or a recombinant TCR at such a locus, while in the meantime, using a TALE base editor to introduce one or several point mutations into immune checkpoints and / or CD52 as previously proposed.

[0083] As shown in the examples, in the above embodiments that combine knockout with targeted gene insertion, such as by using an AAV vector containing the transgene to introduce the transgene by homologous recombination (HDR), accidental transgene trapping (more specifically, called "AAV trapping") when the genome is simultaneously knocked out at another locus is prevented. In this method, nucleases can be used for gene insertion, while at the same time, a TALE base editor is used to inactivate genes (multiple possible) located at other positions in the genome.

[0084] "Rare-cut endonuclease" refers to a sequence-specific endonuclease reagent not found naturally in mammalian cells, and its recognition sequence generally recognizes a sequence in the range of 10 to 50 consecutive base pairs, preferably 12 to 30 bp, more preferably 14 to 20 bp. Such endonuclease reagents are generally, for example, homing endonucleases as described by Arnould S., et al. [International Publication No. 2004 / 067736], zinc finger nucleases (ZFNs) as described by Urnov F., et al. [Highly efficient endogenous human gene correction using designed zinc-finger nucleases (2005) Nature 435:646-651], TALE nucleases as described by Mussolino et al. [A novel TALE nuclease scaffold enables high genome editing activity in combination with low toxicity (2011) Nucl. Acids Res. 39(21):9283-9293] or MegaTAL nucleases as described by Boissel et al. [MegaTALs: a rare-cleaving nuclease architecture for therapeutic genome engineering (2013) Nucleic Acids Research 42(4):2591-2601], etc., nucleic acids encoding "engineered" or "programmable" rare-cut endonucleases.Due to its higher specificity, TALE nuclease has been proven to be a particularly suitable sequence-specific nuclease reagent for therapeutic applications, especially in its heterodimeric form (i.e., it functions as a pair with a "right" monomer (also called "5'" or "forward") and a "left" monomer (also called "3'" or "reverse")), as reported by, for example, Mussolino et al. [TALEN facilitate targeted genome editing in human cells with high specificity and low cytotoxicity (2014) Nucl. Acids Res. 42(10): 6762-6773]. In particular, RNA-guided endonucleases used in conjunction with, for example, Cas9 or Cpf1, as taught by Doudna, J. and Chapentier, E. [The new frontier of genome engineering with CRISPR-Cas9 (2014) Science 346 (6213): 1077], are also rare-cut endonucleases contemplated by the present invention.

[0085] According to a preferred embodiment of the present invention, the endonuclease reagent is transiently expressed intracellularly in the case of conjugates containing polynucleotide(s) and polypeptide(s), such as RNA, more specifically mRNA, protein, or a complex mixing protein and nucleic acid, so-called "ribonucleoprotein". Such conjugates can be formed, for example, using a reagent such as Cas9 or Cpf1 (RNA-guided endonuclease) together with its RNA-guide, as described by Zetsche, B. et al. [Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System (2015) Cell 163(3): 759-771].

[0086] Generally, the electroporation step is used to transfect immune cells with either or both of a nuclease and a TALE base editor, which is usually carried out in a closed chamber containing parallel plate electrodes that generate a pulsed electric field between the electrodes that is substantially uniform throughout the treatment volume and is greater than 100 volts / cm and less than 5,000 volts / cm, as described in International Publication No. WO 2004 / 083379, incorporated herein by reference, particularly on pages 23, line 25 to page 29, line 11. Such an electroporation chamber preferably has a geometric factor (cm-1) defined by the quotient of the square of the electrode gap (cm2) divided by the chamber volume (cm3), the geometric factor being 0.1 cm-1 or less, and the suspension of cells and sequence-specific reagents is in a medium adjusted to have a conductivity in the range of 0.01 to 1.0 millisiemens. Generally, the cell suspension experiences one or more pulsed electric fields. Using this method, the treatment volume of the suspension can be expanded, and the treatment time of the cells in the chamber is substantially uniform. Multiplexing a rare-cut endonuclease and a TALE base editor in immune cells can be carried out by the following protocols reported so far for the nuclease [Poirot et al. (2013) Blood. 122 (21): 1661 and Sachdeva et al. (2019) Nat Commun. 10 (1)].

[0087] The term "exogenous sequence" refers to any nucleotide or nucleic acid sequence that was not initially present at the selected locus. This sequence may be homologous to the genomic sequence, or a copy of the genomic sequence, or it may be a foreign sequence introduced into the cell. The exogenous sequence preferably encodes a polypeptide whose expression confers a therapeutic benefit to sister cells that did not have this exogenous sequence integrated at that locus. The exogenous sequence is generally introduced into the cell as a donor template and integrated into the genome by homologous recombination induced by a rare-cut endonuclease. This donor template can be introduced into the cell by transduction in the form of a viral vector, such as AAV, or, for example, as a polynucleotide, such as a single-stranded oligonucleotide (ssODN), as described in WO 2021224395.

[0088] By this method, immune cells comprising and / or co-expressing rare-cut endonucleases and TALE base editors as described herein can be obtained as a population of cells or intermediate production cells for generating engineered therapeutic cells or cell compositions.

[0089] According to some aspects of the invention, the present TALE base editor is used in gene therapy for in vivo gene correction or inactivation of defective gene expression. In particular, the TALE base editor according to the invention can be directed in vivo to liver cells so as to target cccDNA (covalently closed circular DNA), which is the resistant form of these viruses that remains within hepatocytes of viral genomes, such as hepadnaviruses, in particular HBV (hepatitis B virus).

[0090] The encapsulation of mRNA or polypeptides into nanocarriers, such as liposomes, polymers, and inorganic nanoparticles, has already shown great potential for the delivery of gene editing reagents to hepatocytes [Witzigmann, D. et al. (2020) Lipid nanoparticle technology for therapeutic gene regulation in the liver. Advanced Drug Delivery Reviews, 159:344-363].

[0091] Various types of biodegradable delivery capsules composed in the form of RNA reagents can be manufactured according to the biodegradable matrix contained and the structures of the monomer-forming core hydrophobic domain and polar domain. Delivery specificity can be improved by linking a targeting domain to the proximal polar domain of the nanocarrier so that the delivery capsule can bind to surface antigens of different cell types. The delivery capsule is particularly suitable for intravenous injection for targeting endogenous gene sequences into cells. Such a delivery capsule according to the present invention is useful for delivering a TALE base editor intracellularly in RNA form, particularly for co-delivery of messenger RNAs encoding right and left heterodimeric TALE base editors.

[0092] This application more particularly claims a pharmaceutical composition comprising the biodegradable delivery capsule of the present invention for a treatment comprising a TALE base editor according to the present invention. Such a treatment can be part of a gene therapy where specific gene sequences in anti-infective therapy have to be knocked out or repaired by targeting genes such as the ApoC3, transthyretin (TTR), ANGPTL3, and PCSK9 genes that are useful for treating or preventing, respectively, atherosclerosis, transthyretin (TTR)-mediated amyloidosis (ATTR), hyperlipidemia, and hypercholesterolemia.

[0093] In a preferred embodiment, the present invention provides a list of target window sequences (Table 11) for the ApoC3 gene to which a TALE base editor is particularly accessible for introducing specific mutations while reducing the risk of off-target mutations across the entire human genome.

[0094] According to a more specific embodiment, the present invention provides a method for introducing mutations into TRAC, CD52, PD1, B2m, and ApoC3 by targeting any of the target sequences shown in Tables 7 to 11, respectively, using a TALE base editor as described herein.

[0095] In particular, the present invention includes a method in which a TALE base editor binds to a genomic sequence contained in a gene encoding TRAC selected from any one of SEQ ID NOs: 366 to 407 as shown in Table 7.

[0096] In particular, the present invention includes a method in which a TALE base editor binds to a genomic sequence contained in a gene encoding CD52 selected from any one of SEQ ID NOs: 408 to 422 as shown in Table 8.

[0097] In particular, the present invention includes a method in which a TALE base editor binds to a genomic sequence contained in a gene encoding PD1 selected from any one of SEQ ID NOs: 423 to 466 as shown in Table 9.

[0098] In particular, the present invention includes a method in which a TALE base editor binds to a genomic sequence contained in a gene encoding B2m selected from any one of SEQ ID NOs: 467 to 501 as shown in Table 10.

[0099] In particular, the present invention includes a method in which a TALE base editor binds to a genomic sequence contained in a gene encoding ApoC3 selected from any one of SEQ ID NOs: 502 to 523 as shown in Table 11.

[0100] According to a further aspect of the present invention, mutations can be directly induced into intracellular RNA transcripts by a TALE base editor. This RNA editing method combines the introduction into cells of single-stranded DNA, such as ssODN and heterodimeric TALE base editors as described herein, where the target RNA transcript hybridizes with the single-stranded DNA to form a double-stranded nucleic acid, which is bound by the heterodimeric TALE base editor, resulting in the mutation being directly introduced at the transcript level at the desired C (or G) position in the target sequence.

[0101] As a further embodiment of the present invention, there is a method of correcting genetic deficiencies, particularly dominant alleles with dysfunction, by combining targeted gene integration, such as homologous recombination and inactivation of endogenous genes by sequence-specific base editors, such as TALE base editors, as previously described herein. The principles and schematic diagrams are illustrated in FIGS. 12-17 provided herein by way of example. Such gene therapy methods may consist of using a sequence-specific nuclease to insert a functional copy of a gene or a portion thereof or its modified sequence, in combination with the introduction into cells of a sequence-specific base editor reagent used to inactivate the remaining endogenous sequence that has not been replaced or modified. In some examples, the modified sequence integrated at the endogenous locus is rewritten with respect to the original endogenous sequence by using alternative codons. The sequence-specific base editor that recognizes the remaining intact endogenous allele sequence, preferably the one with the defect causing the genetic disease, can be introduced into cells by various means known to those skilled in the art, such as purified protein, mRNA or viral or non-viral expression vectors.

[0102] According to a preferred embodiment, gene therapy involves combining a site-specific endonuclease for performing targeted gene integration, such as a TALE nuclease, zinc finger nuclease, meganuclease, or RNA-guided endonuclease, with a sequence-specific base editor, such as a TALE base editor described heretofore. The site-specific endonuclease is co-transfected with a DNA template encoding a functional allele sequence, such as an AAV vector or single-stranded DNA, and is designed to facilitate its integration by homologous recombination.

[0103] According to a preferred embodiment, the site-specific endonuclease and the sequence-specific base editor are introduced into the cell sequentially or simultaneously, for example, by co-transfection. Co-transfection by electroporation of the mRNAs encoding both reagents is preferred, but other technical solutions are possible, such as combining viral vectorization, electroporation, nanoparticles, ribonucleotide, or purified protein transfection.

[0104] According to a preferred embodiment, the introduction of the site-specific endonuclease and the sequence-specific base editor is carried out ex vivo, for example, in blood immune cells, preferably in primary immune cells, such as in HSCs or their progeny.

[0105] According to a preferred embodiment, the sequence integrated into the genome for the purpose of correcting a genetic defect is "rewritten", which generally means that an alternative genetic code is used by alternative codon usage different from that of the endogenous allele. Thereby, the integrated rewritten sequence is not recognized by the sequence-specific base editor for the corresponding endogenous allele sequence(s).

[0106] According to a preferred embodiment, the functional gene sequence for the purpose of correcting gene deficiency can be that of an exon or a part thereof, which can be introduced into the genome, for example, according to the strategy "Artex" described in FIG. 15 and International Publication No. 2021 / 224416 incorporated by reference.

[0107] According to a preferred embodiment, the gene therapy method of the present invention targets a dysfunctional allele that causes a disease selected from those listed in Table 19.

[0108] It is also possible to consider that a variant of the above method improves its efficiency by changing different parameters, for example, one of the following: · The sequence to be incorporated for treatment can be inserted at any preferred locus in the genome, not necessarily at the locus of the defective allele. · The sequence to be incorporated for treatment may be inserted upstream of the mutation associated with the disease without a promoter. In such a case, the base editor used is preferably designed to edit the exon downstream of the therapeutic insert. · There may be included a plurality of sequence-specific base editors that target different exons of one incomplete gene.

[0109] Therefore, the present invention relates to a treatment method including one or several of the following steps: · The step of introducing and / or expressing a transgene in a cell inserted at an endogenous locus to correct a gene deficiency; · The step of introducing and / or expressing in the cell a sequence-specific base editor that targets the endogenous sequence(s) of the allele(s) causing the gene deficiency and inactivates its expression.

[0110] The above steps can be performed simultaneously or sequentially. The introduction of the transgene can be carried out by different means known in the art, whether viral or non-viral, for example, by introducing a DNA template encoding the transgene in combination with a site-specific rare-cut endonuclease.

[0111] The advantage of this method is the combination of a site-specific nuclease and a base editor, which can be simultaneously introduced into cells, for example, by electroporation, without the risk of one interacting with the other. In contrast to using multiple nucleases capable of creating chromosomal deletions or rearrangements, the combined, simultaneous use of a site-specific nuclease and a sequence-specific base editor, particularly TALE nucleases and TALE base editors, appears to be safe and without known negative interactions.

[0112] This gene therapy method is not limited to the combined use of TALE base editors and TALE nucleases as described in the examples, but can be carried out using other site-specific endonuclease reagents, such as RNA-guided endonucleases (e.g., Cas9, Cas12...), and other types of sequence-specific base editors, such as catalytically dead Cas9 (dCas9) or nickase Cas9 (nCas9) fused to a deaminase, which are guided to the target locus by a single guide (sgRNA), and any combination thereof.

[0113] Preferably, the transgene sequence is written or has a distinct gene sequence with respect to the endogenous allele causing the gene defect such that the sequence-specific base editor can easily distinguish between the endogenous defective allele and the transgene correcting the gene defect.

[0114] In some embodiments, one or several of the following steps can be performed sequentially or simultaneously: · introducing or expressing a rare-cut endonuclease targeting an endogenous locus into a cell containing a defective gene sequence that causes gene deletion; · introducing into the cell a DNA template that corrects the gene deletion by gene integration at the endogenous locus targeted by the rare-cut endonuclease; · introducing or expressing a base editor, preferably a TALE base editor, such as those described herein, to inactivate at least one endogenous allele that causes the gene deletion.

[0115] The above method occurs simultaneously so as to inactivate all alleles that are presumably involved in the gene deletion, while at the same time providing an exogenous functional copy of such alleles, and is thus particularly suitable for genetic deletions caused by dominant alleles. A non-limiting list of such genetic deletions is provided in Table 19. The method of the present invention is particularly suitable for engineering curative HSCs or T cells ex vivo, considering that it is administered to a patient for treating gene deletions, particularly for treating ADPS1 and STAT3.

[0116] One aspect of the present invention is engineered curative cells, such as HSCs or their progeny, obtainable in and / or involved in the above gene therapy method, which usually contain a transgene for correcting a gene deletion, and the transgene is generally a corrected and / or rewritten version of a defective endogenous allele that causes the gene deletion, and the endogenous allele that causes the gene deletion is inactivated (mutated) by at least one base editor.

[0117] Such engineered therapeutic cells, such as HSCs or their progeny, that can be obtained and / or are involved in the above-described gene therapy method typically include (1) a transgene for correcting a gene defect, generally a corrected and / or rewritten version of the defective endogenous allele that causes the gene defect; (2) a base editor or a transgene sequence encoding the same for inactivating the endogenous allele that causes the gene defect; and optionally, (3) a rare-cut endonuclease or a transgene sequence encoding the same for integrating the transgene at a selected endogenous locus.

[0118] Although the invention has been described generally, further understanding can be obtained by reference to certain specific examples, which are provided herein for illustrative purposes only and are not intended to limit the scope of the claimed invention.

Examples

[0119] Example 1: Materials and Methods T Cell Culture Cryopreserved human PBMCs were obtained from ALLCELLS. PBMCs were cultured in X-vivo-15 medium (Lonza Group) containing 20 ng / ml of human IL-2 (Miltenyi Biotec) and 5% human serum AB (Seralab). Human T cell activator TransAct (Miltenyi Biotec) was used to activate T cells with 25 μl of TransAct per 1 million CD3+ cells on the day after thawing PBMCs. TransAct was maintained in the culture medium for 72 hours.

[0120] TALE Nuclease and TALEB Production TALENs (fused TALE Nter (Delta152)-Repeat 15,5-Cter (40)-Fok1 nuclease domain) and TALE base editors (left TALE binding domain Nter (Delta152)-Repeat 15,5-Cter (40)-DddAtoxsp-Nter-UGI and right TALE binding domain Nter (Delta152)-Repeat 15,5-Cter (40)-DddAtoxsp-Cter) heterodimers, as exemplified in FIG. 1, were assembled using standard molecular biology and / or microbiology techniques, such as enzymatic restriction digestion, ligation, bacterial transformation and plasmid DNA extraction (NEB 10-beta competent E. coli for ccdB selection or NEB stable competent E. coli for blue / white screening) and plasmid DNA extraction. TALE DNA targeting arrays were assembled and cloned into their respective TALEN backbones (pCLS32783) and / or TALE base editor backbones (pCLS35714 and pCLS35715).

[0121] Small-scale mRNA production Plasmids of 37 TALE base editors and 37 matching TALE nucleases derived from the above backbones containing the T7 promoter and polyA sequence were produced as non-cloned after assembly (transformants were seeded directly for culture and plasmid preparation). The plasmids were then linearized using SapI (NEB) and mRNA was produced by in vitro transcription (NEB HiScribe ARCA, NEB).

[0122] Small-scale TALE nuclease and TALE base editor testing (37 endogenous targets and TRAC / CD52 multiplex engineering) T cells activated for 3 days using TransAct (Miltenyi Biotec) were transferred to fresh complete medium containing 20 ng / ml of human IL-2 (Miltenyi Biotec) and 5% human serum AB (Seralab) 10 - 12 hours prior to transfection. The harvested cells were washed once with warmed PBS. 1E6 cells washed with PBS were pelleted and resuspended in 20 μl of Lonza P3 primary cell buffer (Lonza). mRNA of TALE nuclease or TALE base editor at 1 μg / arm / 1 million cells was mixed with the cells, and then the cell mixture was electroporated using a Lonza 4D-Nucleofector under the EO115 program for stimulated human T cells. After electroporation, 80 μl of warmed complete medium was added to the cuvette to dilute the electroporation buffer, and then the mixture was carefully transferred to 400 μl of pre-warmed complete medium in a 48-well plate. Cells transfected with TALE nuclease were incubated at 30 °C overnight during culture and then returned to a 37 °C incubator. Cells transfected with TALE base editor were incubated at 37 °C throughout the process. Cells were harvested on day 6 post-transfection for gDNA extraction and NGS analysis.

[0123] Large-scale TALE nuclease and TALEB mRNA production (CD52-targeted base editor) The plasmid encoding the TRAC TALE nuclease contained a T7 promoter and a polyA sequence. TALE nuclease mRNA was produced from the TRAC TALE nuclease plasmid by Trilink. Sequences targeted by the TRAC TALE nuclease (17 bp recognition site, in uppercase, separated by a 15 bp spacer): 5’-TTCCTCCTACTCACCATcagcctcctggttatGGTACAGGTAAGAGCAA-3’ (SEQ ID NO: 31)

[0124] TALE nuclease mRNA was produced from the CD52 TALE nuclease plasmid by Trilink. Sequences targeted by CD52 TALE nuclease (17 bp recognition sites, in uppercase letters, separated by a 15 bp spacer): 5’-TTCCTCCTACTCACCATcagcctcctggttatGGTACAGGTAAGAGCAACGCCTGGCA-3’ (SEQ ID NO: 32)

[0125] Plasmids encoding the TALE base editor T-25 and the CD52 TALE base editor contained the T7 promoter and the polyA sequence. Before in vitro mRNA synthesis, the plasmid with verified sequence was linearized using SapI (NEB). mRNA was produced using the NEB HiScribe™ T7 Quick High Yield RNA Synthesis Kit (NEB). The 5’ capping reaction was carried out using the ScriptCap™ m7G Capping System (Cellscript). Antarctic phosphatase (NEB) was used to treat the capped mRNA, and final purification was performed using Mag-Bind TotalPure NGS beads (Omega bio-tek) and Invitrogen DynaMag-2 Magnet (ThermoFisher).

[0126] ssODN repair template transfection A pool of ssODNs targeting the TRAC locus (SEQ ID NOs: 33 - 69, see Table 1) was ordered from Integrated DNA Technologies (IDT) and resuspended in ddH2O at 50 pmol / μl.

[0127] T cells activated for 3 days using TransACT were transferred to fresh complete medium containing 20 ng / ml of human IL-2 (Miltenyi Biotec) and 5% human serum AB (Seralab) 10 - 12 hours before transfection.

[0128] The recovered cells were washed once with warmed PBS. 1E6 cells washed with PBS were pelleted and resuspended in 20 μl of Lonza P3 primary cell buffer (Lonza). 200 pmol of the ssODN pool and 1 μg / arm of the TRAC TALE nuclease were mixed with the cells, and then the cell mixture was electroporated using a Lonza 4D-Nucleofector under the EO115 program for stimulated human T cells. After electroporation, 80 μl of warmed complete medium was added to the cuvette to dilute the electroporation buffer, and then the mixture was carefully transferred to 400 ml of pre-warmed complete medium in a 48-well plate. The cells transfected with ssODN and TALE nuclease were incubated at 30 °C until 24 hours after TALE nuclease transfection, and then returned to 37 °C.

[0129] Cells with ssODN KI were cultured for 2 days and then recovered for TALE base editor treatment. The recovered cells were washed once with warmed PBS. 1E6 cells washed with PBS were pelleted and resuspended in 20 μl of Lonza P3 primary cell buffer (Lonza). 1 μg / arm of the TALE base editor T-25 was mixed with the cells, and then the cell mixture was electroporated using a Lonza 4D-Nucleofector under the EO115 program for stimulated human T cells. After electroporation, 80 μl of warmed complete medium was added to the cuvette to dilute the electroporation buffer, and then the mixture was carefully transferred to 400 ml of pre-warmed complete medium in a 48-well plate. The cells transfected with the TALE base editor were incubated at 37 °C for an additional 2 days and then recovered for gDNA extraction and NGS analysis.

[0130] Large-scale CD52 TALE base editor test T cells activated for 3 days using TransACT were transferred to fresh complete medium containing 20 ng / ml of human IL-2 (Miltenyi Biotec) and 5% human serum AB (Seralab) 10 - 12 hours before transfection.

[0131] The recovered cells were washed twice with Cytoporation Medium T (BTXpress, 47 - 0002). 5E6 washed cells were pelleted and resuspended in 180 μl of Cytoporation Medium T. mRNA of the TALE base editor at 2 μg / arm / 1 million cells was mixed with the cells to a final volume of 200 μl, and then the cell / mRNA mixture was electroporated using BTX Pulse Agile in a 0.4 cm gap cuvette. After electroporation, 180 μl of warmed complete medium was added to the cuvette to dilute the electroporation buffer, and then the mixture was carefully transferred to 2 ml of pre-warmed complete medium in a 12-well plate. Cells transfected with the TALE base editor were incubated at 37 °C throughout the process. Cells were harvested on day 6 post-transfection for gDNA extraction and NGS analysis.

[0132] Genomic DNA Extraction Cells were harvested and washed once with PBS. Genomic DNA extraction was performed using the Mag-Bind Blood & Tissue DNA HDQ kit (Omega Bio-Tek) according to the manufacturer's instructions.

[0133] Targeted PCR and NGS Using Phusion High-Fidelity PCR Master Mix (NEB), 100 μg of genomic DNA per reaction was used in a 50 μl reaction. The PCR conditions were set as 1 cycle at 98°C for 30 seconds; 30 cycles at 98°C for 10 seconds, 60°C for 30 seconds, and 72°C for 30 seconds; 1 cycle at 72°C for 5 minutes; and hold at 4°C. The PCR products were then purified with Omega NGS beads (at a ratio of 1:1.2) and eluted in 30 μl of 10 mM Tris buffer pH 7.4. Then, a second PCR incorporating NGS indices was performed with the purified product obtained from the first PCR. 15 μl of the first PCR product was set in a 50 μl reaction using Phusion High-Fidelity PCR Master Mix (NEB). The PCR conditions were set as 1 cycle at 98°C for 30 seconds; 8 cycles at 98°C for 10 seconds, 62°C for 30 seconds, and 72°C for 30 seconds; 1 cycle at 72°C for 5 minutes; and hold at 4°C. The purified PCR products were sequenced on a MiSeq (Illumina) with a 2×250 nano V2 cartridge.

[0134] Flow cytometry TRAC KO was monitored using an anti-TCRa / b antibody (Biolegend, #306732, clone IP26, BV605). CD52 KO was monitored using an anti-52 antibody (BD Biosciences, #563609, clone 4C8, AlexaFlour488). Flow cytometry was performed on a BD FACSCanto (BD Biosciences), and data analysis was processed using FlowJo. Cell populations were first gated for lymphocytes (SSC-A vs FSC-A) and singlets (FSC-H vs FSC-A). The lymphocyte gate was further analyzed for CD52 expression and -TCRa / b expression from this gated population.

[0135] In-silico off-site prediction To evaluate the potential off-target editing of CD52 TALE base editors, the inventors computationally generated a list of potential off-site targets for these base editors. The list was generated as follows. TALE base editors have two binding sequences of 17 bp separated by a spacer. These binding sequences always start with T. Therefore, the inventors first selected as potential targets all genomic sequences that start with T, end with A, have a size between 27 bp and 67 bp (both inclusive), and allow a spacer in the range of 10 - 40 bp. Next, the number of mismatches between the potential target and the binding sequences of the actual TALE base editor target was counted. Potential targets were removed if the total number was more than 8. Finally, all potential targets (editing windows) lacking a G in the left half of the spacer or lacking a C in the right half of the spacer were discarded.

[0136] Off-site and translocation multiplex amplicon sequencing rhAmp primers were designed at on-target and / or off-target sites established by in-silico off-site prediction. Locus-specific forward and reverse primers were obtained from Integrated DNA Technologies (IDT), either as ready-to-use pools or plated individually, and used according to the IDT protocol for RNase H2-dependent multiplex assay amplification (1 cycle at 95°C for 10 minutes; 14 cycles at 95°C for 15 seconds followed by 65°C for 8 minutes; 1 cycle at 99.5°C for 15 minutes; hold at 4°C), followed by universal PCR for adding indices (i5 or i7) for NGS (1 cycle at 95°C for 3 minutes; 24 cycles at 95°C for 15 seconds followed by 60°C for 30 seconds and 72°C for 30 seconds; 1 cycle at 72°C for 1 minute; hold at 4°C). Purified PCR amplicons were sequenced on a NextSeq (Illumina) with a NextSeq 500 / 550 Mid Output kit (150 cycles) cartridge.

[0137] Example 2: Comparison of TALE nuclease and TALEB efficiency To define the key determinants of efficient TALEB editing (C-to-T conversion), the inventors first selected a subset of 37 TALE nucleases that showed high activity (median = 82% and s.d. = 12) in primary T cells using the split-DddaTox strategy described previously (Figure 2A). These 37 target sequences (SEQ ID NOs: 33-69 in Table 1) were carefully selected to target regions with different chromatin states in T cells. The spacer sequence, the sequence between the two TALE binding regions, was also kept constant at 15 bp as previously shown to optimize TALE nucleases [Juillerat, A. et al. Comprehensive analysis of the specificity of transcription activator-like effector nucleases (2014) Nucleic Acids Research, 42(8):5390-5402]. Since previous studies have demonstrated a strong editing preference in the 5'-TC-3' context, the spacer sequences contained various numbers of evenly distributed C, G, TC, or GA (Mok et al. A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing (2020) Nature 583, 631-637). Thirty-seven TALE base editors with DddAtox split and uracil glycosylase inhibitor (UGI), replacing the FokI catalytic domain, were produced as described in Example 1. The G1397 split was used since this fusion showed better editing activity. The maximum editing within the spacer of a given TALE base editor was compared to the indel frequency generated by the corresponding TALE nuclease counterpart (Figure 2B). The complete lack of correlation between the two datasets (TALE nuclease vs. TALE base editing frequency) (Spearman correlation = 0.16, p-value = 0.33) suggests that the key determinant of efficient editing may be the positioning of the target cytosine within the spacer.Indeed, analysis of the editing efficiency from the perspective of the function of the position within the spacer revealed a 4-5 bp editing window defined by both the upper and lower strands (Figure 3).

[0138] Interestingly, for 35 out of 37 base editors, only low-frequency (<0.5%) indels (small insertions and deletions) were observed (indel frequency: median = 0.06% and s.d. = 0.17). Indels at the target site moderately correlated with the editing frequency within the spacer (Spearman correlation = 0.44, p-value = 0.007)) (Figure 4-1A). Furthermore, the inventors measured low by-product (C to A / G) editing within the editing window and overall showed a very high final purity of the edited cell population (Figure 4-1B and Figure 4-2C).

[0139] [Table 1]

[0140] Example 3: Screening and rules for optimal base editing To more comprehensively examine DddA-derived cytosine base editors, a medium- to high-throughput format screening in a defined genomic context was designed by generating a pool of primary T cells containing a predefined TALE base editor target sequence precisely inserted into the TRAC gene. Each of the TALE base editor targets contains a unique TC or GA (target for DddA deaminase) within a spacer sequence flanked by two fixed TALE binding sequences (RVD-L and RVD-R, Figure 5-1A). This setup eliminates editing variability caused by (i) different DNA binding affinities from different TALE array proteins and (ii) the influence of epigenetic factors, such as chromatin relaxation around the artificial base editor target site, allowing for uniform TALE binding to the artificial target site.

[0141] A collection of 30 ssODNs was generated (similar to the inventors' previous collection of TALE base editors targeting endogenous loci) that included the previously described T-25 polynucleotide TALE binding sequences (SEQ ID NO: 57 in Table 1) separated by 15 bp variable spacer arrays as shown below: 5’ TCTAAGAAGTTCCTGCT (variable spacer 15 nucleotides) GAATGTGGTTAGAGACA 3’ (SEQ ID NOs: 70 - 100 in Table 2).

[0142]

Table 2

[0143] Using these ssODNs, a pool of primary T cells with a collection of base editor targets was generated. 30 ssODN oligonucleotides were mixed in equal amounts and transfected into primary T cells by electroporation (200 pmol per million cells) simultaneously with the mRNA encoding the TALE nuclease targeting TRAC (left TALEN monomer of SEQ ID NO: 16 and right TALEN monomer of SEQ ID NO: 17). In the second step, 2 days after transfection of the ssODN pool, the mRNA encoding the T-25 TALE base editor was vectorized by electroporation. Then, for editing analysis, genomic DNA of the cells transfected 2 days after TALE base editor transfection was recovered (A in Figure 5-1). NGS analysis showed that the ssODNs were efficiently and uniformly integrated into the TRAC locus (read count: median = 1667.5, mean = 1686.2, s.d. = 351.7). Control samples treated without the TALE base editor showed low-frequency background mutations, while samples treated with the TALE base editor showed detectable and reproducible levels of C-to-T conversion (B in Figure 5-2 and C in Figure 5-2). Analysis further emphasized an editing window equivalent to that observed using 37 TALE base editors targeting endogenous sequences (D in Figure 5-3), fully validating this pooled approach.

[0144] The ssODN collection was expanded to spacers with various lengths ranging from 5 to 39 bp (i.e., 5, 7, 9, 11... 37, 39 bp). The TCGA quadruplex target sequences were incorporated into the spacers at every other position (A in Figure 6-1). This design containing 191 unique ssODNs (SEQ ID NOs: 103 to 293, in Table 3) enabled the examination of editing efficiency simultaneously in both strands with a single ssODN. Furthermore, to facilitate sequence analysis, unique barcodes were added to each construct (A in Figure 6-1). After filtering the NGS data to remove reads where the barcode collided with the spacer sequence, high and homogeneous representation of each ssODN was obtained (read count: median = 545, mean = 3522.6, s.d. = 7122.5). Similar to the previous collection (15-bp spacer), low-frequency mutations were observed when the TALE base editor was not used, but when the TALE base editor was used, C-to-T conversions were robustly measured in both the plus or minus strands (B in Figure 6-1 and C in Figure 6-2). Analysis of the data showed that spacer lengths in the range of 11 to 17 bp achieved optimal editing and had editing windows of 4 to 5 bp for different spacers (D in Figure 6-2 and E in Figure 6-2).

[0145] To examine the effect of the sequence surrounding the TC context on base editing efficiency, a further collection of ssODNs containing two fixed TALE array protein-binding sites from the T-25 TALE base editor (SEQ ID NO: 57 in Table 1) separated by a 16-bp spacer sequence was designed (SEQ ID NOs: 294 to 357 in Table 4). The spacer sequence consisted of a 10-bp molecular barcode followed by the NTCCNN sequence (the target of the base editor). Cell handling, transfection, and gDNA analysis were performed as previously described.

[0146] After filtering and analysis of the NGS data, the results clearly demonstrated that G or A before TC worked favorably for efficient editing (Figure 7).

[0147]

Table 3

[0148]

Table 4

[0149] Example 4: Application to TALE base editor rules for generating CD52-negative T cells In the context of allogeneic CAR-T therapy, CD52 is often gene-edited to knockout in order to create resistance to alemtuzumab, a CD52-targeted monoclonal antibody used in lymphodepleting regimens. The CD52 gene has only two exons, and exon 2 contains the sequence encoding the mature peptide, so splice site mutations at the intron 1 / exon 2 junction were selected to cause skipping of exon 2, leading to loss of CD52. Therefore, the TALE base editor rules defined above were applied to identify the optimal target, leading to three lead TALE base editors (out of 34 potential base editors, Figure 8A). Three pairs of these TALE base editors (TALEB#1, SEQ ID NO: 20 and SEQ ID NO: 21, TALEB#2, SEQ ID NO: 22 and SEQ ID NO: 23, TALEB#3, SEQ ID NO: 24 and SEQ ID NO: 25) were transfected into primary T cells as mRNA encoding them. Seven days after transfection, phenotypic CD52 knockout was monitored by flow cytometry and splice site editing was measured by NGS. The inventors observed high levels of phenotypic knockout for the three TALE base editors (Figure 8B, TALEB#1, average 81.1% + / - 4.7%, TALEB#2 SA-2 average 83% + / - 3.4% and TALEB#3, average 81.9% + / - 5.3%), which correlated with the editing level (TALEB#1, average 72.6% + / - 1.7%, TALEB#2, average 74.5% + / - 0.6% and TALEB#3, average 74.2% + / - 2.3%, Figure 8C). As predicted from the inventors' previous dataset, very low levels of indels were shown at these sites by NGS data analysis results (TALEB #1, average 0.16% + / - 0.05%, TALEB#2, average 0.28% + / - 0.06%, TALEB#3, average 0.12% + / - 0.02%, mock-transfected, average 0.01% + / - 0.005%; Figure 8C). The polypeptide and polynucleotide target sequences are reported in Table 5.

[0150]

Table 5

[0151] The inventors then attempted to create mutations within the CD52 signal peptide sequence (SEQ ID NO: 365) using a TALE base editor. Mutations in the signal peptide are known to disrupt the processing and translocation of nascent peptides and thus impair surface expression of a particular gene. Accordingly, the inventors designed TALEB: TALE base editor SP (SEQ ID NO: 26 and SEQ ID NO: 27) that could lead to (i) a silent mutation at the Leu23 residue and (ii) several amino acid changes (Gly22Lys, Ser24Leu, and Gly25Lys) in the signal peptide (A of FIG. 9). Changing residues that mutate hydrophobic glycine in the signal peptide to highly charged lysine and polar serine to hydrophobic leucine would significantly affect the ability of the signal peptide to correctly direct translocation. Indeed, 6 days after TALE base editor mRNA transfection (Example 2 SP), CD52-negative cells with an average of 84.2% (+ / - 1.8%) were observed by flow cytometry (B of FIG. 9). NGS sequencing analysis showed that all six positions were mutated at different levels (average editing frequency: G[4]: 73.65+ / - 1%, G[5]: 85.65+ / - 0.7%, C[9]: 11.4+ / - 0.1% C

[11] : 56.5+ / - 0.9%, G

[13] : 0.6+ / - 0.1, G

[14] : 6.5+ / - 0.5%) (C of FIG. 9). Sequence analysis identified 34 different species (including WT) at the protein level and showed that they were present at different ratios (D of FIG. 9).

[0152] Overall, using four CD52 TALEBs, extremely high phenotypic KO (median CD52 negative population: 82.1%) and editing purity (median = 99.7 and s.d. = 0.6) were obtained. To evaluate potential off-target editing among these four CD52 TALEBs, an in-silico list of 276 potential off-site targets was created (Table 6) and monitored using a multiplexed amplicon sequencing assay. No evidence of editing above the control experiment was demonstrated in the targeted amplicon sequencing of these sites (N = 2, independent T cell donors).

[0153]

Table 6

[0154] Finally, since the TALEB CD52 splice site BE generated only minor levels of indels, the inventors hypothesized that multiplex gene editing (i.e., the simultaneous use of base editors and nucleases, such as TALE nucleases) should not generate chromosomal translocations, a phenomenon commonly observed in cells treated with multiple nucleases. As a proof of concept, a TALE nuclease targeting TRAC (SEQ ID NO: 16 and SEQ ID NO: 17) was combined with either a CD52 TALE nuclease (SEQ ID NO: 18 and SEQ ID NO: 19) or a base editor TALE-BE SP (SEQ ID NO: 26 and SEQ ID NO: 27). High, similar levels of phenotypic double gene KO were detected by flow cytometry in both cells treated with TALE nuclease / TALE nuclease and TALE nuclease / TALE base editor (79% and 75%, respectively, Figure 10), but translocations between the two targeted loci (as measured by multiplex amplicon sequencing) were observed only in the TALE nuclease / TALE nuclease-treated samples (479 reads out of 224,406 for the TALE nuclease / TALE nuclease sample and 0 reads out of 144,323 for the TALE nuclease / BE sample, N = 1, single T cell donor, see figure in Figure 11).

[0155] Discussion Base editing corresponds to one of the latest gene editing technologies. Recently, it was demonstrated that the TALE scaffold is compatible with the creation of a new class of cytosine base editors derived from DddA. In the above experimental study, to investigate the determinants of editing by TALE base editors, screening of several base editors targeting various endogenous loci was performed along with the development of a simple robust medium-throughput approach. This throughput screening strategy utilized highly efficient and accurate TALE nuclease-mediated ssODN knock-in in primary T cells, enabling the evaluation of TALE base editor editing efficiency at hundreds of different targets in cellulo. Since all base editor artificial target sequences were inserted into the same predefined locus in the genome, this method made it possible to focus on how variations in target / spacer sequences might affect TALE base editors while excluding factors such as DNA binding affinity or epigenetic variations. The experimental results showed an optimal 13-17 bp spacer length window for editing, and for the best editing activity, the G1397C-bearing arm of the TALE base editor was positioned 4-7 bp downstream in the 3' direction of the target TC.

[0156] The extremely accurate introduction of the intended mutations (high-purity end products) is an essential requirement for applications such as gene correction. However, the generation of DSBs by base editors, especially since CRISPR / Cas base nucleases have recently been associated with major on-target genomic instability or chromosomal abnormalities [Weisheit, , et al. (2020). Detection of Deleterious On-Target Effects after HDR-Mediated CRISPR Editing. Cell Rep. 31. Boutin, J., et al. (2022). ON-Target Adverse Events of CRISPR-Cas9 Nuclease: More Chaotic than Expected. Cris. J. 5, 19-30], can pose the greatest concerns. In this study, dozens of these molecular tools were examined by TALE base editors, and even at high editing frequencies ( >80% in bulk populations), only minor byproduct mutations (C~A / G) were detected, and more importantly, low indels were generated. However, careful design of the base editor positioning made it possible to prevent or minimize bystander mutations.

[0157] Base editors have been used to edit or mutate conserved genetic elements, such as enhancers [Zeng, J., et al. (2020). Therapeutic base editing of human hematopoietic stem cells. Nat. Med. 26, 535-541], start codons, splice sites [Kluesner, M. G et al. (2021). CRISPR-Cas9 cytidine and adenosine base editing of splice-sites mediates highly-efficient disruption of proteins in primary and immortalized cells. Nat. Commun. 12:1-12], branch points, and conserved active sites [Hanna, R. E., et al. (2021). Massively parallel assessment of human variants with base editor screens. Cell 184, 1064-1080]. It is estimated that approximately 46,000 (46,608) splice sites in the genome are likely to be targeted by the TALE base editors according to the present invention, affecting 15,279 different transcripts corresponding to 76.57% of all transcripts in the human genome, indicating that overall, splice site editing could be a viable approach for gene knockout by TALE base editors. To demonstrate the feasibility of such an approach, a highly efficient TALE base editor targeting the conserved G at the intron 1 / exon 2 junction of the CD52 gene was designed. As an alternative to splice site editing, targeting signal peptides has also been demonstrated to lead to efficient surface CD52 protein knockout.

[0158] Therefore, base editors, which have hitherto been restricted to knockout or gene correction, correspond to promising molecular tools for multiplex gene manipulation. The feasibility of efficient multiplex gene manipulation using a combination of two different molecular tools, nucleases and base editors, has been demonstrated herein. Such multiplex / multi-tool strategies exhibit several advantages. First, they prevent the generation of translocations, which are often observed in the simultaneous use of several (>1) nucleases. Second, they enable the possibility of more than multiple knockouts while still allowing gene knock-in at nuclease target sites, and overall, they expand the possible scope of applications while better controlling the results of the engineered cell population (e.g., the absence of translocations). The precise position rules determined here for TALE base editors enable a lower frequency of unwanted indel generation and an increased reach to additional cellular compartments beyond traditional nuclear targets. Thereby, the potential scope of the TALE-based multiplex / multi-tool strategy is extended beyond the capabilities of most other non-TALE editing tools.

[0159] Example 5: Application to gene therapy for correcting exon 24 of the PIK3CD gene that causes combined immunodeficiency ADPS1 The method of the present invention described herein aims to improve the efficiency and safety of TALEN-mediated therapeutic gene insertion in long-term hematopoietic stem cells (LT-HSCs) of individuals suffering from dominant-negative genetic diseases. The treatment consists of TALEN-mediated insertion of a therapeutic repair matrix (cDNA of the mutated gene) in the intron or exon of the defective gene, followed by TALE base editor-mediated inactivation of the same defective gene. The TALE base editor treatment proposed by this method could theoretically increase the frequency of cells carrying a normal phenotype without creating additional genomic harmful events for the simultaneous creation of double-strand breaks. Overall, the inactivation of the remaining defective gene is thought to improve the therapeutic outcome of gene therapy interventions.

[0160] APDS1 is a combined immunodeficiency caused by a gain-of-function mutation (E1021K) that occurs in exon 24 of the PIK3CD gene. This indication can benefit from the TALEN / TALE base editor-mediated targeted repair approach, the principle of which is described in Figures 12-16 (inactivation downstream of the original exon by using the Artex integration of the rewritten PIK3CD correction sequence + base editor). Such a TALEN / TALE base editor-mediated targeted repair / inactivation approach for exon 24 of PIK3CD is exemplified below in Figures 20 and 21. Treatment of APDS1 cells with TALEN targeting intron 1 (between exons 1 and 2) facilitates the insertion of a recoded therapeutic cDNA matrix that retains the correct version of the PIK3CD cDNA (derived from exons 2 to 24). Co-treatment with a TALE base editor targeting exon 3 (see selection of TALE base editor target sites in Table 12) creates a stop codon downstream of the therapeutic cassette insertion site, thus preventing the mutated allele from being expressed.

[0161] Example 6: Effect of spacer length on the editing efficiency of C0, C11, C40 C from C to T The TALE base editor heterodimer is a double-stranded bacterial deaminase characterized by the fusion of 1) a catalytic domain divided into two inactive halves that, when reconstituted, catalyzes the conversion of cytosine (C) to thymine (T), 2) a transcription activator-like effector domain (TALE) for DNA binding, and 3) a uracil glycosylase inhibitor (UGI) (Mok B. Y. et al., Nature 2020). These TALE base editors have been used for several applications, including the generation of mutations in mitochondrial DNA in mitochondria (Mok B. Y. et al., Nature 2020) or in chloroplasts (Beum-Chang Kang, et al., Nature Plants 2021;). Despite these successful applications, the editing rules and target sequence specificities of TALE base editors remain limited. Therefore, more detailed and comprehensive studies are needed to create further generations of TALE base editors. However, such progress is difficult. In vitro studies require purified recombinant TALE base editors, and cell-based approaches are tedious because they require multiple different TALEs targeting various loci to exclude confounding effects such as epigenetic factors or modifications.

[0162] To define the key determinants of efficient TALE base editing (C to T conversion) from the perspective of the function of the 15 bp spacer length / reported preferred 5'-TC position within the editing window (Figure 1), the inventors designed a medium to high-throughput format screening in a defined genomic context (Figure 5-1) by generating a pool of primary T cells containing a predefined TALE base editor target sequence precisely inserted into the TRAC gene. Each of the TALE base editor targets contains a unique TC or GA (target for DddA deaminase) within a spacer sequence flanked by two fixed TALE binding sequences (RVD-L and RVD-R, Figure 22). This setup excludes editing variability caused by different DNA binding affinities from different TALE array proteins and epigenetic factors, such as the effect of chromatin relaxation around the artificial BE target site, enabling uniform TALE base editor binding to the artificial target site.

[0163] To examine whether the length of the linker connecting the split heads of both arms of the TAL array potentially affects the movement of the DddA head split and thus the possibility of changing target specificity, STAT3 TALE base editors with different TALE C-terminal lengths called C40, C11, and C0 skeletons were constructed (Table 13, Figure 23).

[0164]

Table 7

[0165] Effect of spacer length (C0, C11, and C40) on C to T editing efficiency To evaluate the differences associated with spacer lengths within the STAT3 target array, a collection of ssODNs containing two fixed TALE array protein binding sites from STAT3 TALE base editors separated by spacers of various lengths as shown in FIGS. 24-1 to 24-4 spanning 5 to 17 bp (i.e., 5, 7, 9, 11... 17 bp) was constructed. The TCGA quadruplex target array was incorporated into the spacer at every other position to generate a pool of primary T cells carrying a collection of BE targets. Furthermore, to facilitate sequence analysis, a barcode unique to each construct was added. The 37 unique ssODNs obtained (Table 15) were mixed in equal amounts and transfected into primary T cells by electroporation (200 pmol per million cells) simultaneously with the mRNA encoding the TALE nuclease targeting TRAC (SEQ ID NOs: 16 and 17). In the second step, two days after transfection of the TRAC TALEN and ssODN pool, the mRNA encoding the STAT3 TALE base editor (mixed linker lengths) was co-electroporated.

[0166]

Table 8

[0167] Subsequently, for editing analysis, genomic DNA of the cells transfected two days after TALE base editor transfection was recovered. The NGS analysis data compiled and presented in the figures of FIGS. 24-1 to 24-4 (C11 / C11 and C40 / C40 TALEB heterodimers) and FIG. 25 (mixed C11 / C40 and C40 / C11 TALEB heterodimers) showed the following: · The spacer was best edited when it was between 11 bp (iv) and 15 bp (vi) in length, with a maximum at 13 bp (v). · Editing was generally better when using C11 / C11, followed immediately by C40 / C40. · The C / G position and spacer size were still the parameters most affecting editing efficiency.

[0168] Effect of the context around TC: Spacer length of 15 bp As shown in Figures 26 and 27 and detailed in Table 16, a library containing 256 unique ssODNs was designed to evaluate the effect of the surrounding context within a spacer length of 15 bp on TALEB editing efficiency. PBMCs from two donors were transfected with three different pools of oligos inserted into the TRAC TALEN and the TRAC locus, followed by either STAT3 BE C40 / C40 or C11 / C11 transfection for editing of the cells using oligo KI. Genomic DNA (gDNA) was prepared from cells treated with the three oligo pools and samples were sent for sequencing on a MiSeq.

[0169] Data analysis from bioinformatics determined the contribution of each surrounding base to the efficiency of C editing, as shown in Figure 28 (A and B), and this was found to be similar for both constructs: · At position M2: A ≤ T << G < C. · At position M1: T ≤ C < A << G. · At position 1: T ≤ G < A << C · At position 2: T < C < G << A

[0170] The inventors examined multiple edits where C follows the central TC (from TCC to TTT), and editing analysis showed that the C40 structure is more tolerant than C11 (Figure 29).

[0171] These results suggest for the first time that for gene editing projects where a single point mutation (C->T) is desired, particularly for target sequences showing a spacer of 15 bp, the C11 structure is best suited for such foci. Such target sequences have the following general formula: 5’-T 0 -N left -N y -RTC-N X -N right -A 0 -3’, or 5’-T0 -N left -N x -GAY-N y -N right -A 0 -3’ [wherein, N can be A, T, C or G, R can be G or A, preferably G, Y can be C or T, N left can be a polynucleotide sequence containing 9 to 20 nucleotides, and each individual nucleotide can be A, T, C or G, N right can be a polynucleotide sequence containing 9 to 20 nucleotides, and each individual nucleotide can be A, T, C or G, G is the complementary base of C] by preferably, the following formula: 5’-T 0 -N left -N y -RTCC-N X -N right -A 0 -3’, or 5’-T 0 -N left -N x -GGAY-N y -N right -A 0 -3’ by more preferably, the following formula 5’-T 0 -N left -N y -GTCC-N X -N right -A 0 -3’, or 5’-T 0 -N left -N x -GGAC-N y -N right -A 0 -3’ [wherein, x = 2 to 6 y = 6 to 10 It can be defined by x + y = 11] It can be defined by

[0172] Influence of the context around TC: 13bp spacer length For a spacer length of 13bp, another library containing 256 unique ssODNs was designed (as detailed in Table 17). PBMCs from two donors were transfected with three different pools of oligos inserted into the TRAC TALEN and TRAC locus, followed by either STAT3 BE C40 / C40 or C11 / C11 transfection for cell editing using oligo KI. Genomic DNA (gDNA) was prepared from cells treated with eight oligo pools and samples were sent for sequencing on MiSeq. Data analysis from bioinformatics presented in Figures 30 A and B determined the contribution of each surrounding base to the efficiency of C editing, which was found to be similar for both constructs.

[0173] · At position M2: T < A << G = C. This position is not adjacent to TC but seems to be important. · At position M1: T < C < A < G · At position 1: T < G < A < C · At position 2: T < A = C < G. This position seems to be less important when using C11 in the same spacer.

[0174] The inventors examined multiple edits (from TCC to TTT) where C follows the central TC, and editing analysis showed that editing of both Ts at 13bp is more permissive for both constructs (Figure 31).

[0175] TALEB appears to be more tolerable, surprisingly, when targeting an array with a 13bp spacer than when using an array with a 15bp spacer. These results suggest that when multiple edits are desired in a gene editing project, the design of the TALE base editor should preferably be designed for genomic sequences with a 13bp spacer. Such target sequences are represented by the following general formula: 5’-T 0 -N left -N y -RTC-N X -N right -A 0 -3’, or 5’-T 0 -N left -N x -GAY-N y -N right -A 0 -3’ [wherein, N can be A, T, C or G, R can be G or A, Y can be C or T, N left can be a polynucleotide sequence containing 9 to 20 nucleotides, and each individual nucleotide can be A, T, C or G, N right can be a polynucleotide sequence containing 9 to 20 nucleotides, and each individual nucleotide can be A, T, C or G, G is the complementary base of C] Preferably, the following formula: 5’-T 0 -N left -N y -RTCC-N X -N right -A 0 -3’, or 5’-T 0 -N left -N x -GGAY-N y -N right -A 0 -3’ by More preferably, 5'-T 0 -N left -N y -GTCC-N X -N right -A 0 -3', or 5'-T 0 -N left -N x -GGAC-N y -N right -A 0 -3' [wherein, x = 2 to 4 y = 6 to 8 x + y = 9] can be defined by.

[0176] As shown in FIGS. 28(A and B), data analysis from bioinformatics aimed at determining the contribution of each surrounding base to the efficiency of C editing showed surprisingly similar results for the surrounding bases of TC for determining high editing targets containing different spacers. However, regardless of the spacer, the C11 TALEB scaffold showed stronger specificity for their target sequences.

[0177] Materials and Methods T Cell Culture Cryopreserved human PBMCs were obtained from ALLCELLS. PBMCs were cultured in X-vivo-15 medium (Lonza Group) containing 20 ng / ml of human IL-2 (Miltenyi Biotec) and 5% human serum AB (Seralab). Human T cell activator TransAct (Miltenyi Biotec) was used to activate T cells with 25 μl of TransAct per 1 million CD3+ cells on the day after thawing PBMCs. TransAct was maintained in the culture medium for 72 hours.

[0178] TALE Nuclease or TALE Base Editor Production The skeletons of TALEN (pCLS32783) and TALE base editors (pCLS35714, pCLS35715, pCLS37473, and pCLS37474, Table 13) were assembled using standard molecular biology and / or microbiology techniques, such as enzymatic restriction digestion, ligation, bacterial transformation, and plasmid DNA extraction. TALE DNA targeting arrays were assembled using standard molecular biology and / or microbiology techniques, such as enzymatic restriction digestion, ligation, bacterial transformation (NEB 10-beta competent E. coli for ccdB selection or NEB stable competent E. coli for blue / white screening), and plasmid DNA extraction, and cloned into the TALEN and / or TALE base editor skeletons.

[0179] Large-scale TALE nuclease and TALE base editor mRNA production (STAT3-targeting TALEB) The plasmid encoding the TRAC TALE nuclease contained a T7 promoter and a polyA sequence. TALE nuclease mRNA was produced from the TRAC TALE nuclease plasmid by Trilin. The sequence targeted by the TRAC TALE nuclease (17bp recognition site, uppercase, separated by a 15bp spacer).

[0180] The plasmid encoding the STAT3 TALE base editor contained a T7 promoter and a polyA sequence. Before in vitro mRNA synthesis, the plasmid with verified sequence was linearized using SapI (NEB). mRNA was produced using the NEB HiScribe™ T7 Quick High Yield RNA Synthesis Kit (NEB). The 5’ capping reaction was performed using the ScriptCap™ m7G Capping System (Cellscript). Antarctic phosphatase (NEB) was used to treat the capped mRNA, and final purification was carried out using Mag-Bind TotalPure NGS beads (Omega bio-tek) and Invitrogen DynaMag-2 Magnet (ThermoFisher).

[0181] ssODN repair template transfection The ssODN pool targeting the TRAC locus (Tables 15, 16, and 17) was ordered from Integrated DNA Technologies (IDT) and resuspended in ddH2O at 50 pmol / μl.

[0182] T cells activated for 3 days using TransACT were transferred to fresh complete medium containing 20 ng / ml of human IL-2 (Miltenyi Biotec) and 5% human serum AB (Seralab) 10 - 12 hours before transfection.

[0183] The recovered cells were washed once with warmed PBS. 1E6 cells washed with PBS were pelleted and resuspended in 20 μl of Lonza P3 primary cell buffer (Lonza). 200 pmol of the ssODN pool and 1 mg / arm of the TRAC TALE nuclease were mixed with the cells, and then the cell mixture was electroporated using a Lonza 4D-Nucleofector under the EO115 program for stimulated human T cells. After electroporation, 80 μl of warmed complete medium was added to the cuvette to dilute the electroporation buffer, and then the mixture was carefully transferred to 400 ml of pre-warmed complete medium in a 48-well plate. The cells transfected with ssODN and TALE nuclease were incubated at 30 °C for 24 hours after TALE nuclease transfection and then returned to 37 °C.

[0184] Cells with ssODN KI were cultured for 2 days and then recovered for TALEB treatment. The recovered cells were washed once with warmed PBS. 1E6 cells washed with PBS were pelleted and resuspended in 20 μl of Lonza P3 primary cell buffer (Lonza). 1 mg / arm of STAT3 TALEB (C0, C11 or C40) was mixed with the cells, and then the cell mixture was electroporated using a Lonza 4D-Nucleofector under the EO115 program for stimulated human T cells. After electroporation, 80 μl of warmed complete medium was added to the cuvette to dilute the electroporation buffer, and then the mixture was carefully transferred to 400 ml of pre-warmed complete medium in a 48-well plate. The cells transfected with the TALE base editor were incubated at 37 °C for an additional 2 days and then recovered for gDNA extraction and NGS analysis.

[0185] Genomic DNA extraction The cells were recovered and washed once with PBS. Genomic DNA extraction was performed using a Mag-Bind Blood & Tissue DNA HDQ kit (Omega Bio-Tek) according to the manufacturer's instructions.

[0186] Targeted PCR and NGS Using Phusion High-Fidelity PCR Master Mix (NEB), 100 μg of genomic DNA per reaction was used in a 50 μl reaction. The PCR conditions were set as 1 cycle at 98°C for 30 seconds; 30 cycles at 98°C for 10 seconds, 60°C for 30 seconds, and 72°C for 30 seconds; 1 cycle at 72°C for 5 minutes; and hold at 4°C. The PCR products were then purified with Omega NGS beads (at a ratio of 1:1.2) and eluted in 30 μl of 10 mM Tris buffer pH 7.4. Then, a second PCR incorporating NGS indices was performed with the purified product obtained from the first PCR. 15 μl of the first PCR product was set in a 50 μl reaction using Phusion High-Fidelity PCR Master Mix (NEB). The PCR conditions were set as 1 cycle at 98°C for 30 seconds; 8 cycles at 98°C for 10 seconds, 62°C for 30 seconds, and 72°C for 30 seconds; 1 cycle at 72°C for 5 minutes; and hold at 4°C. The purified PCR products were sequenced on a MiSeq (Illumina) with a 2 × 250 nano V2 cartridge.

[0187] Example 7: TALEB according to the present invention prevents AAV trapping On day 0, frozen human peripheral blood mononuclear cells (PBMCs) from AllCells (Alameda, California 94502) were thawed, washed, counted, and resuspended in OpTmizer medium (Gibco: A1048501) supplemented with 5% AB serum (GeminiBio: 100 - 318) and 20 ng / mL of recombinant human IL-2 (Miltenyi: 130 - 097 - 743). The cells were then transferred to an incubator set at 37°C and 5% CO2.

[0188] On the first day, PBMCs were counted, analyzed by flow cytometry to evaluate the percentage of CD3+ cells, centrifuged, and resuspended in Optimizer medium supplemented with 5% AB serum, 20 ng / mL of human IL-2, and Transact beads CD3 CD28 (Miltenyi: 130-111-160). Then, the cells were transferred to an incubator set at 37 °C and 5% CO2.

[0189] On the fourth day, the T cells were subcultured in fresh OpTimizer medium supplemented with 5% AB serum and 20 ng / ml of IL-2. Then, the plates were transferred to an incubator set at 37 °C and 5% CO2.

[0190] On the fifth day, the cells were co-electroporated using AgilePulse technology with 1 μg of mRNA encoding the left and right arms of either TRAC TALEN (SEQ ID NOs: 562 and 563) or B2M TALEN (SEQ ID NOs: 564 and 565) or TRAC TALEB (SEQ ID NOs: 566 and 567) targeting the TRAC genomic sequence SEQ ID NO: 561. During transfection, the cells were incubated in fresh OpTimizer medium at 37 °C for 15 minutes and then transferred to 30 °C for an additional 15 minutes. Then, the cells were counted and concentrated to 8E6 cells / mL, and as previously reported [Jo et al (2022) Nat Commun 13(1) and Sachdeva et al. (2019) Nat Commun. 10 (1)], were transduced or not transduced with AAV6 particles encoding HLA-E (SEQ ID NO: 560) at 50,000 vg / cell for targeted integration at the B2M locus (SEQ ID NO: 559). The modified cells were cultured overnight at 30 °C and the next day, they were subcultured in fresh OpTimizer medium supplemented with 5% AB serum and 20 ng / ml of IL-2. Then, the cells were transferred to an incubator set at 37 °C and 5% CO2.

[0191] On the eighth day, the modified T cells were harvested and analyzed by flow cytometry using anti-TCRab, anti-HLA-ABC, and anti-HLA-E antibodies.

[0192] The sequences of the reagents used in these experiments are reported in Table 18.

[0193] As shown in Figure 32-1A, approximately 80%, 60%, and 10% of the cells were TCRab-negative when treated with TRAC TALEN, TRAC TALEB, and B2M TALEN, respectively.

[0194] These results demonstrated that both TRAC TALEN and the TALE base editor are highly effective. Furthermore, when transduced with AAV6 particles, 50% HLA-E positive cells could be detected when cells were transfected with B2M TALEN, demonstrating a high targeting efficiency. When cells were transfected with TRAC TALEN and transduced with AAV6 particles (used as template DNA designed for insertion of the HLAE coding sequence at the B2M locus by homologous recombination), approximately 0.5% HLA-E positive cells could be detected. These HLA-E positive cells are not artifacts as shown in Figure 32-2B, and these results demonstrate that the DNA template is not originally designed to be inserted at the TRAC locus, but the AAV6 construct can be trapped at the TRAC locus. Importantly, when cells were transfected with the TRAC TALE base editor and transduced with AAV6 particles, no HLA-E positive cells could be detected (Figure 32-1A and 32-2B), demonstrating that such trapping can be eliminated when using the TALE base editor. Therefore, the combination of TALEN and TALEB was found to be highly efficient and rapid in ensuring a higher degree of genomic integrity, especially when performing multiple gene editing in therapeutic immune cells, particularly when combining gene editing consisting of knocking-in into the transgene using TALE nuclease according to the present invention and knocking out the endogenous gene using TALEB.

[0195]

Table 9

[0196]

Table 10

[0197]

Table 11

[0198]

Table 12

[0199]

Table 13

[0200]

Table 14

[0201]

Table 15

[0202]

Table 16

[0203]

Table 17

[0204]

Table 18

[0205]

Table 19

Claims

1. A method for designing and producing a TALE base editor heterodimer that converts a specific C to A and / or the position of its complementary G to T in a double-stranded nucleic acid sequence, i) In the nucleic acid sequence, 5'-T 0 -N left -N y -RTC-N X -N right -A 0 -3', and 5’-T 0 -N left -N x -GAY-N y -N right -A 0 -3’ [In the formula, N can be A, T, C, or G. R can be G or A, Y can be C or T, N left This can be a polynucleotide sequence containing 9 to 20 nucleotides, where each individual nucleotide can be A, T, C, or G. N right This can be a polynucleotide sequence containing 9 to 20 nucleotides, where each individual nucleotide can be A, T, C, or G. G is the complementary base of C, x = 2 to 6, [y = 6 to 10, preferably x + y ≥ 11, and more preferably x + y = 12] A step of identifying a target sequence selected from, ii) Each N left and N right A step of synthesizing polynucleotide sequences that encode left and right TALE-linked polypeptides to be bound to the polynucleotide sequence, iii) A step of fusing the polynucleotide sequence encoding the TALE-binding polypeptide on the left to the polynucleotide encoding the N-terminal split DddAtox, iv) A step of fusing the polynucleotide sequence encoding the TALE-binding polypeptide shown on the right to the polynucleotide encoding the C-terminal split DddAtox, v) The step of fusing a polynucleotide sequence encoding UGI (uracil glycosylase inhibitor) to at least one polynucleotide sequence encoding the polynucleotide sequence obtained from ii) and iii). Methods that include...

2. The method according to claim 1, wherein the left and right TALE-conjugated polypeptides include an 11-amino acid or 40-amino acid C-terminal domain from SEQ ID NO:

270.

3. The method according to claim 1, wherein the left and right TALE-binding polypeptides include AvrBs3-like repeats with D (aspartic acid) residues at positions 4 and 32 of either the canonical sequence of AvrBs3.

4. At least one of the aforementioned AvrBs3 repeats is 【Chemistry 1】 It includes one polypeptide sequence selected from the group consisting of, X 12 X 13 The method according to claim 3, wherein the amino acid is one that forms two variable residues.

5. The C-terminal domain of the TALE-binding polypeptide is 【Chemistry 2】 It consists of a polypeptide sequence of 40 to 80 residues having at least 85% identity with, or X 1 , X 2 and X 3 The method according to claim 1, wherein the residue is a K (lysine), H (histidine), or R (arginine) residue.

6. A method for introducing mutations into the genome of a cell, comprising the step of introducing or expressing a TALE base editor into the cell, the TALE base editor comprising a heterodimer fusion of left and right TALE-binding polypeptides having a C-terminal domain of 1 to 50 amino acids with the C-terminal and N-terminal split DddAtox, respectively, wherein the heterodimer TALE base editor 【Transformation 3】 [In the formula, N can be A, T, C, or G. R can be G or A, Y can be C or T, N left It can be a polynucleotide sequence, where each individual nucleotide can be A, T, C, or G. N right It can be a polynucleotide sequence, where each individual nucleotide can be A, T, C, or G. G is the complementary base of C, x = 2 to 6, y = 6 to 10 Preferably x + y ≥ 11, and more preferably x + y = 12. A method for binding to a selected genome sequence.

7. The method according to claim 6, wherein the left and right TALE-conjugated polypeptides have C-terminal domains of 1 to 50 amino acids, preferably 8 to 40 amino acids, and more preferably 10 to 30 amino acids.

8. The method according to claim 6, wherein the left and right TALE-conjugated polypeptides each have a C-terminal domain of 11 or 40 amino acids.

9. The method according to claim 6, wherein the cells are hematopoietic stem cells.

10. The method according to claim 6, wherein the cells are immune cells.

11. The method according to claim 10, wherein the immune cells are T cells or NK cells.

12. The method according to claim 6, wherein the immune cells are conjugated with a chimeric antigen receptor (CAR) or recombinant TCR.

13. The method according to claim 6, wherein the TALE base editor binds to a genomic sequence contained in a gene encoding TRAC.

14. The method according to claim 13, wherein the TALE base editor binds to a genome sequence selected from any one of sequence numbers 366 to 407.

15. The method according to claim 6, wherein the TALE base editor binds to a genomic sequence contained in a gene encoding a target of an immunosuppressant drug.

16. The method according to claim 15, wherein the TALE base editor binds to a genomic sequence contained in a gene encoding CD52.

17. The method according to claim 16, wherein the TALE base editor binds to a genome sequence selected from any one of sequence numbers 408 to 422.

18. The method according to claim 6, wherein the TALE base editor binds to a genomic sequence contained in a gene encoding an immune checkpoint.

19. The method according to claim 18, wherein the TALE base editor binds to a genomic sequence in the PD1 gene selected from any one of sequence numbers 423 to 466.

20. The method according to claim 6, wherein the TALE base editor binds to a genomic sequence contained in a gene encoding B2M.

21. The method according to claim 20, wherein the TALE base editor binds to a genomic sequence in a B2M gene selected from either SEQ ID NO: 467 or SEQ ID NO:

501.

22. The method according to claim 6, wherein the genome sequence is targeted in the liver and gene therapy is performed.

23. The method according to claim 22, wherein the genome sequence encodes a defective ApoC3 protein.

24. The method according to claim 22, wherein the TALE base editor binds to a genomic sequence in ApoC3 selected from either SEQ ID NO: 502 or SEQ ID NO: 523.