Optimized Base Editor
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2026-04-06
AI Technical Summary
Current adenine base editors (ABEs) lack high average editing efficiency, wide editing windows, and low off-target effects, limiting their application to a wide range of host cells, particularly in plants.
The development of a systematically optimized Cas12a-ABE by modifying nuclear localization signals, linkers, adenosine deaminase domains, and crRNA components in iterative cycles, resulting in a monomeric structure with enhanced versatility and efficiency.
The optimized Cas12a-ABE achieves comparable or improved editing efficiency compared to known ABEs, with a simpler domain structure, wider applicability, and reduced off-target effects, making it suitable for diverse host cells, including plants.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to adenine base editors (ABEs) and components thereof. The present invention also relates to complexes comprising an adenine base editor (ABE) and a guide RNA in a functionally associated form. The present invention further relates to a nucleic acid molecule encoding an ABE / guide RNA, an expression construct or vector comprising a nucleic acid sequence encoding an adenine base editor and / or a nucleic acid sequence encoding a guide RNA. The present invention further relates to a cell comprising an adenine base editor and a method for adenine base editing of a target site in a genome of interest in at least one cell of a prokaryotic or eukaryotic organism, including bacterial and archaeal organisms. In addition, the present invention relates to various methods, kits and uses related to the provided ABEs. [Background technology]
[0002] Base editors, on the one hand, represent a very useful biotechnological tool to generate precise nucleotide substitutions at specific DNA target sites, especially for site-specific eukaryotic and prokaryotic (including bacterial and archaeal) genome editing of complex genomes where high precision is paramount. Currently, there are two main types: cytidine / cytosine (CBEs) and adenine / adenosine base editors (ABEs). CBEs are usually created by fusing a cytidine deaminase domain to a catalytically inactive Cas9 (either dead Cas9(D10A / H840A) or nickase Cas9(D10A)). Various cytidine deaminases have been used for base editing, such as APOBEC1(A1), A3A, A3B, PmCDA1, AID and their derivatives (Rees and Liu, 2018. Nat. Rev. Genet.; doi:10.1038 / s41576-018-0059-1). CBE catalyzes the deamination of cytidine to uracil on the non-targeted DNA strand, ultimately resulting in a CG to TA mutation (for CBE, see Komor et al., Nature 533, 420-460, 2016; Komor et al., 2017, Science Advances, doi:10.1126 / sciadv.aao4774). With regard to Cas9 variants suitable for base editing, nCas9 is believed to be more active than dCas9 because nicking of the targeted strand allows the non-targeted strand to be used as a template for mismatch-mediated repair (e.g., Eid et al., Biochem J. 2018 Jun 15;475(11):1955-1964).
[0003] ABE is derived from TadA evolved from Escherichia coli (Gaudelli et al., Nature, 551:464-471, 2017) and catalyzes adenine to inosine, which is repaired as guanine, resulting in an AT to GC transition. Similar to CBE, the use of D10A Cas9 nickase increases editing frequency and the editing window is similar (Gaudelli et al., 2017, supra; Koblan et al., Nature Biotech, 36, 843-846, 2018). Unlike CBE, the repair product is more precise and fewer indels are observed (Rees and Liu, Nat Rev Genet, 19(12):770-788, 2018).
[0004] CBE and ABE are currently used in a wide range of species, with Cas9 being the primary nuclease platform, but it has limited target scope because the PAM and deamination window control the target space. Many groups have utilized alternative Cas9 PAM variants to overcome this limitation (e.g., Tan et al., Nat Comm, 110, 439, 2019).
[0005] Besides Cas9, Cas12a (previously called Cpf1, CRISPR class II type V nuclease; see Zetsche et al., Cell, 163(3):759-7741, 2015) represents another programmable DNA endonuclease guided by single guide RNA (gRNA; sgRNA), which on the one hand represents an important tool for genome editing in higher eukaryotic cells, including plant cells (Bandyopadhyay et al., Front Plant Sci, 11:58411, 2020). On the other hand, various Cas12a variants with altered and enhanced PAM specificity have been provided (Gao et al., Nat Biotechnol, 35:789-792, 2017; Toth et al., Nucleic Acids Res., 48:3722-3733, 2020). Moreover, temperature-resistant variants of Cas12a have already been described (Schindele and Puchta, 2020 Plant Biotechnology Journal, 18, 1118-1120). It has also been reported that ABEs based on Cas9 variants in combination with TadA8e and TadA9 have certain off-target effects that prevent the widespread and targeted use of these ABEs in plants (Li et al., Genome Biology 2022, 23:51, https: / / Doi.org / 10.1186 / s13059-022-02618-w).
[0006] Although Cas12a-derived base editors have been reported occasionally, no systematic reports on the activity of base editors using Cas12a are currently available, and the reported functions of Cas12a-derived editors are usually considered to be much lower than those of Cas9-based base editors, taking into account the fact that highly active nickase variants of Cas12a (such as the nCas9 D10A mutant) are not available. In particular, suitable Cas12a-derived base editors with high activity and specificity suitable for plant applications, let alone ABEs, are still not available (cf. Molla et al., Nature Plants, 7: 1166-1187, 2021, see especially p. 1173).
[0007] WO 2022 / 020407A1 describes Cas12a-based ABEs that are functional in plants. These ABEs have a heterodimeric structure with respect to their adenosine deaminase domains, one of which is an evolved / mutated adenosine deaminase domain.
[0008] Considering the specific PAM target space of Cas12a compared to Cas9, Cas12a base editors will be of great interest not only for basic science, but also for precise genome editing in therapy, applications in unicellular organisms (prokaryotes and eukaryotes, including yeast) and genome editing in plants. The Cas12a system may be very useful for gene inactivation, since Cas12a nuclease cleaves DNA at the distal region associated with the respective PAM sequence, which is therefore not as critical for target binding and cleavage, and the cleaved and subsequently repaired DNA sequence may be re-cleaved as the critical recognition motif is maintained.
[0009] Typically, optimization of genome editing tools such as base editors requires time-consuming testing of numerous constructs and targets, as generally applicable high-throughput platforms for testing and modifying novel base editors are still not available (Komor et al., 2017, supra; Gaudelli et al., supra; Gao et al., supra). Furthermore, findings on one base editor system, e.g., nCas9 base CBE, cannot be simply extrapolated when trying to define new base editors based on another base nuclease, such as Cas12a. As detailed above, CBEs and ABEs also differ significantly in structure, applicability and specificity.
[0010] Currently, there is still no optimized ABE with high average editing efficiency, wide editing window, low off-target effects, and broad applicability to various types of host cells. Meanwhile, those skilled in the art are aware of several TadA variants, orthologues, and mutants that have been identified in silico, evolved, and tested for ABE functionality (see Zhang et al., Nature Communications, 2023, 14:414, https: / / doi.org / 10.1038 / s41467-023-36003-3 and supplementary data). Nevertheless, only focusing on specific Cas9-based ABEs, and not even testing Cas12a-based ABEs or TadA9 alone or in combination with ABEs, Zhang et al. 2023 shows the significant difficulties associated with identifying functional ABEs, since all functional parts of the ABE fusion need to be optimized in terms of the overall structure and individual parts (proteins and linkers) to provide a functional ABE.
[0011] Even if certain ABEs are already available for very specific purposes, the development of new ABEs with improved functionality is extremely challenging: the development, testing and validation of such large fusion proteins must start from scratch, as even the smallest modifications can result in a complete loss of function, despite the large size of such fusion proteins. Summary of the Invention [Problem to be solved by the invention]
[0012] Therefore, using a systematic approach, herein referred to as ITER (Iterative Testing of Editing Reagents), the objective of the present invention is to de novo develop and iteratively optimize Cas12a-ABE by modifying various nuclear localization signals (NLS), linkers, adenosine deaminase domains and crRNA components in several iterative cycles to provide a new ABE tool applicable in a wide range of target cells (prokaryotes and eukaryotes), which shows targeted and high specific activity for improving genome editing technologies in general, especially in plants.
[0013] Another object of the present invention is to provide a highly versatile Cas12a-based ABE that functions in plants, has a relatively simple domain structure, and is suitable for a variety of applications. The Cas-12a-based ABE according to the present invention has a monomeric structure for its adenosine deaminase domain. Compared to known Cas-12a-based ABEs that function in plants, this simpler structure reduces the complexity and may allow for increased versatility in the use of each ABE according to the present invention. Furthermore, the relatively simple structure of the Cas12a system results in a relatively small molecular weight and size, which is beneficial in terms of efficiency in many methods of gene delivery and transfection.
[0014] It is a further object of the present invention to provide a Cas12a-based ABE that functions in plants, where the relatively simple domain structure of the Cas12a-based ABE still allows for at least similar, and preferably even better, editing efficiency compared to known ABEs that function in plants. [Brief description of the drawings]
[0015] [Figure 1a] Figure 1a-b are schematic diagrams of the tested Cas12a base editor (BE) expression construct (a) and guide RNA (gRNA) expression construct (b). Figure 1c is the fluorescent reporter system used to measure ABE activity in wheat protoplasts. Upon A:T to G:C base editing, mutating the stop codon to a Gln(Q) codon restores a functional GFP coding sequence. [Figure 1b] Figure 1a-b are schematic diagrams of the tested Cas12a base editor (BE) expression construct (a) and guide RNA (gRNA) expression construct (b). Figure 1c is the fluorescent reporter system used to measure ABE activity in wheat protoplasts. Upon A:T to G:C base editing, mutating the stop codon to a Gln(Q) codon restores a functional GFP coding sequence. [Figure 1c] Figure 1a-b are schematic diagrams of the tested Cas12a base editor (BE) expression construct (a) and guide RNA (gRNA) expression construct (b). Figure 1c is the fluorescent reporter system used to measure ABE activity in wheat protoplasts. Upon A:T to G:C base editing, mutating the stop codon to a Gln(Q) codon restores a functional GFP coding sequence. [Diagram 2] Editing efficiency measured by GFP recovery (GFP cells / mCherry cells [%]) determined during iterative testing combining various Cas12a base editor constructs with various gRNA constructs. As described in Figure 1, BE and gRNA constructs are represented by numbers (1-12) and letters (a-h), respectively. [Diagram 3]Figure 3a-b shows the editing efficiency in wheat protoplasts measured by GFP recovery (GFP cells / mCherry cells [%]) determined during iterative testing combining different Cas12a base editor constructs with different gRNA constructs. Figure 3c shows a side-by-side comparison of the key BE-gRNA constructs obtained along the optimization path. Components leading to increased activity are displayed on the right side of the panel. BE and gRNA constructs are represented by numbers (1-12) and letters (a-h), respectively, as described in Figure 1. [Figure 4] Figure 4a is a schematic of the Cas12a base editor (BE) protein and guide RNA (gRNA) constructs tested. Figure 4b is the editing efficiency of maize protoplasts as measured by GFP recovery (GFP cells / mCherry cells [%]) determined during iterative testing of different Cas12a base editor constructs in combination with different gRNA constructs. Components leading to increased activity are indicated on the right side of the panel. [Diagram 5] Base editing efficiency of six Cas12a-ABE configurations in simplex, measured by the ratio of A:T to G:C converted reads in wheat. See Figure 1a and b, where the various BE configurations are labeled with numbers and letters. The bar graphs display the A:T to G:C conversion rates at individual on-target target sites. The x-axis indicates the targeted adenines at various positions along the protospacer, with the base adjacent to the PAM at position 1. Seven separate bar graphs are shown for each targeted adenine, representing from left to right: (i) negative control (no BE), (ii) v6a, (iii) v9a, (iv) v6h, (v) v9h, (vi) v11h, (vii) v12h. Editing rates were calculated from two or three independent biological replicates, shown as points on the bar graphs. The violin plots represent the pooled efficiency at all target sites of the individual base editor constructs. Significance was calculated by Kruskal-Wallis test with Dunn post-hoc test (P<0.05). [Figure 6]Base editing efficiency of six Cas12a-ABE configurations in multiplex, measured by the ratio of A:T to G:C converted reads in wheat. See Figure 1a and b, where the various BE configurations are labeled with numbers and letters. The bar graphs display the A:T to G:C conversion rates at individual on-target target sites. The x-axis indicates the targeted adenines at various positions along the protospacer, with the base adjacent to the PAM at position 1. Seven separate bar graphs are shown for each targeted adenine, representing from left to right: (i) negative control (no BE), (ii) v6a, (iii) v9a, (iv) v6h, (v) v9h, (vi) v11h, (vii) v12h. Editing rates were calculated from three independent biological replicates, shown as points on the bar graphs. The violin plots represent the pooled efficiency at all target sites of the individual base editor constructs. Significance was calculated by Kruskal-Wallis test with Dunn post-hoc test (P<0.05). [Figure 7] Figure 7a-b: Frequency of T0 wheat plants with A:T to G:C base edits at independent positions of the target site measured by NGS (n>153) for each of the two base editors. Heterozygous mutations (25%-75% editing rate) and homozygous mutations (>75%) are displayed. Two Cas12a-ABEs (v9h and v11h) are compared with TS60-A (a) and TS112-A (b). Figure 7c-d: Frequency of individual T0 wheat plants with heterozygous or homozygous A:T to G:C conversions at individual positions (n>153) of the target site. Frequencies for heterozygous (HZ) and homozygous (HM) T0 plants are shown in d) and combined frequencies in c). TaTS60 and TaTS112 values for the three subgenomes are shown. Asterisks in (a–d) indicate significant differences in efficiency between v9h and v11h, as measured by Z-score test of two ratios (*: p<0.05; **: p<0.01; ***: p<0.001). [Figure 8]Figure 8a is the frequency of genotypes generated by Cas12a-ABE v9h and v11h. A>G indicates base editing measured at the target site position. Aa and aa indicate heterozygous and homozygous base editing on subgenome A, respectively. Figure 8b is the percentage of T0 wheat plants with at least one mutation in one of the subgenome TS60 and / or TS112 loci. Asterisks indicate significant differences in editing efficiency between v9h and v11h as measured by two-ratio Z-score test (*: p<0.05; **: p<0.01; ***: p<0.001). [Figure 9] Comparison of editing rates measured by ddPCR drop-off or NGS A:T to G:C conversion rates for TaTS60-A (a) or TaTS112-A (b) in individual TO wheat plants (n>54) for two base editors (v9h and v11h). The gradient grey scale represents editing efficiency. [Figure 10] Base editing efficiency of six Cas12a-ABE constructs in multiplex as measured by the ratio of A:T to G:C converted reads in maize. See Figure 1a and b, where the various BE constructs are labeled with numbers and letters. The bar graphs display the A:T to G:C conversion rates at individual on-target target sites. The x-axis indicates the targeted adenines at various positions along the protospacer, with the base adjacent to the PAM at position 1. Seven separate bar graphs are shown for each targeted adenine, representing from left to right: (i) negative control (no BE), (ii) v6a, (iii) v9a, (iv) v6h, (v) v9h, (vi) v11h, (vii) v12h. Editing rates were calculated from three independent biological replicates shown as points on the bar graphs. The violin plots represent the pooled efficiency at all target sites of the individual base editor constructs. Significance in (a–c) was calculated by Kruskal-Wallis test with Dunn post hoc test (P<0.05). [Figure 11]Figure 11a-c: Frequency of T0 corn plants with A:T to G:C base edits at independent positions of target sites Zm-TS3 (a), Zm-TS4 (b) and Zm-TS8 (c) as measured by Sanger sequencing for base editors v9h (n=25), v11h (n=27) and v12h (n=21). Heterozygous mutations (editing rates 25%-75%) and homozygous mutations (>75%) are displayed. Figure 11d-e: Frequency of individual T0 corn plants with heterozygous or homozygous A:T to G:C conversions at individual sites of target sites Zm-TS1, Zm-TS3, Zm-TS4 and Zm-TS8. Activity of three base editors: v9h (n=25), v11h (n=27) and v12h (n=21) are compared. Frequencies for heterozygous (HZ) and homozygous (HM) TO plants are shown in e) and combined frequencies in d). Asterisks in (c-e) indicate significant differences in efficiency between v12h and v9h or v11h as measured by Z-score test of two ratios (*: p<0.05). [Figure 12] Distribution of plants (wheat) with A to G edits in the T1 generation without the Cas12-ABE transgene. Twelve independent T1 lines were analyzed. WT: wild type; HZ: heterozygous; HM: homozygous. [Figure 13a] ABE activity of TadA9>(GGGSS)6×>dLbCas12a-D156R in canola and soybean protoplasts. a) GFP fluorescence from the defective GFP gene in canola protoplasts reflecting editing of the TAG stop codon to a functional CAG codon compared to the fluorescence from a functional GFP gene (GFP control). b) Levels of ABE activity at the endogenous target site measured by NGS (canola protoplasts) or ddPCR (soybean protoplasts). [Figure 13b]ABE activity of TadA9>(GGGSS)6×>dLbCas12a-D156R in canola and soybean protoplasts. a) GFP fluorescence from the defective GFP gene in canola protoplasts reflecting editing of the TAG stop codon to a functional CAG codon compared to the fluorescence from a functional GFP gene (GFP control). b) Levels of ABE activity at the endogenous target site measured by NGS (canola protoplasts) or ddPCR (soybean protoplasts). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] A brief description of the sequence
[0017] [Table 1]
[0018] [Table 2]
[0019] [Table 3]
[0020] [Table 4]
[0021] [Table 5]
[0022] [Table 6]
[0023] [Table 7]
[0024] [Table 8]
[0025] [Table 9]
[0026] [Table 10]
[0027] [Table 11]
[0028] [Table 12]
[0029] "Identity" and / or "homology", when used in reference to a comparison of two or more nucleic acid or amino acid molecules, means that the sequences of the molecules share a certain degree of sequence similarity and that the sequences are partially identical.
[0030] Enzyme variants can be defined by their sequence identity when compared to the parent enzyme. Sequence identity is usually indicated as "% sequence identity" or "% identity". In a first step, to determine the percent identity between two amino acid sequences, a pairwise sequence alignment is made between the two sequences, and the two sequences are aligned over their entire length (i.e., pairwise global alignment). The alignment is made using a program that implements the Needleman and Wunsch algorithm (J. Mol. Biol. (1979) 48, p. 443-453), preferably "NEEDLE" (European Molecular Biology Open Software Suite (EMBOSS)) with the program default parameters (gap open=10.0, gap extension=0.5 and matrix=EBLOSUM62). The preferred alignment for the purposes of the present invention is the alignment that allows the determination of maximum sequence identity.
[0031] The following example is intended to illustrate two nucleotide sequences, but the same calculations apply to protein sequences. Sequence A: AAGATACTG; length: 9 bases Sequence B: GATCTGA; length: 7 bases
[0032] Therefore, the shorter sequence is sequence B.
[0033] Producing a pairwise global alignment showing both sequences over their full length gives: Sequence A: AAGATACTG- Array B:--GAT-CTGA
[0034] The symbol "I" in the alignment indicates identical residues (meaning bases in the case of DNA and amino acids in the case of proteins). The number of identical residues is 6.
[0035] The symbol "-" in the alignment indicates a gap. The number of gaps introduced by the alignment in sequence B is 1. The number of gaps introduced by the alignment at the edge of sequence B is 2 and at the edge of sequence A is 1.
[0036] The alignment length, showing sequences aligned over the full length, is 10.
[0037] According to the present invention, when a pairwise alignment is generated showing a shorter sequence over its entire length, the result is: Sequence A: GATACTG- Array B: GAT-CTGA
[0038] According to the present invention, generating a pairwise alignment showing sequence A over its entire length results in: Sequence A: AAGATACTG Sequence B: --GAT-CTG
[0039] According to the present invention, generating a pairwise alignment showing sequence B over its entire length results in: Sequence A: GATACTG- Array B: GAT-CTGA
[0040] The alignment length, which shows the shorter sequence over its entire length, is 8 (there is one gap included in the alignment length of the shorter sequence).
[0041] Therefore, the alignment length showing sequence A over its entire length is 9 (meaning that sequence A is a sequence of the invention).
[0042] Therefore, the alignment length showing sequence B over its entire length is 8 (meaning that sequence B is a sequence of the invention).
[0043] After aligning the two sequences, in a second step, an identity value is determined from the resulting alignment. For the purposes of this description, the percent identity is calculated by %-identity=(identical residues / length of the alignment region showing each of the sequences of the invention over its entire length)×100. Thus, the sequence identity related to the comparison of two amino acid sequences according to this embodiment is calculated by dividing the number of identical residues by the length of the alignment region showing each of the sequences of the invention over its entire length. Multiplying this value by 100 gives the "% identity". According to the example given above, the % identity is (6 / 9)*100=66.7% when sequence A is the sequence of the invention, and (6 / 8)*100=75% when sequence B is the sequence of the invention.
[0044] Indel is a term for random insertion or deletion of bases in the genome of an organism associated with the repair of DSBs by NHEJ. It is classified among small genetic variations of 1-10000 base pairs in length. As used herein, indel refers to random insertion or deletion of bases in or immediately adjacent to a target site (e.g., less than 1000bp, 900bp, 800bp, 700bp, 600bp, 500bp, 400bp, 300bp, 250bp, 200bp, 150bp, 100bp, 50bp, 40bp, 30bp, 25bp, 20bp, 15bp, 10bp or 5bp upstream and / or downstream).
[0045] In a first aspect, the present invention provides a method for the preparation of a cascade of polypeptides comprising, in order, the following structural elements: a.) at least one N-terminal NLS sequence, b.) an adenosine deaminase domain selected from the TadA9 domain and the TadA8 domain, preferably the TadA9 domain, or a functional variant thereof of said domains, c.) at least one linker domain, d.) dCas12a or a functional fragment thereof or aCas12a or a functional fragment thereof, comprising at least one or more mutations, said at least one or more mutations resulting in increased activity and / or enhanced temperature tolerance, preferably at least one of said mutations being D156 of SEQ ID NO: 14, 15 or 16, E174 of SEQ ID NO: 17, 18 or 19 and E175 of SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27 or 28, respectively. 184, wherein the at least one mutation results in increased activity and / or enhanced temperature tolerance, and in particular the at least one mutation in the dCas12a orthologue or homologue corresponds to a D to R, E to R, or K to D / E mutation at a homologous position in any one of the deadCas12a variants of SEQ ID NOs: 14 to 43 as reference sequences, respectively; e.) an adenine base editor (ABE) that may include at least one C-terminal NLS sequence, wherein the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence may be the same or different.
[0046] Those skilled in the art are familiar with the nomenclature and structure of TadA molecules and their classification (see Gaudelli et al., 2017 supra; Gaudelli et al., 2020; https: / / doi.org / 10.1101 / 2020.03.13.990630). For example, TadA8e is known to be derived from TadA-7.10 (see SEQ ID NO: 114) by introducing eight amino acid changes. TadA8.20 (see SEQ ID NO: 115) is also derived from TadA-7.10, but contains only five amino acid changes that differ from those of TadA8e. TadA9 was derived from TadA8e, for example, by introducing two amino acid mutations (V82S and Q154R) from TadA8.20. As used herein, a particular class of TadA, e.g., TadA8e, TadA9 or TadA-7.10, refers to a molecule derived from Escherichia coli TadA and having the characteristic mutations (also called signature mutations) of the respective TadA subclass. Still, as the skilled artisan will recognize, there may be certain additional mutations, insertions or deletions at positions other than those that characterize the class, e.g., truncated N- or C-termini, mutations at sites different from the characteristic positions of the class, etc. Such variants having at least 80%, at least 85%, at least 90%, preferably at least 95% sequence identity at the amino acid level with the corresponding TadA molecule are also considered to fall into the same class. For example, a TadA9 molecule having all the characteristic positions of the class as the TadA9 sequence of SEQ ID NO: 58, but with certain variations (e.g., 4%), would still be considered a TadA9 molecule, so long as it has the overall deaminase function of TadA9 and the characteristic positions of the class as described in the art, e.g., as shown in SEQ ID NO: 117. For example, TadA8e (e.g., SEQ ID NO: 57, 116) or TadA9 (e.g., SEQ ID NO: 58) have signature mutations at positions 81 and 153, respectively, which allow one of skill in the art to identify the TadA class. Furthermore, there may be additional mutations that affect properties of TadA other than its deaminase function.For example, in one embodiment, TadA, including TadA8e and TadA9, may comprise the mutation V105W at position 105 according to SEQ ID NOs: 57 and 58 to reduce off-target activity and / or N107Q / S according to SEQ ID NOs: 57 and 58 to further reduce cytosine deaminase activity (Jeong et al., 2021, Nature Biotechnology, https: / / doi.org / 10.1038 / s41587-021-00943-2). Furthermore, in a further embodiment, TadA, including TadA8e and TadA9, may comprise the mutation F147A at position 147 according to SEQ ID NOs: 57 and 58 to narrow the editing range (cf. Li et al., 2023, https: / / doi.org / 10.1016 / j.omtn.2022.12.001). Even with these additional mutations, one of skill in the art would still recognize TadA8e and TadA9 as belonging to the TadA8e and TadA9 classes, respectively. Based on the above, a "functional variant" or "functional fragment" in the context of TadA or in the context of any dCas12, nCas12a or ABE disclosed and claimed herein refers to a TadA, dCas12a, nCas12a or ABE that has the same class-characterizing (or signature) positions as the TadA from which it is derived, but the functional variant may be a shorter variant, e.g. a truncated variant that still includes the relevant catalytic active site and the class-characterizing positions, or in another embodiment or aspect, for example, a functional variant may be a molecule that has high (>80%, preferably at least 90%, more preferably at least 95%) sequence identity at the amino acid level with the TadA molecule from which it is derived and that includes the particular mutation, but the variant still includes the class-characterizing positions.
[0047] The terms "protein," "polypeptide," and "amino acid sequence" are used interchangeably herein, e.g., with reference to adenine base editors.
[0048] The terms "adenine" and "adenosine" are used interchangeably herein, e.g., in reference to nucleic acid or base editors.
[0049] The terms "cytosine" and "cytidine" are used interchangeably herein, e.g., in reference to nucleic acid or base editors.
[0050] The term "in order" as used herein in relation to a polypeptide / protein describes that each (sub)element (also referred to herein as a domain, portion or (sub)part) is present throughout the polypeptide / protein in a specified order from the N-terminus to the C-terminus of the amino acid sequence that makes up the polypeptide / protein. The term "in order" also means that any additional intervening sequences, linkers, etc. may be present between the portions that are present in the given order. When applied to "in order", it means in the direction from the 5' to the 3' end of each nucleic acid sequence.
[0051] The term "structural element" as used herein, for example in the context of a protein or an adenine base editor, refers to a region of a polypeptide chain of a protein that represents a distinct functional entity.
[0052] The term "NLS sequence" as used herein describes a nuclear localization signal, which is a part of a protein that facilitates the transport of the respective protein to the cell nucleus by nuclear transport. Typical features of nuclear localization signals (e.g., the presence of positively charged amino acids such as lysine and arginine) are known to those skilled in the art. The mechanism of nuclear transport is also known to those skilled in the art.
[0053] The term "increased activity and / or enhanced temperature tolerance" as used herein, i.e. in relation to adenine base editors (ABEs), describes an increased enzymatic activity and / or enhanced temperature tolerance in an active Cas12a, which can be induced by at least one or more mutations in the coding sequence of the active Cas12a, which at least one or more mutations in the coding sequence cause at least one or more amino acid exchanges in the amino acid sequence of the active Cas12a. If the ABE comprises a Cas12, or dCas12a, or nCas12, having at least one or more mutations that result in an increased activity and / or enhanced temperature tolerance as described above, then the increased activity and / or enhanced temperature tolerance can be directly transferred to the ABE.
[0054] An adenine base editor according to the invention may comprise at least one N-terminal NLS sequence, preferably one N-terminal NLS sequence that corresponds to SEQ ID NO:52 or is a triple SV40 NLS sequence (3xSV40) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:52.
[0055] In one embodiment, an adenine base editor according to the invention may comprise at least one N-terminal NLS sequence, preferably one N-terminal NLS sequence that corresponds to SEQ ID NO:53 or is a bipartite SV40 NLS sequence (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:53.
[0056] In one embodiment, an adenine base editor according to the invention may comprise at least one N-terminal NLS sequence, preferably one which is the SV40 NLS sequence (SV40) corresponding to SEQ ID NO:54 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:54.
[0057] In one embodiment, an adenine base editor according to the invention may comprise at least one N-terminal NLS sequence, preferably one N-terminal NLS sequence that is a Flag-tagged SV40 nuclear localization signal sequence corresponding to SEQ ID NO:55 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% identity to SEQ ID NO:55.
[0058] In one embodiment, an adenine base editor according to the invention may comprise at least one N-terminal NLS sequence, preferably one N-terminal NLS sequence which is the nuclear localization signal, nucNLS, of the nucleoplasmin gene of Xenopus laevis corresponding to SEQ ID NO:56 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:56.
[0059] In yet another embodiment, an adenine base editor according to the invention comprises an SV40 NLS sequence (BP) corresponding to SEQ ID NO:52, or a triple SV40 NLS sequence (3xSV40) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:52 and SEQ ID NO:53, or a bipartite SV40 NLS sequence (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:53 and SEQ ID NO:54, or an SV40 NLS sequence (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:54. NLS sequence (SV40) and a Flag-tagged SV40 nuclear localization signal sequence corresponding to SEQ ID NO:55 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% identity to SEQ ID NO:55 and nucNLS, the nuclear localization signal of the nucleoplasmin gene of Xenopus laevis corresponding to SEQ ID NO:56 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:56, or a combination thereof.
[0060] An adenine base editor according to the invention may comprise at least one C-terminal NLS sequence, preferably one C-terminal NLS sequence that corresponds to SEQ ID NO:52 or is a triple SV40 NLS sequence (3xSV40) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:52.
[0061] In one embodiment, an adenine base editor according to the invention may comprise at least one C-terminal NLS sequence, preferably one C-terminal NLS sequence that corresponds to SEQ ID NO:53 or is a bipartite SV40 NLS sequence (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:53.
[0062] In one embodiment, an adenine base editor according to the invention may comprise at least one C-terminal NLS sequence, preferably one C-terminal NLS sequence that corresponds to SEQ ID NO:54 or has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:54 (SV40).
[0063] In one embodiment, an adenine base editor according to the invention may comprise at least one C-terminal NLS sequence, preferably one C-terminal NLS sequence that is a Flag-tagged SV40 nuclear localization signal sequence corresponding to SEQ ID NO:55 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% identity to SEQ ID NO:55.
[0064] In one embodiment, an adenine base editor according to the invention may comprise at least one C-terminal NLS sequence, preferably one C-terminal NLS sequence which is the nuclear localization signal, nucNLS, of the nucleoplasmin gene of Xenopus laevis corresponding to SEQ ID NO:56 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:56.
[0065] In one embodiment, an adenine base editor according to the invention comprises an SV40 NLS sequence (BP) corresponding to SEQ ID NO:52, or a triple SV40 NLS sequence (3xSV40) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:52 and SEQ ID NO:53, or a bipartite SV40 NLS sequence (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:53 and SEQ ID NO:54, or an SV40 NLS sequence (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:54. NLS sequence (SV40) and a Flag-tagged SV40 nuclear localization signal sequence corresponding to SEQ ID NO:55 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% identity to SEQ ID NO:55 and nucNLS, the nuclear localization signal of the nucleoplasmin gene of Xenopus laevis corresponding to SEQ ID NO:56 or having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:56, or a combination thereof.
[0066] Particularly preferred, in certain embodiments, an adenine base editor according to the invention may comprise one or more N-terminal NLS sequences and one or more C-terminal NLS sequences.
[0067] Particularly preferably, an adenine base editor according to the invention may comprise one or more N-terminal NLS sequences and one or more C-terminal NLS sequences that correspond to SEQ ID NO:53 or are bipartite SV40 NLS sequences (BP) having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:53.
[0068] The term "domain" as used herein describes a region of a polypeptide chain of a protein that is self-stabilizing and preferably folds independently from the rest of the protein.
[0069] As used herein, the term "adenosine deaminase domain" describes a portion of a protein and / or fusion protein that promotes the deamination of adenosine to inosine by replacement of the amino group with a keto group catalyzed by the respective protein and / or fusion protein.
[0070] Suitable adenosine deaminase domains are disclosed herein or known to those of skill in the art (Huang et al., 2021, Nature Protocols, 16, 1089-1128; doi:10.1038 / s41596-020-00450-9).
[0071] An adenine base editor according to the invention may comprise an adenosine deaminase domain that is a TadA8 domain, preferably a TadA8e domain corresponding to SEQ ID NO:57 or to a sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:57.
[0072] In one embodiment, an adenine base editor according to the invention may comprise an adenosine deaminase domain which is a TadA9 domain, preferably a TadA9 domain corresponding to SEQ ID NO:58 or to a sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:58.
[0073] The term "linker domain" as used herein describes a part of a fusion protein that connects two functional domains of the respective fusion protein, thus facilitating the prevention of undesirable effects such as misfolding of the respective fusion protein. In particular, the linker can ensure proper spacing between different structural elements or entities so that each can properly exert its function within the fusion.
[0074] An adenine base editor according to the invention may comprise at least one linker domain, preferably an XTEN 32aa linker domain, especially preferably an XTEN 32aa linker domain corresponding to SEQ ID NO:48 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NO:48.
[0075] In one embodiment, an adenine base editor according to the invention may comprise at least one linker domain, preferably an XTEN 48aa linker domain, especially preferably an XTEN 48aa linker domain corresponding to SEQ ID NO:49 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NO:49.
[0076] In one embodiment, an adenine base editor according to the invention may comprise one or more linker domains selected from the group consisting of sequences corresponding to SEQ ID NO: 48, SEQ ID NO: 49, SEQ ID NO: 50 or SEQ ID NO: 51, or one or more linker domains corresponding to a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NOs: 48, 49, 50 or 51.
[0077] In one embodiment, an adenine base editor according to the present invention may comprise at least one linker domain, preferably a GGGGS linker domain corresponding to SEQ ID NO:50.
[0078] Particularly preferably, an adenine base editor according to the present invention may comprise at least one linker domain, preferably a hexaGGGGS linker domain corresponding to SEQ ID NO:51.
[0079] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0080] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, a hexaGGGGS linker corresponding to SEQ ID NO:51, a dCas12a according to SEQ ID NO:29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0081] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 30, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0082] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, a hexaGGGGS linker corresponding to SEQ ID NO:51, a dCas12a according to SEQ ID NO:31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0083] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0084] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0085] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:30, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0086] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0087] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0088] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, a hexaGGGGS linker corresponding to SEQ ID NO:51, a dCas12a according to SEQ ID NO:29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0089] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 30, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0090] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, a hexaGGGGS linker corresponding to SEQ ID NO:51, a dCas12a according to SEQ ID NO:31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0091] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0092] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0093] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:30, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0094] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0095] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0096] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0097] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0098] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0099] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0100] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to SEQ ID NO: 29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0101] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0102] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0103] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0104] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0105] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0106] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a hexaGGGGS linker corresponding to SEQ ID NO: 51, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0107] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0108] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO:52.
[0109] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 32aa linker corresponding to SEQ ID NO: 48, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0110] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:523, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 32aa linker corresponding to SEQ ID NO:48, a dCas12a according to SEQ ID NO:31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO:52.
[0111] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0112] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0113] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 30 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0114] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0115] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0116] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0117] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:30, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0118] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA9 domain corresponding to SEQ ID NO:58, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0119] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0120] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0121] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 30 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0122] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0123] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53.
[0124] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:29 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0125] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:30, and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0126] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:31 and a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53.
[0127] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0128] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0129] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0130] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0131] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0132] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to SEQ ID NO: 29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0133] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0134] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA9 domain corresponding to SEQ ID NO: 58, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0135] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0136] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0137] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0138] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, a GGGGS linker corresponding to SEQ ID NO: 50, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0139] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to any one of SEQ ID NOs: 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, or 43, and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0140] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO:53, a TadA8e domain corresponding to SEQ ID NO:57, an XTEN 48aa linker corresponding to SEQ ID NO:49, a dCas12a according to SEQ ID NO:29 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO:52.
[0141] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to SEQ ID NO: 30 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0142] In one embodiment, an adenine base editor according to the invention may comprise, in order, a bipartite SV40 nuclear localization signal corresponding to SEQ ID NO: 53, a TadA8e domain corresponding to SEQ ID NO: 57, an XTEN 48aa linker corresponding to SEQ ID NO: 49, a dCas12a according to SEQ ID NO: 31 and a triple SV40 nuclear localization signal corresponding to SEQ ID NO: 52.
[0143] The term "dCas12a" as used herein describes any mutant of an orthologue of Cas12a that has at least one mutation that at least significantly reduces or eliminates the DNase activity of the corresponding wild-type Cas12a enzyme. Such DNase-dead mutants of Cas12a nuclease (dead Cas12a) or functional fragments thereof may contain one, two, three or more mutations, particularly preferably one or two mutations, that at least reduce or even eliminate their respective DNase activity. The preferably one, two, three or more mutations, particularly preferably one or two mutations that render the respective nuclease activity non-functional, are located in the nuclease active site, for example the RuvC site of the respective Cas12a nuclease or functional fragment thereof.
[0144] The term "functional fragment" as used herein defines a subdomain of the enzyme or protein used, particularly Cas12a or TadA, that is capable of folding and exerting at least one function of the full-length protein from which it is derived, but only includes at least one functional domain or fragment of the full-length protein. Functional fragments may also include N- or C-terminally truncated versions of the corresponding full-length protein. In either case, the functional fragment is smaller than the corresponding full-length protein, and therefore has lower steric requirements. In the context of ABEs, as they represent multi-domain proteins, the term "functional fragment" or "functional variant" refers to an ABE that has substantially the same overall structure with respect to the location of the CRISPR effector and TadA deaminase and at least one linker, but includes certain additional mutations or domains, which, however, do not affect the overall ABE base editor activity measurable by monitoring a given ABE resulting in editing activity at at least one target site of interest.
[0145] As used herein, the term "genome" describes all of the genetic information of an organism, consisting of nucleotide sequences which may exist in the form of deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA).
[0146] As used herein, the term "target site" describes a nucleotide sequence, typically a DNA sequence, at which base editing can be performed using the base editors described herein. Typically, the target site is part of a genome.
[0147] The term "PAM" as used herein describes a short nucleotide sequence, typically a DNA sequence, typically about 2-6 base pairs in length, that is located within or near a given target site. Different types of nucleases recognize and bind to one or more specific PAM sequences. In the case of Cas9 nucleases, Cas-9-mediated DNA cleavage occurs at the PAM-proximal sequence region. In the case of Cas12a nucleases, Cas12a-mediated DNA cleavage occurs at a more distal region with respect to the respective PAM sequence.
[0148] The terms "PAM" and "PAM sequence" are used interchangeably herein.
[0149] Preferably, an adenine base editor according to the present invention may comprise dCas12a or a functional fragment thereof, wherein the RNA processing activity of dCas12a or a functional fragment thereof is unaffected by at least one mutation that significantly reduces or eliminates at least the DNase activity of the corresponding wild-type Cas12a enzyme.
[0150] An adenine base editor according to the invention is dCas12a, or a functional fragment thereof, that has one, two, three or more mutations that render the nuclease activity non-functional, particularly preferably one of the one or two mutations has at least a 75% similarity to any of SEQ ID NOs: 1, 13, 15, 30, 44, 45, 46 or 47, or a sequence corresponding to SEQ ID NOs: 1, 13, 15, 30, 44, 45, 46 or 47, or a functional fragment thereof. , 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to a Cas12a ortholog or homolog or functional fragment thereof at a position homologous to D832.
[0151] In one embodiment, an adenine base editor according to the invention may comprise dCas12a or a functional fragment thereof, wherein the nuclease activity is non-functional, with one, two, three or more mutations, particularly preferably one of the one or two mutations, corresponding to a mutation in a Cas12a orthologue or homologue or a functional fragment thereof at a position homologous to D908 of a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of SEQ ID NOs: 2, 18 or 33 or a sequence corresponding to SEQ ID NOs: 2, 18 or 33 or a functional fragment thereof.
[0152] In one embodiment, an adenine base editor according to the invention is dCas12a, or a functional fragment thereof, comprising one, two, three or more mutations that render the nuclease activity non-functional, particularly preferably one of the one or two mutations at least similar to any of SEQ ID NO: 3, 4, 5, 21, 24, 27, 36, 39 or 42, or a sequence corresponding to SEQ ID NO: 3, 4, 5, 21, 24, 27, 36, 39 or 42, or a functional fragment thereof. In some embodiments, the dCas12a or functional fragment thereof corresponds to a mutation in a Cas12a ortholog or homolog, or a functional fragment thereof, at a position homologous to D917 of a sequence having 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity.
[0153] Preferably, an adenine base editor according to the invention may comprise dCas12a or a functional fragment thereof, which has one, two, three or more mutations that render the nuclease activity non-functional, particularly preferably one of the one or two mutations corresponds to a D to A mutation in a Cas12a orthologue or homologue.
[0154] In one embodiment, an adenine base editor according to the invention is dCas12a, or a functional fragment thereof, comprising one, two, three or more mutations that render the nuclease activity non-functional, particularly preferably one of the one or two mutations at least similar to any of SEQ ID NOs: 1, 13, 14, 29, 44, 45, 46 or 47, or a sequence corresponding to SEQ ID NOs: 1, 13, 14, 29, 44, 45, 46 or 47, or a functional fragment thereof. or at least 99% sequence identity to a sequence having 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to a Cas12a ortholog or homolog, or a functional fragment thereof.
[0155] In one embodiment, an adenine base editor according to the invention may comprise dCas12a or a functional fragment thereof, wherein the nuclease activity is non-functional, with one, two, three or more mutations, particularly preferably one of the one or two mutations, corresponding to a mutation in a Cas12a orthologue or homologue or a functional fragment thereof at a position homologous to E993 of a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of SEQ ID NO: 2, 17 or 32 or a sequence corresponding to SEQ ID NO: 2, 17 or 32 or a functional fragment thereof.
[0156] In one embodiment, an adenine base editor according to the invention is dCas12a, or a functional fragment thereof, comprising one, two, three or more mutations that render the nuclease activity non-functional, particularly preferably one of the one or two mutations at least similar to any of SEQ ID NO: 3, 4, 5, 20, 23, 26, 35, 38 or 41, or a sequence corresponding to SEQ ID NO: 3, 4, 5, 20, 23, 26, 35, 38 or 41, or a functional fragment thereof. In some embodiments, the dCas12a or functional fragment thereof corresponds to a mutation in a Cas12a ortholog or homolog, or a functional fragment thereof, at a position homologous to E1006 of a sequence having 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity.
[0157] In one embodiment, an adenine base editor according to the invention may comprise dCas12a or a functional fragment thereof, which has one, two, three or more mutations that render its nuclease activity non-functional, particularly preferably one of the one or two mutations corresponds to an E to A mutation in a Cas12a orthologue or homologue.
[0158] In another embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, wherein the dCas12a or a functional fragment thereof comprises at least one or more mutations that confer increased activity and / or enhanced temperature tolerance.
[0159] In one embodiment, an adenine base editor according to the invention may comprise dCas12a or a functional fragment thereof, which comprises at least one or more mutations that, when present in Cas12a, result in increased activity and / or enhanced temperature tolerance.
[0160] In one embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, where at least one of the one or more mutations that result in enhanced temperature tolerance corresponds to a mutation in a dCas12a orthologue or homologue or functional fragment thereof at a position homologous to D156 of SEQ ID NO: 1, 13, 14, 15, 44, 45, 46, or 47, or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% sequence identity to any of SEQ ID NOs: 1, 13, 14, 15, 44, 45, 46, or 47.
[0161] In one embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, where one of the at least one or more mutations that result in enhanced temperature tolerance corresponds to a mutation in a dCas12a orthologue or homologue, or a functional fragment thereof, at a position homologous to E174 of a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% sequence identity to any of SEQ ID NOs: 2, 17, or 18, or a sequence corresponding to SEQ ID NOs: 2, 17, or 18.
[0162] In one embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, where at least one of the one or more mutations that result in enhanced temperature tolerance corresponds to a mutation in a dCas12a orthologue or homologue or functional fragment thereof at a position homologous to E184 of any of SEQ ID NO:3, 4, 5, 20, 21, 23, 24, 26, or 27, or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% sequence identity to any of SEQ ID NOs:3, 4, 5, 20, 21, 23, 24, 26, or 27.
[0163] In one embodiment, an adenine base editor according to the invention may comprise dCas12a or a functional fragment thereof, wherein one of the at least one or more mutations that result in increased activity and / or enhanced temperature tolerance corresponds to a D to R mutation.
[0164] In one embodiment, an adenine base editor according to the invention may comprise nCas12a, or a functional fragment thereof, comprising at least one or more mutations that, when present in Cas12a, result in increased activity and / or enhanced temperature tolerance.
[0165] In one embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, or nCas12a or a functional fragment thereof, comprising at least one or more mutations that result in increased activity and / or enhanced temperature tolerance, wherein one of the at least one or more mutations corresponds to a mutation in a dCas12a ortholog or homolog at a position homologous to D156 of SEQ ID NO: 14, 15, or 16, and wherein at least one mutation in the dCas12a ortholog or homolog corresponds to a D to R mutation at the homologous position.
[0166] In another embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, wherein one of the at least one or more mutations that result in increased activity and / or enhanced temperature tolerance is an E to R mutation. In one embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, or nCas12a or a functional fragment thereof, comprising at least one or more mutations that result in increased activity and / or enhanced temperature tolerance, wherein at least one of the at least one or more mutations corresponds to a mutation in a dCas12a orthologue or homolog at a position homologous to E174 of SEQ ID NO: 17, 18, or 19, and wherein at least one mutation in the dCas12a orthologue or homolog corresponds to an E to R mutation at the homologous position.
[0167] In one embodiment, an adenine base editor according to the invention can comprise dCas12a or a functional fragment thereof, or nCas12a or a functional fragment thereof, comprising at least one or more mutations that result in increased activity and / or enhanced temperature tolerance, wherein one of the at least one or more mutations corresponds to a mutation in a dCas12a orthologue or homolog at a position homologous to E184 of SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27, or 28, and wherein at least one mutation in the dCas12a orthologue or homolog corresponds to an E to R mutation at the homologous position.
[0168] In yet another embodiment, an adenine base editor according to the present invention may comprise dCas12a or a functional fragment thereof, wherein at least one of the one or more mutations that result in enhanced temperature tolerance is a K to D / E mutation in direct comparison with any one of the deadCas12a variants of SEQ ID NOs: 14-43, respectively, as a reference sequence.
[0169] In one embodiment, an adenine base editor according to the invention comprises dCas12a or a functional fragment thereof having one or more mutations, preferably 1 to 5 mutations, particularly preferably 2 to 4 mutations, conferring altered PAM specificity compared to the respective wild-type Cas12a nuclease or functional fragment thereof, preferably wherein the altered PAM specificity results in the recognition of one or more PAM sequences selected from the group consisting of TYCV, TATV, TACV, CTCV, CCCV, TTYN, VTTV and TRTV.
[0170] In one embodiment, one or more mutations, preferably one to five mutations, particularly preferably one of two to four mutations, that confer altered PAM specificity compared to the respective wild-type Cas12a nuclease or functional fragment thereof, is a mutation in a dCas12a orthologue or homologue or a functional fragment thereof at a position homologous to G532 of SEQ ID NO: 1, 13, 14, 15, 29 or 30, or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NO: 1, 13, 14, 15, 29 or 30.
[0171] In one embodiment, one or more mutations, preferably one to five mutations, particularly preferably one of two to four mutations, that confer altered PAM specificity compared to the respective wild-type Cas12a nuclease or functional fragment thereof correspond to a mutation in a dCas12a orthologue or homologue or a functional fragment thereof at a position homologous to K595 of SEQ ID NO: 1, 13, 14, 15, 29 or 30 or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of SEQ ID NOs: 1, 13, 14, 15, 29 or 30.
[0172] In one embodiment, one or more mutations, preferably one to five mutations, particularly preferably one of two to four mutations, that confer altered PAM specificity compared to the respective wild-type Cas12a nuclease or functional fragment thereof correspond to a mutation in a dCas12a orthologue or homologue or a functional fragment thereof at a position homologous to K583 of SEQ ID NO: 1, 13, 14, 15, 29 or 30 or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of SEQ ID NOs: 1, 13, 14, 15, 29 or 30 corresponding thereto.
[0173] In one embodiment, one or more mutations, preferably one to five mutations, particularly preferably one of two to four mutations, that confer altered PAM specificity compared to the respective wild-type Cas12a nuclease or functional fragment thereof, correspond to a mutation in a dCas12a orthologue or homologue or a functional fragment thereof at a position homologous to Y542 of SEQ ID NO: 1, 13, 14, 15, 29 or 30, or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NO: 1, 13, 14, 15, 29 or 30.
[0174] Particularly preferably, each of the one or more mutations, preferably 1 to 5 mutations, especially preferably 2 to 4 mutations, that confer an altered PAM specificity compared to the respective wild-type Cas12a nuclease or a functional fragment thereof, is selected from the group consisting of G to R, K to R, K to V and Y to R mutations.
[0175] Particularly preferably, the one or more mutations, preferably 1 to 5 mutations, especially preferably 2 to 4 mutations, conferring an altered PAM specificity compared to the respective wild-type Cas12a nuclease or a functional fragment thereof, are individually selected from the group consisting of G532R, K595R, K538R and Y542R, for any of the sequences according to SEQ ID NO: 1, 13, 14, 15, 29 or 30 or sequences having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NO: 1, 13, 14, 15, 29 or 30.
[0176] As used herein, the term "nCas12a" describes a mutant of Cas12a or a functional fragment thereof that exhibits nickase activity and is therefore preferably capable of inducing a single-strand break (a "nick") with comparable or the same specificity as the respective wild-type Cas12a nuclease inducing a double-strand break.
[0177] In the art, nCas12a variants have been described (e.g., WO 2017 / 127807, WO 2019 / 233990A1, and WO 2018 / 176009). However, to date, no nCas12a has been reported that has in vivo functionality. Therefore, further developments are expected. In view of the fact that the nCas9 nickase, which is much easier to make and use given the individual nuclease domains of wild-type Cas9, is highly suitable as a CBE and ABE element, any nCas12a can be used as part of the ABE disclosed herein in place of dCas12a in a similar manner.
[0178] In one embodiment, an adenine base editor (ABE) may comprise, in order, the following structural elements: a.) at least one N-terminal NLS sequence, b.) an adenosine deaminase domain selected from a TadA8 or TadA9 domain or a functional variant thereof, c.) at least one linker domain, wherein the at least one linker domain comprises a hexaGGGGS linker according to SEQ ID NO: 51, d.) dCas12a or a functional fragment thereof or nCas12a or a functional fragment thereof, e.) at least one C-terminal NLS sequence, wherein the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence may be the same or different.
[0179] In a preferred embodiment, an adenine base editor (ABE) may comprise, in order, the following structural elements: a.) at least one N-terminal NLS sequence, b.) an adenosine deaminase domain selected from a TadA8 or TadA9 domain or a functional variant thereof, c.) at least one linker domain, wherein the at least one linker domain comprises a hexaGGGGS linker according to SEQ ID NO: 51, d.) dCas12a or a functional fragment thereof or nCas12a or a functional fragment thereof, e.) at least one C-terminal NLS sequence, wherein the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence are identical. In one embodiment, dCas12a or nCas12a or a functional fragment thereof may comprise at least one or more additional mutations as defined above, wherein one of the at least one or more additional mutations that result in enhanced temperature tolerance corresponds to a mutation in a dCas12a orthologue or homologue at position D156 of SEQ ID NO: 14, 15 or 16, or at position E174 of SEQ ID NO: 17, 18 or 19, or at position E184 of SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27 or 28, or at a homologous position in a Cas12a orthologue or homologue, and preferably, one of the at least one or more additional mutations that result in enhanced temperature tolerance corresponds to a mutation in a dCas12a orthologue or homologue at position D156 of SEQ ID NO: 14, 15 or 16 as a reference sequence, or at a homologous position in a Cas12a orthologue or homologue. or one of the at least one or more additional mutations conferring temperature resistance corresponds to E174R, as compared to SEQ ID NO: 17, 18 or 19 as a reference sequence, or at a homologous position within a Cas12a orthologue or homolog; or one of the at least one or more additional mutations conferring temperature resistance corresponds to E184R, as compared to SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27 or 28 as a reference sequence, or at a homologous position within a Cas12a orthologue or homolog; more preferably, the at least one or more additional mutations correspond to at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124, 125, 126, 127 or 130 as a reference sequence, or at a homologous position within a Cas12a orthologue or homolog;or (iii) D156R, D832A and E925A, or at least one or more additional mutations corresponding to: (i) D156R and D832A, or (ii) D156R and E925A, or (iii) D156R, D832A and E925A, compared to a sequence with 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity, or at a homologous position within a Cas12a orthologue or homologue; compared to SEQ ID NO:2 as a reference sequence, or compared to a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to the corresponding reference sequence, or to a relative sequence within a Cas12a orthologue or homologue. At the same positions, (iv) E174R and D908A, or (v) E174R and E993A, or (vi) E174R, D908A, and E993A, or at least one or more additional mutations are at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135%, 136%, 137%, 138%, 139%, 140%, 141%, 142%, 143%, 144%, 145%, 146%, 147%, 148%, 149%, 150%, 151%, 152%, 153%, 154%, 155%, 156%, 157%, 158%, 159%, 160%, 161%, 162%, 163%, 164%, 165%, 166%, 167%, %, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity, or at homologous positions within a Cas12a orthologue or homologue, (viii) E184R and D917A, or (ix) E184R and E1006A, or (x) E184R, D917A, and E1006A.
[0180] In one embodiment, the at least one N-terminal NLS sequence and / or the at least one C-terminal NLS sequence is selected from triple SV40 NLS (SEQ ID NO:52), bipartite SV40 NLS (SEQ ID NO:53), SV40 NLS (SEQ ID NO:54), FNLS (SEQ ID NO:55) or nucNLS (SEQ ID NO:56), preferably the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence are at least one bipartite SV40 NLS (SEQ ID NO:53) or a functional homologue thereof or a sequence having at least 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:53.
[0181] In another embodiment, the adenosine deaminase domain can be a TadA8e domain according to SEQ ID NO:57 or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:57.
[0182] In another embodiment, the adenosine deaminase domain is a TadA9 domain according to SEQ ID NO:58 or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:58.
[0183] In another aspect, the invention relates to a complex comprising an adenine base editor as described herein and a functionally related form of a guide RNA or a sequence encoding a guide RNA, wherein the guide RNA is specific for dCas12a or nCas12a as defined herein, and optionally the guide RNA is expressed from a construct comprising a truncated tRNA at the 5' end and at least one direct repeat structure 5'- and 3'- of a sequence of the spacer RNA or a sequence encoding the spacer RNA.
[0184] The term "complex" as used herein describes an adenine base editor in functional association with at least one guide RNA. It is known to those skilled in the art that the nuclease domain of a given adenine base editor is usually non-covalently and reversibly associated with the respective guide RNA.
[0185] In the present disclosure, the terms guide RNA and crRNA are used interchangeably. Those skilled in the art are aware of the fact that naturally occurring CRISPR nucleases and cognate guide RNAs are compatible with each other. Furthermore, those skilled in the art are aware that different CRISPR / Cas effectors are guided by different types of guide RNAs.
[0186] For example, certain CRISPR nucleases, such as Cas9, use dual heteroduplex guide RNAs (crRNA::tracrRNA), which can be combined as a single guide RNA when used in molecular biology. Other CRISPR nucleases, including class 2 type V CRISPR nuclease Cas12a and its variants, i.e., dCas12a or nCas12a, use a single crRNA RNA as a guide molecule. Guide RNA, as used herein, is a general term that describes any type of RNA that guides a CRISPR nuclease or its variants. Thus, as such, the term guide RNA, when used in connection with Cas12a effectors, refers to any suitable crRNA-based construct that is suitable for interacting with a crRNA or Cas12a variant or an ABE or a fusion protein comprising it to guide it to a target site of interest that includes a suitable PAM.
[0187] Advantageously, a construct or nucleic acid molecule for the expression of a guide RNA comprising at least one direct repeat structure at the 5'- and 3'-position of the spacer RNA sequence or encoding the spacer RNA facilitates accurate 3'-end processing of the guide RNA transcript by the still intact RNA processing activity of dCas12a or a functional fragment thereof and / or nCas12a or a functional fragment thereof, allowing the production of accurate guide RNA molecules.
[0188] The term "spacer RNA" as used herein describes an RNA sequence that is complementary to a particular target region and thus facilitates (i) localization of the respective target region by a complex comprising an adenine base editor as described herein and a functionally associated guide RNA, and (ii) binding of the complex to the respective target region.
[0189] Preferably, the guide RNA can be expressed from a construct comprising a sequence encoding a spacer RNA, and the spacer RNA and the sequence encoding the spacer RNA are 18 to 30 nucleotides long, preferably 20 to 27 nucleotides long, particularly preferably 21 to 25 nucleotides long, and particularly preferably 22 to 24 nucleotides long.
[0190] The guide RNA may be expressed from a construct that includes a T-extended terminator 3' of a direct repeat structure located 3' of a sequence encoding a spacer RNA. Preferably, the T-extended terminator consists of 3 to 15, preferably 4 to 10, and particularly preferably 5 to 8 thymine (T) residues.
[0191] The terms "thymine" and "thymidine" are used interchangeably herein, e.g., in reference to nucleic acid and / or base editors.
[0192] As known to those of skill in the art, Pol III terminates transcription at heterogeneous positions within the T-extension terminator.
[0193] Advantageously, the expression constructs and / or nucleic acids described herein comprising a T-extended terminator located 3'-of a direct repeat structure located 3'-of a spacer RNA coding sequence overcome the drawback of heterogeneous transcription termination by pol III as dCas12a or a functional fragment thereof or nCas12a or a functional fragment thereof, and utilize the still intact RNA processing activity during processing of the guide RNA to cleave the direct repeat structure located 3'-of the spacer RNA coding sequence together with the poly-U tail transcribed from the T-extended terminator.
[0194] The guide RNA can be expressed from a construct comprising at least one polymerase (pol) III promoter. Those skilled in the art are aware of the pol III promoters typically used in the art. Preferably, the at least one pol III promoter is individually selected from the sequence corresponding to SEQ ID NO: 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 or 79 or the sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity with any of the sequences corresponding to SEQ ID NO: 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 or 79.
[0195] Particularly preferably, the guide RNA may be expressed from a construct comprising at least one pol III promoter corresponding to SEQ ID NO:69 or a sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:69.
[0196] In one embodiment, the guide RNA may be encoded by a scaffold provided by any one of SEQ ID NOs: 59, 60 or 61 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity to at least one of the corresponding reference sequences of SEQ ID NOs: 59, 60 or 61, respectively.
[0197] The term "scaffold" as used herein describes a DNA sequence that comprises all the elements required for transcription of a guide RNA in a manner that allows functional binding of a complex comprising an adenine base editor as described herein and the respective guide RNA.
[0198] The positions marked with "n" in any of the sequences corresponding to SEQ ID NO: 59, 60 or 61 represent a variable region encoding the spacer RNA and may be 18 to 30, preferably 20 to 27, particularly preferably 21 to 25, and especially preferably 22 to 24 nucleotides in length, and each position may be any nucleotide individually selected from the group consisting of A, G, C and T.
[0199] In a further aspect, the present invention relates to a nucleic acid molecule encoding an adenine base editor as described herein and / or a nucleic acid molecule encoding a guide RNA as described herein. According to all embodiments relating to a nucleic acid molecule encoding an adenine base editor as described herein and / or a nucleic acid molecule encoding a guide RNA as described herein, each respective nucleic acid molecule may be codon optimized for expression in a particular species of interest. The particular species of interest may be a plant species, a bacterial species, a fungal species, an archaeal species or an animal species.
[0200] The one or more specific prokaryotic species of interest may be Gluconobacter oxydans, Gluconobacter asaii, Achromobacter delmarvae, Achromobacter viscosus, Achromobacter lacticum, Agrobacterium tumefaciens, Agrobacterium tumefaciens, Agrobacterium radiobacter, Alcaligenes faecalis, Arthrobacter citreus, citreus, Arthrobacter tumescens, Arthrobacter paraffineus, Arthrobacter hydrocarboglutamicus, Arthrobacter oxydans, Aureobacterium saperdae, Azotobacter indicus, Brevibacterium ammoniagenes, Brevibacterium divaricatum, Brevibacterium lactofermentum, Brevibacterium flavum flavum, Brevibacterium globosum, Brevibacterium fuscum, Brevibacterium ketoglutamicumketoglutamicum, Brevibacterium helcolum, Brevibacterium pusillum, Brevibacterium testaceum, Brevibacterium roseum, Brevibacterium immariophilium, Brevibacterium linens, Brevibacterium protopharmiae, Corynebacterium acetophilum, Corynebacterium glutamicum, Corynebacterium callunae callunae, Corynebacterium acetoacidophilum, Corynebacterium acetoglutamicum, Enterobacter aerogenes, Erwinia amylovora, Erwinia carotovora, Erwinia herbicola, Erwinia chrysanthemi, Flavobacterium peregrinum, Flavobacterium fucatum, Flavobacterium aurantinum aurantinum, Flavobacterium rhenanum, Flavobacterium sewanense, Flavobacterium brevebreve, Flavobacterium meningosepticum, Micrococcus sp. CCM825, Morganella morganii, Nocardia opaca, Nocardia rugosa, Planococcus eucinatus, Proteus rettgeri, Propionibacterium shermanii, Pseudomonas synxantha, Pseudomonas azotoformans, Pseudomonas fluorescens jluorescens, Pseudomonas ovalis, Pseudomonas stutzeri, Pseudomonas acidovolans, Pseudomonas mucidolens, Pseudomonas testosteroni, Pseudomonas aeruginosa, Rhodococcus erythropolis, Rhodococcus rhodochrous, Rhodococcus sp. ATCC15592, Rhodococcus sp. ATCC19070, Sporosarcina urea ureae, Staphylococcus aureus, Vibrio metschnikovii, Vibrio tyrogenes, Actinomadura madurae, Actinomyces violaceochromogenesviolaceochromogenes, Kitasatosporia parulosa, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces flavelus, Streptomyces griseolus, Streptomyces lividans, Streptomyces olivaceus, Streptomyces tanashiensis, Streptomyces virginiae, Streptomyces antibioticus, Streptomyces cacaoi cacaoi, Streptomyces lavendulae, Streptomyces viridochromogenes, Aeromonas salmonicida, Bacillus pumilus, Bacillus circulans, Bacillus thiaminolyticus, Escherichia freundii, Microbacterium ammoniaphilum, Serratia marcescens, Salmonella typhimurium, Salmonella shotmuleri schottmulleri, Xanthomonas citri, Synechocystis sp., Synechococcus elongatuselongatus, Thermosynechococcus elongatus, Microcystis aeruginosa, Nostoc sp., N. commune, N. sphaericum, Nostoc punctiforme, Spirulina platensis, Lyngbya majuscula, L. lagerheimii, Phormidium tenue, Anabaena sp. and Leptolyngbya sp.
[0201] The one or more particular eukaryotic microbial species of interest may be Saccharomyces cerevisiae, Hansenula species, such as Hansenula polymorpha, Schizosaccharomyces species, such as Schizosaccharomyces pombe, Kluyveromyces species, such as Kluyveromyces lactis and Kluyveromyces marxianus, Yarrowia species, such as Yarrowia lipolytica, Pichia species, such as Pichia methanolica, methanolica, Pichia stipites and Pichia pastoris, Zygosaccharomyces species such as Zygosaccharomyces rouxii and Zygosaccharomyces bailii, Candida species such as Candida boidinii, Candida utilis, Candida freyschussii, Candida glabrata and Candida sonorensis. sonorensis, Schwanniomyces species such as Schwanniomyces occidentalis, Arxula species such as Arxula adeninivorans, Ogataea species such as Ogataea minuta, Klebsiella species such as Klebsiella pneumoniae.pneumonia, Aspergillus species such as Aspergillus niger, and Myceliophthora thermophila.
[0202] The one or more specific plant species of interest may be maple spp., Actinidia spp., Abelmoschus spp., Agave sisalana, Agropyron spp., Agrostis stolonifera, Allium spp., Amaranthus spp., Ammophila arenaria, Ananas comosus, Annona spp., celery (Apium graveolens), Arachis spp., Artocarpus spp., Asparagus spp., and the like. officinalis, Avena spp. (e.g. Avena sativa, Avena fatua, Avena byzantina, Avena fatua var.sativa, Avena hybrida), star fruit (Averrhoa carambola), Bambusa sp., Benincasa hispida, Brazil nut (Bertholletia excelsea), sugar beet (Beta vulgaris), Brassica spp. (e.g. Brassica napus, Brassica rapa subsp. ssp. [canola, rapeseed, turnip rape], Cadaba farinosa, Camellia sinensis, Canna indica, Cannabis sativa, Capsicum spp., Carex elata, Carica papaya, Carissa macrocarpa, Carya spp.), safflower (Carthamus tinctorius), chestnut species (Castanea spp.), kapok (Ceiba pentandra), endive (Cichorium endivia), cinnamon species (Cinnamomum spp.), watermelon (Citrullus lanatus), citrus species (Citrus spp.), coconut species (Cocos spp.), coffee species (Coffea spp.), taro (Colocasia esculenta), cola species (Cola spp.), coriander (Coriandrum sativum), hazel species (Corylus spp.), hawthorn species (Crataegus spp.), saffron (Crocus sativus), pumpkin species (Cucurbita spp.), Cucumis spp., Cynara spp., Daucus carota, Desmodium spp., Dimocarpus longan, Dioscorea spp., Diospyros spp., Echinochloa spp., Elaeis (e.g. Elaeis guineensis, Elaeis oleifera), Eleusine coracana, Eragrostis tef, Erianthus spp., Eriobotrya japonica, Eucalyptus spp. sp.), Pitanga (Eugenia uniflora), Buckwheat (Fagopyrum spp.), Beech (Fagus spp.), Tall fescue (Festuca arundinacea), Fig (Ficus carica), Fortunella spp., Fragaria spp., Ginkgo (Ginkgo biloba), Glycine spp.) (e.g. Glycine max, Soja hispida or Soja max), cotton (Gossypium hirsutum), Helianthus spp. (e.g. Helianthus annuus, Hemerocallis fulva), Hibiscus spp., Hordeum spp. (e.g. Hordeum vulgare), sweet potato (Ipomoea batatas), Juglans spp., lettuce (Lactuca sativa), Lathyrus spp., lentil (Lens culinaris), flax (Linum usitatissimum), litchi (Litchi chinensis, Lotus spp., Luffa acutangula, Lupinus spp., Luzula sylvatica, Lycopersicon spp. (e.g. Lycopersicon esculentum, Lycopersicon lycopersicum, Lycopersicon pyriforme), Macrotyloma spp., Malus spp., Acerola (Malpighia emarginata), Mammea americana, Mangifera indica, Manihot spp. spp.), Sapodilla (Manilkara zapota), Medicago sativa, Melilotus spp., Mentha spp., Miscanthus sinensis, Momordica spp., Morus nigra, Musa spp., Nicotiana spp., Olea spp.), Opuntia spp., Ornithopus spp., Oryza spp. (e.g., Oryza sativa, Oryza latifolia), Panicum miliaceum, Panicum virgatum, Passiflora edulis, Pastinaca sativa, Pennisetum sp., Persea spp., Parsley (Petroselinum crispum), Phalaris arundinacea, Phaseolus spp., Timothy grass (Phleum pratense, date palms (Phoenix spp.), common reeds (Phragmites australis), nightshades (Physalis spp.), pines (Pinus spp.), pistachios (Pistacia vera), peas (Pisum spp.), Poa spp., populus spp., Prosopis spp., cherry blossoms (Prunus spp.), Psidium spp., pomegranates (Punica granatum), pears (Pyrus communis), oaks (Quercus spp.), radishes (Raphanus sativus), and rhubarb (Rheum rhabarbarum, Ribes spp., Ricinus communis, Rubus spp., Saccharum spp., Salix sp., Sambucus spp., Secale cereale, Sesamum spp., Sinapis sp., Solanum spp.) (e.g. potato (Solanum tuberosum), Solanum integrifolium or tomato (Solanum lycopersicum)), sorghum (Sorghum bicolor), spinach (Spinacia spp.), myrtaceae (Syzygium spp.), Tagetes spp., tamarind (Tamarindus indica), cacao (Theobroma cacao), Trifolium spp., gamagrass (Tripsacum dactyloides), Triticosecale rimpaui, wheat (Triticum spp.) (e.g. wheat (Triticum aestivum), durum wheat (Triticum durum), riveted wheat (Triticum turgidum, Triticum hybernum, Triticum macha, Triticum sativum, Triticum monococcum or Triticum vulgare, Tropaeolum minus, Tropaeolum majus, Vaccinium spp., Vicia spp., Vigna spp., Viola odorata, Vitis spp., Zea mays, Zizania palustris and Ziziphus spp.
[0203] In yet another aspect, the present invention also relates to an expression construct or vector comprising a nucleic acid sequence as described herein, wherein the nucleic acid sequence encoding an adenine base editor and / or the nucleic acid sequence encoding a guide RNA are present on the same expression construct or vector or on at least two separate expression constructs or vectors, optionally wherein an expression construct or vector encoding a guide RNA is present, and wherein the guide RNA is expressed from an RNA polymerase III promoter or an RNA polymerase II promoter, preferably wherein the promoter is selected from the U3, U6, H1 and ubiquitin promoter.
[0204] An expression construct or vector comprising a nucleic acid sequence as described herein in which a nucleic acid sequence encoding an adenine base editor and / or a nucleic acid sequence encoding a guide RNA is present may comprise at least one pol III promoter. A person skilled in the art is aware of pol III promoters typically used in the art.
[0205] Preferably, at least one pol III promoter is individually selected from a sequence corresponding to SEQ ID NO: 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 or 79 or a sequence having 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to any of the sequences corresponding to SEQ ID NO: 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 or 79.
[0206] Preferably, an expression construct or vector comprising a nucleic acid sequence as described herein in which a nucleic acid sequence encoding an adenine base editor and / or a nucleic acid sequence encoding a guide RNA is present comprises at least one pol III promoter corresponding to SEQ ID NO:69 or a sequence with 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99% sequence identity to SEQ ID NO:69.
[0207] In a further aspect, the invention also relates to a cell comprising an adenine base editor as described herein, or comprising a nucleic acid sequence encoding a complex as described herein, or comprising an expression construct or vector as described herein.
[0208] In another aspect, the invention provides a method for adenine base editing of a target site in a genome of interest in at least one cell of a prokaryotic or eukaryotic organism, comprising the steps of: (a) providing to the at least one cell a nucleic acid molecule or expression construct encoding at least one adenine base editor or at least one complex described herein or the same as described herein; (b) optionally allowing functional expression and / or assembly of the complexes defined herein into a functionally relevant form; and (c) infecting the genome of interest of the at least one cell with a complex comprising at least one adenine base editor or at least one complex described herein. (d) optionally selecting the at least one modified cell; and (e) obtaining at least one cell that comprises at least one adenine base edit at the target site, wherein the method further relates to processes for modifying the genetic identity of a human germline, the use of human embryos for industrial or commercial purposes, as well as processes for modifying the genetic identity of an animal that may cause suffering to the animal without providing substantial medical benefit to the human or animal, and also to animals resulting from such processes, and further to the treatment of the human or animal body by therapy or surgery.
[0209] Optionally, the method comprises: (f) regenerating at least one population of edited cells, tissue, organ, material, or whole organism from the at least one edited cell. Includes.
[0210] Preferably, the at least one cell may be derived from a plant, algae, yeast or fungal organism, preferably the at least one cell is a plant cell, preferably a cell of Acer spp., Actinidia spp., Abelmoschus spp., Agave sisalana, Agropyron spp., Agrostis stolonifera, Allium spp., Amaranthus spp., Ammophila arenaria, Ananas comosus, Annona spp., Apium graveolens, Arachis spp. spp.), breadfruit spp. (Artocarpus spp.), asparagus (Asparagus officinalis), oats spp. (e.g. Avena sativa, Avena fatua, Avena byzantina, Avena fatua var.sativa, Avena hybrida), star fruit (Averrhoa carambola), bamboo spp. (Bambusa sp.), wax gourd (Benincasa hispida), Brazil nuts (Bertholletia excelsea), sugar beet (Beta vulgaris), Brassica spp. (e.g. Brassica napus), napus, Brassica rapa ssp. (canola, rapeseed, turnip rape), Cadaba farinosa, Camellia sinensis, Canna indica, Cannabis sativa, Capsicum spp.), Carex elata, Papaya (Carica papaya), Carissa macrocarpa, Pecan (Carya spp.), Safflower (Carthamus tinctorius), Chestnut (Castanea spp.), Kapok (Ceiba pentandra), Endive (Cichorium endivia), Cinnamomum spp., Watermelon (Citrullus lanatus), Citrus spp., Cocos spp., Coffea spp., Taro (Colocasia esculenta), Cola spp., Corchorus sp., Coriander (Coriandrum sativum, Corylus spp., Crataegus spp., Crocus sativus, Cucurbita spp., Cucumis spp., Cynara spp., Daucus carota, Desmodium spp., Dimocarpus longan, Dioscorea spp., Diospyros spp., Echinochloa spp., Elaeis spp. (e.g. Elaeis guineensis, Elaeis oleracea), Elaeis oleracea (e.g. ... oleifera), finger millet (Eleusine coracana), teff (Eragrostis tef), Erianthus sp., loquat (Eriobotrya japonica), eucalyptus sp., pitanga (Eugenia uniflora), buckwheat (Fagopyrum spp.), beech (Fagus spp.), fescue (Festuca arundinacea), fig (Ficus carica), fortunella spp.), Fragaria spp., Ginkgo biloba, Glycine spp. (e.g. Glycine max, Soja hispida or Soja max), cotton (Gossypium hirsutum), Helianthus spp. (e.g. Helianthus annuus, Hemerocallis fulva), Hibiscus spp., Hordeum spp. (e.g. Hordeum vulgare), sweet potato (Ipomoea batatas), Juglans spp., lettuce (Lactuca sativa, Lathyrus spp., Lens culinaris, Linum usitatissimum, Litchi chinensis, Lotus spp., Luffa acutangula, Lupinus spp., Luzula sylvatica, Lycopersicon spp. (e.g. Lycopersicon esculentum, Lycopersicon lycopersicum, Lycopersicon pyriforme), Macrotyloma spp., Malus spp., Malpighia emarginata, Mammea americana, Mango (Mangifera indica), Cassava spp., Sapodilla (Manilkara zapota), Medicago sativa, Melilotus spp., Mentha spp., Miscanthus sinensis, Momordica spp.), Morus nigra, Musa spp., Nicotiana spp., Olea spp., Opuntia spp., Ornithopus spp., Oryza spp. (e.g. Oryza sativa, Oryza latifolia), Panicum miliaceum, Panicum virgatum, Passiflora edulis, Pastinaca sativa, Pennisetum spp., Persea spp., Parsley spp. crispum, Reed canary grass (Phalaris arundinacea), Phaseolus spp., Timothy grass (Phleum pratense), Date palm (Phoenix spp.), Common reed (Phragmites australis), Physalis spp., Pinus spp., Pistachio (Pistacia vera), Pea spp., Poa spp., Populus spp., Prosopis spp., Cherry spp., Psidium spp., Pomegranate (Punica granatum), Pear (Pyrus communis, Quercus spp., Raphanus sativus, Rheum rhabarbarum, Ribes spp., Ricinus communis, Rubus spp., Saccharum spp., Salix spp., Sambucus spp., Secale cereale, Sesamum spp., Sinapis spp., Solanum spp.) (e.g. potato (Solanum tuberosum), Solanum integrifolium or tomato (Solanum lycopersicum)), sorghum (Sorghum bicolor), spinach (Spinacia spp.), myrtaceae (Syzygium spp.), Tagetes spp., tamarind (Tamarindus indica), cacao (Theobroma cacao), Trifolium spp., gamagrass (Tripsacum dactyloides), Triticosecale rimpaui, wheat (Triticum spp.) (e.g. wheat (Triticum aestivum), durum wheat (Triticum durum), riveted wheat (Triticum turgidum, Triticum hybernum, Triticum macha, Triticum sativum, Triticum monococcum or Triticum vulgare, Tropaeolum minus, Tropaeolum majus, Vaccinium spp., Vicia spp., Vigna spp., Viola odorata, Vitis spp., Zea mays, Zizania palustris or Ziziphus spp. The plant cell may be derived from a plant cell belonging to the superfamily Viridiplantae, in particular monocotyledonous and dicotyledonous plants, including fodder or forage legumes, ornamental plants, food crops, trees or shrubs, selected from the list including:
[0211] The term "plant" as used herein includes whole plants, ancestors and progeny of plants, and plant parts including seeds, shoots, stems, leaves, roots (including tubers), flowers, and tissues and organs. Further disclosed in relation to plants are plant cells, suspension cultures, callus tissue, embryos, meristematic regions, gametophytes, sporophytes, pollen, and microspores from the plants that can be obtained, analyzed, and processed according to the disclosure provided herein.
[0212] Plants that are particularly useful in the method of the present invention include all plants belonging to the superfamily Viridiplantae, in particular monocotyledons and dicotyledons. In one embodiment, the method of the present invention relates to the use of fodder or forage legumes, ornamental plants, food crops, trees or shrubs. For example, the method of the present invention relates to the use of crop plants, such as the crop plants listed below. The plants include maple spp., Actinidia spp., Abelmoschus spp., Agave sisalana, Agropyron spp., Agrostis stolonifera, Allium spp., Amaranthus spp., Ammophila arenaria, pineapple (Ananas comosus), Annona spp., celery (Apium graveolens), Arachis spp., Artocarpus spp., Asparagus (Asparagus officinalis), Avena spp. spp.) (e.g. Avena sativa, Avena fatua, Avena byzantina, Avena fatua var.sativa, Avena hybrida), star fruit (Averrhoa carambola), Bambusa sp., Benincasa hispida, Brazil nut (Bertholletia excelsea), sugar beet (Beta vulgaris), Brassica spp. (e.g. Brassica napus, Brassica rapa ssp.) [canola, rapeseed, turnip rape]), Cadaba farinosa, Camellia sinensis, Canna indica, Cannabis sativa, Capsicum spp., Carex elata, Carica papaya, Carissa macrocarpa, Carya spp., Carthamus tinctorius, Castanea spp., Kapok (Ceiba pentandra), Cichorium endivia, Cinnamomum spp., Citrullus lanatus, Citrus spp. spp.), Cocos spp., Coffea spp., Taro (Colocasia esculenta), Cola spp., Corchorus spp., Coriandrum sativum, Corylus spp., Crataegus spp., Saffron (Crocus sativus), Cucurbita spp., Cucumis spp., Cynara spp., Daucus carota, Desmodium spp., Longan (Dimocarpus longan), Dioscorea spp. spp.), Diospyros spp., Echinochloa spp., Elaeis spp. (e.g. Elaeis guineensis, Elaeis oleifera), Eleusine coracana, Eragrostis tef, Erianthus spp., Eriobotrya japonica, Eucalyptus spp.), Eugenia uniflora, buckwheat species (Fagopyrum spp.), beech species (Fagus spp.), fescue (Festuca arundinacea), figs (Ficus carica), fortunella species (Fortunella spp.), Fragaria spp., Ginkgo biloba, Glycine spp. (e.g. Glycine max, Soja hispida or Soja max), cotton (Gossypium hirsutum), Helianthus spp. (e.g. Helianthus annuus, Hemerocallis fulva), Hibiscus spp. spp.), Hordeum spp. (e.g. barley (Hordeum vulgare)), sweet potato (Ipomoea batatas), Juglans spp., lettuce (Lactuca sativa), Lathyrus spp., lentil (Lens culinaris), flax (Linum usitatissimum), Litchi (Litchi chinensis), Lotus spp., Luffa acutangula, Lupinus spp., Luzula sylvatica, Lycopersicon spp. (e.g. tomato (Lycopersicon esculentum, Lycopersicon lycopersicum, Lycopersicon pyriforme, Macrotyloma spp., Malus spp., Acerola (Malpighia emarginata), Mammea americana, Mango (Mangifera indica), Manihot spp.), Sapodilla (Manilkara zapota), Medicago sativa, Melilotus spp., Mentha spp., Miscanthus sinensis, Momordica spp., Morus nigra, Musa spp., Nicotiana spp., Olea spp., Opuntia spp., Ornithopus spp., Oryza spp. (e.g. Oryza sativa, Oryza latifolia), Panicum spp. miliaceum, Panicum virgatum, Passiflora edulis, Parsnip (Pastinaca sativa), Pennisetum sp., Persea spp., Parsley (Petroselinum crispum), Reed cane grass (Phalaris arundinacea), Bean (Phaseolus spp.), Timothy grass (Phleum pratense), Date palm (Phoenix spp.), Common reed (Phragmites australis), Nightshade (Physalis spp.), Pine (Pinus spp.), Pistachio (Pistacia vera), Pea (Pisum spp.), Poa spp., Populus spp., Prosopis spp., Prunus spp., Psidium spp., Pomegranate (Punica granatum), Pyrus communis, Quercus spp., Radish (Raphanus sativus), Rheum rhabarbarum, Ribes spp., Castor bean (Ricinus communis), Rubus spp.), sugarcane spp., Salix spp., Sambucus spp., Secale cereale, Sesamum spp., Sinapis spp., Solanum spp. (e.g. potato (Solanum tuberosum), Solanum integrifolium or tomato (Solanum lycopersicum)), sorghum (Sorghum bicolor), spinach spp., myrtle spp., Syzygium spp., Tagetes spp., tamarind (Tamarindus indica), cocoa (Theobroma cacao), cacao, Trifolium spp., Tripsacum dactyloides, Triticosecale rimpaui, Triticum spp. (e.g. Triticum aestivum, Triticum durum, Triticum turgidum, Triticum hybernum, Triticum macha, Triticum sativum, Triticum monococcum or Triticum vulgare), Tropaeolum minus, Tropaeolum majus, Vaccinium spp. spp.), broad bean spp., cowpea spp., sweet violet spp., Vitis spp., Zea mays, wild rice spp., and Ziziphus spp., among others.
[0213] Preferred plants are Abelmoschus spp., Allium spp., celery (Apium graveolens), Asparagus officinalis, Avena spp. (e.g. Avena sativa, Avena fatua, Avena byzantina, Avena fatua var. sativa, Avena hybrida), sugar beet (Beta vulgaris), Brassica spp. (e.g. Brassica napus, Brassica rapa subsp.), spp. (canola, rapeseed, turnip rape)], Capsicum spp., Citrullus lanatus, Cucumis spp., Cynara spp., Daucus carota, Glycine spp. (e.g. Glycine max, Soja hispida or Soja max), Gossypium hirsutum, Helianthus spp. (e.g. Helianthus annuus), Hordeum spp. (e.g. Hordeum vulgare), Lactuca sativa, Medicago sativa, Oryza spp. (e.g. Oryza sativa, Oryza latifolia), Pennisetum spp., Saccharum spp., Secale cereale, Solanum spp.) (e.g. potato (Solanum tuberosum), Solanum integrifolium or tomato (Solanum lycopersicum)), sorghum (Sorghum bicolor), spinach species (Spinacia spp.), wheat species (Triticum spp.) (e.g. wheat (Triticum aestivum), durum (Triticum durum), riveted wheat (Triticum turgidum), Triticum hybernum, Mach wheat (Triticum macha), Triticum sativum, einkorn (Triticum monococcum) or Triticum vulgare), maize (Zea mays).
[0214] Particularly preferred plants are Brassica spp. (e.g. Brassica napus, Brassica rapa spp. (canola, rapeseed, turnip rape)], Capsicum spp., Glycine spp. (e.g. Glycine max, Soja hispida or Soja max), Gossypium hirsutum, Helianthus spp. (e.g. Helianthus annuus), Oryza spp. (e.g. Oryza sativa, Oryza latifolia), latifolia), Solanum spp. (e.g. potato (Solanum tuberosum), Solanum integrifolium or tomato (Solanum lycopersicum)), Triticum spp. (e.g. wheat (Triticum aestivum), durum (Triticum turgidum), hybernum (Triticum hybernum), macha (Triticum sativum), einkorn (Triticum monococcum or Triticum vulgare), and Zea mays.
[0215] In a further aspect, the present invention also relates to edited cells or tissues, organs, materials (e.g. materials from leaves or germ cells or parts of organs or materials from parts of seeds, e.g. in crushed form) or whole organisms obtained or obtainable by the methods described herein. All methods and uses disclosed herein specifically exclude processes for modifying the germline genetic identity of humans, the use of human embryos for industrial or commercial purposes, processes for cloning humans and further processes for modifying the genetic identity of animals and also animals resulting from such processes that may cause suffering to the animals without providing substantial medical benefit to the human or animal, and the edited cells do not include human germline cells or human embryos. However, the methods and uses disclosed herein specifically refer to uses and methods using non-embryo and non-germline human cells, e.g. primary human cells such as macrophages, T cells, etc., that are edited ex vivo under in vitro conditions.
[0216] Furthermore, in all aspects and embodiments disclosed herein, the method or use described herein, insofar as it refers to a plant cell, includes that said at least one plant cell, tissue, organ, plant or seed is not obtained by an essentially biological process. Instead, said at least one plant cell, tissue, organ, plant or seed is obtained by at least one step of artificial intervention in the form of using ABE as disclosed herein, which does not exist in nature per se and affects plant cells, by modifying and / or introducing steps of a technical nature that affect sexual crossing and selection. Such steps include genome editing steps, for example to exchange bases or nucleotides of interest, chemical treatments, factors or genes or gene products, including chromosome doubling, removal of chromosomes, introduction of exogenous genes or genetic material into the plant genome (nuclear, mitochondrial or plastid genome), etc., or any combination thereof.
[0217] In yet another aspect, the invention also relates to a kit comprising (a) an adenine base editor described herein, and / or a complex described herein, and / or a nucleic acid molecule described herein, and / or an expression construct described herein, and / or a cell described herein, and (b) a container containing reaction components including buffers, and, optionally, (c) instructions for use.
[0218] According to all embodiments of the kits described herein, the reaction components, including buffers, provide suitable reaction conditions to promote activity of an adenine base editor described herein, and / or a complex described herein, and / or a nucleic acid sequence described herein, and / or an expression construct described herein, and / or a cell described herein.
[0219] In yet another aspect, there is provided a method of obtaining a plant or a seed thereof, or a progeny regenerated from the plant or seed, which may comprise propagating a trait introduced by at least one adenosine base editor described herein to at least one genomic target site, i.e., the site of the plant or seed thereof to be modified and / or the site where the ABE interacts. In one embodiment, the genome of the plant or seed modified by at least one targeted edit mediated by at least one adenosine base editor described herein can thus be used to modify the genome of the progeny in a targeted manner by specifically combining with the genome of the original polyploid plant, such that at least one targeted edit is present in at least one allele of the progeny.
[0220] In one aspect, there may be provided a use of an adenine base editor, complex, or expression construct or vector described herein for adenine base editing of a target site in a genome of interest in at least one cell of a prokaryotic or eukaryotic organism, including bacterial and archaeal organisms. EXAMPLES
[0221] Example 1: Molecular methods Example 1.1: Cloning PCR was performed using Q5® High Fidelity DNA Polymerase (M0491, NEB) with DNA oligonucleotides from Integrated DNA Technologies (IDT). PCR products were gel purified using a gel purification kit (Zymo Research, no. D4002). To generate entry vectors, DNA fragments were inserted into BsaI-digested GreenGate empty entry vectors via Gibson assembly (2x NEBuilder Hifi DNA Assembly Mix, NEB) or restriction ligation with T4 DNA ligase (NEB). Base editor, gRNA and fluorescent reporter vectors were constructed using Golden Gate cloning (30 cycles (37°C, 5 min, 16°C, 5 min), 50°C for 5 min, 80°C for 5 min) using BsaI or BbsI. Vectors were transformed into DH5α E. coli or One Shot™ ccdB Survival™ competent cells (Thermo Fisher Scientific) by heat shock transformation. Depending on the selectable marker, cells were cultured at 100 μg mL -1 of carbenicillin, 100 μg mL -1 of spectinomycin, 25 μg mL -1 of kanamycin or 40 μg mL -1 The plasmids were isolated (GeneJET Plasmid Miniprep Kit, Thermo Fisher Scientific) and verified by restriction enzyme digestion and / or Sanger sequencing (Eurofins, Mix2seq).
[0222] Example 1.2: Entry clone TadA7.10d was synthesized on a BioXP3200 DNA synthesis platform (Codex DNA) based on the published sequence (Gaudelli et al., Nature 551, 464-471, 2017), and TadA8e (#138489) and Tad8.20m (#136300) were ordered from Addgene.
[0223] The LbCas12a(D832A) sequence was codon-optimized for wheat and then synthesized (Twist Biosciences). The three synthesized fragments were cloned into the entry vector using Gibson assembly. The LbCas12a(D156R-D832A) variant was generated by site-directed mutagenesis PCR via Gibson assembly. The 3xSV40-NLS and BPstar-NLS sequences were previously published (Richter at el., Nature Biotechnology 38, 883-891, 2020; Alok et al., Frontiers in Plant Science 11, 264, 2020) and cloned by annealing oligos and subsequent ligation. Nuc-UGI-SV40 was amplified from A3A-PBE (Addgene #119768); cloned into the entry vector via Gibson assembly. The CaMV terminator was isolated from the PABE-7 plasmid (Addgene #115628) and cloned into the entry vector via Gibson assembly.
[0224] Example 2: Plant growth conditions Wheat seeds (Fielder) were treated by successive washings with sterile water for 3 min, isopropanol for 45 s, sterile water for 3 min and 6% sodium hypochlorite (Chem-lab nv) for 10 min. Sterile seeds were washed six times with sterile water in a laminar flow cabinet and sown on a sterile growth medium containing 1 / 2MS pH 5.7 (Duchefa Biochemie, M0221.0050), 2.5 mM MES (Duchefa Biochemie, M1503.0100) and 0.5% plant tissue agar (NEOGEN, NCM0250A). Two seeds were sown per sterile 175 mL cylindrical container (Greiner Bio-one, #960162) and stratified for 3 days at 4°C in the dark. Plants were grown under SpectraluxPlus NL 36 W / 840 Plus (Radium Lampenwerk) fluorescent lights at 25°C with long days (16 h light / 8 h dark).
[0225] B104 corn seeds were sown directly onto Jiffy substrate (Jiffy Products International, No. 32170138). Seed germination was carried out for 5 days under long-day conditions (16 h light / 8 h dark) at 25°C and 55% relative humidity under light provided by a metal halide lamp equipped with high pressure sodium vapor (RNP-T / LR / 400W / S / 230 / E40; Radium) and a quartz burner (HRI-BT / 400W / D230 / E40; Radium). Seedlings were transferred to the dark for 8 days before isolating protoplasts.
[0226] Example 3: Protoplast isolation and transfection Example 3.1: Isolation of wheat protoplasts Wheat leaves were harvested 7–8 days after emergence (DAG). Approximately 40–50 second leaves were cut into strips of 0.5–1 mm latitude with a sharp razor blade and the leaf strips were incubated in 0.6 M D-mannitol (Sigma-Aldrich, M1902) for 10 min in the dark. The mannitol was removed and 25 mL of cell wall enzyme solution (20 mM MES, 1.5% Cellulase R10 (C8001.0010), 0.75% Macerozyme R10 (M8002.0010), 0.6 M D-mannitol and 10 mM KCl, 0.1% BSA and 10 mM CaCl2) was added to the protoplasts and incubated for 8 h in the dark at 25 °C with shaking at 40 rpm. After enzymatic digestion, 25 mL of W5 solution (2 mM MES pH 5.7, 154 mM NaCl, 125 mM CaCl2, 0.5 mM KCl) was added to release the protoplasts. The mixture was filtered through a sterile 40 μm cell strainer (Corning, #431750) and the protoplasts were harvested by centrifugation at 80 g (slow acceleration and brake) for 3 min at room temperature. The supernatant was discarded and the protoplasts were resuspended in 6 mL of W5 solution and incubated on ice for 30 min. Protoplast yield was determined by adding MMGTa solution (4 mM MES pH 5.7, 0.4 M mannitol, 15 mM MgCl2) onto the cell pellet to obtain a yield of 1 × 106 cells mL -1 was measured using a Neubauer chamber before reaching a concentration of
[0227] Example 3.2: Transfection of wheat protoplasts The protoplasts were then incubated on ice for approximately 30 min before transfection. 12 μg of total plasmid DNA was added to MMGTa to a total volume of 20 μL in a 1 mL strip tube (National Scientific Supply Co, TN0946-08B). 100 μL of protoplasts (105 cells) and 110 μL of PEG solution (0.2 M mannitol, 100 mM CaCl2), 40% PEG (Sigma 81240) were added to the DNA using a multichannel pipette and mixed immediately by gently inverting the strips. Eight transfections were processed in parallel for each strip. The protoplasts were incubated for 15-20 min and transfection was stopped by adding W5 solution. After centrifugation at 80 g (slow acceleration and braking) for 3 min, the supernatant was discarded and the protoplast pellet was resuspended in 1 mL of W5 solution. The cells were then transferred to a 24-well plate (VWR734-2325EU catalogue) and incubated in the dark at 25°C for 42-46 hours.
[0228] Example 3.3: Isolation of maize protoplasts Etiolated maize leaves were harvested at 12 or 13 DAG. The central portion of the second or third leaf was cut into 0.5 mm strips. The strips were then infiltrated with 25 mL of cell wall enzyme solution (0.6 M D-mannitol, 10 mM MES, 1.5% cellulose, 0.3% Macerozyme R10, 0.1% BSA and 1 mM CaCl2) using a vacuum (50 mmMg) for 30 min in the dark, and then incubated at 25°C for 2 h with shaking (40 rpm). The solution containing the protoplasts was filtered using a sterile 40 μm cell strainer (Corning) and collected by centrifugation at 100 g (slow acceleration and brake) for 3 min. The supernatant was removed and the protoplasts were washed with ice-cold 0.6 M D-mannitol by centrifugation at 100 g (slow acceleration and brake) for 2 min. The cells were then resuspended in 5 mL of 0.6 M D-mannitol and incubated in the dark for 30 min. Protoplasts were resuspended in MMGZm solution (0.6 M mannitol, 15 mM MgCl2, 4 mM MES) and counted using a Neubauer chamber to yield a total of 1 × 106 cells mL -1 The concentration was adjusted to .
[0229] Example 3.4: Transfection of maize protoplasts 20 μg of total plasmid DNA was added to MMGZm to a total volume of 20 μL in a 1 mL strip tube. 100 μL of protoplasts (105 cells) and 110 μL of PEG (0.2 M mannitol, 100 mM CaCl2, 40% PEG (Sigma 81240)) solution were added to the DNA using a multichannel pipette and mixed immediately by gently inverting the strip. For each strip, eight transfections were processed in parallel. The cells were then incubated in the dark for 10-15 min and transfection was stopped by adding W5 solution. After centrifugation for 2 min at 100 g (slow acceleration and braking), the supernatant was discarded and the protoplast pellet was resuspended in 1 mL of W5 solution. The cells were then transferred to a 24-well plate (VWR) using a tip with wide holes and incubated for 2 days in the dark at 25 °C with shaking (20 rpm).
[0230] Example 4: High Content Image Analysis Two days after transfection, 50 μL of protoplasts were transferred into 96-well Cell Carrier Ultra plates (#6055302) and imaged with an Opera Phenix® high content screen system (PerkinElmer). Image acquisition was performed using a 20x water immersion objective in confocal mode, acquiring seven Z-planes and nine fields of view per well, covering four image channels: bright field, chlorophyll, GFP, and mCherry. Raw images were transferred to a Columbus™ image data storage and analysis system for automated image processing and quantification.
[0231] After flat-field correction and smoothing of the chlorophyll channel, single wheat cells were segmented and selected as protoplasts based on circularity. The mCherry and GFP signals were used to identify nuclei and non-transformed protoplasts were excluded based on the lack of nuclear mCherry signal. The intensity of mCherry and GFP in the nuclei of transformed protoplasts was used to identify and quantitate GFP expressing transformed protoplasts.
[0232] In the case of maize, the chlorophyll channel could not be used for cell division because the plants were etiolated. The analysis was focused directly on the transformed protoplast nuclei and segmented based on the mCherry and GFP channels. The same analysis as above was used to identify and quantitate the transformed nuclei expressing GFP.
[0233] Results were exported as tables and all calculations and image processing were performed on an in-house cluster (VIB). The time required from the start of imaging to obtaining processed results takes 3-4 h for a 96-well plate. Code for the wheat and maize analysis workflows is available in the Supplementary Data.
[0234] Example 5: FACS Images were acquired using a BD Biosciences FACS imaging-compatible prototype cell sorter equipped with an optical module that allows multicolor fluorescent imaging of fast-flowing cells in a stream, enabled by BD CellView® imaging technology, which is based on fluorescent imaging using radio frequency tagging emission (FIRE).
[0235] Two days after transfection, 500 μL of protoplast solution was used for sorting. The gating strategy for GFP was first established in cells expressing pZmUBI-GFP-NLS (p02243), and similar settings were used for all experiments in wheat and maize. A quality check was performed by running the sorted cell fraction on the instrument and imaged using the imaging system integrated into the FACS instrument. For both wheat and maize, a 130 μm nozzle was used to sort 1,000–5,000 cells into a 1.5 mL Eppendorf tube containing 10 μL of dilution buffer from the Phire Tissue Direct PCR Master Mix kit (Thermo Fisher Scientific, F160L).
[0236] Example 6: Genotyping and NGS analysis For genotyping individual wheat transformed plants, leaf pieces (0.5–1 cm) were collected in 1 mL tubes on a 96-well plate (VWR, 732-3716) and flash frozen in liquid nitrogen. The tissue was ground into a powder by adding two metal beads (3 mm) and shaking the plate at 20 Hz for 1 min (Retsch, Mixer Mill MM 400). 400 μL of extraction buffer (100 mM Tris-HCl pH 8.0, 500 mM NaCl, 50 mM EDTA, 0.7% SDS) was added to each sample and incubated at 60 °C for 30 min. Samples were centrifuged and 300 μL of the supernatant was mixed with 300 μL of isopropanol for DNA precipitation. Samples were then centrifuged and the supernatant was removed. The pellet was washed with 70% ethanol, dried at room temperature and dissolved in 100 μL of 10 mM Tris-HCl pH=8.0.
[0237] For sorted material, 2 μL of solution containing sorted cells was used as template in a total reaction volume of 20 μL for amplicon PCR using the Phire Plant Direct PCR Kit (Thermo Fisher Scientific, F160L) according to the manufacturer's recommendations.
[0238] For dCas12-BE and nuclease activity LbCas12a, the efficiency of base editing and indels was measured using NGS. Amplicons of 210–260 bp were designed to amplify the target site. For pooling and demultiplexing of amplicon reads after sequencing, 6-base indexes were added to the forward and reverse primers. 5 μL of Phire PCR reaction was verified on a 2% agarose gel with a low molecular weight ladder (NEB, no. N3233S). 15 μL of PCR product was pooled and purified using a PCR purification kit (Zymo Research Co., D4013). Depending on the amplicon, special gel purification was performed to specifically isolate the PCR band of the target site (Zymo Research, D4002). DNA concentration was measured using Qubit (Invitrogen) according to the manufacturer's protocol and was 2 ng μL. -1 Paired-end sequencing was performed using Eurofins NGSelect amplicons (5M reads 2 × 150 bp). Reads were demultiplexed using Je-demultiplex and individual fastq files were obtained using the Galaxy workflow (https: / / usegalaxy.be).
[0239] Base edits were calculated using CRISPResso2Pooled or CRISPRessoBatch. Editing window and read quality were defined as follows: cleavage offset was set to -1, quantification window size was set to 10, quantification window center was set to -12, and minimum average read quality was set to 30. Indels were calculated using CRISPResso2Pooled or CRISPRessoBatch with the following settings, defining the Cas12a cleavage site and read quality: cleavage offset was set to -4, and minimum average read quality was set to 30.
[0240] Example 7: Stable wheat transformation Immature embryos, 2-3 mm in size, were isolated from sterilized ears of wheat cultivars. They were fielded and bombarded using a PDS-1000 / He particle delivery system (Bio-Rad) using the following particle bombardment parameters: gold particle diameter, 0.6 μm; target distance, 6 cm; bombardment pressure, 7.584 kPa; gap distance, 8-10 mm; microcarrier flight distance, 10 mm; vacuum in the bombardment chamber, 27.5''Hg. For each shot, approximately 150 μg of gold particles carrying 570 ng of plasmid DNA were delivered.
[0241] The applied plasmid DNA was a mixture of Cas12a-ABE vectors pCG392 or pCG434, pCG406 and pCG408 (gRNA), and pBAY02032 (selection marker). Vector pBAY02032 contains an eGFP-BAR fusion gene under the control of the 35S promoter. Bombarded immature embryos were transferred to non-selective WLS callus induction medium for approximately 1 week, then transferred to WLS containing 5 mg L-1 phosphinothricin (PPT) for an initial selection round of approximately 3 weeks, followed by 10 mg L-1 phosphinothricin (PPT) for an additional 3 weeks. -1 A second selection round was performed with WLS containing PPT. PPT-resistant calli were selected and cultured at 5 mg L -1 The shoots were transferred to shoot regeneration medium containing PPT.
[0242] Example 8: Replication of Cas12-ABE components in wheat Various components of the Cas12a-ABE underwent extensive iterative testing in wheat protoplasts to develop an optimized Cas12a-ABE construct.
[0243] Isolation of wheat protoplasts was carried out according to Example 3.1: Wheat leaves were harvested 7 or 8 days after germination. Approximately 40-50 second leaves were cut into strips with a latitude of 0.5-1 mm with a sharp razor blade and the leaf strips were incubated in 0.6 M D-mannitol (Sigma-Aldrich) for 10 min in the dark. The mannitol was removed and 25 mL of cell wall enzyme solution (20 mM MES, 1.5% Cellulase R10, 0.75% Macerozyme R10, 0.6 M D-mannitol and 10 mM KCl, 0.1% BSA and 10 mM CaCl2) was added to the protoplasts and incubated for 8 h in the dark at 25 °C with shaking at 40 rpm. After enzymatic digestion, 25 mL of W5 solution (2 mM MES pH 5.7, 154 mM NaCl, 125 mM CaCl2, 0.5 mM KCl) was added to release the protoplasts. The protoplasts were harvested and centrifuged at 80 g for 3 min at room temperature. The supernatant was discarded and the protoplasts were resuspended in 6 mL of W5 solution and incubated on ice for 30 min. The yield of protoplasts was determined by adding MMG solution (4 mM MES pH 5.7, 0.4 M mannitol, 15 mM MgCl2) onto the cell pellet to obtain a yield of 1 × 10 6 cells mL -1 was measured using a Neubauer chamber before reaching a concentration of
[0244] Transfection of wheat protoplasts was carried out according to Example 3.2: Protoplasts were then incubated on ice for ±30 min before transfection. 12 μg total plasmid DNA was added to MMG to a total volume of 20 μL in a 1 mL strip tube (National Scientific Supply Co). 100 μL of protoplasts (=1 × 10 5Cells) and 110 μL of PEG solution (0.2 M mannitol, 100 mM CaCl2, 40% PEG (Sigma 81240)) were added to the DNA using a multichannel pipette and mixed immediately by gently inverting the strips. For each strip, eight transfections were processed in parallel. Protoplasts were incubated for 15-20 min and transfection was stopped by adding W5 solution. After centrifugation at 80 g for 3 min, the supernatant was discarded and the protoplast pellet was resuspended in 1 mL of W5 solution. Cells were then transferred to a 24-well plate and incubated in the dark at 25 °C for 42-46 h. For all of Examples 1.1-1.2, a wheat-codon-optimized version of LbCas12a was used.
[0245] Overall, 12 different expression constructs were constructed that contained nucleic acid sequences encoding the adenine base editors described herein (see FIG. 1a; constructs 1-12). In addition, a total of eight different expression constructs were constructed that contained nucleic acid sequences encoding guide RNAs described herein (see FIG. 1b; constructs a-h). To test ABE activity, wheat protoplasts were co-transfected with three vectors: (1) a vector encoding a mutant GFP gene in which the Gln(Q) codon (CAG) was mutated to a stop codon (TAG) (FIG. 1c), (2) a Cas12a-ABE expression vector with a p35S:mCherry-NLS cassette, and (3) a vector encoding a gRNA targeting the mutant GFP codon. Two days after transfection, the protoplasts were transferred to a 96-well Cell Carrier Ultra plate and imaged with an Opera Phenix® High Content Screen System (Perkin Elmer). Image acquisition was performed using a 20x water immersion objective in confocal mode, acquiring seven Z-planes and nine fields of view per well, covering four image channels: bright field, chlorophyll, EGFP and mCherry. Raw images were transferred to a Columbus™ image data storage and analysis system for automated image processing and quantification. After flat-field correction and smoothing of the chlorophyll channel, single wheat cells were segmented and selected as protoplasts based on their circularity. Editing the TAG codon to a CAG codon restores the GFP coding sequence, resulting in GFP fluorescence. The ratio [%] of GFP cells / mCherry cells was determined as a measure of ABE activity. Thus, a higher ratio of GFP cells / mCherry cells indicates higher ABE activity.
[0246] Example 8.1: Evaluation of Adenosine Deaminase Domains Three different expression constructs containing the nucleic acid sequence encoding the adenine base editor described herein were comparatively analyzed for ABE activity because each of these expression constructs contained a different adenosine deaminase domain as a unique feature: (i) TadA7.10d (construct 1; see FIG. 1a), (ii) TadA8.20 (construct 2; see FIG. 1a), and (iii) TadA8e (construct 3; see FIG. 1a). The ABE activity of each was tested in wheat protoplasts and expressed as GFP cells / mCherry cells [%] (see FIG. 2a). Each of the three different adenine base editor expression constructs was co-transfected with the same guide RNA expression construct (construct a; see FIG. 1b). As a negative control, all three adenine base editor expression constructs were tested without the guide RNA expression construct (see FIG. 2a).
[0247] The results show that ABE activity without guide RNA was undetectable in all cases (see FIG. 2a). The ABE activity (0.6%) with TadA8e as the adenosine deaminase was consistently higher compared to the other two adenosine deaminase domains tested (0.0% and 0.1% for TadA7.10d and TadA8.20, respectively; see FIG. 2a).
[0248] Example 8.2: Evaluation of NLS sequences Five different expression constructs containing nucleic acid sequences encoding the adenine base editors described herein were comparatively analyzed with respect to their ABE activity. Each of these expression constructs contained a different combination of NLS sequence configurations and plant terminator sequences: (i) one 3xSV40 (SEQ ID NO:52) at the 3'- of the dLbCas12a domain (construct 3; see FIG. 1a), (ii) one BP (SEQ ID NO:53) at the 5'- of the adenosine deaminase domain and one 3xSV40 (SEQ ID NO:52) at the 3'- of the dLbCas12a domain (constructs 4 [G7T terminator] and 5 [CaMV terminator]; see FIG. 1a), (iii) one BP (SEQ ID NO:53) at the 5'- of the adenosine deaminase domain and one BP (SEQ ID NO:53) at the 3'- of the dLbCas12a domain (constructs 6 [G7T terminator] and 7 [CaMV terminator]; see FIG. 1a). The ABE activity of each was tested in wheat protoplasts and expressed as GFP cells / mCherry cells [%] (see FIG. 2b). Each of the different adenine base editor expression constructs was co-transfected with the same guide RNA expression construct (construct a; see FIG. 1a). As a negative control, all adenine base editor expression constructs were tested without the guide RNA expression construct (see FIG. 2b).
[0249] The highest ABE activity was found in constructs 6 and 7 (3.1% and 2.7%, respectively; see FIG. 2b), which contain one BP (SEQ ID NO: 53) at the 5'-end of the adenosine deaminase domain and one BP (SEQ ID NO: 53) at the 3'-end of the dLbCas12a domain. Other constructs yielded ABE activities ranging from 0.9% to 1.9% (constructs 3 to 5; see FIG. 2b). The effect of the terminator sequence on ABE activity was not significant.
[0250] Example 8.3: Evaluation of guide RNA systems Adenine base editor expression construct 6 (see Example 1.2) was tested in combination with eight different guide RNA expression constructs (constructs a-h; see FIG. 1b). The respective ABE activities were determined in wheat protoplasts and expressed as GFP cells / mCherry cells [%] (see FIG. 2c). As a negative control, adenine base editor expression construct 6 was tested without guide RNA expression construct (see FIG. 2c).
[0251] Adenine base editor expression construct 6, in combination with guide RNA expression construct h, yielded the highest ABE activity (11.9%; see Figure 2c), which contained (i) a truncated tRNA at the first 5'- of two mature direct repeat sequences, (ii) one mature direct repeat sequence at the 5'- of the spacer RNA coding sequence, (iii) a second mature direct repeat sequence at the 3'- of the spacer RNA coding sequence, and (iv) a poly-T tail (T-extended terminator). In combination with adenine base editor expression construct 6, other guide RNA expression constructs tested yielded ABE activities ranging from 0.3% to 5.5% (constructs c-g; see Figure 2c).
[0252] Example 8.4: Evaluation of Cas12a domains Four different adenine base editor expression constructs (constructs 3, 6, 8 and 9; see FIG. 1a) were tested in combination with a guide RNA expression construct (see FIG. 1b). The respective ABE activities were determined in wheat protoplasts and expressed as GFP cells / mCherry cells [%] (see FIG. 3a). As negative controls, all adenine base editor expression constructs were tested without a guide RNA expression construct (see FIG. 3a).
[0253] Significantly higher ABE activity was determined for adenine base editor expression constructs 8 and 9 (23.4% and 29.4%, respectively; see FIG. 3a) compared to adenine base editor expression constructs 3 and 6 (8.7% and 13.0%, respectively; see FIG. 3a). In contrast to constructs 3 and 6 (both containing dLbCas12a), constructs 8 and 9 contained the D156R mutant of dLbCas12a, which exhibited increased activity and / or enhanced temperature tolerance compared to the wild-type LbCas12a enzyme (Schindele and Puchta, 2020, Plant Biotechnol. J., 18(5), p.1118-1120. doi:https: / / doi.org / 10.1111 / pbi.13275). The highest ABE activity was detected in construct 9 (29.4%), which contained one BP (SEQ ID NO: 53) 3' of the dLbCas12a domain as the C-terminal NLS sequence (see FIG. 1a). In contrast, construct 8 contained one 3xSV40 (SEQ ID NO: 52) 3'- of the dLbCas12a domain as the C-terminal NLS sequence (see FIG. 1a).
[0254] Example 8.5: Evaluation of TadA8e vs. TadA9 and the 32aa linker vs. hexaGGGGS linker Four different adenine base editor expression constructs (constructs 9, 10, 11 and 12; see FIG. 1a) were tested in combination with a guide RNA expression construct (see FIG. 1b). The ABE activity of each was tested in wheat protoplasts and expressed as GFP cells / mCherry cells [%] (see FIG. 3b). As negative controls, all adenine base editor expression constructs were tested without a guide RNA expression construct (see FIG. 3b).
[0255] Overall, both constructs containing TadA9 as the adenosine deaminase domain (constructs 11 and 12; see FIG. 1a) produced higher ABE activity compared to both constructs containing TadA8e as the adenosine deaminase domain (constructs 9 and 10). The highest ABE activity (41.9%) was determined for construct 12 (see FIG. 3b), which contained TadA9 as the adenosine deaminase domain (see FIG. 1a) and a hexaGGGGS linker domain connecting TadA9 to the dLbCas12a (D156R mutant) domain located 3'- of TadA9 (see FIG. 3b). In contrast, construct 11 contained a 32aa linker domain (see FIG. 1a).
[0256] Example 8.6: Comparative analysis of Cas12a-ABE in wheat We compared successive BE structures in single wheat protoplast experiments and confirmed that individual modifications along the optimization path of LbCas12a-ABE lead to increased activity (Figure 3c). Consistent with previous experiments, the introduction of the truncated tRNA DR-DR crRNA system, the LbCas12a(D156R) variant, the TadA9 and the 6×GGGGS linker all led to a significant increase in base editing efficiency (one-way ANOVA, Tukey HSD: P<0.05).
[0257] Example 9: Comparative analysis of Cas12a-ABE in maize Six different combinations of adenine base editor expression constructs and guide RNA expression constructs were comparatively analyzed in maize protoplasts. The respective ABE activities were expressed as GFP cells / mCherry cells [%] (see Figure 4b). As negative controls, all adenine base editor expression constructs were tested without guide RNA expression constructs (see Figure 4b).
[0258] The results show that the introduction of an N-terminal BP and a C-terminal BP (SEQ ID NO: 53) as an NLS in combination with a guide RNA structure comprising a truncated tRNA flanked by mature direct repeat sequences in the 5'- and 3'-directions (construct h; see FIG. 1b) leads to a first significant increase in editing efficiency (v6h, 13.3%, v3a, 0.8%, v6a, 0.2%, see FIGS. 4a and b). Furthermore, the introduction of a mutation in the dLbCas12a domain (D156R) that results in an enhanced temperature resistance further increases the editing efficiency by 2-fold, from 13.3% (construct v6h) to 26.1% (construct v9h; see FIGS. 4a and b). Furthermore, the replacement of the TadA8e adenosine deaminase domain with TadA9 further increases the editing efficiency by almost 2-fold (47.7% in construct v11h, see FIGS. 4a and b). Finally, replacement of the 32aa linker domain with a hexaGGGGS linker domain increases the editing efficiency by a further 20% (67.9% in construct v12h; see Figures 4a and b).
[0259] Example 10: Validation of Cas12a-ABE activity at endogenous target sites in wheat and maize Example 10.1: Cas12a-ABE activity at endogenous target sites in wheat (Triticum aestivum) The base editing activity of the optimized LbCas12a-ABE was determined by measuring A to G substitutions at endogenous wheat genomic sites. Transfected protoplast samples were sorted using a BD Biosciences FACS imaging-compatible prototype cell sorter according to Example 5. Two days after transfection, 500 μL of protoplast solution was used for sorting. The gating strategy for GFP was first established in cells expressing pZmUBI-GFP-NLS, and similar settings were used for all experiments in wheat and maize. For both wheat and maize, a 130 μm nozzle was used to sort 1,000–5,000 cells into 1.5 mL Eppendorf tubes containing 10 μL of dilution buffer from the Phire Tissue Direct PCR Master Mix kit (Thermo Fisher Scientific). 2 μL of the solution containing the sorted cells was used as template in a total reaction volume of 20 μL for amplicon PCR using the Phire Plant Direct PCR Kit, following the manufacturer's recommendations. PCR products were pooled and purified using a PCR purification kit (Zymo Research Co., D4013). Paired-end sequencing was performed using Eurofins NGSelect amplicons (5M reads 2 × 150 bp). Reads were demultiplexed using Je-demultiplex (Girardot et al., 2016 BMC Bioinformatics 17, 419) and individual .fastq files were obtained using the Galaxy workflow (https: / / usegalaxy.be) as previously described (Bollier et al., 2020 BioRxiv 11.13.381046). Base editing was calculated using CRISPResso2Pooled or CRISPRessoBatch (Clement et al., 2019 Nature Biotechnology 37, 224-226).
[0260] Six LbCas12a-ABE and crRNA constructs were selected (numbers and letters follow the nomenclature in Figure 1a and b), and their ABE activities were determined at four sites (Ta-TS60-A, Ta-TS105-B, Ta-TS112-A, and Ta-TS121-D; Table 1) in simplex (one guide RNA, see Figure 5) and multiplex (four guide RNAs, Figure 6). For all four target sites, A to G conversions in the range of 0.5% to 10% within the editing window of A08 to A11 were observed in v9h, v11h, and v12h (see Figures 5 and 6). These findings indicate that (i) the inclusion of a truncated tRNA and two mature direct repeats 5'- and 3'-of the spacer RNA-encoding sequence in the guide RNA construct, and (ii) the D156R mutant of dLbCas12a, and (iii) the incorporation of the TadA9 domain into the LbCas12a-ABE construct improves editing efficiency.
[0261] [Table 13]
[0262] To determine whether the ABE constructs showing higher ABE activity in wheat protoplasts also lead to efficient adenine base editing in wheat plants, activity against the two most efficient wheat genomic target sites (Ta-TS60-A and Ta-TS112-A) was tested in stably transformed wheat plants.
[0263] Wheat transformation was performed according to Example 7: Immature embryos of 2-3 mm in size were isolated from sterilized ears of wheat cultivars. They were fielded and bombarded using a PDS-1000 / He particle delivery system (Bio-Rad) as described by Sparks and Jones (2014; Cereal Genomics: Methods in Molecular biology, vol. 1099, Chapter 17) using the following particle bombardment parameters: gold particle diameter, 0.6 μm; target distance, 6 cm; bombardment pressure, 7.584 kPa; gap distance, 8-10 mm; microcarrier flight distance, 10 mm; vacuum in the bombardment chamber, 27.5''Hg. For each shot, approximately 150 μg of gold particles carrying 570 ng of plasmid DNA were delivered. The applied plasmid DNA was a mixture of vectors pCG392 (ABE v9: SEQ ID NO: 62) or pCG434 (ABE v12: SEQ ID NO: 63) (ABE), pCG406 (crRNA h: SEQ ID NO: 67) and pCG408 (crRNA h: SEQ ID NO: 68) (gRNA) and pBAY02032 (selection marker (SM)). Vector pBAY02032 contains an eGFP-BAR fusion gene under the control of the 35S promoter. Further culture of bombarded immature embryos was performed essentially as described by Ishida et al. (2015; Agrobacterium Protocols: Volume 1, Methods in Molecular Biology, vol. 1223, Chapter 15, 189-198). Bombarded immature embryos were transferred to non-selective WLS callus induction medium for about 1 week, then transferred to WLS containing 5 mg L-1 phosphinothricin (PPT) for a first selection round of about 3 weeks, followed by a second selection round on WLS containing 10 mg L-1 PPT for another 3 weeks. PPT-resistant calli were selected and transferred to shoot regeneration medium containing 5 mg L-1 PPT.
[0264] In two experiments, a total of 689 and 756 immature embryos were bombarded with a mixture of pCG392 base editor, two gRNAs, and SM plasmid DNA, and a total of 140 and 227 embryos were obtained as phosphinothricin (PPT)-resistant shoot regeneration lines. In two experiments, a total of 776 and 766 immature embryos were bombarded with a mixture of plasmid DNA and pCG434 base editor, and a total of 223 and 180 embryos were obtained as PPT-resistant shoot regeneration lines. All plants developed from one immature embryo were treated as a pool. Genomic DNA was extracted from pooled leaf samples for ddPCR analysis. The ddPCR assay was designed using Primer3Plus software with modified settings compatible with the applied master mix. To avoid loss of binding sites, primers and reference probes were designed away from the cleavage site. PCR primers were designed according to the following guidelines: primer length, 17-24 bases; primer melting temperature, 55-60 °C; ideal temperature, 58 °C; difference in melting temperature of two primers, 2 °C or less; primer GC content, 35-65%; amplicon size, 100-250 bases. Drop-off probes were designed to lose their binding site if one or more base substitutions were introduced within the base editing activity window. The sequences of the probes and primers are shown in Table 2.
[0265] [Table 14]
[0266] The 20x ddPCR mix consisted of 18 μM forward and 18 μM reverse primers, 5 μM reference probe and 5 μM drop-off probe. The following reagents were mixed in a 96-well plate to create a 25 μL reaction: 11 μL ddPCR Supermix for probes (without dUTP), 1.1 μL 10x Assay Mix (Bio-Rad Laboratories, CA, USA), 100–250 ng genomic DNA in water and up to 22 μL water. Droplets were generated using a QX100 droplet generator according to the manufacturer's instructions (Bio-Rad Laboratories) and transferred to a 96-well plate for standard PCR in a C1000 thermal cycler with a deep-well block (Bio-Rad Laboratories, Hercules, CA, USA). Thermal cycling consisted of a 10 min activation period at 95 °C, followed by 40 cycles of a two-step thermal profile of denaturation at 95 °C for 30 s, combined annealing and extension at 60 °C for 3 min, and one cycle at 98 °C for 10 min. After PCR, droplets were analyzed in "absolute quantification" mode using a QX100 droplet reader (BioRad Laboratories, Hercules, CA, USA). The ddPCR drop-off level is a measure of the level of base editing in plants. Table 3 summarizes the editing frequencies observed. For Ta-TS60-A, 17–38% of pools have more than 10% editing and 3–10% of pools have more than 50% editing. The target site Ta-TS112-A shows higher levels of editing, with 35–57% of pools having more than 10% editing and 9–27% of pools having more than 50% editing. This indicates that efficient adenine base editing is occurring in some pools.
[0267] [Table 15]
[0268] For almost all shoots that showed base editing, the drop-off levels were approximately 50% or 100%.
[0269] Individual T0 wheat plants were analyzed by NGS to determine the editing frequency and alleles of on-target and off-target loci. DNA was isolated using the Edwards method (Edwards et al., 1991, Nucleic Acids Research 19, 1349). Amplicons were obtained by PCR using Q5 polymerase (Invitrogen), pooled, and purified using a PCR purification kit (Zymo Research Co., D4013). Paired-end sequencing was performed (Eurofins NGSelect amplicon: 5M reads 2 × 150 bp) and base editing was calculated using CRISPResso2 (Clement et al., 2019 Nature Biotechnology 37, 224-226).
[0270] The proportion of wheat plants with A:T to G:C conversions at positions A7 and A9 in the case of TaTS60-A and A8 and A10 in the case of TaTS112-A ranged from 2.6 to 34% for v9h and 11.5 to 46.7% for v11h (n>153 for each BE structure, Fig. 7a-c). Consistent with the results from protoplasts, v11h efficiency was significantly higher than v9h at these positions (1.4-4.7-fold increase, p<0.05 z-score test for the two ratios: Fig. 7a-c). Base editing was also observed at secondary positions (A10, A11, and A15 in the case of TaTS60-A and A6, A7 in the case of TaTS112-A), with no indels detected in any of the lines. Overall, v9h and v11h activity in TS60-A and TS112-A generated panels of 18 and 37 unique genotypes, respectively (Fig. 8a).
[0271] Consistent with previous protoplast results, Ta-TS112 was more active than Ta-TS60: 34–47% of independent events were edited with TS112-A and 16–35% with TS60-A (Figure 7).
[0272] We compared the editing rates obtained by ddPCR and NGS for a subset of individual T0 lines (n>54). A strong correlation was observed between the editing rates obtained by the two methods for both v9h and v11h in TaTS60-A and TaTS112-A (Figure (Figure9a–b),9), indicating that the ddPCR assay reliably predicts A:T to G:C mutation levels.
[0273] We also analyzed off-target editing of the homologous subgenomes by NGS (Figure 7c-d, target sites TaTS60-D, TaTS60-B, TaTS112-D and TaTS112-B) and observed A:T to G:C conversions consistent with previous protoplast results. The majority of mutant lines were scored as heterozygous, with 13% and 29% of lines containing heterozygous on-target edits of the v11h constructs in TS60 and TS112, respectively (Figure 7d). Homozygous lines were also recovered, with 3–18% recovered for major targeted bases. Approximately 30% of plants transformed with v9h had at least one base edit in at least one of the homeologs of TS60 or TS112, compared with an average of 50% for v11h (Figure 8b; p<0.01; two-proportion Z-score test). Furthermore, double mutants of TS60 and TS112 were generated at a rate of 11% for v9h and 19% for v11h (Figure 8b). These results indicate that the v9h and v11h Cas12a-ABEs can efficiently induce base editing in stable wheat lines without inducing indels. Thus, with the specific constructs provided, contrary to the literature (see Li et al., 2022, supra), our ABEs showed robust base editing efficiency without severe off-target effects.
[0274] As confirmation, a subset of T1 progeny was genotyped to confirm that Cas12a-ABE could be used to generate stably transmitted alleles. To exclude continued activity of Cas12-ABE in T1, we first screened for the absence of functional Cas12a-ABE transgenes in 20 v9h and 20 v11h lines. Twelve lines were selected for both constructs and genotyped at the TS60 and TS112 sites via NGS from three or four transgene-free plants, respectively (total of 94 plants). Ten of 12 and 12 of 12 v9h and v11h lines, respectively, contained edits in T1. Mutations were heterozygous or homozygous A-to-G edits at all target sites and subgenomes in both constructs, with no indels detected (Figure 12). Taken together, these results indicate that both Cas12a-ABEs can be used to efficiently introduce heritable multiplex base edits into indel-free wheat.
[0275] Example 10.2: Cas12a-ABE activity at endogenous target sites in maize (Zea mays) Similar experiments as described above for wheat were also performed on maize protoplasts (see Figure 10). The base editing activity of the optimized LbCas12a-ABE was determined by measuring A to G substitutions at endogenous maize genome sites. Six LbCas12a-ABE and crRNA constructs were selected (numbers and letters follow the nomenclature in Figures 1a and b), and their respective ABE activities were determined at four sites (Zm-TS1, Zm-TS3, Zm-TS4 and Zm-TS8; Table 1) in a multiplex (four guide RNAs, Figure 10). Similar to the results determined for wheat, v9h, v11h and v12h showed the best performance with editing efficiencies ranging from 0.5 to 20% (see Figure 10).
[0276] The activities of three Cas12a-ABE constructs (v9h, v11h, and v12h) were then assessed and compared at four endogenous sites in a multiplex of maize T0 plants. A vector containing a WUS-BBM cassette and an array of four crRNAs targeting Zm-TS1, Zm-TS3, Zm-TS4 and Zm-TS8 was constructed and used to transfect the B104 Public Maize Inbred with Improved Tissue Culture and Use of Morphogenic Regulators. Frontiers in Plant Science. 13.) and transformed into Agrobacterium tumefaciens strain HA105.
[0277] The four target sites Zm-TS1, Zm-TS3, Zm-TS4 and Zm-TS8 and the molecular tools used are shown in Table 4 below, and SEQ ID NOs: 103 to 105 show the specific maize targeting ABE constructs designed, produced and used to target Zm-TS1, Zm-TS3, Zm-TS4 and Zm-TS8.
[0278] Maize transformation was performed from immature maize embryos as described by Aesaert et al., 2022. Leaf samples were taken from individual regenerating shoots and subjected to DNA isolation and genotyping as for the wheat analysis.
[0279] For v9h and v11h, A:T to G:C conversions were detected at two of the four sites (Zm-TS4-, Zm-TS8) in up to 44% of edited plants (Fig. 11b, d). In contrast, three of the four sites were edited for v12h (Zm-TS3, Zm-TS4 and Zm-TS8), with up to 52.4% of the plants edited (Fig. 11c, d). Consistently, v12h also showed significantly higher activity in Zm-S8 (Fig. 11c-e; p<0.05: two-proportion z-score test). Most plants were shown to be heterozygous for the mutation, but homozygous mutations for v11h and v12h were also observed (see Fig. 11b, c, e). Similar to the above results in wheat stable plants, none of the regenerated T0 plants showed any indels. Analysis of progeny plants from one v9h, two v11h and one v12h maize lines showed that both the transgene insertion and the A to G edit were inherited by the next generation.
[0280] These results support that the optimized Cas12-ABE can stably introduce A:T to G:C mutations at endogenous sites in different monocotyledonous plant species and therefore the optimized ABE can be widely and successfully applied to various plant species.
[0281] [Table 16]
[0282] Example 11: LbCas12a-ABE activity in soybean and rapeseed To test the activity of LbCas12a-ABE in dicotyledonous plants, additional experiments using rapeseed (Brassica napus) and soybean (Glycine max) protoplasts were performed. Rapeseed protoplasts were isolated from leaves of 4- to 7-week-old plants grown axenically. Healthy leaves were cut into fine strips with a sharp razor blade. The strips were infiltrated with a cell wall lytic enzyme solution (1.5% Cellulase R10 and 0.75% Macerozyme R10 in 10 mM KCl and 0.6 M mannitol, pH 7.5) and incubated overnight in the dark with gentle shaking (40 rpm) at 24 °C. After enzyme digestion, the released protoplasts were collected by filtering the mixture through a 40 μm nylon mesh and resuspended in W5 solution. The resuspended protoplasts were placed on ice and allowed to settle by gravity, after which the cell pellet was resuspended in MMG. For transformation, 200 μL of cells (2.5 × 105) were mixed with 20 μg of plasmid DNA and 220 μL of freshly prepared polyethylene glycol (PEG) solution. The mixture was incubated in the dark for 15–20 min. After removing the PEG solution, the protoplasts were resuspended in 2 mL of W5 solution and incubated at 24 °C. Soybean protoplasts were isolated from unifoliate leaves of 6-day-old seedlings and transfected essentially as described for rapeseed. After removing the PEG solution, the protoplasts were resuspended in 2 mL of WI solution.
[0283] Cas12a-ABE activity was first evaluated using a GFP reporter system similar to that described in wheat. To this end, rapeseed protoplasts were co-transfected with three vectors: (1) a vector encoding a mutant GFP gene containing an early stop codon (SEQ ID NO:118), (2) a Cas12a-ABE expression construct containing TadA9 as the adenosine deaminase domain and a hexaGGGGS linker connecting TadA9 to the dLbCas12a(D156R) module located 3' of TadA9 (construct 12 in Fig. 1a; SEQ ID NO:119), and (3) a vector encoding a gRNA targeting the dGFP reporter and containing two mature direct repeats at the 5' and 3' ends of the spacer (SEQ ID NO:120). The Cas12a-ABE vector contained the Arabidopsis ubiquitin 10 promoter for constitutive expression, while expression of the gRNA was driven by the polymerase III type promoter of the Arabidopsis U6 snRNA gene. Editing the TAG stop codon to the original CAG codon restored the GFP coding sequence and resulted in GFP fluorescence. As a positive control, protoplasts were transfected with a construct containing wild-type GFP behind the strong Cauliflower Mosaic Virus (CaMV) 35S promoter. As a negative control, the Cas12a-ABE fusion protein was tested without gRNA. Fluorescence imaging 2 days after transfection revealed approximately 35% GFP fluorescent cells in the positive control and 4.2% in the Cas12a-ABE (see Figure 13a). Importantly, no GFP-positive cells were observed in the absence of gRNA (data not shown).
[0284] To confirm Cas12a-ABE activity at the endogenous target site, the TadA9>(GGGSS)6×>dLbCas12a-D156R expression construct was co-transfected into rapeseed or soybean protoplasts with expression constructs of Cas12a gRNAs targeting the BnFAD2 (SEQ ID NO: 121), BnALS3 (SEQ ID NO: 122) or GmFAD2 (SEQ ID NO: 123) genes, respectively. Transfected rapeseed protoplasts were cultured in alginate and editing efficiency was determined by deep amplicon sequencing 14 days after transfection. Conversely, soybean protoplasts were incubated in WI solution for 72 hours and analyzed by droplet digital PCR. As shown in Figure 13b, expression of Cas12a-ABE resulted in relatively high editing efficiency at the FAD2 target site, with up to 8.4% of sequence reads showing A to G substitutions and less than 0.025% showing indel formation. Low but significant levels of base editing were observed when targeting the BnALS3 or GmFAD2 genes (average 0.52% and 1.58%, respectively). Together with the wheat and maize data, these results indicate that TadA9-containing ABEs are active in both monocotyledonous and dicotyledonous plants.
Claims
1. In order, the following structural elements: a.) At least one N-terminal NLS sequence, b.) TadA9 adenosine deaminase domain or its functional variant, c.) At least one linker domain, preferably comprising or consisting of a hexaGGGGS linker according to Sequence ID No. 51, d.) dCas12a or a functional fragment thereof or nCas12a or a functional fragment thereof, comprising at least one mutation, wherein the at least one mutation results in increased activity and / or enhanced temperature tolerance, and one of the at least one mutations corresponds to a mutation in the dCas12a ortholog or homolog at a position homologous to D156 in SEQ ID NOs. 14, 15 or 16, E174 in SEQ ID NOs. 17, 18 or 19, and E184 in SEQ ID NOs. 20-28, respectively, wherein the at least one mutation results in increased activity and / or enhanced temperature tolerance, and in particular, the at least one mutation in the dCas12a ortholog or homolog corresponds to a D-to-R, E-to-R, or K-to-D / E mutation at a position homologous to SEQ ID NOs. 14-43 as a reference, e.) At least one C-terminal NLS sequence an adenine base editor (ABE) comprising, wherein the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence may be the same or different.
2. In order, the following structural elements: a.) At least one N-terminal NLS sequence, b.) An adenosine deaminase domain or a functional variant thereof selected from the TadA8, preferably TadA8e, or TadA9 domain, c.) At least one linker domain comprising or consisting of a hexaGGGGS linker according to Sequence ID No. 51, d.) dCas12a or its functional fragment or nCas12a or its functional fragment, e.) At least one C-terminal NLS sequence an adenine base editor (ABE) comprising, wherein the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence may be the same or different.
3. The dCas12a or nCas12a or the functional fragment thereof comprises at least one additional mutation as described in claim 1, one of which results in increased activity and / or enhanced temperature tolerance, and which corresponds to a mutation in the dCas12a ortholog or homolog at position D156 of SEQ ID NO: 14, 15 or 16, or position E174 of SEQ ID NO: 17, 18 or 19, or position E184 of SEQ ID NO: 20, 21, 22, 23, 24, 25, 26, 27 or 28, or at a position homologous to a homologous position in the Cas12a ortholog or homolog. Preferably, one of the at least one additional mutations that results in increased activity and / or temperature tolerance corresponds to D156R at a homologous position in the Cas12a ortholog or homolog compared to sequence number 14, 15, or 16 as the reference sequence, or to E174R at a homologous position in the Cas12a ortholog or homolog compared to sequence number 17, 18, or 19 as the reference sequence, or to E174R at a homologous position in the Cas12a ortholog or homolog, or one of the at least one additional mutations that results in increased activity and / or temperature tolerance corresponds to E184R at a homologous position in the Cas12a ortholog or homolog compared to sequence number 20, 21, 22, 23, 24, 25, 26, 27, or 28 as the reference sequence, More preferably, the at least one additional mutation corresponds to (i) D156R and D832A, or (ii) D156R and E925A, or (iii) D156R and D832A and E925A, or at homologous positions within the Cas12a ortholog or homolog, when compared to SEQ ID NO: 1 as a reference sequence, or when compared to a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the corresponding reference sequence, or at homologous positions within the Cas12a ortholog or homolog, or The at least one additional mutation described above corresponds to (iv) E174R and D908A, or (v) E174R and E993A, or (vi) E174R, D908A, and E993A, or at homologous positions within the Cas12a ortholog or homolog, when compared to SEQ ID NO: 2 as the reference sequence, or when compared to sequences having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the corresponding reference sequence, or at homologous positions within the Cas12a ortholog or homolog, to (iv) E174R and D908A, or (v) E174R and E993A, or (vi) E174R, D908A, and E993A, or The adenine base editor according to claim 1, wherein the at least one additional mutation corresponds to (viiii)E184R and D917A, or (ix)E184R and E1006A, or (x)E184R and D917A and E1006A, when compared to SEQ ID NO: 3 as a reference sequence, or when compared to a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the corresponding reference sequence, or at homologous positions within a Cas12a ortholog or homolog, respectively.
4. The adenine base editor according to claim 1, wherein the at least one N-terminal NLS sequence and / or the at least one C-terminal NLS sequence is selected from the triple SV40 NLS of SEQ ID NO: 52, the binocular SV40 NLS of SEQ ID NO: 53, the SV40 NLS of SEQ ID NO: 54, the FNLS of SEQ ID NO: 55, or the nucNLS of SEQ ID NO: 56, and preferably the at least one N-terminal NLS sequence and the at least one C-terminal NLS sequence are sequences having at least one binocular SV40 NLS of SEQ ID NO: 53 or its functional homolog, or sequences having at least 95%, 96%, 97%, 98%, or at least 99% sequence identity with SEQ ID NO:
53.
5. The adenosine deaminase domain is a TadA8e domain according to Sequence ID No. 57 or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% sequence identity with Sequence ID No. 57, or the adenosine deaminase domain is a TadA8e domain according to Sequence ID No. 57 or a sequence having at least 75%, 76%, 77%, 78%, 79%, 80 The adenine base editor according to claim 1, wherein the ndeaminase domain is a sequence having at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% sequence identity with TadA9 or sequence number 58 according to SEQ ID NO:
58.
6. A complex comprising the adenine base editor described in claim 1 and a functionally related form of guide RNA or a nucleic acid molecule encoding the guide RNA, wherein the guide RNA is specific to dCas12a or nCas12a described in claim 1, and optionally, the guide RNA is expressed from a construct comprising a truncated tRNA at the 5' end and at least one direct repeat structure at the 5'- and 3'- of the spacer RNA sequence or the sequence encoding the spacer RNA.
7. The complex according to claim 6, wherein the guide RNA is encoded by a scaffold structure provided by a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the sequence numbers 59, 60, or 61 or the corresponding reference sequences of each sequence number.
8. A nucleic acid molecule encoding the adenine base editor according to claim 1.
9. An expression construct or vector comprising a nucleic acid sequence according to the nucleic acid molecule described in claim 8, wherein the nucleic acid sequence encoding the adenine base editor and / or the nucleic acid sequence encoding the guide RNA (i) are present on the same expression construct or vector, or (ii) the nucleic acid sequence encoding the adenine base editor and / or the nucleic acid sequence encoding the guide RNA are present on at least two separate expression constructs or vectors, and optionally, an expression construct or vector encoding a guide RNA is present, the guide RNA is expressed from an RNA polymerase III promoter or an RNA polymerase II promoter, preferably the promoter is selected from U3, U6, H1 and ubiquitin promoters.
10. A cell comprising the adenine base editor described in claim 1.
11. A method for adenine base editing of a target site in the genome of interest in at least one cell of a prokaryote or eukaryote, including bacteria and archaea, (a) A step of providing at least one adenine base editor according to claim 1 to the at least one cell, (b) A step that optionally enables the construction of the complex described in claim 6 into a functional expression and / or a functionally related form, (c) The step of contacting the target genome of at least one cell with at least one functionally related form of a complex comprising at least one adenine base editor or at least one complex as described in claim 1 to obtain at least one modified cell, (d) A step of selectively selecting at least one of the modified cells, (e) A step of obtaining at least one cell having at least one adenine base edit in the target site This includes, but excludes, processes that alter the genetic identity of the human germline, the use of human embryos for industrial or commercial purposes, and processes that alter the genetic identity of animals, which may cause suffering to animals without providing substantial medical benefit to humans or animals, as well as animals resulting from such processes, and further excludes the treatment of the human or animal body by therapy, and optionally, (f) The process of regenerating at least one population of edited cells, tissues, organs, materials or an entire organism from the at least one edited cell. A method that includes this.
12. The method according to claim 11, wherein the at least one cell is derived from a plant, algae, yeast, or fungal organism, and preferably the at least one cell is a plant cell, preferably a plant cell belonging to the subkingdom Viridiplanta of the superfamily, particularly monocots and dicots including legumes used for fodder or animal feed, ornamental plants, food crops, trees, or shrubs.
13. Edited cells or tissues, organs, materials or whole organisms obtained or obtainable by the method of claim 11.
14. (a) comprising the adenine base editor described in claim 1, and (b) A container for containing the reaction components, including a buffer. This includes, and optionally, (c) Instructions for use A kit that includes this.
15. Use of the adenine base editor, complex, or expression construct or vector according to claim 1 for adenine base editing of a target site in the genome of interest in at least one cell of a prokaryote or eukaryote, including bacteria and archaea, for use to alter the genetic identity of a human germline, for use in human embryos for industrial or commercial purposes, and for use to exclude a process of altering the genetic identity of an animal that may cause suffering to an animal without providing substantial medical benefit to a human or animal, and also exclude an animal resulting from such a process, and further exclude the treatment of a human or animal body by therapy.