Cjcas9 mutants with relaxed pam compatibility
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2026-04-08
AI Technical Summary
The challenge lies in the limited packaging capacity of adeno-associated virus (AAV) vectors for delivering full-length Cas9 orthologues, such as SpCas9, SaCas9, and A/me2Cas9, which restricts their use in genome editing due to size constraints, and the stringent PAM recognition of smaller orthologues like CjCas9, limiting their application, especially for base and prime editing.
Development of the evoCjCas9 variant through phage-assisted continuous directed evolution, which relaxes PAM compatibility, allowing recognition of N4AH and N5HA sequences, and combines a compact size with enhanced nuclease activity, enabling efficient genome editing and base/prime editing using single AAV vectors.
evoCjCas9 achieves higher editing rates comparable to SpCas9 variants, with broader targeting range and reduced PAM stringency, facilitating efficient in vivo genome editing and base/prime editing in various cell types, including postmitotic primary cells.
Smart Images

Figure IMGF000041_0001 
Figure IMGF000006_0001 
Figure 00000049_0000
Abstract
Description
[0001] CjCAS9 mutants with relaxed PAM compatibility
[0002] This application claims the benefit of priority of European applications EP23175382.3 filed on 25 May 2023, and EP23203923.0 filed 16 October 2023, both of which are incorporated herein by reference.
[0003] Field
[0004] The present invention relates to variants of Campylobacter jejuni CAS9 variants having relaxed PAM requirements, and fusion proteins comprising same, as well as nucleic acids encoding such variants.
[0005] Background
[0006] Genome editing with RNA-guided programmable nucleases has transformed biomedical research and holds promise for therapeutic application in patients with genetic diseases. Cas9 endonucleases of the type II CRISPR immune defence system are the most widely used RNA- guided nucleases for genome editing. They are paired with programmable single guide (sg)RNAs, which contain a spacer sequence that targets the complementary protospacer on genomic DNA in case a downstream protospacer adjacent motif (PAM) is present. Binding of Cas9 endonucleases to target sites subsequently induces double strand (ds)DNA breaks, which lead to the installation of insertion / deletion (indel) mutations or precise editing from DNA donor templates via homologous recombination. More recently, the fusion of deaminases or reverse transcriptases (RTs) to catalytically impaired Cas9 led to the establishment of base editors (BEs) and prime editors (PEs). Importantly, these Cas9-based genome editing tools work independent of dsDNA break formation and homologous recombination and therefore enable precise editing in all cell types, including postmitotic primary cells.
[0007] Cas9 of Streptococcus pyogenes (SpCas9) was the first RNA-programmable endonuclease employed for genome editing, and due to its high activity and relatively simple PAM sequence (NGG) remains the most widespread Cas9 ortholog used to date. The large size of SpCas9 (1 ,368 aa), however, presents a challenge for efficient in vivo delivery, in particular if adeno- associated virus (AAV) vectors are used. AAVs are among the most promising nucleic acid delivery vehicles for the clinics, as they show low rates of genome integration and immunogenicity, and as they cover a broad range of tissue tropisms. Nevertheless, their limited cargo capacity of ~5 kb (including ITRs) prohibits the packaging of full-length SpCas9 together with its sgRNA on a single AAV vector. One possible solution to this problem is the replacement of SpCas9 with more compact Cas9 orthologues. Recent studies therefore employed Cas9 of Staphylococcus aureus (SaCas9 - 1 ,053 aa) and Neisseria Meningitidis (A / me2Cas9 - 1 .080 aa) and packaged them as full-length nucleases or BEs together with their sgRNAs on single AAV vectors. Nevertheless, these constructs were still close to the packaging limit of AAV, and for BEs tissue-specific promoters had to be replaced by ubiquitous, minimal promoters (EF-1a or the nuclear RNA (U1a) promoter).
[0008] With 984 aa, Campylobacter jejuni (Cy)Cas9 is the smallest known Cas9 orthologue, making it a particularly attractive candidate for in vivo genome editing applications. However, similar to other smaller Cas9 orthologues, CjCas9 recognizes a relatively complex PAM sequence (N3VRYAC), which occurs less frequently in the genome than the NGG motif of SpCas9. This strongly restrains the application of CjCas9, in particular for base- and prime editing where Cas9 has to bind DNA in a precisely defined distance to the targeted base.
[0009] The EMBL database discloses a Campylobacter jejuni CAS9 variant under accession number EAI4715943. The Uniprot database discloses a Campylobacter jejuni CAS9 variant under accession numbers A0A690ANS8 and A0A3STLQ3.
[0010] Nakagawa et al. report on an engineered Campylobacter jejuni Cas9 variant with enhanced activity and broader targeting range (Communications Biology, vol. 5, no. 1 , 8 March 2022 (2022-03-08).
[0011] Based on the above-mentioned state of the art, the objective of the present invention is to provide C / CAS9 variants with lower PAM recognition stringency. This objective is attained by the subjectmatter of the independent claims of the present specification, with further advantageous embodiments described in the dependent claims, examples, figures and general description of this specification.
[0012] Summary of the Invention
[0013] The inventors employed phage-assisted continuous directed evolution to broaden the PAM compatibility of CjCas9, the smallest Cas9 orthologue characterized to date. The identified variant, termed evoCjCas9, recognizes N4AH and N5HA PAM sequences, which occur ten-fold more frequently in the genome than the canonical N3VRYAC PAM site. evoCjCas9 also exhibits a higher nuclease activity than wild type CjCas9 on canonical PAMs, with editing rates that are comparable to commonly used SpCas9 variants. When fused to deaminases or reverse transcriptases, evoCjCas9 enables robust base- and prime editing, with the small size of evoCjCas9 base editors also allowing efficient installation of A-to-G and C-to-T transition mutations in the murine liver using single adeno-associated virus (AAV) vector systems. By combining a compact size with a broad targeting range, evoCjCas9 is a powerful addition to the CRISPR-Cas9 toolbox, ideally suited for translational genome editing using single-vector systems.
[0014] A first aspect of the invention relates to a Campylobacter jejuni Cas9 (CjCas9) variant, particularly a CjCas variant described by SEQ ID NO 001 , which is characterized by the mutations E789K.
[0015] An alternative first aspect of the invention relates to a CjCas9 variant, particularly a CjCas variant described by SEQ ID NO 001 , which is characterized by the mutations S951 G. Another aspect of the invention relates to a polypeptide comprising a C / CAS9 variant according to the invention that further comprises a nucleotide deaminase activity.
[0016] The invention further relates to polynucleotide sequences encoding the variant polypeptides provided by the invention.
[0017] Terms and definitions
[0018] For purposes of interpreting this specification, the following definitions will apply and whenever appropriate, terms used in the singular will also include the plural and vice versa. In the event that any definition set forth below conflicts with any document incorporated herein by reference, the definition set forth shall control.
[0019] The terms “comprising”, “having”, “containing”, and “including”, and other similar forms, and grammatical equivalents thereof, as used herein, are intended to be equivalent in meaning and to be open-ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items. For example, an article “comprising” components A, B, and C can consist of (i.e., contain only) components A, B, and C, or can contain not only components A, B, and C but also one or more other components. As such, it is intended and understood that “comprises” and similar forms thereof, and grammatical equivalents thereof, include disclosure of embodiments of “consisting essentially of or “consisting of.”
[0020] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0021] Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X.”
[0022] As used herein, including in the appended claims, the singular forms “a”, “or” and “the” include plural referents unless the context clearly dictates otherwise.
[0023] "And / or" where used herein is to be taken as specific recitation of each of the two specified features or components with or without the other. Thus, the term "and / or" as used in a phrase such as "A and / or B" herein is intended to include "A and B," "A or B," "A" (alone), and "B" (alone). Likewise, the term "and / or" as used in a phrase such as "A, B, and / or C" is intended to encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone). Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art (e.g., in cell culture, molecular genetics, nucleic acid chemistry, hybridization techniques and biochemistry, organic synthesis). Standard techniques are used for molecular, genetic, and biochemical methods (see generally, Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed. (2012) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. and Ausubel et al., Short Protocols in Molecular Biology (2002) 5th Ed, John Wiley & Sons, Inc.) and chemical methods.
[0024] Any patent document cited herein shall be deemed incorporated by reference herein in its entirety.
[0025] Sequences
[0026] Sequences similar or homologous (e.g., at least about 85% sequence identity) to the sequences disclosed herein are also part of the invention. In some embodiments, the sequence identity at the amino acid level can be about 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher. At the nucleic acid level, the sequence identity can be about 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher. Alternatively, substantial identity exists when the nucleic acid segments will hybridize under selective hybridization conditions (e.g., very high stringency hybridization conditions), to the complement of the strand. The nucleic acids may be present in whole cells, in a cell lysate, or in a partially purified or substantially pure form.
[0027] In the context of the present specification, the terms sequence identity and percentage of sequence identity refer to a single quantitative parameter representing the result of a sequence comparison determined by comparing two aligned sequences position by position. Methods for alignment of sequences for comparison are well-known in the art. Alignment of sequences for comparison may be conducted by the local homology algorithm of Smith and Waterman, Adv. Appl. Math. 2:482 (1981 ), by the global alignment algorithm of Needleman and Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson and Lipman, Proc. Nat. Acad. Sci. 85:2444 (1988) or by computerized implementations of these algorithms, including, but not limited to: CLUSTAL, GAP, BESTFIT, BLAST, FASTA and TFASTA. Software for performing BLAST analyses is publicly available, e.g., through the National Center for Biotechnology-Information (http: / / blast.ncbi.nlm.nih.gov / ).
[0028] One example for comparison of amino acid sequences is the BLASTP algorithm that uses the default settings: Expect threshold: 10; Word size: 3; Max matches in a query range: 0; Matrix: BLOSUM62; Gap Costs: Existence 11 , Extension 1 ; Compositional adjustments: Conditional compositional score matrix adjustment. One such example for comparison of nucleic acid sequences is the BLASTN algorithm that uses the default settings: Expect threshold: 10; Word size: 28; Max matches in a query range: 0; Match / Mismatch Scores: 1 .-2; Gap costs: Linear. Unless stated otherwise, sequence identity values provided herein refer to the value obtained using the BLAST suite of programs (Altschul et al., J. Mol. Biol. 215:403-410 (1990)) using the above identified default parameters for protein and nucleic acid comparison, respectively. Reference to identical sequences without specification of a percentage value implies 100% identical sequences (i.e. the same sequence).
[0029] Nucleic acid sequence positions are indicated according to the IUPAC code: The term having substantially the same biological activity in the context of the present invention relates to either one or both main functions of a nucleic acid sequence editing protein, i.e. Indel, Base or Prime editing) in percent of edited sequence determined via High-throughput sequencing (HTS, see Figures 4b and the description relating thereto.
[0030] General Biochemistry: Peptides, Amino Acid Sequences The term polypeptide in the context of the present specification relates to a molecule consisting of 50 or more amino acids that form a linear chain wherein the amino acids are connected by peptide bonds. The amino acid sequence of a polypeptide may represent the amino acid sequence of a whole (as found physiologically) protein or fragments thereof. The term "polypeptides" and "protein" are used interchangeably herein and include proteins and fragments thereof. Polypeptides are disclosed herein as amino acid residue sequences.
[0031] Amino acid residue sequences are given from amino to carboxyl terminus. Capital letters for sequence positions refer to L-amino acids in the one-letter code (Stryer, Biochemistry, 3rded. p. 21 ). Lower case letters for amino acid sequence positions refer to the corresponding D- or (2R)- amino acids. Sequences are written left to right in the direction from the amino to the carboxy terminus. In accordance with standard nomenclature, amino acid residue sequences are denominated by either a three letter or a single letter code as indicated as follows: Alanine (Ala, A), Arginine (Arg, R), Asparagine (Asn, N), Aspartic Acid (Asp, D), Cysteine (Cys, C), Glutamine (Gin, Q), Glutamic Acid (Glu, E), Glycine (Gly, G), Histidine (His, H), Isoleucine (lie, I), Leucine (Leu, L), Lysine (Lys, K), Methionine (Met, M), Phenylalanine (Phe, F), Proline (Pro, P), Serine (Ser, S), Threonine (Thr, T), Tryptophan (Trp, W), Tyrosine (Tyr, Y), and Valine (Vai, V).
[0032] The term variant refers to a polypeptide that differs from a reference polypeptide, but retains essential properties. A typical variant of a polypeptide differs in its primary amino acid sequence from another, reference polypeptide. Generally, differences are limited so that the sequences of the reference polypeptide and the variant are closely similar overall and, in many regions, identical. A variant and reference polypeptide may differ in amino acid sequence by one or more modifications (e.g., substitutions, additions, and / or deletions). A substituted or inserted amino acid residue may or may not be one encoded by the genetic code. A variant of a polypeptide may be naturally occurring such as an allelic variant, or it may be a variant that is not known to occur naturally.
[0033] The term homologue in the context of the present specification relates to a functional polypeptide having a sequence identity of 85 % or more with SEQ ID NO 002.
[0034] General Molecular Biology: Nucleic Acid Sequences, Expression
[0035] The term transgene in the context of the present specification relates to a gene or genetic material that has been transferred from one organism to another. In the present context, the term may also refer to transfer of the natural or physiologically intact variant of a genetic sequence into tissue of a patient where it is missing. It may further refer to transfer of a natural encoded sequence the expression of which is driven by a promoter absent or silenced in the targeted tissue.
[0036] The term recombinant in the context of the present specification relates to a nucleic acid, which is the product of one or several steps of cloning, restriction and / or ligation and which is different from the naturally occurring nucleic acid. A recombinant virus particle comprises a recombinant nucleic acid.
[0037] The terms gene expression or expression, or alternatively the term gene product, may refer to either of, or both of, the processes - and products thereof - of generation of nucleic acids (RNA) or the generation of a peptide or polypeptide, also referred to transcription and translation, respectively, or any of the intermediate processes that regulate the processing of genetic information to yield polypeptide products. The term gene expression may also be applied to the transcription and processing of a RNA gene product, for example a regulatory RNA or a structural (e.g. ribosomal) RNA. If an expressed polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell. Expression may be assayed both on the level of transcription and translation, in other words mRNA and / or protein product.
[0038] The term sgRNA (single guide RNA) in the context of the present specification relates to an RNA molecule capable of sequence-specific repression of gene expression via the CRISPR (clustered regularly interspaced short palindromic repeats) mechanism.
[0039] The term nucleic acid expression vector in the context of the present specification relates to a plasmid, a viral genome or an RNA, which is used to transfect (in case of a plasmid or an RNA) or transduce (in case of a viral genome) a target cell with a certain gene of interest, or -in the case of an RNA construct being transfected- to translate the corresponding protein of interest from a transfected mRNA. For vectors operating on the level of transcription and subsequent translation, the gene of interest is under control of a promoter sequence and the promoter sequence is operational inside the target cell, thus, the gene of interest is transcribed either constitutively or in response to a stimulus or dependent on the cell’s status. In certain embodiments, the viral genome is packaged into a capsid to become a viral vector, which is able to transduce the target cell.
[0040] Detailed Description of the Invention
[0041] The invention generally relates to polypeptide sequences derived of Campylobacter jejuni Cas9 comprising one mutation selected from E789K and S951G, or both E789K and S951G, the numeration following the numeration of SEQ ID NO 001 .
[0042] A first aspect of the invention relates to a Campylobacter jejuni Cas9 (CjCas9) variant, particularly a CjCas variant described by SEQ ID NO 001 , which is characterized by the mutations E789K.
[0043] An alternative first aspect of the invention relates to a CjCas9 variant, particularly a CjCas variant described by SEQ ID NO 001 , which is characterized by the mutations S951 G.
[0044] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G.
[0045] The position of the mutation is calculated relative to its position in SEQ ID NO 001 , i.e. if a homologue with substantially identical sequence, but shorter or longer N-terminal sequence is used, the given numeration refers to the aligning positions of SEQ ID NO 001 . The CjCas9 variant may be identical in all other positions to SEQ ID NO 001. Alternatively, the CjCas9 variant may comprise one, two, three or more of the mutations further described herein, and may otherwise be may be identical in all other positions to SEQ ID NO 001 . Yet alternatively, the CjCas9 variant may be only equal or more than (>) 85%, >90%, >92%, %, >95%, >96%, %, >97%, >98%, identical to SEQ ID NO 001 , and comprise at least one, particularly both of the mutations E789K and S951 G, and optionally further L58Y, N821 K, and D900K. In addition to Campylobacter jejuni Cas9 orthologues, Neisseria meningitidis Cas9 orthologues share a high degree of identity with SEQ ID NO 001 and are expected to provide the same benefits as the sequence of CjCAS9-1 of SEQ ID NO 001 (see Chen et al, The CRISPR Journal Vol. 5(3); https: / / doi.Org / 10.1089 / crispr.2021 .0143.
[0046] Any of the Cas variants disclosed herein may be associated to a sgRNA molecule guiding its activity to a specific sequence.
[0047] D900
[0048] In certain embodiments, the CjCas9 variant according to the invention further comprises a mutation D900K or D900R. In certain particular embodiments, the CjCas9 variant further comprises the mutation D900K. This is described in WO2022045169A1 (Univ Tokyo; priority date 25 Aug 2020; also Nakagawa et al., Communications Biology Vol. 5, Article number: 211 (2022)).
[0049] In certain embodiments, the CJCas9 variant according to the invention comprises the mutations E789K and D900K.
[0050] In certain embodiments, the CJCas9 variant according to the invention comprises the mutations S951G and D900K.
[0051] In certain particular embodiments, the CJCas9 variant according to the invention comprises both E789K and S951G, and D900K.
[0052] L58
[0053] In certain embodiments, the CJCas9 variant according to the invention further comprises a substitution of L58 by Y or Q, particularly L58Y.
[0054] The inventors found that the evolution described in the examples also produced L58Q. The most favourable results in later testing were delivered by L58Y. The authors of the publication that first cites this position as interesting (Nakagawa et al., ibid.) suggest substituting leucine at amino acid number 58 by one of tyrosine, tryptophan, phenylalanine, or isoleucine.
[0055] In certain embodiments, L58 is substituted by an amino acid selected from the group consisting of tyrosine, tryptophan, phenylalanine, and isoleucine. In certain particular embodiments, L58 is substituted by tyrosine.
[0056] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations E789K and L58Y.
[0057] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations S951G and L58Y.
[0058] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G, and L58Y. In certain embodiments, the CjCas9 variant according to the invention comprises the mutations E789K, L58Y and D900K.
[0059] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations S951G, L58Y and D900K.
[0060] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G, L58Y and D900K.
[0061] N821
[0062] In certain embodiments, the CjCas9 variant according to the invention further comprises a substitution of N821 , particularly N821 K.
[0063] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations E789K and N821 K.
[0064] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations S951G and N821 K.
[0065] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G, and N821 K.
[0066] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations E789K, L58Y and N821 K.
[0067] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations S951G, L58Y and N821 K.
[0068] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G, L58Y and N821 K.
[0069] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations E789K, D900K and N821 K.
[0070] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations S951G, D900K and N821 K.
[0071] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G, D900K and N821 K.
[0072] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations E789K, L58Y, D900K and N821 K.
[0073] In certain embodiments, the CjCas9 variant according to the invention comprises the mutations S951G, L58Y, D900K and N821 K.
[0074] In certain particular embodiments, the CjCas9 variant according to the invention comprises both E789K and S951G, L58Y, D900K and N821 K. Fig. 2 shows the improvement, relative to the WT sequence, of the variant comprising a full set of muations (L58Y, E789K, N821 K, D900K, S951G). Single mutations tested for PAM relaxation that resulted in a significant effect included E789K, E871 K, D891 N, E946K and S951G, which all resulted in PAM relaxation. The effect was particularly significant for a single E789K or S951 G mutation (data not shown).
[0075] In certain embodiments, the CJCas9 variant according to the invention comprising any of the combinations recited above, further comprises one or more mutations selected from the group of G323V, D376Y, F434C, K622N, D663E and T671A.
[0076] The evolution performed in arriving at the present invention selected for mutations beneficial in CRISPR activation (CRISPRa). This is an application often used in research. Further work performed in the examples centred on Cas9 variants able to cut DNA (which is a function that cannot be screened in PACE). The mutations G323V, D376Y, F434C, K622N, D663E and T671A impacted nuclease activity, and thus were not further pursued. The inventors however do expect these mutations to be advantageous in applications where a nuclease deficient Cas9 is required (e.g. CRISPR activation & inhibition, epigenome editing).
[0077] In certain embodiments, the CJCas9 variant according to the invention comprising any of the combinations recited above, further comprises one or more mutations selected from the group of E871 K, D891 N and E946K.
[0078] In certain embodiments, the CJCas9 variant according to the invention comprises the sequence of SEQ ID NO 002, or a sequence having at least 95% identity to SEQ ID NO 002 and bears the mutations L58Y, E789K, N821 K, D900K, S951G. The variant according to these embodiments defined by identity to SEQ ID NO 002 is characterized by a less stringent PAM requirement (defined as functional PAM sites in a random sequence) than SEQ ID NO 001.
[0079] A functional PAM can be defined as a PAM site that allows clearly distinguishable editing efficiency (comprising at least 50% efficiency of a protein most efficiently recognized PAM).
[0080] PAM stringency can be indirectly quantified by calculating the probability of PAM occurrence within a perfectly random DNA sequence of length PAM and considering only one strand direction (see Example 1 ). This parameter is quantified in Fig. 3d as PAM availability in %. For example, an NGN PAM would correspond to 25%. For evoCJCas9, the inventors determined this value to be 19.34%. For further information, see the caption of Fig. 3d and Example 1 .
[0081] The determination whether a potential PAM sequence is actually recognized by the protein is a parameter for which different methods are available. The value for this parameter used in the context of this specification (at least 50% efficiency compared to the most active PAM site of a Cas9) may be determined experimentally in a number of ways (e.g. HT-PAMDA or HTS (both were performed in our manuscript). The method described in Example 1 is used when in doubt. Once it has been defined when a PAM sequence is considered as recognized, one can (mathematically) determine a value for PAM relaxation. The inventors quantified this as the probability of PAM occurrence within a (perfectly) random (ss)DNA sequence of length PAM.
[0082] Editing enzymes
[0083] Another aspect of the invention relates to a polypeptide comprising a C / CAS9 variant according to any one of the preceding embodiments of the first aspect of the invention, wherein the polypeptide further comprises a nucleotide deaminase activity.
[0084] In certain embodiments, the deaminase is an adenosine deaminase activity.
[0085] A great number of deaminases offer themselves for linking to the CAS9 variant of the invention to construct a base editing enzyme. The inventors hypothesize, without wanting to be bound by theory, that deaminases characterized by a high kinetic rate of activity may have a favourable effect in combining with the CAS9 variant of the invention, as a higher kinetic rate is expected to compensate for the presumably shorter residency time resulting from the reduction in PAM stringency.
[0086] In certain embodiments, the deaminase is tadA, an essential tRNA-specific adenosine deaminase from Escherichia coli.
[0087] In certain embodiments, the deaminase is tadA 7.10, an evolved essential tRNA-specific adenosine deaminase from Escherichia coli, see Richter et al., Nat Biotechnol. 2020 Jul; 38(7): 883-891.
[0088] In certain embodiments, the deaminase is tadA 8, an evolved essential tRNA-specific adenosine deaminase from Escherichia coli, see Yan, Molecular Plant 14, 722-731 (2021 ).
[0089] In certain embodiments, the deaminase is a cytidine deaminase activity.
[0090] In certain embodiments, the deaminase is APOBEC1 , a human cytidine deaminase (Uniprot P41238).
[0091] In certain embodiments, the deaminase is TadCBEd, a human cytidine deaminase evolved from tadA. See Neugebauer et al. Nat Biotechnol (2022). https: / / doi.org / 10.1038 / s41587-022-01533-6.
[0092] In certain embodiments, the polypeptide comprising a CjCAS9 variant according to the preceding aspect further comprises a reverse transcriptase activity.
[0093] In certain embodiments, the polypeptide comprising a CjCAS9 variant according to the preceding aspect does not comprise an RNAseH activity.
[0094] Polynucleotide sequences
[0095] Another aspect of the invention relates to a polynucleotide sequence encoding a CJCas9 variant or a polypeptide comprising a CjCAS9 variant, as specified in any one of the preceding aspects and embodiments.
[0096] The polynucleotide sequence according to this aspect of the invention may be a DNA sequence. The DNA sequence may be a “naked” DNA sequence, for example an expression plasmid sequence or a microchromosome. The DNA sequence may alternatively be a DNA viral sequence.
[0097] In other embodiments, the polynucleotide sequence according to this aspect of the invention may be a RNA sequence.
[0098] The RNA sequence may be a “naked” RNA sequence, for example an mRNA sequence comprised of “natural” or synthetic RNA. The RNA sequence may alternatively be a RNA viral sequence.
[0099] The polynucleotide sequence encoding a CJCas9 variant may be present in a recombinant cell under control of a promoter sequence operable in that cell to provide transgene expression of the variant or editing enzyme that the variant may be part of.
[0100] The polynucleotide sequence encoding a CJCas9 variant may be present in a viral vector capable of specifically entering a target cell to deliver the variant or editing enzyme that the variant may be part of.
[0101] Another aspect of the invention relates to a nucleic acid expression vector comprising a polynucleotide sequence according to the above specification under control of a promoter operable in a target cell. In particular embodiments, the expression vector is an AAV virus particle.
[0102] In a more particular embodiment, the expression vector is an AAV9 virus particle.
[0103] The inventors illustrated the invention in AAV variant AAV9, which in their hands was easily produced and efficient in the experiments performed. The skilled person is aware how to select an AAV variant based on its specificity for a particular target tissue. The inventors predict that the invention can be easily practiced in any AAV variant available depending on the application at hand.
[0104] The invention further encompasses the following items:
[0105] 1 . A CJCas9 variant of SEQ ID NO 001 , characterized by at least one of the following mutations: E789K and / or S951 G.
[0106] 2. The CJCas9 variant according to item 1 , further comprising the mutation D900K or D900R, particularly D900K.
[0107] 3. The CJCas9 variant according to any one of the preceding items, further comprising a substitution of L58 by Y or Q, particularly L58Y.
[0108] 4. The CJCas9 variant according to any one of the preceding items, further comprising a substitution of N821 , particularly N821 K.
[0109] 5. The CJCas9 variant according to any one of the preceding items, comprising both E789K and S951G.
[0110] 6. The CJCas9 variant according to any one of the preceding items, further comprising a substitution selected from the group of G323V, D376Y, F434C, K622N, D663E and T671A. 7. The CJCas9 variant according to any one of the preceding items, further comprising a substitution selected from E871 K, D891 N and E946K.
[0111] 8. The CJCas9 variant according to any one of the preceding items, comprising the sequence of SEQ ID NO 002.
[0112] 9. The CJCas9 variant according to any one of the preceding items 1 to 8, comprising a sequence having at least 95% identity, particularly > 97%, more particular > 98% identity to SEQ ID NO 002, wherein the sequence bears the mutations E789K and S951G, and wherein the cjCas9 variant is characterized by having a less stringent PAM requirement than SEQ ID NO 001 .
[0113] 10. The CJCas9 variant according to item 9, wherein the sequence further bears the mutation L58Y.
[0114] 11 . The CJCas9 variant according to item 9 or 10, wherein the sequence further bears the mutation N821 K.
[0115] 12. The CJCas9 variant according to item 9, wherein the sequence bears all the mutations L58Y, E789K, N821 K, D900K, S951 G.
[0116] 13. A polypeptide comprising a C / CAS9 variant according to any one of the preceding items, wherein the polypeptide further comprises a nucleotide deaminase activity.
[0117] 14. The polypeptide according to item 13, wherein the deaminase is selected from: an adenosine deaminase activity; or a cytidine deaminase activity.
[0118] 15. A polypeptide comprising a CjCAS9 variant according to any one of the preceding items 13 to 14, wherein the polypeptide further comprises a reverse transcriptase activity.
[0119] 16. The polypeptide according to item 15, wherein the polypeptide does not comprise an RNAseH activity.
[0120] 17. A polynucleotide sequence encoding a CJCas9 variant or a polypeptide comprising a CjCAS9 variant, as specified in any one of the preceding items.
[0121] 18. An expression vector comprising a polynucleotide according to item 17, under control of a promoter.
[0122] 19. The expression vector according to item 18, wherein the expression vector is an AAV virus particle.
[0123] Wherever alternatives for single separable features are laid out herein as “embodiments”, it is to be understood that such alternatives may be combined freely to form discrete embodiments of the invention disclosed herein. The invention is further illustrated by the following examples and figures, from which further embodiments and advantages can be drawn. These examples are meant to illustrate the invention but not to limit its scope.
[0124] Description of the Figures
[0125] Fig. 1 | Phage-assisted continuous and non-continuous evolution (PACE) to broaden the PAM compatibility of CJCas9. a, Schematics of the PACE experiment. Selection phages (SP) infect E. coli host cells and use the hosts machinery to replicate and translate required phage genes. Instead of gene III, SP carries the ω-subunit-dCJCas9 fusion gene. Cas9 recognition of a PAM and protospacer sequence on the accessory plasmid (AP) allows recruitment of endogenous E. coli RNA polymerase via the ω-subunit and subsequent translation of gene III (pill). Efficient recognition of the PAM and protospacer sequence results in infectious phages that are able to sustain in the lagoon, whereas weak or lack of recognition results in phage wash-out. The mutation rate during phage replication is controlled and elevated via arabinose-inducible expression of mutagenesis genes from a mutagenesis plasmid (DP6). b, Illustrative overview of 19 rounds of PANCE using accessory plasmids containing different PAM sites (according to IUPAC notation). Phages were increasingly diluted from one round to the next. Grey indicates presence of phages, white indicates absence. Reference sequences are given as SEQ ID NO 054 and 055. c, Long- read sequencing of individual phages (n = 811 ) at 8 different time points during PACE. For each time point, the mutation frequency and amino acid position is shown. Amino acid positions substituted in more than 50% of phage genotypes are indicated.
[0126] Figure 2 | Characterization of the PAM compatibility of evoCjCas9. a, HT-PAMDA characterization of CJCas9 and evoCJCas9 illustrating their PAM preference at PAM positions 4 to 8. The logw rate constant represents the mean of two replicates against two distinct spacer sequences, b, Protein structure of CJCas9 with evoCjCas9 mutations introduced and highlighted (left, based on PDB: 5X2H). Detailed view illustrating differences between CJCas9 (top panels) and evoCjCas9 (bottom panels) residues interacting with PAM nucleotides of the target strand.
[0127] Figure 3 | Indel formation rates of evoCjCas9 compared to CjCas9 and other commonly used RNA-guided endonucleases, a, Illustration and size comparison of the expression vectors encoding for different Cas enzymes, b, Observed indel frequencies for CJCas9 and evoCjCas9 on 8 different PAM sites, c, Observed indel frequencies for different commonly used RNA-guided endonucleases and PAM relaxed Cas variants on different target sites with the most optimal PAM for each nuclease. Colors represent different timepoints, d, PAM frequency of the different RNA- guided endonucleases, illustrated as the probability of encountering a PAM on a randomly selected target site in a random DNA sequence (in %). Calculations are based on TTGAT for TnpB; TTTR for CasMINIv3.1 ; PAM sequences with a kinetic rate constant >10-3for CJCas9 variants; N2GRRT for SaCas9; N3RRT for SaKKH; N2GG for Saur / Cas9; N2RG for SauriKKH;
[0128] N4CC for Nme2Cas9; N4C for eNme2-C.NR; NGG for SpCas9; NG for SpG and NR for SpRY. e, Mean indel frequencies of CJCas9-variants at on-target sites (0 MM) and potential off-target sites containing up to 5 mismatches (MM), f, Indel frequencies of evoCJCas9 at sites with dinucleotide mismatches normalized to the on-target site. Dinucleotides highlighted in red represent mismatch positions. PAM at position 23-30. Number above bars indicate amount of different target sites, g, Relative indel frequencies evaluated by HTS at off-target sites detected by CHANGE-seq for two sgRNAs. Boxplots in (b, c, e) represent the 25th, 50th, and 75thpercentiles. Whiskers indicate 5 and 95 percentiles. Numbers (n) above plots indicate number of datapoints. Bars in (f, g) represent mean, with error bars representing standard error of the mean, b, c, e, f, g, Indels refer to insertion, deletions, or substitutions. The sequences shown in panel g relate to the SEQ ID NOs: 28-53, 56.
[0129] Figure 4 | Base- and Prime editing with CjCas9 and evoCjCas9. a, Illustration of expression vectors encoding for different CjCas9 BEs. b, Observed nucleotide transitions within the target sites with canonical and non-canonical PAMs. Bars represent mean of three independent biological replicates on N4ACAC, N4ATAC, N4GTAC, and N4GCAC (canonical) or N4AAAC, N4CAAC, N4GAAC and N4TAAC (non-canonical) PAM sites. Number of target sites (n) is indicated within each plot. Protospacer positions from -9 to 22 are shown, c, Illustration of expression vectors encoding for CjCas9 prime editors with the full length (top) or RnaseH depleted M-MLV reverse transcriptase (bottom), d, Comparison of CjCas9 or evoCjCas9 PEs installing an A to G transition at the AAVS1 locus, e, Mean editing efficiencies of CjCas9 or evoCjCas9-PEΔRnHon a selection of loci with canonical or non-canonical PAM sites. The loci of the target sites and type of edits are indicated. Bars (b, d, e) represent mean of biological replicates in HEK293T cells, with error bars indicating standard error of the mean, f, Mean prime editing efficiencies and indel rates on 64 (canonical) and 172 (non-canonical) integrated target sites. Boxplots represent the 25th, 50th, and 75thpercentiles. Whiskers indicate 5 and 95 percentiles. CjCas9 in blue, evoCjCas9 in red (b, d - f).
[0130] Figure 5 | In vivo genome editing with compact evoCjCas9 adenine and cytosine BE. a, Schematic representation of single AAV vectors used for evoCjCas9 base editing. Elements are not illustrated to scale, b, Illustration of the experimental workflow for in vivo adenine or cytosine base editing at the Pcsk9 locus in the liver and of the Gpr6 locus in the brain, c, d, A-to-G editing at the targeted Pcsk9 splice site with different AAV concentrations, analyzed in isolated hepatocytes, whole liver samples, the tail (c, same legend as in h) and other tissues (d). e, Plasma PCSK9 levels relative to untreated control as determined by ELISA. ***P - 0.0006, **P - 0.0029, ***P = 0.0001 , ****P < 0.0001 (left to right), f, Plasma LDL cholesterol levels. P = 0.2461 , *P = 0.0278, *P - 0.027 (left to right), g, A-to-G editing at the targeted site in the Gpr6 locus (non- canonical CCGCCAAC PAM) in isolated brain tissues for CjCas9 and evoCjCas9. h, C-to-T editing at the targeted Pcsk9 site with evoCjCas9 cytosine base editor (eAID) in isolated hepatocytes, whole liver samples and the tail, i, Plasma PCSK9 levels relative to untreated control as determined by ELISA. *P - 0.0286. j, Plasma LDL cholesterol levels. **P - 0.0095. Means were compared using one-way ANOVA with Dunnett correction (e,f) or one-tailed t-test (Mann-Whitney) (i, j) . Bars (c-j) represent mean with error bars representing standard error of the mean. Individual data points as black dots, vg, vector genomes, chi, chimeric Intron, ICV, intracerebroventricular.
[0131] Figure 6 | Base editing with CjCas9 and evoCjCas9. Observed nucleotide transitions of adenine BEs (top) and cytosine BEs (bottom panel) with CjCas9 and evoCjCas9 on 8 different PAM sites (with >20 target sites per PAM and a total of 254 target sites for each architecture). The BE architecture is indicated in rightmost panels. Bars represent mean of three independent biological replicates, with error bars representing standard error of the mean.
[0132] Figure 7| Contribution of consensus mutations in CjCas9 on PAM relaxation. HT-PAMDA characterization of enCjCas9 (top panel) with introduced PAM relaxing mutations (lower panels) illustrating their PAM preference at PAM positions 5 to 8. The logw rate constant represents the mean of two replicates against two distinct spacer sequences.
[0133] Figure 8| Indel formation with CjCas9 and evoCjCas9 on endogenous target sites, a, b, Indel frequencies of CjCas9 and evoCjCas9 on 14 non-canonical and 10 canonical PAM sites in HEK293T cells with selection harvested 9 days (a) or without selection 4 days (b) post transfection. Bars (a, b) represent mean with error bars representing standard error of the mean. Individual data points as black dots.
[0134] Figure 9| Indel formation rates with CjCas9 and evoCjCas9 at target sites with different PAMs. Mean indel frequency in self-targeting libraries on canonical (a) and non-canonical (b) PAM sites. Bars represent mean with error bars representing standard error of the mean, with individual data points of biological triplicates shown as black dots (a) or indicated above error bars (b). The in vitro PAMDA assay (Fig. 2a) revealed few motifs recognized by CjCas9 that lie outside of the simplified N3VRYAC motif (e.g. N3MACAN), albeit with low efficiency. We define sequences that fall outside of N3VRYAC as non-canonical motifs for the purpose of simplification.
[0135] Figure 10| Prime editing rates with CjCas9 and evoCjCas9 on endogenous target sites. Prime editing and corresponding unintended editing frequencies of CjCas9 and evoCjCas9-PEΔRnHon 3 canonical and 7 non-canonical target sites in HEK293T cells. Bars represent mean with error bars representing standard error of the mean. Individual data points as black dots.
[0136] Figure 111 Cytosine base editing with CjCas9 and evoCjCas9 on sites in the Pcsk9 locus. Editing frequency (C-to-T) with eAID or TadCDd CBEs with one or two C-terminal UGI domains in N2a cells. Bars represent mean with error bars representing standard error of the mean. Individual data points as black dots.
[0137] Figure 12| In vivo genome editing with the evoCjCas9-TadCDd cytosine BE. Observed editing at the targeted Pcsk9 site with an AAV construct containing the evoCjCas9-TadCDd base editor. Bars represent mean with error bars representing standard error of the mean, with individual data points shown as black dots. Examples
[0138] The examples show phage-assisted continuous directed evolution (PACE) to generate evoCJCas9. This CjCas9 variant obtained thereby recognizes N4AH and N5HA PAM sequences, which occur ten-times more frequently in the genome than the canonical N3VRYAC PAM sequence. evoCjCas9 also possesses a higher nuclease activity than wild type CjCas9 and enables robust base- and prime editing at canonical and non-canonical PAM sites. Proof-of-concept is provided for tissuespecific in vivo base editing with evoCjCas9 using a one-vector AAV system.
[0139] Example 1: PAM recognition assay
[0140] We define recognition of potential PAMs by performing high-throughput PAM detection assay as recently described (Walton et. al., 2021 , Nat. Protoc. (https: / / doi.org / 10.1038 / s41596-020-00465- 2)). In brief, Cas9 is cloned in pCMV-T7-SpCas9-P2A-EGFP (Addgene #139987) and expressed in HEK293T (ATCC CRL-3216) cells for 48 hours. Whole-cell lysates are collected and normalized to a concentration corresponding to 150 nM fluorescein dye. Target-specific sgRNAs are in vitro transcribed using HiScribe T7 High Yield RNA Synthesis Kit. Substrate libraries with different target sites and PAM libraries are cloned into p11-LacY-wtxq (Addgene #69056). gRNA (1.1 μM) and normalized cell lysate (83 nM fluorescein) are complexed for 10 minutes at 37 °C. The ribonucleoprotein mixture (0.5 μM gRNA, 37.5 nM fluorescein lysate) is added to the substrate library (2.5 nM) and the reaction is stopped at different time intervals (1 , 8, and 32 minutes). For all time points, substrate libraries are individually PCR amplified using time point-specific barcodes. Samples are pooled and sequenced on a high-throughput sequencing platform (e.g. Illumina). A Cas9 protein is characterized on two different target sites, with the PAM library spanning positions 4 to 8 (Target 1 and Target 2). Two replicates for each target sequence are performed on different days, and sequencing data is processed using the published python script (Walton et. al., 2021 , Nat. Protoc.). Derived depletion rate constants for each PAM are averaged. We define a Cas9 to exhibit activity towards a PAM when the depletion rate constant is equal to or larger than 0.001 .
[0141] Based on the number of PAMs a Cas9 recognizes, the PAM availability in % is calculated as follows (see Fig. 3d):
[0142] (number of PAMs a Cas9 exhibits activity on) I (number of possible PAMs (= 4 to the power of total possible PAM positions) * 100
[0143] At the example of evoCJCas9: 198 / (4^5)*100 = 19.33%.
[0144] At the example of wild-type CJCas9: 19 / (4^5)*100 = 1 .86%
[0145] At the example of enCJCas9: 35 / (4^5)*100 = 3.42%
[0146] At the example of SpCas9 (assuming NGG PAM being recognized): 4 / (4^3)*100 = 6.25%
[0147] Alternatively, PAM recognition can be expressed as a parameter describing every how many bases in a random DNA sequence a Cas9 would encounter a PAM on which it exhibits activity: (number of possible PAMs) I (number of PAMs a Cas9 exhibits activity on)
[0148] At the example of evoCJCas9: (4^5) / 198 = 5.17
[0149] At the example of wild-type CJCas9: (4^5) / 19 = 53.89
[0150] At the example of enCJCas9: (4^5) / 35 = 29.26
[0151] At the example of SpCas9 (assuming NGG PAM being recognized): (4^3) / 4 = 16
[0152] Example 2: Results
[0153] Phage-assisted continuous and non-continuous directed evolution of CjCas9
[0154] PACE is a technique for directed protein evolution, where the gene of interest is continuously mutated and transferred from bacterial host cell to host cell via a modified M13 filamentous bacteriophage in an activity of interest-dependent manner. In our PACE setup the essential gene III on the M13 selection phage (SP) is replaced by a catalytically dead CJCas9 (dCJCas9), which is C-terminally fused to the ω subunit of bacterial RNA polymerase. SP infection of E. coli host cells carrying the accessory plasmid (AP) results in complexion of dCJCas9 with an sgRNA that directs the complex to the promoter region of gene III (Fig. 1a). Recognition of the protospacer and PAM sequence leads to stable binding of the dCJCas9-ω complex, recruitment of the RNA polymerase and subsequent expression of gene III. This completes the life cycle of the phage and enables the assembly of functional and infectious phage particles. The additional presence of the drift plasmid (DP) in E. coli host cells ensures mutagenesis of the phage genome during replication, and continuous addition and draining of media with fresh host cells leads to a steady phage dilution and sets the evolutionary pressure. By installing a library of non-canonical PAM sequences on the AP, we further ensured linkage of phage propagation to dCjCas9 variants that recognize these novel PAMs in our setup (Fig. 1a). Since the first four positions of the canonical CJCas9 PAM show only minor nucleotide preferences (NNNV), we focused our efforts on the relaxation of the last 4 PAM positions (RYAC). Our first attempt to start PACE with a PAM library that is fully randomized at these 4 positions, however, was not successful as the selection was too stringent and phages could not be propagated (determined by RT-qPCR; CT > 30). Therefore, we next tried to pre-select CJCas9 variants with relaxed PAM recognition using phage-assisted non-continuous evolution (PANCE). In this setup we used the same evolution logic, but reduced the dilution pressure by propagating phages via serial phage dilution instead of dilution in a continuous flow. However, also the moderate dilution rates used in PANCE resulted in infrequent gene III activation and loss of phages, as determined by RT-qPCR (CT > 30). To reduce stringency further, we next performed PANCE in an arrayed fashion where we individually replaced nucleotides at PAM positions 5-8 on the AP in all possible combinations (15 different nucleotide compositions per position, resulting in 56 different combinations plus 4 canonical N4ACAC PAMs - Fig. 1 b). In this setup we identified phages to withstand 200-fold dilution steps for PAM combinations close to the canonical PAM, while more unfavoured PAM sequences (e.g. T, G and C at positions 5, 6 and 7, respectively) led to a phage wash-out unless the AP was mixed with a fraction of plasmids with canonical PAM nucleotides (Fig. 1b). Interestingly, lagoons with all nucleotide variations at position 8 supported phage propagation, while positions 6 and 7 possessed more stringent PAM requirements (Fig. 1 b). After 12 rounds of dilution (a total of 5*1023-fold dilution), 45 lagoons maintained detectable phage genomes. These were mixed in a 1 :1 ratio with wild-type ω -dCJCas9-phages and used to seed a lagoon of a PACE apparatus with a continuous inflow of fresh host cells containing an AP with a randomized PAM library for positions 6-8. The pre-relaxation with PANCE allowed phages to sustain a gradual increase of flow rates from 0.5 to 3 lagoon volumes per hour over 10 days, with a phage wash-out starting to occur at higher flow rates. In the initial phase of PACE (first 100 hours), we modulated the selection pressure by fluctuating the dilution rates between 0.5 and 1 lagoon volume per hour (LV / h) in order to facilitate crossing of potential fitness valleys, followed by a more stringent selection with a dilution rate that was increased to 3 LV / h. To assess the shift of phage genotypes during PACE, we collected phage DNA at different timepoints during the experiment and analysed it by nanopore long read sequencing. Evaluation of the sequencing data revealed a mutation bias of G-C to A-T conversions (51.3%), with A-T to T-A conversions being the least frequently observed mutations (2.0%, n=762). Importantly, amino acid changes occurred most often in the RuvC III domain, phosphate lock loop domain, wedge domain and PAM interacting domain of CjCas9 (Fig. 1c). After 108 hours of PACE, phages with the consensus mutations L58Q, G323V, D376Y, F434C, K622N, D663E, T671A, E789K, N821 K, D891 N, S918T, S932N and S951G started to occur in the lagoon. These mutations were further enriched until termination of the experiment, where a consensus genotype made up the bulk of the phage.
[0155] Establishment of an evolved CjCas9 variant with broad PAM recognition
[0156] In our PACE and PANCE evolution logic we selected for PAM recognition via CJCas9 binding, potentially leading to mutations that impair the nuclease activity of Cas9. Hence, we first assessed the effect of the consensus mutations individually in a kinetic in vitro cleavage assay on a randomized PAM library, with the aim of finding a set of mutations that lead to PAM relaxation without impairing the nuclease activity. Of the 13 consensus mutations, E789K, E871 K, D891 N and E946K had a strong contribution to PAM relaxation compared to wild type CJCas9 at position 5 and 7 (Fig. 2a), with all 4 mutations leading to the recognition of non-canonical C and T bases at position 5 and non-canonical C, G and T bases at position 7. Since E789K resulted in a higher cleavage activity at most non-canonical PAMs than the other three variants, and additionally allowed pronounced recognition of N5ACA PAMs, we decided to select this mutation for our evolved (evo)CjCas9 variant. Interestingly, mapping E789K onto the cryogenic electron microscopy (cryo- EM) structure of CjCas9 (Protein Data Bank (PDB): 5X2H) revealed that it is located in the phosphate lock loop domain and interacts with the phosphate backbone of PAM position 1 . The positively charged lysine residue may therefore facilitate distortion of the target DNA strand, which in turn allows recognition of non-canonical bases at downstream PAM positions.
[0157] The mutation S951 G occurred with the highest frequency of all PAM-relaxing consensus mutations (after 170h of PACE 93.9% of phage genomes contained S951 G) and was the only mutation that allowed relaxation of PAM position 6, leading to the recognition of a non-canonical A nucleobase in addition to the canonical C and T. We therefore also selected this mutation for introduction into evoCjCas9. Notably, S951 is located in the PAM interacting domain, but instead of interacting with nucleobases at position 6 of the PAM it interacts with the nucleobases on the opposing DNA target strand via hydrogen bonding (Fig. 2b). We therefore speculate that the shorter side chain of glycine compared to serine reduces the interactions with A nucleobases at this position, leading to the detection of non-canonical T at position 6 of the PAM.
[0158] Substitutions of residues L58 and N821 were also frequently observed at the end of our PACE experiment (98.6% and 76.7%, respectively), but did not substantially contribute to PAM relaxation of CjCas9 in our PAM detection assay. Nevertheless, L58Y has previously been described to increase the stability of the CjCas9-sgRNA complex by forming a stacking interaction with U48 in the sgRNA, thereby enhancing the DNA cleavage activity of CjCas9 on canonical PAMs (Fig. 2a). Since N821 is in close physical distance and interacts with the same nucleotide of the sgRNA via hydrogen bonding (Fig. 2b), we speculate that the positively charged lysine residue at this position could further stabilize the CjCas9-sgRNA complex and enhance cleavage activity, prompting us to introduce L58Y and N821 K into evoCjCas9.
[0159] Mutations G323V, D376Y, F434C, K622N, D663E and T671A were also enriched during PACE, but had a detrimental effect on the catalytic activity of CjCas9 nuclease in our cleavage assay. Thus, while these mutations may pose a fitness advantage regarding target recognition and binding of CjCas9, they would prohibit applications where cutting or nicking of DNA is required. We therefore did not integrate these mutations into evoCjCas9.
[0160] Taken together, we found that the mutations L58Y, E789K, N821 K and S951 G increase PAM flexibility or cleavage activity. Combined with D900K, a mutation that was not enriched in our PACE experiment but was previously described to enhance CJCas9 cleavage activity (Nakagawa, R. et al. Engineered Campylobacter jejuni Cas9 variant with enhanced activity and broader targeting range. Commun. Biol. 5, 1-8 (2022)), these mutations were introduced into CJCas9 to generate evoCjCas9. Compared to wild type CJCas9 and the previously established enhanced (en)CJCas9 variant with the substitutions L58Y and D900K, evoCjCas9 shows a substantial relaxation in PAM recognition (Fig. 2a); it recognizes all nucleotides at the first three PAM positions and NAHNN or NNHAN sequences at PAM positions 4-8. This allows evoCJCas9 to detect 10 times more target sequences than wild type CJCas9, which requires a VRYAC motif at PAM positions 4-8 (198 versus 19 PAM sequences with a kinetic rate constant >1 O'3, Fig. 2a). evoCjCasd shows high cleavage activity at canonical and non-canonical PAMs in human cells
[0161] To comprehensively characterize the activity of evoCJCas9, we compared indel frequencies of CJCas9 and evoCJCas9 in mammalian cells at multiple sites with canonical and non-canonical PAM sequences. We therefore generated a lentiviral library containing 267 CJCas9 sgRNA constructs paired with their target sites (self-targeting library), with the different target sites comprising a set of canonical (N4ACAC, N4ATAC, N4GTAC and N4GCAC) and non-canonical (N4AAAC, N4CAAC, N4GAAC and N4TAAC) PAM sites. After integrating the library into the genome of HEK293T cells, plasmids expressing wild type CjCas9 or evoCjCas9 were transfected (Fig. 3a) and after selection genomic DNA was isolated and analysed by high throughput sequencing (HTS). Interestingly, at canonical PAM sites evoCJCas9 was already 2-fold more efficient than wild type CjCas9 in generating indels (on average 54.0% versus 24.7%, Fig. 3b). Moreover, evoCJCas9 also reached average editing rates of 48.6% at non-canonical PAM sites, at which wild type CjCas9 did not show substantial editing (Fig. 3b).
[0162] Since other PAM relaxed Cas9 variants often come with the trade-off of reduced activity, we next benchmarked indel formation rates of commonly used Cas9 orthologs (CJCas9, enCjCas9, SaCas9, SauriCas9, SpCas9 - Fig. 3a) and their PAM-relaxed variants (evoCjCas9, SaKKH, SauriCas9-KKH, SpRY, SpG). We designed a self-targeting library consisting of 45 target sites, with each target site containing the optimal protospacer length and PAM sequence for each orthologue (22 bp and N3AACAC for CjCas9, 22 bp and N2GRRT for SaCas9, 22 bp and N2GG for SauriCas9, 19 bp and NGG for SpCas9). The library was cloned into a lentiviral plasmid and after virus production integrated into the genome of HEK293T cells. The cell pool was subsequently transfected with plasmids expressing the different Cas9 variants, and after selection genomic DNA was extracted for analysis by HTS. Confirming previous studies, we found that PAM relaxed SaCas9-, SauriCas9- and SpCas9 variants showed lower activity than the wild type orthologs on canonical PAMs (Fig. 3c). In contrary, evoCjCas9 showed substantially higher activity than wild type CjCas9 on the canonical N3AACAC PAM sequence, which could be explained by the activity enhancing mutations that are also present in enCjCas9 (Fig. 3c). Cross-comparison to other Cas9 orthologs further revealed that evoCJCas9 works with similar efficiency as SpG, a frequently used SpCas9 variant detecting NGN PAM sequences that occur with similar frequencies than evoCJCas9 PAM sequences. Moreover, in our assay evoCJCas9 outperforms SpRY, the most PAM relaxed SpCas9 variant that has been established so far (Fig. 3c, d). evoCjCasd has a low tolerance for mismatches between the spacer and target site
[0163] Off-target editing at genomic sites that share sequence similarities to the targeted locus are limiting clinical application of CRISPR-Cas nucleases. We therefore next assessed the mismatch tolerance of wild type and evoCjCas9. To this end, we designed a self-targeting library with 687 members, consisting of 37 target sites and up to 19 potential off-target sites with 1-5 mismatches. The library was stably integrated into HEK293T cells using lentiviral vectors and cells were transfected with plasmids expressing wild type- or evoCjCas9. HTS analysis of the target sites revealed that both variants have a similar mismatch tolerance; the mean off-target (1-5 mismatches) to on-target ratio for CjCas9 and evoCjCas9 was 0.36 and 0.39, respectively, with both variants not tolerating more than 2 mismatches between the spacer and protospacer (Fig. 3e). To further investigate how the location of the mismatches within the protospacer influences off-target editing activities, we next designed a dual-mismatch permutation library for 10 CjCas9 target sites. As expected from previous observations with SpCas9, mismatches in the seed region of the sgRNA (position 14-22) were more detrimental to target recognition than PAM distal mismatches (Fig. 3f).
[0164] Efficient base- and prime editing with evoCjCasd
[0165] The broad PAM recognition and high activity of evoCJCas9 makes it an ideal nuclease for generating size optimized BEs and PEs. To first assess base editing with evoCJCas9, we fused evoCJCas9- or wild type CJCas9 variants that nick the target strand (D8A) C-terminally to either an adenosine deaminase (ABE8e; Richter, M. F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol. 38, 883-891 (2020); WO2021151756A1 , incorporated herein by reference) or to a cytosine deaminase (hsAIDmax (Komor et al. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424 (2016)), eAID (Liu, Z. etal. Improved base editorfor efficient editing in GC contexts in rabbits with an optimized AID-Cas9 fusion. 33, 9210-9219 (2019)), BE4max and AncBE4max (Koblan, L. W. et al. Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat. Biotechnol. 36, 843-848 (2018))) (Fig. 4a). Cytosine BEs were furthermore N-terminally fused to an uracil glycosylase inhibitor (UGI) to block the conversion of the installed U back to a C. We next designed a self-targeting library consisting of 254 target sites with canonical and non-canonical PAM sites for base editing with CJCas9-BEs. The library was stably integrated into HEK293T cells using lentiviral vectors, which were subsequently transfected with plasmids expressing the different base editors. In line with our results obtained with the evoCjCas9 nuclease, evoCjCas9 BEs already showed higher editing efficiency than wild type CjCas9 BEs at canonical PAMs. In addition, evoCjCas9 BEs allowed base editing at non-canonical PAMs, where wild type CjCas9 BEs were largely inactive (Fig. 4b). Similar to previously reported BEs established from other compact Cas9 orthologues, CjCas9 BEs had a broader editing window than the typical editing window of SpCas9 BEs. We reasoned that this might be caused by a less constrained R-loop formation, and next replaced the HNH domain of evoCjCas9 with a heterodimer of TadA-TadA8e to constrain the deaminase as close to the R-loop as possible. However, while the editing window of HNHxevoCjCas9 was indeed more constrained, similar to reports with HNHxSpCas9 BEs, editing rates were also substantially lower than for CjCas9 BEs with N-terminally fused TadA8e (Fig. 4b, last panel). Finally, we also assessed if CjCas9 base editors show preferences for certain trinucleotide motifs (bases flanking the targeted A or C). Consistent with recent reports for SpCas9 Bes, we found that evoCjCas9 BEs using the rAPOBECI deaminase show a preference for editing of C bases that are preceded by a T, while evoCjCas9 BEs using AID show a preference for editing of C bases that are preceded by A or G. In addition, we observed that evoCjCas9 BEs using the TadA8e deaminase largely disfavor target bases preceding A.
[0166] Prime editing is a more versatile genome editing technology than base editing that enables the introduction of all 12 possible base conversions as well as multi base replacements and small insertions or deletions. To generate miniature PE systems based on CJCas9, we first introduced the H559A mutation into wild type CJCas9 and evoCJCas9 to obtain variants that only nick the nontarget strand. These variants were then N-terminally fused to either the full length M-MLV reverse transcriptase (PE) or an M-MLV reverse transcriptase that lacks the RNaseH domain (PEΔRnH) (Fig. 4c). When we tested these variants by introducing an A to G mutation into the AAVS1 locus, we observed that wild type CJCas9 PEs only achieved 6% editing while evoCJCas9 PEs achieved up to 26% editing, with no significant difference between full length PE and PEΔRnH(Fig. 4d). We next generated a self-targeting library containing 236 prime editing guide (peg)RNAs paired with their target sites. These target sites include human disease loci with canonical- and non-canonical PAMs, which are corrected back to the wild type sequence by their corresponding pegRNAs. The library was integrated into HEK293T cells using lentiviral vectors, and cells were subsequently transfected with plasmids expressing CJCas9-PEΔRnHor evoCJCas9-PEΔRnH. In line with the results from the AAVS1 locus, evoCJCas9-PEΔRnHalready achieved higher editing rates (up to 30.1 %) than CJCas9-PEΔRnH(up to 24.6%) at target sites with canonical PAM sequences. In addition, evoCJCas9-PEΔRnHalso enabled editing at target sites with non-canonical PAMs (up to 22%), whereas wild type CJCas9-PEΔRnHwas unable to edit these sites (Fig. 4e;). The same pattern was also observed when the library was integrated in K562 cells, albeit with lower overall editing rates. Further analysis of CjCas9 prime editing outcomes on the library revealed that single nucleotide replacements were more efficient than insertions or deletions, with A to C and C to T conversions being the most efficient edits. Notably, within our library we also observed multiple loci with low or no prime editing, suggesting that at these sites our pegRNA design was not optimal (Fig. 4f). We therefore next assessed if we could identify pegRNA designs that correlate with editing rates in our library (n=238). We found that the most influential features contributing to CjCas9 prime editing were the melting temperature of the primer binding sequence (PBS) (R=0.43) and the reverse transcriptase template (RTT) (R=0.26, suggesting that extending the pegRNA length downstream of the scaffold sequence may benefit editing rates.
[0167] In vivo base editing with single AAV constructs
[0168] The small size of evoCJCas9 compared to other Cas9 orthologs offers major advantages for in vivo genome editing, and in principle facilitates packaging of evoCJCas9 nucleases and evoCJCas9 BEs together with their sgRNAs on single AAV vectors. To test whether this approach is feasible and enables efficient editing in vivo, we cloned an AAV construct that expresses evoCJCas9-ABE8e under the hepatocyte specific P3 promoter and an sgRNA targeting the AG splice acceptor site of the Pcsk9 intron 3 under the U6 promoter (Fig. 5a). Pcsk9 is a secreted protein that is expressed in the liver and that acts as a negative regulator of the low-density lipoprotein (LDL) receptor; reducing Pcsk9 levels by installing splice site mutations is therefore a potential therapeutic approach for reducing blood LDL levels. We packaged the evoCJCas9 BE expression construct spanning 4988 bp including ITRs into hepatotropic AAV9 capsids, which were systemically administered into 5-week-old C57BL / 6J mice at a dose of 2.5 x 1013or 5 x 1013vector genomes (vg) / kg body weight (Fig. 5b). 6 weeks post injection we isolated liver tissue, hepatocytes and serum from treated mice to analyze editing rates and the effect on serum Pcsk9 and LDL levels. HTS on genomic DNA isolated from liver tissue revealed 22% and 21 % editing in animals treated with the lower- and higher dose, respectively (Fig. 5c). In isolated hepatocytes, editing rates were substantially higher, with 38% and 39%, confirming hepatocyte-specific expression of evoCjCas9- ABE8e from the P3 promoter. Analysis of the serum of treated mice further revealed a 71 % reduction of Pcsk9 levels and a 35% and 43% reduction of LDL levels in animals injected with a dose of 2.5 x 1013vg / kg and 5 x 1013vg / kg, respectively (Fig. 5d, e). Together, these data demonstrate that evoCjCas9 can be employed for efficient and tissue specific in vivo base editing from single AAV vectors at clinically relevant doses.
[0169] DISCUSSION
[0170] The use of compact Cas9 orthologs instead of the 1 ,368 aa long SpCas9 nuclease is an attractive alternative for translational genome editing, as their small size facilitates in vivo delivery. With 984 aa, CjCas9 is the smallest known Cas9 ortholog, which in principle fits as a nuclease or a BE fusion protein together with its sgRNA on a single AAV vector. However, similar to other small-sized Cas9 orthologs, wild type CjCas9 recognizes a relatively long PAM, reducing the number of targetable sites. Using phage-assisted evolution (PACE and PANCE), we identified a set of amino acid substitutions that led to PAM relaxation of CjCas9.
[0171] The novel variant, termed evoCjCas9, recognizes N4AH and N5HA PAM sequences, which occur 10 times more frequently in the genome than the canonical N3VRYAC PAM site. In addition, evoCjCas9 possesses a nuclease activity that is higher than that of wild type CjCas9 (2-fold), and in range with the nuclease activity of PAM-relaxed SpCas9 variants such as SpG. Moreover, the requirement of a 22 to 24 nucleotide long spacer sequence with a mismatch tolerance of less than 3 bases makes evoCjCas9 a highly precise targeted nuclease.
[0172] By fusing deaminases or the M-MLV reverse transcriptase to D8A or H559A evoCjCas9 nickase variants we also constructed compact base- and prime editors. Analyzing evoCjCas9 BEs, we observed average editing rates above 20% for various cytidine and adenine BE architectures, albeit with wider editing windows than SpCas9 BEs. Similar observations have been made with BEs that were based on other compact Cas9 proteins, such as SaCas9 and Nme2Cas9, suggesting differences in the R-loop accessibility between SpCas9 and smaller Cas9 orthologs. While with evoCjCas9 PEs we also achieved >20% editing at certain target sites, on average the editing rates were lower than for evoCjCas9 BEs. Optimizing linker architectures or the CjCas9 pegRNA design may therefore be required to obtain robust editing at a larger number of loci.
[0173] Due to the small size of evoCjCas9 it can be packaged as an adenine- or cytosine BE on a single AAV vector. While adenine base editing from single AAV vectors has already been demonstrated previously with Cy3Cas9, SaCas9 and A / me2Cas9, to our knowledge this study provides the first demonstration for cytosine base editing from a single AAV vector. Furthermore, the compact size of CJCas9 would also allow for expression of ABEs using tissue-specific promoters other than the demonstrated P3 and hSyn promoters. Examples include the hUPII (bladder; urothelium), ProA7 (eye; cone cells), Ksp-cadherin (kidney; renal tubules), a-MHC (cardiac muscle), CK8e or SPc5- 12 (skeletal muscle) promoters.
[0174] Taken together, we used PACE and PANCE to develop evoCjCas9, a compact Cas9 nuclease with a broad targeting range and high nuclease activity. evoCJCas9 possesses a high specificity and enables base- and prime editing in mammalian cells. Due to its small size, we could also develop a single vector AAV system for in vivo adenine base editing from tissue-specific promoters, demonstrating the potential of evoCJCas9 for future therapeutic applications.
[0175] METHODS
[0176] General methods and cloning. PCRs were performed using the Q5 High-Fidelity DNA polymerase. All expression vectors were assembled using NEBuilder HiFi DNA assembly. Plasmids expressing sgRNAs were cloned using the T4 DNA ligase (New England Biolabs). Plasmids used in mammalian tissue culture were purified using NucleoBond Xtra Midi kits (Macherey-Nagel).
[0177] Phage-assisted continuous and non-continuous evolution. PANCE and PACE were performed as recently described (Hu, J. H. et al. Evolved Cas9 variants with broad PAM compatibility and high DNA specificity. Nature 556, 57-63 (2018); Miller, S. M., Wang, T. & Liu, D. R. Phage-assisted continuous and non-continuous evolution. Nat. Protoc. 15, 4101-4127 (2020)). The M13 filamentous selection phage encodes a catalytically dead CJCas9 fused via an Ala-Ala- Linker to the ω subunit of bacterial RNA polymerase. The accessory plasmid harbors the PAM and protospacer sequence in reverse, 114 nt upstream of gene III. To increase the mutagenesis rate, DP6 (Addgene #140446) was used and induced with L-arabinose (10 mM) one hour before selection phage addition. All experiments were performed using the S2060 bacterial strain (Addgene #105064), cultured in Davis Rich Media. PANCE was performed using PAM libraries with all possible nucleobase combinations for either PAM positions 5, 6, 7, or 8. PANCE lagoons were contained in deep well plates at 1 mL volume and selection phages were added once cultures reached the log phase (approx. OD 0.4 - 0.6). Phages were incubated overnight and harvested via centrifugation and sterile filtration (0.2 pm, Sarstedt) before adding a dilution to the next round. Phage titers were determined using plaque assays or RT-qPCR for higher throughput. qPCR results with CT > 30 were considered to resemble the absence of phages. Similarly, PACE was performed using a pool of PANCE-derived phage variants after round 12, mixed 1 :1 with wild type CJCas9 containing phages. As accessory plasmid, a PAM library for positions 6-8 was used. S2060 host cells were continuously cultured at OD 0.5 in a custom-built miniature bioreactor. Host cells were pumped through a lagoon and the mutagenesis plasmid DP6 was induced with L-arabinose (10mM) before the cells reached the lagoon. The flow through the lagoon was controlled via a microcontroller and adjusted to increase the selection pressure over time. Phage titers in the bioreactor (no phages detected) and lagoon were monitored using RT-qPCR and plaque assays. Regular time points were collected and analysed using long-read Nanopore sequencing (Oxford Nanopore Technologies).
[0178] Phylogenetic analysis of directed evolution outcomes. Phage genotypes emerging during PACE were sequenced using a previously established method. Briefly, from collected phages at different time points ω subunit-CjCas9 variants were PCR amplified using unique molecular identifiers (UMIs) and timepoint specific barcodes and were subcloned. Per timepoint, a few hundred colonies were selected and clonally amplified. Plasmids were purified and ω subunit- CjCas9 was extracted using unique restriction sites. Nanopore adapters were ligated, and variants were sequenced using the manufacturer's protocol. Timepoints were demultiplexed and consensus reads for each UMI were built using a python script.
[0179] High-throughput PAM detection assay. HT-PAMDA was performed as recently described (Walton et al. Scalable characterization of the PAM requirements of CRISPR-Cas enzymes using HT-PAMDA. Nat. Protoc. 16, 1511-1547 (2021 ). Cas9-variants were cloned in pCMV-T7-SpCas9- P2A-EGFP (Addgene #139987) and expressed in HEK293T cells for 48 hours. Whole-cell lysates were collected and normalized to a concentration corresponding to 150 nM fluorescein dye. Targetspecific sgRNAs were in vitro transcribed using HiScribe T7 High Yield RNA Synthesis Kit. Substrate libraries with different target sites and PAM libraries were cloned into p11-LacY-wtxq (Addgene #69056). gRNA (1.1 μM) and normalized cell lysate (83 nM fluorescein) were complexed for 10 minutes at 37 °C. The RNP mixture (0.5 μM gRNA, 37.5 nM fluorescein lysate) was added to the substrate library (2.5 nM) and the reaction was stopped after different time intervals (1 , 8, and 32 minutes). For all time points, substrate libraries were individually PCR amplified using timepoint-specific barcodes, followed by an amplification using protein variant-specific barcodes. Samples were pooled and sequenced on a NovaSeq 6000 (Illumina). Cas9-variants were characterized on two different target sites, with the PAM library spanning either positions 1 to 5 (Spacer 3) or 3 to 8 (Spacer 1 and 2). Two replicates were performed on different days. Sequencing data were processed using the provided python script. Rate constants were calculated from normalized read counts for all characterized CJCas9 variants.
[0180] Target library design using sgRNAs. Self-targeting constructs were designed with guide sequence preceding 5’-G nucleotide (not counted to total guide length) and ordered as singlestranded DNA oligo pools (Twist Bioscience). Oligo pools were amplified according to the manufacturer’s protocols using NEBNext Ultra II polymerase. The amplified pool was cloned into BsmB I -digested Lenti_gRNA-Puro (Addgene #84752) and electroporated into Endura electrocompetent cells (Lucigen). Colonies were harvested (> 2000x coverage) and plasmids were purified.
[0181] Target library screen using pegRNAs. A recently developed prime editing prediction tool (PRIDICT, unpublished) was adapted to assist in the design of the pegRNAs for the library. The library was cloned at > 2000x coverage as described above. Upon harvest, plasmids were digested using BsmBI, purified, and the CJCas9 scaffold sequence was inserted via golden gate assembly. Purified ligation reactions were electroporated, colonies were harvested (> 2000x coverage) and plasmids were purified.
[0182] Cell culture and high-throughput sequencing. All cell lines were cultured at 37°C and 5% CO2 in cell culture incubators. HEK293T cells (ATCC CRL-3216) were maintained in Dulbecco’s modified Eagle’s medium (DMEM) plus GlutaMAX (Thermo Fisher Scientific), and K562 (ATCC CCL-243) cells were maintained in RPMI-1640 medium (Thermo Fisher Scientific), both supplemented with 10% (v / v) fetal bovine serum (FBS; Sigma-Aldrich) and 1 % penicillin / streptomycin (Thermo Fisher Scientific). Cells were maintained at confluency below 90% and depending on library size seeded on 6-well / 48-well cell culture plates or T175 cell culture flasks (Greiner). For base editor target library screens, HEK293T cells were transfected at 70% confluency (day 0) using 10 pl Lipofectamine 2000 (Thermo Fisher Scientific), 3 pg of base editor expression plasmid and 1 pg of Tol2-helper plasmid (Addgene #31828) according to the recommendation of the manufacturer.
[0183] Prime editing library screens in HEK293T were performed in T25 flasks by transfection of 2.64 μg of editor plasmid with OptiMEM (total volume of 75.14 μl) and 21.7 μl polyethylenimine (PEI, 1 mg / ml) (Day 0). PE screens in K562 were performed in 48-well plates (16x lipofection per editor or control) by seeding 100,000 K562 cells in 300 μl RPMI++ and transfection of 1000 ng editor plasmid with OptiMEM (total volume of 25 μl) mixed with 1.5 μl of Lipofectamine 2000 (Thermo Fisher Scientific) and 23.5 μl of OptiMEM. Controls for HEK293T and K562 were transfected without editor plasmid. Arrayed editing experiments in HEK293T were performed in 48-well plates by seeding 130,000 cells 6h before transfection. Transfection was performed with 750 ng of editor plasmid and 250 ng of gRNA plasmid with 1 μl of Lipofectamine 2000 as described above for K562 (day 0). On day 1 , HEK293T (detached using 1x TrypLE (Gibco)) or K562 cells were transferred into blasticidin (7.5 mg / mL) containing media (no selection for controls) and maintained for 6 (PE) or 9 (BE) more days (unless otherwise noted).
[0184] Genomic DNA was isolated by direct lysis and locus-specific primers were used to generate targeted amplicons (NEBNext Ultra II) for deep sequencing, maintaining > 500x coverage. Cells were tested negative for Mycoplasma contamination and were authenticated by the supplier by STR analysis.
[0185] Lentiviral integration of self-targeting libraries. For lentivirus production, HEK293T cells were seeded at 15x106cells in a T175 flask (Greiner) and transfected at 70% confluency using polyethylenimine (PEI). Briefly, 152 μl PEI (0.1 mg / ml) was mixed with 506 μl Opti-MEM containing 2.6 μg VSV-G (Addgene #12259), 5.2 μg PAX2 (Addgene #12260) and 10.4 μg lentiviral vector plasmid. The mix was incubated for 20 min at room temperature prior to transfection. One day post- transfection, media was replaced and three days post-transfection, the supernatant was harvested and purified by filtration (0.4 μm, Sarstedt). A dilution series of produced lentivirus was used to transduce HEK293T cells. 24 hours post-transduction, cells were selected with 2.5 μg / ml puromycin. Lentiviral concentration with a 50% survival rate 3 days post-transduction was used for subsequent library cell generation. Library cells were generated accordingly, with cell numbers being sufficient for > 2000x coverage (post-selection). 3 days post-transduction, puromycin- containing media was replaced, and 5 days post-transduction library cells were detached and cryopreserved.
[0186] Expression and purification of CJCas9 and evoCJCas9. DNA sequences encoding CJCas9 and evoCJCas9 were inserted into the 1S plasmid (Addgene #29659) using ligation-independent cloning, resulting in a construct comprising a N-terminal His6-SUMO tag followed by a tobacco etch virus (TEV) protease cleavage site. The constructs were expressed in E. coli BL21 Rosetta2 (DE3) cells (Novagen). Cells were lysed in 20 mM Tris-HCI pH 8.0, 1 M NaCI, 20 mM imidazole, 1 pg / mL pepstatin, 200 μg / mL 4-(2-Aminoethyl) benzenesulfonyl fluoride hydrochloride (AEBSF), 1 mM Tris-(2-Carboxyethyl)phosphine hydrochloride (TCEP) by ultrasonication. Clarified lysate was loaded on a 5 mL Ni-NTA (Sigma-Aldrich) affinity column. The column was washed with 20 mM Tris-HCI pH 8.0, 1 M NaCI, 20 mM imidazole, 1 mM TCEP and bound protein was eluted with 20 mM Tris-HCI pH 8.0, 300 mM NaCI, 300 mM imidazole, 1 mM TCEP. The eluted protein was dialyzed, overnight at 4°C, against a buffer yielding a final imidazole concentration of 40 mM imidazole, in the presence of 1 mM TCEP and TEV protease to remove the His6-SUMO tag. Cleaved protein was passed through a 5 mL Ni-NTA (Sigma-Aldrich) affinity column to remove the cleaved affinity tag and uncleaved protein; the column was washed with 20 mM Tris-HCI pH 8.0, 300 mM NaCI, 40 mM imidazole, 1 mM TCEP, and the flow-through and wash fractions were pooled for subsequent purification using a 5 mL HiTrap HP Heparin column (Cytiva), eluting with a linear gradient up to 1 M NaCI. The elution fractions were pooled, concentrated in the presence of 1 mM Dithiothreitol (DTT) and further purified by size exclusion chromatography using a Superdex 200 (16 / 600) column (Cytiva) in 20 mM Tris-HCI pH 8.0, 150 mM NaCI, 1 mM DTT, yielding pure, monodispersed proteins. Aliquots were flash-frozen in liquid nitrogen and stored at -80 °C.
[0187] CHANGE-seq. The sgRNAs were first tested for functionality by digesting a PCR amplicon of the genomic target site. The library was prepared as previously described44. Data were processed using the CHANGE-seq analysis pipeline (https: / / github.com / tsailabSJ / changeseq) with the following parameters: ‘read_threshold: 4, window_size: 3, mapq_threshold: 50, start_threshold: 1 , Gap_threshold: 3, mismatch_threshold: 10, search_radius: 30, merged_analysis: True. The respective target sites were deep sequenced and covered by at least 10,000 reads per site.
[0188] Adeno-associated virus production. Vectors (AAV2 serotype 9) were produced by the Viral Vector Facility of the Neuroscience Center Zurich. Briefly, AAV vectors were ultracentrifuged and diafiltered. Physical titers (vg / mL) were determined using a Qubit 3.0 fluorometer as previously described. Identity of the packaged genomes of each AAV vector were confirmed as size verification of PCR amplicon (gel electrophoresis) and Sanger sequencing.
[0189] Animal studies. Mouse experiments were performed in accordance with protocols approved by the Kantonales Veterinaramt Zurich. Mice were housed in a pathogen-free animal facility at the Institute of Pharmacology and Toxicology at UZH Zurich and kept in a temperature- and humidity- controlled room (21 °C, 50% RH) on a 12-h light / dark cycle. Mice were fasted for 3-4 h before blood was collected from the inferior vena cava before liver perfusion. Mice were injected with 0.5- 1 x 1012AAV vector genomes per mouse. Injection volumes were 120-150 μl. Only male C57BL / 6J animals at an age of 5-6 weeks were used.
[0190] Primary hepatocyte isolation. Mice were euthanized using CO2 and immediately perfused with Hank’s balanced salt solution (Thermo Fisher Scientific) plus 0.5 mM EDTA via the inferior vena cava and a subsequent incision in the portal vein. During this step, one liver lobe was squeezed off via a thread to inhibit perfusion of this lobe to collect whole liver samples for whole liver lysates. After blanching of the liver, mice were perfused with digestion medium (low-glucose DMEM plus 1 x penicillin-streptomycin (Thermo Fisher Scientific), 15 mM HEPES and freshly added Liberase (Roche)) for 5 min. Livers were isolated in cold isolation medium (low-glucose DMEM supplemented with 10% (vol / vol) FBS plus 1 x penicillin-streptomycin (Thermo Fisher Scientific) and GlutaMax (Thermo Fisher Scientific)), and the liver was gently dissociated to yield a cell suspension that was passed through a 100-pm filter. The suspension was then centrifuged at 50g for 2 min and washed with isolation medium 2-3 times until the supernatant was clear. The primary hepatocytes were pelleted for direct lysis for NGS sample preparation.
[0191] ELISAs. Mouse PCSK9 levels were determined by using Mouse Proprotein Convertase 9 / PCSK9 Quantikine ELISA Kit (R&D Systems, cat. no. MPC900) according to the manufacturer’s instructions. Absorbance was measured at 450 nm and background at 540 nm; the latter was subtracted for quantification. Total cholesterol, triglyceride and high-density lipoprotein (HDL) from all mouse samples were measured as routine parameters at the Division of Clinical Chemistry and Biochemistry at the University Children’s Hospital Zurich using Alinity ci-series. LDL levels were calculated by using the Friedewald formula.
[0192] HTS data analysis. Libraries were sequenced on a MiSeq or NovaSeq 6000 (Illumina, 150bp, paired-end). Amplicon sequences were analyzed using custom Python scripts or aligned to their reference sequences using CRIPResso2. In brief, editing rates for BE and PE libraries were determined by calculating the ratio of edited bases (BE) or sequence-stretches (PE; 2bp before nick until 5bp after flap) and subsequently correcting editing rates with background rates from control samples as previously described (PRIDICT, unpublished).
[0193] Sequences
[0194] CJCasQ (wild-type) (SEQ ID NO: 1):
[0195] MARILAFD I GI S S I GWAFSENDELKDCGVRI FTKVENPKTGE SLALPRRLARSARKRLARRKARLNHLKHL IANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEK GAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQRE FGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRI INLLNNLKNTEGI LYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEI AKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVA INEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKE
[0196] QNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFD
[0197] DSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLND
[0198] TRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAID
[0199] AVI IAYANNS IVKAFSDFKKEQESNSAELYAKKI SELDYKNKRKFFEPFSGFRQKVLDKIDE IFVSKPERK
[0200] KPSGALHEETFRKEEEFYQSYGGKEGVLKALELGKIRKVNGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTM
[0201] DFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKDMQEPEFVYYNAFTSSTVSLIVSK
[0202] HDNKFETLSKNQKILFKNANEKEVIAKSIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKK evoCjCas9 (L58Y, E789K, N821K, D900K, S951G) (SEQ ID NO: 2):
[0203] MARILAFDIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHLKHL
[0204] IANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEK
[0205] GAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQRE
[0206] FGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGI
[0207] LYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEI
[0208] AKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVA
[0209] INEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKE
[0210] QNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFD
[0211] DSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLND
[0212] TRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAID
[0213] AVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERK
[0214] KPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTM
[0215] DFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLIVSK
[0216] HDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKK evoCjCas9(D8A)-ABE8e (Nuclear localization signal, TadA8e deaminase, XTEN-Linker)
[0217] (SEQ ID NO: 3):
[0218] MKRTADGSEFESPKKKRKVSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHD
[0219] PTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRWFGVRNSKRGAAGSLMNVLNYP
[0220] GMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGS
[0221] SGGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHL
[0222] KHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDD
[0223] KEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKK
[0224] QREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNT
[0225] EGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDL
[0226] NEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNL
[0227] KVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKI
[0228] EKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSR
[0229] SFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRN
[0230] LNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHH AIDAVI IAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKP
[0231] ERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPI
[0232] YTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLI
[0233] VSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKKSGGSKRT
[0234] ADGSEFEPKKKRKV evoCJCas9(D8A)-ABE8eWQ (Nuclear localization signal, TadA8e (V106W, N108Q) deaminase, XTEN-Linker) (SEQ ID NO: 4):
[0235] MKRTAD GSE FE S PKKKRKVSE VE FSHE YWRHALTLAKRARDERE VPVGAVEVLNNRVT GE GWNRAI GLHD
[0236] PTAHAEimLRQGGIAMQNYRlADATT / YVTFEPCWCAGAMIHSRIGRW'FGWRQSKRGAAGSIMNVLNYP GMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGS
[0237] SGGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHL
[0238] KHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDD
[0239] KEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKK QREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRI INLLNNLKNT
[0240] EGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDL
[0241] NEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNL
[0242] KVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKI
[0243] EKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSR
[0244] SFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRN
[0245] LNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHH
[0246] AIDAVI IAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKP
[0247] ERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPI
[0248] YTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLI
[0249] VSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKKSGGSKRT
[0250] ADGSEFEPKKKRKV evoCJCas9(D8A)-ABEmax (Nuclear localization signal, TadA-TadA7.10 deaminase, XTEN-
[0251] Linker) (SEQ ID NO: 5):
[0252] MKRTAD GSE FE S PKKKRKVSE VE FSHE YWMRHALTI AKRA HDERE WVX3AVTVHNMRVI GE GWKI'RP I GRHD
[0253] PTAHAEIMAIAQGGLWQNYRIMDATLEVTLEPCVMCAGAMIHSRIGRWFC3ARDAKTGAAGSLMDVLHHP
[0254] GMNHRVEITEGIIADECAALLSDFFRMRRQEIKAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGS
[0255] SGGSSEVEFSHEYWMRHALTIJMSWSRDEREVWGAV'IAIAINRVIGEGWKRAIGLHDPTAHAEIMALRQGGL
[0256] WQNYRLIDATLYWFEPCWCAGAMIHSRIGWVFCWRNAKTGaAGSIMDVLHYPCa-INHKVEITEGILMD
[0257] ECAALLCYFFRMPRQVFNAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSARILAFAIGIS
[0258] SIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHLKHLIANEFKLNYEDY
[0259] QSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKAIKQNEE
[0260] KLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEE
[0261] VLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRI INLLNNLKNTEGILYTKDDLNALLN
[0262] EVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITLIKDEIK LKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKKDFLPAF
[0263] NETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYKAKKDAE
[0264] LECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNKVLVFTK
[0265] QNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLNDTRYIARLVLNYT
[0266] KDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAIDAVI IAYANNS IV
[0267] KAFSDFKKEQESNSAELYAKKI SELDYKNKRKFFEPFSGFRQKVLDKIDE IFVSKPERKKPSGALHKETFR
[0268] KEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTMDFALKVLPNKAV
[0269] ARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLIVSKHDNKFETLSKNQ
[0270] KILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKKSGGSKRTADGSEFEPKKKRKV evoCjCas9(D8A)-HNHxABEmaxHD (Nuclear localization signal, TadA-TadA7.10 deaminase,
[0271] Linker sequences) (SEQ ID NO: 6):
[0272] MKRTADGSEFESPKKKRKVGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLA
[0273] RSARKRYARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVI
[0274] LHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYER
[0275] CIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFM
[0276] FVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKY
[0277] KE FIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTP
[0278] LMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKIN
[0279] IELGGS SEVEFSHE YHMRHALTLAKRAWDEREVPVGAVLVHNNRVI GEGWNRPI GRHDPTAHAE IMALRQG
[0280] GLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRWFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGIL
[0281] ADECAALLSDFFRMRRQEIKAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSEVEFSHEY
[0282] WMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLY
[0283] VTFEPCVMCAGAMIHSRIGRWFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMP
[0284] RQVFNAQKKAQSSTDSGGYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTW
[0285] GFSAKDRNNHLHHAIDAVI IAYANNS IVKAFSDFKKEQESNSAELYAKKI SELDYKNKRKFFEPFSGFRQK
[0286] VLDKIDEIFVSKPERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIF
[0287] KHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFV
[0288] YYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQ
[0289] REDFKKSGGSKRTADGSEFEPKKKRKV evoCJCas9(D8A)-HNHxABE8eHD (Nuclear localization signal, TadA-TadA8e deaminase,
[0290] Linker sequences) (SEQ ID NO: 7):
[0291] MKRTADGSEFESPKKKRKVGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLA
[0292] RSARKRYARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVI
[0293] LHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYER
[0294] CIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFM
[0295] FVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKY
[0296] KEFIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTP LMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKIN
[0297] IELGGS SEVEFSHE YWMRHALTIAKRAWDEREVPVGAVLVHNNRVI GEGWNRPI GRHDPTAHAE IMALRQG
[0298] GLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRWFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGIL
[0299] ADECAALLSDFFRMRRQEIKAQKKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSEVEFSHEY
[0300] WMRHALTLAKRARDEREVPVGAVLVLNNRVI GEGWNRAI GLHDPTAHAE IMALRQGGLVMQNYRL IDATLY
[0301] VTFEPCVMCAGAMIHSRIGRWFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMP
[0302] RQVFNAQKKAQSSTNSGGYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTW
[0303] GFSAKDRNNHLHHAIDAVI IAYANNS IVKAFSDFKKEQESNSAELYAKKI SELDYKNKRKFFEPFSGFRQK
[0304] VLDKIDEIFVSKPERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIF
[0305] KHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFV
[0306] YYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQ
[0307] REDFKKSGGSKRTADGSEFEPKKKRKV evoCjCas9(D8A)-BE4max (Nuclear localization signal, APOBEC1 deaminase, XTEN-Linker,
[0308] SGGsS-Linker, Uracil Glycosyiase Inhibitor) (SEQ ID NO: 8):
[0309] MKRTADGSEFESPKKKRKVSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRH
[0310] TSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHHAD
[0311] PRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLYVLELYCIILGLPPCLNI
[0312] LRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGGSSGGSSGSETPGTSESATPESSGGSSGGSARILA
[0313] FAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHLKHLIANEFK
[0314] LNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKA
[0315] IKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFS
[0316] KKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKDD
[0317] LNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITL
[0318] IKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKK
[0319] DFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYK
[0320] AKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNK
[0321] VLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLNDTRYIAR
[0322] LVLNYTKD YLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAIDAVIIAY
[0323] ANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERKKPSGAL
[0324] HKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTMDFALKV
[0325] LPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLIVSKHDNKFE
[0326] TLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKKSGGSGGSGGSTNLSDI
[0327] IEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGE
[0328] NKIKMLSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVM
[0329] LLTSDAPEYKPWALVIQDSNGENKIKMLSGGSKRTADGSEFEPKKKRKV evoCjCas9(D8A)-AncBE4max (Nuclear localization signal, Ancestral APOBEC1 deaminase,
[0330] XTEN-Linker, SGGsS-Linker, Uracil Glycosylase Inhibitor) (SEQ ID NO: 9);
[0331] MKRTADGSEFESPKKKRKVSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEIKWGTSHKIWRH
[0332] SSKNTTKHVEVNFIEKFTSERHFCPSTSCSITWFLSWSPCGECSKAITEFLSQHPNVTLVIYVARLYHHMD
[0333] QQNRQGLRDLVNSGVTIQIMTAPEYDYCWRNFVNYPPGKEAHWPRYPPLWMKLYALELHAGILGLPPCLNI
[0334] LRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKSGGSSGGSSGSETPGTSESATPESSGGSSGGSARILA
[0335] FAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHLKHLIANEFK
[0336] LNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKA
[0337] IKQNEEKLANYQSVGEYLYKE YFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFS
[0338] KKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKDD
[0339] LNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITL
[0340] IKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKK
[0341] DFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYK
[0342] AKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNK
[0343] VLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLNDTRYIAR
[0344] LVLNYTKD YLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAIDAVIIAY
[0345] ANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERKKPSGAL
[0346] HKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPIYTMDFALKV
[0347] LPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLIVSKHDNKFE
[0348] TLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKKSGGSGGSGGSTNLSDI
[0349] IEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGE
[0350] NKIKMLSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVM
[0351] LLTSDAPEYKPWALVIQDSNGENKIKMLSGGSKRTADGSEFEPKKKRKV evoCjCas9(D8A)-AID (Nuclear localization signal, Activation Induced Deaminase, XTEN-
[0352] Linker, SGGsS-Linker, Uracil Glycosylase Inhibitor) (SEQ ID NO: 10);
[0353] MKRTADGSEFESPKKKRKVDSLLMNRRKFLYQFKNVRWAKGRRETYLCYWKRRDSATSFSLDFGYLRNKN
[0354] GCHVELLFLRYISDWDLDPGRCYRVTWFTSWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPE
[0355] GLRRLHRAGVQIAIMTFKDYFYCWNTFVENHERTFKAWEGLHENSVRLSRQLRRILLPLYEVDDLRDAFRT
[0356] LGLSGGSSGGSSGSETPGTSESATPESSGGSSGGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVEN
[0357] PKTGESLALPRRLARSARKRYARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRA
[0358] LNELLSKQDFARVILHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSK
[0359] EFTNVRNKKESYERCIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFT
[0360] DEKRAPKNSPLAFMFVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEF
[0361] KGEKGTYFIEFKKYKEFIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKD
[0362] HLNISFKALKLVTPLMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVL
[0363] NALLKKYGKVHKINIELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQK
[0364] EFCAYSGEKIKISDLQDEKMLEIDHIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIE
[0365] VLAKNLPTKKQKRILDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSK
[0366] VHVEAKSGMLTSALRHTWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISEL DYKNKRKFFEPFSGFRQKVLDKIDEIFVSKPERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIR
[0367] KVKGKI VKNGDMFRVD I FKHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLY
[0368] KDSLILIQTKKMQEPEFVYYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVF
[0369] EKYIVSALGEVTKAEFRQREDFKKSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPE
[0370] SDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSGGSGGSTNLSDIIEKETGKQL
[0371] VIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGG
[0372] SKRTADGSEFEPKKKRKV evoCJCas9(D8A)-eAID (Nuclear localization signal, Enhanced Activation Induced
[0373] Deaminase (Anuclear export signal), XTEN-Linker, SGGsS-Linker, Uracil Glycosyiase
[0374] Inhibitor) (SEQ ID NO: 11);
[0375] MKRTADGSEFESPKKKRKVDSLLMNRREFLYQFKNVRWAKGRRETYLCYVVKRRDSATSFSLDFGYLRNKN
[0376] GCHVELLFLRYISDWDLDPGRCYRVTWFISWSPCYDCARHVADFLRGNPNLSLRIFTARLYFCEDRKAEPE
[0377] GLRRLHRAGVQIAIMTFKDYFYCWNTFVENHGRTFKAWEGLHENSVRLSRQLRRILLPSGGSSGGSSGSET
[0378] PGTSESATPESSGGSSGGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARS
[0379] ARKRYARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILH
[0380] IAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCI
[0381] AQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFV
[0382] ALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKE
[0383] FIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLM
[0384] LEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIE
[0385] LAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQ
[0386] DEKMLEIDHIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILD
[0387] KNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRH
[0388] TWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFR
[0389] QKVLDKIDEIFVSKPERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVD
[0390] IFKHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPE
[0391] FVYYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEF
[0392] RQREDFKKSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDEN
[0393] VMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEV
[0394] IGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSKRTADGSEFEPKKKR
[0395] KV evoCjCas9(D8A)-Tad-CDd (Nuclear localization signal, TadCBEd, XTEN-Linker, SGGsS-
[0396] Linker, Uracil Glycosyiase Inhibitor) (SEQ ID NO: 12);
[0397] MKRTADGSEFESPKKKRKVSEVEFSHEYWMRHALTLAKRARDERKAPVGAVLVLNNRVIGEGWNRAIGLHD
[0398] PTAHAEIIALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMINSRIGRWFGVRNSKRGAAGSLMNVLNYP
[0399] GMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSINSGGSSGGSSGSETPGTSESATPESSGGS
[0400] SGGSARILAFAIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLARSARKRYARRKARLNHL
[0401] KHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVILHIAKRRGYDDIKNSDD KEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYERCIAQSFLKDELKLIFKK
[0402] QREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFMFVALTRIINLLNNLKNT
[0403] EGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKYKEFIKALGEHNLSQDDL
[0404] NEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTPLMLEGKKYDEACNELNL
[0405] KVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKINIELAREVGKNHSQRAKI
[0406] EKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISDLQDEKMLEIDHIYPYSR
[0407] SFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRILDKNYKDKEQKNFKDRN
[0408] LNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSALRHTWGFSAKDRNNHLHH
[0409] AIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSGFRQKVLDKIDEIFVSKP
[0410] ERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFRVDIFKHKKTNKFYAVPI
[0411] YTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQEPEFVYYNAFTSSTVSLI
[0412] VSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKAEFRQREDFKKSGGSGGS
[0413] GGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWA
[0414] LVIQDSNGENKIKMLSGGSGGSGGSTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAY
[0415] DESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKMLSGGSKRTADGSEFEPKKKRKV evoCjCas9(H559A)-PEmax (Nuclear localization signal, Linker sequences, M-MLV Reverse
[0416] Transcriptase) (SEQ ID NO: 13);
[0417] MKRTADGSEFESPKKKRKVGSARILAFDIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLA
[0418] RSARKRYARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVI
[0419] LHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYER
[0420] CIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFM
[0421] FVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKY
[0422] KEFIKALGEHNLSQDDLNEIAKDITLIKDEIKLKKALAKYDLNQNQIDSLSKLEFKDHLNISFKALKLVTP
[0423] LMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKIN IELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISD
[0424] LQDEKMLEIDAIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRI
[0425] LDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSAL
[0426] RHTWGFSAKDRNNHLHHAIDAVIIAYANNSIVKAFSDFKKEQESNSAELYAKKISELDYKNKRKFFEPFSG
[0427] FRQKVLDKIDEIFVSKPERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFR
[0428] VDIFKHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQE
[0429] PEFVYYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKA
[0430] EFRQREDFKKSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLS
[0431] DFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPL
[0432] LPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFA
[0433] FEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGT
[0434] RALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCR
[0435] LFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQE IKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLT
[0436] QKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLS
[0437] NARMTHYQALLLDTDRVQFGPWALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADHTWYTD GSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHG
[0438] E I YRRRGWLT SEGKE IKNKDE I LALLKALFLPKRLS I I HCPGHQKGHSAE ARGNRMADQAARKAAI TETPD
[0439] TSTLLIENSSPSGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD evoCjCas9(H559A)-PEmax-ARnaseH (Nuclear localization signal, Linker sequences, M-MLV
[0440] Reverse Transcriptase (ARnaseH)) (SEQ ID NO: 14):
[0441] MKRTADGSEFESPKKKRKVGSARILAFDIGISSIGWAFSENDELKDCGVRIFTKVENPKTGESLALPRRLA
[0442] RSARKRYARRKARLNHLKHLIANEFKLNYEDYQSFDESLAKAYKGSLISPYELRFRALNELLSKQDFARVI
[0443] LHIAKRRGYDDIKNSDDKEKGAILKAIKQNEEKLANYQSVGEYLYKEYFQKFKENSKEFTNVRNKKESYER
[0444] CIAQSFLKDELKLIFKKQREFGFSFSKKFEEEVLSVAFYKRALKDFSHLVGNCSFFTDEKRAPKNSPLAFM
[0445] FVALTRIINLLNNLKNTEGILYTKDDLNALLNEVLKNGTLTYKQTKKLLGLSDDYEFKGEKGTYFIEFKKY
[0446] KE FIKALGEHNLSQDDLNE I AKD I TL IKDE IKLKKALAKYDLNQNQ IDSLSKLE FKDHLNI S FKALKLVTP
[0447] LMLEGKKYDEACNELNLKVAINEDKKDFLPAFNETYYKDEVTNPWLRAIKEYRKVLNALLKKYGKVHKIN
[0448] IELAREVGKNHSQRAKIEKEQNENYKAKKDAELECEKLGLKINSKNILKLRLFKEQKEFCAYSGEKIKISD
[0449] LQDEKMLEIDAIYPYSRSFDDSYMNKVLVFTKQNQEKLNQTPFEAFGNDSAKWQKIEVLAKNLPTKKQKRI
[0450] LDKNYKDKEQKNFKDRNLNDTRYIARLVLNYTKDYLDFLPLSDDENTKLNDTQKGSKVHVEAKSGMLTSAL
[0451] RHTWGFSAKDRNNHLHHAIDAVI I AY ANNS IVKAFSDFKKEQESNSAELYAKKI SELDYKNKRKFFEPFSG
[0452] FRQKVLDKIDEIFVSKPERKKPSGALHKETFRKEEEFYQSYGGKEGVLKALELGKIRKVKGKIVKNGDMFR
[0453] VDIFKHKKTNKFYAVPIYTMDFALKVLPNKAVARSKKGEIKDWILMDENYEFCFSLYKDSLILIQTKKMQE
[0454] PEFVYYNAFTSSTVSLIVSKHDNKFETLSKNQKILFKNANEKEVIAKGIGIQNLKVFEKYIVSALGEVTKA
[0455] EFRQREDFKKSGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGSTLNIEDEYRLHETSKEPDVSLGSTWLS
[0456] DFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPL
[0457] LPVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFA
[0458] FEWRDPEMGISGQLTWTRLPQGFKNSPTLFNEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGT
[0459] RALLQTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGKAGFCR
[0460] LFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQE IKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLT
[0461] QKLGPWRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLS
[0462] NARMTHYQALLLDTDRVQFGPWALNPSGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD
[0463] AAV constructs :
[0464] P3-evoCJCas9(D8A)-ABE8e (ITR, Promoter (P3), chimeric intron (chl), Kozak, Nuclear localization signal, TadASe, Linker-sequence, evoCJCas9, poly-A, evoCJCas9-scaffold RNA,
[0465] Spacer Sequence (21-23bp), Promoter (U6), ITR) (SEQ ID NO: 15):
[0466] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGC
[0467] CTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCTACCGGCGCGCCG
[0468] GGGGAGGCTGCTGGTGAATATTAACCAAGGTCACCCCAGTTATCGGAGGAGCAAACAGGGGCTAAGTCCAC
[0469] ACGCGTGGTACCGTCTGTCTGCACATTTCGTAGAGCGAGTGTTCCGATACTCTAATCTCCCTAGGCAAGGT
[0470] TCATATTTGTGTAGGTTACTTATTCTCCTTTTGTTGACTAAGTCAATAATCAGAATCAGCAGGTTTGGAGT
[0471] CAGCTTGGCAGGGATCAGCAGCCTGGGTTGGAAGGAGGGGGTATAAAAGCCCCTTCACCAGGAGAAGCCGT
[0472] CACACAGATCCACAAGCTCCTGGGGTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATAGAAACT GGGCTTGTCGAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCACTTTG
[0473] CCTTTCTCTCCACAGGGCCGCCACCATGAAACGGACAGCCGACGGAAGCGAGTTCGAGTCACCAAAGAAGA
[0474] AGCGGAAAGTCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGCCAAGAGG
[0475] GCACGGGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCGAGGGCTG
[0476] GAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCTGAGACAGGGCGGCCTGG
[0477] TCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTGCGCCGGC
[0478] GCCATGATCCACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCGCAGGCTC
[0479] CCTGATGAACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTGGCAGATG
[0480] AATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAAGGCCCAG
[0481] AGCTCCATCAACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGGCACAA
[0482] GCGAGAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGCCCGCATCCTGG
[0483] CCTTTGCCATCGGCATCAGCAGCATCGGATGGGCCTTCAGCGAGAATGACGAGCTGAAGGACTGCGGTGTG
[0484] AGAATCTTTACAAAGGTCGAGAACCCCAAGACCGGCGAAAGCCTGGCGCTGCCAAGACGGCTGGCCCGGTC
[0485] TGCTCGGAAGAGATACGCCCGGAGAAAAGCCAGACTGAATCACCTGAAACACCTGATCGCCAACGAGTTCA
[0486] AACTGAACTACGAGGACTACCAGAGCTTCGATGAGAGCTTGGCTAAGGCCTACAAAGGCAGCCTGATCAGC
[0487] CCCTACGAGCTGAGATTCAGAGCTCTGAATGAGCTCCTGAGCAAGCAGGACTTCGCTAGAGTGATCCTGCA
[0488] CATCGCCAAGCGAAGAGGCTACGACGACATCAAGAACTCTGATGACAAGGAGAAGGGCGCCATTCTGAAAG
[0489] CCATCAAGCAGAACGAAGAGAAACTGGCCAATTACCAGAGCGTGGGCGAGTACCTGTACAAGGAGTACTTC
[0490] CAGAAGTTCAAAGAAAATAGCAAAGAGTTCACCAACGTGCGGAACAAGAAGGAGTCTTATGAAAGATGTAT
[0491] CGCCCAGAGCTTCCTGAAAGACGAGCTGAAACTGATCTTCAAGAAGCAAAGAGAGTTTGGCTTCAGCTTCA
[0492] GCAAAAAATTTGAAGAAGAGGTGCTGTCTGTCGCCTTCTACAAGAGGGCTCTGAAGGACTTCAGCCACCTG
[0493] GTGGGCAATTGCAGCTTTTTCACAGACGAAAAGCGGGCCCCTAAGAACAGCCCTCTGGCCTTCATGTTCGT
[0494] GGCTCTGACCAGAATCATCAACCTGCTGAACAACCTGAAAAATACCGAGGGCATCCTCTATACCAAGGACG
[0495] ATCTGAACGCCCTGCTGAATGAGGTGCTCAAGAATGGCACCCTGACCTACAAGCAAACAAAAAAACTGCTG
[0496] GGCCTGTCTGACGACTACGAGTTTAAAGGCGAGAAGGGCACCTATTTCATCGAATTTAAGAAGTACAAGGA
[0497] ATTCATCAAGGCACTGGGCGAACACAACCTGTCTCAGGACGACCTGAACGAGATCGCCAAGGACATCACCC
[0498] TGATCAAGGATGAGATCAAGCTGAAAAAAGCTCTGGCCAAGTACGACCTGAATCAAAACCAGATCGACAGC
[0499] TTAAGCAAGCTGGAATTTAAGGATCACCTGAACATCTCCTTTAAGGCCCTGAAGCTGGTGACCCCACTGAT
[0500] GCTGGAAGGAAAGAAGTACGACGAAGCATGCAACGAGCTTAACCTGAAGGTTGCTATCAACGAGGATAAGA
[0501] AGGATTTCCTGCCTGCCTTTAACGAGACATACTACAAGGATGAGGTGACCAACCCCGTGGTGCTGAGAGCT
[0502] ATCAAAGAGTACAGAAAGGTGCTGAACGCCCTGCTGAAGAAGTACGGCAAGGTCCACAAGATCAATATTGA
[0503] GCTGGCCCGGGAGGTTGGAAAGAACCACTCTCAGAGAGCAAAGATCGAGAAGGAGCAAAACGAGAACTACA
[0504] AAGCGAAGAAGGACGCCGAACTGGAGTGCGAGAAGCTTGGCCTGAAGATCAACTCTAAGAATATCCTGAAA
[0505] CTCAGACTTTTCAAAGAACAGAAGGAATTCTGTGCCTACAGCGGCGAGAAGATCAAAATTTCTGACCTGCA
[0506] GGATGAAAAGATGCTGGAGATCGACCACATCTACCCTTACTCCAGAAGCTTCGACGACAGCTATATGAACA
[0507] AAGTGCTGGTGTTCACAAAGCAGAACCAGGAGAAGCTGAATCAGACCCCTTTCGAGGCCTTCGGCAATGAC
[0508] TCCGCCAAGTGGCAGAAAATCGAGGTGCTGGCCAAAAACCTGCCAACCAAGAAACAGAAGAGAATCCTCGA
[0509] CAAGAACTACAAGGACAAGGAACAGAAGAACTTCAAGGATCGGAACCTGAACGACACCCGGTACATCGCCA
[0510] GGCTGGTGTTAAATTACACCAAGGACTACCTGGATTTCCTGCCCCTGAGCGACGACGAGAACACCAAGCTG
[0511] AACGACACACAGAAGGGCAGCAAGGTGCACGTGGAAGCCAAGAGCGGCATGCTGACCAGCGCACTGCGCCA CACTTGGGGCTTCAGCGCTAAGGACCGGAACAACCACCTGCATCACGCCATCGATGCCGTGATCATAGCCT
[0512] ACGCCAACAACTCAATCGTGAAAGCTTTCAGTGACTTTAAGAAAGAACAGGAGAGCAACTCTGCCGAACTG
[0513] TACGCCAAGAAAATTAGCGAGCTGGACTACAAGAACAAACGGAAGTTCTTCGAACCTTTCTCAGGATTTAG
[0514] ACAGAAGGTGCTGGATAAGATCGATGAAATCTTCGTGTCCAAGCCCGAGAGAAAGAAGCCTAGCGGAGCCC
[0515] TGCACAAGGAAACCTTCAGAAAGGAAGAAGAGTTCTACCAGTCTTATGGGGGCAAAGAAGGCGTGCTGAAG
[0516] GCCCTGGAACTCGGCAAGATCAGAAAGGTGAAGGGAAAGATTGTGAAGAACGGCGACATGTTCAGAGTGGA
[0517] CATCTTCAAGCACAAGAAGACCAACAAGTTCTATGCCGTGCCTATCTACACAATGGATTTCGCCCTGAAAG
[0518] TGCTGCCTAACAAGGCTGTCGCCAGATCTAAGAAGGGCGAGATCAAGGACTGGATCCTGATGGACGAAAAC
[0519] TACGAGTTCTGCTTCAGCCTGTACAAGGACAGCCTGATCCTGATCCAAACAAAGAAGATGCAGGAGCCTGA
[0520] ATTCGTGTACTACAACGCCTTCACAAGCAGCACCGTGAGCCTGATCGTGTCTAAACATGATAACAAGTTCG
[0521] AAACCCTGTCCAAGAACCAAAAGATCCTGTTCAAGAACGCCAACGAGAAGGAAGTGATCGCCAAGGGCATC
[0522] GGCATTCAGAATCTGAAGGTGTTCGAGAAATATATCGTGTCCGCTCTGGGAGAGGTTACAAAGGCCGAGTT
[0523] TCGGCAGAGAGAAGATTTTAAGAAGTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCA
[0524] AGAAGAAGAGGAAAGTCTAATAGATCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCC
[0525] CCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATC
[0526] GCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGG
[0527] AAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAGCGGCCCAAAAAAAGCGGTTTTAG
[0528] GGGATTGTAACCCCGCAGAGTCCCGCAAACTCTTTATTCTAGTCCCTTTTCAGGGACTAGAACNNNNNNN
[0529] NNNNNNNNNNNNNNNGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATCGAAATACTTTC
[0530] AAGTTACGGTAAGCATATGATAGTCCATTTTAAAACATAATTTTAAAACTGCAAACTACCCAAGAAATTAT
[0531] TACTTTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTCT
[0532] AACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCAGGAACCCCTAGTGATG
[0533] GAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCC
[0534] GGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG
[0535] MSyn-evoCJCas9(D8A)-ABE8e (ITR, Promoter (hSyn), Kozak, Nuclear localization signal,
[0536] TadASe, Linker-sequence, evoCJCas9, poly-A, evoCJCas9-scaffold RNA, Spacer Sequence
[0537] (21-23bp), Promoter (U6), ITR) (SEQ ID NO: 16):
[0538] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGC
[0539] CTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCTGAGGGCCCTGCG
[0540] TATGAGTGCAAGTGGGTTTTAGGACCAGGATGAGGCGGGGTGGGGGTGCCTACCTGACGACCGACCCCGAC
[0541] CCACTGGACAAGCACCCAACCCCCATTCCCCAAATTGCGCATCCCCTATCAGAGAGGGGGAGGGGAAACAG
[0542] GATGCGGCGAGGCGCGTGCGCACTGCCAGCTTCAGCACCGCGGACAGTGCCTTCGCCCCCGCCTGGCGGCG
[0543] CGCGCCACCGCCGCCTCAGCACTGAAGGCGCGCTGACGTCACTCGCCGGTCCCCCGCAAACTCCCCTTCCC
[0544] GGCCACCTTGGTCGCGTCCGCGCCGCCGCCGGCCCAGCCGGACCGCACCACGCGAGGCGCGAGATAGGGGG
[0545] GCACGGGCGCGACCATCTGCGCTGCGGCGCCGGCGACTCAGCGCTGCCTCAGTCTGCGGTGGGCAGCGGAG
[0546] GAGTCGTGTCGTGCCTGAGAGCGCAGTCGAGAGGCCGCCACCATGAAACGGACAGCCGACGGAAGCGAGTT
[0547] CGAGTCACCAAAGAAGAAGCGGAAAGTCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCC
[0548] TGACCCTGGCCAAGAGGGCACGGGATGAGAGGGAGGTGCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGA
[0549] GTGATCGGCGAGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATGGCCCT
[0550] CTCAGAAGAAGGCCCAGAGCTCCATCAACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGA
[0551] GACACCTGGCACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGG
[0552] TCAGCCCGCATCCTGGCCTTTGCCATCGGCATCAGCAGCATCGGATGGGCCTTCAGCGAGAATGACGAGCT
[0553] GAAGGACTGCGGTGTGAGAATCTTTACAAAGGTCGAGAACCCCAAGACCGGCGAAAGCCTGGCGCTGCCAA
[0554] GACGGCTGGCCCGGTCTGCTCGGAAGAGATACGCCCGGAGAAAAGCCAGACTGAATCACCTGAAACACCTG
[0555] ATCGCCAACGAGTTCAAACTGAACTACGAGGACTACCAGAGCTTCGATGAGAGCTTGGCTAAGGCCTACAA
[0556] AGGCAGCCTGATCAGCCCCTACGAGCTGAGATTCAGAGCTCTGAATGAGCTCCTGAGCAAGCAGGACTTCG
[0557] CTAGAGTGATCCTGCACATCGCCAAGCGAAGAGGCTACGACGACATCAAGAACTCTGATGACAAGGAGAAG
[0558] GGCGCCATTCTGAAAGCCATCAAGCAGAACGAAGAGAAACTGGCCAATTACCAGAGCGTGGGCGAGTACCT
[0559] GTACAAGGAGTACTTCCAGAAGTTCAAAGAAAATAGCAAAGAGTTCACCAACGTGCGGAACAAGAAGGAGT
[0560] CTTATGAAAGATGTATCGCCCAGAGCTTCCTGAAAGACGAGCTGAAACTGATCTTCAAGAAGCAAAGAGAG
[0561] TTTGGCTTCAGCTTCAGCAAAAAATTTGAAGAAGAGGTGCTGTCTGTCGCCTTCTACAAGAGGGCTCTGAA
[0562] GGACTTCAGCCACCTGGTGGGCAATTGCAGCTTTTTCACAGACGAAAAGCGGGCCCCTAAGAACAGCCCTC
[0563] TGGCCTTCATGTTCGTGGCTCTGACCAGAATCATCAACCTGCTGAACAACCTGAAAAATACCGAGGGCATC
[0564] CTCTATACCAAGGACGATCTGAACGCCCTGCTGAATGAGGTGCTCAAGAATGGCACCCTGACCTACAAGCA
[0565] AACAAAAAAACTGCTGGGCCTGTCTGACGACTACGAGTTTAAAGGCGAGAAGGGCACCTATTTCATCGAAT
[0566] TTAAGAAGTACAAGGAATTCATCAAGGCACTGGGCGAACACAACCTGTCTCAGGACGACCTGAACGAGATC
[0567] GCCAAGGACATCACCCTGATCAAGGATGAGATCAAGCTGAAAAAAGCTCTGGCCAAGTACGACCTGAATCA
[0568] AAACCAGATCGACAGCTTAAGCAAGCTGGAATTTAAGGATCACCTGAACATCTCCTTTAAGGCCCTGAAGC
[0569] TGGTGACCCCACTGATGCTGGAAGGAAAGAAGTACGACGAAGCATGCAACGAGCTTAACCTGAAGGTTGCT
[0570] ATCAACGAGGATAAGAAGGATTTCCTGCCTGCCTTTAACGAGACATACTACAAGGATGAGGTGACCAACCC
[0571] CGTGGTGCTGAGAGCTATCAAAGAGTACAGAAAGGTGCTGAACGCCCTGCTGAAGAAGTACGGCAAGGTCC
[0572] ACAAGATCAATATTGAGCTGGCCCGGGAGGTTGGAAAGAACCACTCTCAGAGAGCAAAGATCGAGAAGGAG
[0573] CAAAACGAGAACTACAAAGCGAAGAAGGACGCCGAACTGGAGTGCGAGAAGCTTGGCCTGAAGATCAACTC
[0574] TAAGAATATCCTGAAACTCAGACTTTTCAAAGAACAGAAGGAATTCTGTGCCTACAGCGGCGAGAAGATCA
[0575] AAATTTCTGACCTGCAGGATGAAAAGATGCTGGAGATCGACCACATCTACCCTTACTCCAGAAGCTTCGAC
[0576] GACAGCTATATGAACAAAGTGCTGGTGTTCACAAAGCAGAACCAGGAGAAGCTGAATCAGACCCCTTTCGA
[0577] GGCCTTCGGCAATGACTCCGCCAAGTGGCAGAAAATCGAGGTGCTGGCCAAAAACCTGCCAACCAAGAAAC
[0578] AGAAGAGAATCCTCGACAAGAACTACAAGGACAAGGAACAGAAGAACTTCAAGGATCGGAACCTGAACGAC
[0579] ACCCGGTACATCGCCAGGCTGGTGTTAAATTACACCAAGGACTACCTGGATTTCCTGCCCCTGAGCGACGA
[0580] CGAGAACACCAAGCTGAACGACACACAGAAGGGCAGCAAGGTGCACGTGGAAGCCAAGAGCGGCATGCTGA
[0581] CCAGCGCACTGCGCCACACTTGGGGCTTCAGCGCTAAGGACCGGAACAACCACCTGCATCACGCCATCGAT
[0582] GCCGTGATCATAGCCTACGCCAACAACTCAATCGTGAAAGCTTTCAGTGACTTTAAGAAAGAACAGGAGAG
[0583] CAACTCTGCCGAACTGTACGCCAAGAAAATTAGCGAGCTGGACTACAAGAACAAACGGAAGTTCTTCGAAC
[0584] CTTTCTCAGGATTTAGACAGAAGGTGCTGGATAAGATCGATGAAATCTTCGTGTCCAAGCCCGAGAGAAAG
[0585] AAGCCTAGCGGAGCCCTGCACAAGGAAACCTTCAGAAAGGAAGAAGAGTTCTACCAGTCTTATGGGGGCAA AGAAGGCGTGCTGAAGGCCCTGGAACTCGGCAAGATCAGAAAGGTGAAGGGAAAGATTGTGAAGAACGGCG
[0586] ACATGTTCAGAGTGGACATCTTCAAGCACAAGAAGACCAACAAGTTCTATGCCGTGCCTATCTACACAATG
[0587] GATTTCGCCCTGAAAGTGCTGCCTAACAAGGCTGTCGCCAGATCTAAGAAGGGCGAGATCAAGGACTGGAT
[0588] CCTGATGGACGAAAACTACGAGTTCTGCTTCAGCCTGTACAAGGACAGCCTGATCCTGATCCAAACAAAGA
[0589] AGATGCAGGAGCCTGAATTCGTGTACTACAACGCCTTCACAAGCAGCACCGTGAGCCTGATCGTGTCTAAA
[0590] CATGATAACAAGTTCGAAACCCTGTCCAAGAACCAAAAGATCCTGTTCAAGAACGCCAACGAGAAGGAAGT
[0591] GATCGCCAAGGGCATCGGCATTCAGAATCTGAAGGTGTTCGAGAAATATATCGTGTCCGCTCTGGGAGAGG
[0592] TTACAAAGGCCGAGTTTCGGCAGAGAGAAGATTTTAAGAAGTCTGGCGGCTCAAAAAGAACCGCCGACGGC
[0593] AGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAATAGATCTCGACTGTGCCTTCTAGTTGCCAGCCATC
[0594] TGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAA
[0595] ATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGC
[0596] AAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATGGCTCGAGCGGCCCA
[0597] AAAAAAGCGGTTTTAGGGGATTGTAACCCCGCAGAGTCCCGCAAACTCTTTATTCTAGTCCCTTTTCAGGG
[0598] ACTAGAACNNNNNNNNNNNNNNNNNNNNNNNNGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAG
[0599] AAATCGAAATACTTTCAAGTTACGGTAAGCATATGATAGTCCATTTTAAAACATAATTTTAAAACTGCAAA
[0600] CTACCCAAGAAATTATTACTTTCTACGTCACGTATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAA
[0601] TTCTAATTATCTCTCTAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCA
[0602] GGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAA
[0603] AGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAG
[0604] P3-evoCJCas9(D8A)-eAID-CBE ( ITR, Promoter (P3), Kozak, Nuclear localization signal, enhanced activation induced deaminase (eAID), Linker-sequence, evoCJCas9, Uracil
[0605] Glycosylase Inhibitor, poly-A, evoCJCas9-scaffold RNA, Spacer Sequence (21-23bp),
[0606] Promoter (U6), ITR) (SEQ ID NO: 17):
[0607] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGC CTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCTACCGGCGCGCCG
[0608] GGGGAGGCTGCTGGTGAATATTAACCAAGGTCACCCCAGTTATCGGAGGAGCAAACAGGGGCTAAGTCCAC
[0609] ACGCGTGGTACCGTCTGTCTGCACATTTCGTAGAGCGAGTGTTCCGATACTCTAATCTCCCTAGGCAAGGT
[0610] TCATATTTGTGTAGGTTACTTATTCTCCTTTTGTTGACTAAGTCAATAATCAGAATCAGCAGGTTTGGAGT
[0611] CAGCTTGGCAGGGATCAGCAGCCTGGGTTGGAAGGAGGGGGTATAAAAGCCCCTTCACCAGGAGAAGCCCT
[0612] CACACAGATCCACAAGCTCCTGGGCCCACCATAGAA CGGACAGCCGACGGAAGCGAGTTCGAGTCACCA
[0613] AAGAAGAAGCGGAAAGTCAGTGATAGCCTCTTGATGAAT AGACGCGAATTCCTGTATCAGTTTAAAAACGT
[0614] GAGATGGGCAAAAGGCCGACGAGAGACATATCTGTGCTATGTCGTTAAGCGCAGAGATTCAGCCACCAGTT
[0615] TCTCTCTCGACTTCGGCTACCTGCGGAACAAGAATGGTTGCCATGTTGAGCTCCTGTTCCTGAGGTATATC
[0616] AGCGACTGGGATTTGGACCCAGGGCGGTGCTATAGGGTGACATGGTTTATTTCCTGGTCACCTTGTTATGA
[0617] CTGCGCGCGGCATGTTGCCGATTTTCTGAGAGGGAACCCTAACCTGTCTCTGAGGATCTTCACCGCGCGAC
[0618] TGTACTTCTGTGAGGACCGGAAAGCCGAACCCGAGGGACTGAGACGCCTCCACAGAGCGGGTGTGCAGATT
[0619] GCCATAATGACCTTTAAGGA.CTACTTCTACTGCTGGAACACCTTCGTCGAAAATCACGGCCGGACTTTCAA
[0620] GGCTTGGGAAGGATTGCATGAAAACAGCGTCAGGCTTTCCAGGCAGCTTCGCCGCATTCTTCTCCCGTCTG
[0621] GCGGATCTAGCGGAGGATCCTCTGGCAGCGAGACACCAGGAACAAGCGAGTCAGCAACA CCAGAGAGCAGTGGCGGCAGCAGCGGCGGGTCAGCCCGCATCCTGGCCTTTGCCATCGGCATCA
[0622] GCAGCATCGGATGGGCCTTCAGCGAGAATGACGAGCTGAAGGACTGCGGTGTGAGAATCTTTACAAAGGTC
[0623] GAGAACCCCAAGACCGGCGAAAGCCTGGCGCTGCCAAGACGGCTGGCCCGGTCTGCTCGGAAGAGATACGC
[0624] CCGGAGAAAAGCCAGACTGAATCACCTGAAACACCTGATCGCCAACGAGTTCAAACTGAACTACGAGGACT
[0625] ACCAGAGCTTCGATGAGAGCTTGGCTAAGGCCTACAAAGGCAGCCTGATCAGCCCCTACGAGCTGAGATTC
[0626] AGAGCTCTGAATGAGCTCCTGAGCAAGCAGGACTTCGCTAGAGTGATCCTGCACATCGCCAAGCGAAGAGG
[0627] CTACGACGACATCAAGAACTCTGATGACAAGGAGAAGGGCGCCATTCTGAAAGCCATCAAGCAGAACGAAG
[0628] AGAAACTGGCCAATTACCAGAGCGTGGGCGAGTACCTGTACAAGGAGTACTTCCAGAAGTTCAAAGAAAAT
[0629] AGCAAAGAGTTCACCAACGTGCGGAACAAGAAGGAGTCTTATGAAAGATGTATCGCCCAGAGCTTCCTGAA
[0630] AGACGAGCTGAAACTGATCTTCAAGAAGCAAAGAGAGTTTGGCTTCAGCTTCAGCAAAAAATTTGAAGAAG
[0631] AGGTGCTGTCTGTCGCCTTCTACAAGAGGGCTCTGAAGGACTTCAGCCACCTGGTGGGCAATTGCAGCTTT
[0632] TTCACAGACGAAAAGCGGGCCCCTAAGAACAGCCCTCTGGCCTTCATGTTCGTGGCTCTGACCAGAATCAT
[0633] CAACCTGCTGAACAACCTGAAAAATACCGAGGGCATCCTCTATACCAAGGACGATCTGAACGCCCTGCTGA
[0634] ATGAGGTGCTCAAGAATGGCACCCTGACCTACAAGCAAACAAAAAAACTGCTGGGCCTGTCTGACGACTAC
[0635] GAGTTTAAAGGCGAGAAGGGCACCTATTTCATCGAATTTAAGAAGTACAAGGAATTCATCAAGGCACTGGG
[0636] CGAACACAACCTGTCTCAGGACGACCTGAACGAGATCGCCAAGGACATCACCCTGATCAAGGATGAGATCA
[0637] AGCTGAAAAAAGCTCTGGCCAAGTACGACCTGAATCAAAACCAGATCGACAGCTTAAGCAAGCTGGAATTT
[0638] AAGGATCACCTGAACATCTCCTTTAAGGCCCTGAAGCTGGTGACCCCACTGATGCTGGAAGGAAAGAAGTA
[0639] CGACGAAGCATGCAACGAGCTTAACCTGAAGGTTGCTATCAACGAGGATAAGAAGGATTTCCTGCCTGCCT
[0640] TTAACGAGACATACTACAAGGATGAGGTGACCAACCCCGTGGTGCTGAGAGCTATCAAAGAGTACAGAAAG
[0641] GTGCTGAACGCCCTGCTGAAGAAGTACGGCAAGGTCCACAAGATCAATATTGAGCTGGCCCGGGAGGTTGG
[0642] AAAGAACCACTCTCAGAGAGCAAAGATCGAGAAGGAGCAAAACGAGAACTACAAAGCGAAGAAGGACGCCG
[0643] AACTGGAGTGCGAGAAGCTTGGCCTGAAGATCAACTCTAAGAATATCCTGAAACTCAGACTTTTCAAAGAA
[0644] CAGAAGGAATTCTGTGCCTACAGCGGCGAGAAGATCAAAATTTCTGACCTGCAGGATGAAAAGATGCTGGA
[0645] GATCGACCACATCTACCCTTACTCCAGAAGCTTCGACGACAGCTATATGAACAAAGTGCTGGTGTTCACAA
[0646] AGCAGAACCAGGAGAAGCTGAATCAGACCCCTTTCGAGGCCTTCGGCAATGACTCCGCCAAGTGGCAGAAA
[0647] ATCGAGGTGCTGGCCAAAAACCTGCCAACCAAGAAACAGAAGAGAATCCTCGACAAGAACTACAAGGACAA
[0648] GGAACAGAAGAACTTCAAGGATCGGAACCTGAACGACACCCGGTACATCGCCAGGCTGGTGTTAAATTACA
[0649] CCAAGGACTACCTGGATTTCCTGCCCCTGAGCGACGACGAGAACACCAAGCTGAACGACACACAGAAGGGC
[0650] AGCAAGGTGCACGTGGAAGCCAAGAGCGGCATGCTGACCAGCGCACTGCGCCACACTTGGGGCTTCAGCGC
[0651] TAAGGACCGGAACAACCACCTGCATCACGCCATCGATGCCGTGATCATAGCCTACGCCAACAACTCAATCG
[0652] TGAAAGCTTTCAGTGACTTTAAGAAAGAACAGGAGAGCAACTCTGCCGAACTGTACGCCAAGAAAATTAGC
[0653] GAGCTGGACTACAAGAACAAACGGAAGTTCTTCGAACCTTTCTCAGGATTTAGACAGAAGGTGCTGGATAA
[0654] GATCGATGAAATCTTCGTGTCCAAGCCCGAGAGAAAGAAGCCTAGCGGAGCCCTGCACAAGGAAACCTTCA
[0655] GAAAGGAAGAAGAGTTCTACCAGTCTTATGGGGGCAAAGAAGGCGTGCTGAAGGCCCTGGAACTCGGCAAG
[0656] ATCAGAAAGGTGAAGGGAAAGATTGTGAAGAACGGCGACATGTTCAGAGTGGACATCTTCAAGCACAAGAA
[0657] GACCAACAAGTTCTATGCCGTGCCTATCTACACAATGGATTTCGCCCTGAAAGTGCTGCCTAACAAGGCTG
[0658] TCGCCAGATCTAAGAAGGGCGAGATCAAGGACTGGATCCTGATGGACGAAAACTACGAGTTCTGCTTCAGC
[0659] CTGTACAAGGACAGCCTGATCCTGATCCAAACAAAGAAGATGCAGGAGCCTGAATTCGTGTACTACAACGC
[0660] CTTCACAAGCAGCACCGTGAGCCTGATCGTGTCTAAACATGATAACAAGTTCGAAACCCTGTCCAAGAACC AAAAGATCCTGTTCAAGAACGCCAACGAGAAGGAAGTGATCGCCAAGGGCATCGGCATTCAGAATCTGAAG
[0661] GTGTTCGAGAAATATATCGTGTCCGCTCTGGGAGAGGTTACAAAGGCCGAGTTTCGGCAGAGAGAAGATTT
[0662] TAAGAAGAGCGGCGGGAGCGGCGGGAGCGGGGGGAGCACTAATCTGAGCGACATCATTGA
[0663] GAAGGAGACTGGGAAACAGCTGGTCATTCAGGAGTCCATCCTGATGCTGCCTGAGGAGGT
[0664] GGAGGAAGTGATCGGCAACAAGCCAGAGTCTGACATCCTGGTGCACACCGCCTACGACG
[0665] AGTCCACAGATGAGAATGTGATGCTGCTGACCTCTGACGCCCCCGAGTATAAGCCTTGGG
[0666] CCCTGGTCATCCAGGATTCTAACGGCGAGAATAAGATCAAGATGCTGTCTGGCGGCTCAAAAA
[0667] GAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGTCTAATAGATCTCGACTGTGCCTTCT
[0668] AGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGT
[0669] CCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGG
[0670] TGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGGGCTCTATG
[0671] GCTCGAGCGGCCCAAAAAAAGCGGTTTTAGGGGATTGTAACCCCGCAGAGTCCCGCAAACTCTTTATTCTA
[0672] GTCCCTTTTCAGGGACTAGAACNNNNNNNNNNNNNNNNNNNNNNGGTGTTTCGTCCTTTCCACAAGA
[0673] TATATAAAGCCAAGAAATCGAAATACTTTCAAGTTACGGTAAGCATATGATAGTCCATTTTAAAACATAAT
[0674] TTTAAAACTGCAAACTACCCAAGAAATTATTACTTTCTACGTCACGTATTTTGTACTAATATCTTTGTGTT
[0675] TACAGTCAAATTAATTCTAATTATCTCTCTAACAGCCTTGTATCGTATATGCAAATATGAAGGAATCATGG
[0676] GAAATAGGCCCTCAGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGA
[0677] GGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCA
[0678] G
[0679] P3-evoCJCas9(D8A)-TadCD-CBE (ITR, Promoter (P3), Kozak, Nuclear localization signal,
[0680] Tad-cytosine deaminase (TadCD), Linker-sequence, evoCJCas9, Uracil Glycosylase
[0681] Inhibitor, poly-A, evoCJCas9-scaffold RNA, Spacer Sequence (21-23bp), Promoter (U6),
[0682] ITR) (SEQ ID NO: 18):
[0683] CTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCAAAGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGC
[0684] CTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCTACCGGCGCGCCG
[0685] GGGGAGGCTGCTGGTGAATATTAACCAAGGTCACCCCAGTTATCGGAGGAGCAAACAGGGGCTAAGTCCAC
[0686] ACGCGTGGTACCGTCTGTCTGCACATTTCGTAGAGCGAGTGTTCCGATACTCTAATCTCCCTAGGCAAGGT
[0687] TCATATTTGTGTAGGTTACTTATTCTCCTTTTGTTGACTAAGTCAATAATCAGAATCAGCAGGTTTGGAGT
[0688] CAGCTTGGCAGGGATCAGCAGCCTGGGTTGGAAGGAGGGGGTATAAAAGCCCCTTCACCAGGAGAAGCCGT
[0689] CACACAGATCCACAAGCTCCTGGGCCGCCACCATGAAACGGACAGCCGACGGAAGCGAGTTCGAGTCACCA
[0690] AAGAAGAAGCGGAAAGTCTCTGAGGTGGAGTTTTCCCACGAGTACTGGATGAGACATGCCCTGACCCTGGC
[0691] CAAGAGGGCACGGGATGAGAGGAAGGCCCCTGTGGGAGCCGTGCTGGTGCTGAACAATAGAGTGATCGGCG
[0692] AGGGCTGGAACAGAGCCATCGGCCTGCACGACCCAACAGCCCATGCCGAAATTATCGCCCTGAGACAGGGC
[0693] GGCCTGGTCATGCAGAACTACAGACTGATTGACGCCACCCTGTACGTGACATTCGAGCCTTGCGTGATGTG
[0694] CGCCGGCGCCATGATCAACTCTAGGATCGGCCGCGTGGTGTTTGGCGTGAGGAACTCAAAAAGAGGCGCCG
[0695] CAGGCTCCCTGATGAACGTGCTGAACTACCCCGGCATGAATCACCGCGTCGAAATTACCGAGGGAATCCTG
[0696] GCAGATGAATGTGCCGCCCTGCTGTGCGATTTCTATCGGATGCCTAGACAGGTGTTCAATGCTCAGAAGAA
[0697] GGCCCAGAGCTCCATCAACTCCGGAGGATCTAGCGGAGGCTCCTCTGGCTCTGAGACACCTGG
[0698] CACAAGCGAGAGCGCAACACCTGAAAGCAGCGGGGGCAGCAGCGGGGGGTCAGCCCGCA TCCTGGCCTTTGCCATCGGCATCAGCAGCATCGGATGGGCCTTCAGCGAGAATGACGAGCTGAAGGACTGC
[0699] GGTGTGAGAATCTTTACAAAGGTCGAGAACCCCAAGACCGGCGAAAGCCTGGCGCTGCCAAGACGGCTGGC
[0700] CCGGTCTGCTCGGAAGAGATACGCCCGGAGAAAAGCCAGACTGAATCACCTGAAACACCTGATCGCCAACG
[0701] AGTTCAAACTGAACTACGAGGACTACCAGAGCTTCGATGAGAGCTTGGCTAAGGCCTACAAAGGCAGCCTG
[0702] ATCAGCCCCTACGAGCTGAGATTCAGAGCTCTGAATGAGCTCCTGAGCAAGCAGGACTTCGCTAGAGTGAT
[0703] CCTGCACATCGCCAAGCGAAGAGGCTACGACGACATCAAGAACTCTGATGACAAGGAGAAGGGCGCCATTC
[0704] TGAAAGCCATCAAGCAGAACGAAGAGAAACTGGCCAATTACCAGAGCGTGGGCGAGTACCTGTACAAGGAG
[0705] TACTTCCAGAAGTTCAAAGAAAATAGCAAAGAGTTCACCAACGTGCGGAACAAGAAGGAGTCTTATGAAAG
[0706] ATGTATCGCCCAGAGCTTCCTGAAAGACGAGCTGAAACTGATCTTCAAGAAGCAAAGAGAGTTTGGCTTCA
[0707] GCTTCAGCAAAAAATTTGAAGAAGAGGTGCTGTCTGTCGCCTTCTACAAGAGGGCTCTGAAGGACTTCAGC
[0708] CACCTGGTGGGCAATTGCAGCTTTTTCACAGACGAAAAGCGGGCCCCTAAGAACAGCCCTCTGGCCTTCAT
[0709] GTTCGTGGCTCTGACCAGAATCATCAACCTGCTGAACAACCTGAAAAATACCGAGGGCATCCTCTATACCA
[0710] AGGACGATCTGAACGCCCTGCTGAATGAGGTGCTCAAGAATGGCACCCTGACCTACAAGCAAACAAAAAAA
[0711] CTGCTGGGCCTGTCTGACGACTACGAGTTTAAAGGCGAGAAGGGCACCTATTTCATCGAATTTAAGAAGTA
[0712] CAAGGAATTCATCAAGGCACTGGGCGAACACAACCTGTCTCAGGACGACCTGAACGAGATCGCCAAGGACA
[0713] TCACCCTGATCAAGGATGAGATCAAGCTGAAAAAAGCTCTGGCCAAGTACGACCTGAATCAAAACCAGATC
[0714] GACAGCTTAAGCAAGCTGGAATTTAAGGATCACCTGAACATCTCCTTTAAGGCCCTGAAGCTGGTGACCCC
[0715] ACTGATGCTGGAAGGAAAGAAGTACGACGAAGCATGCAACGAGCTTAACCTGAAGGTTGCTATCAACGAGG
[0716] ATAAGAAGGATTTCCTGCCTGCCTTTAACGAGACATACTACAAGGATGAGGTGACCAACCCCGTGGTGCTG
[0717] AGAGCTATCAAAGAGTACAGAAAGGTGCTGAACGCCCTGCTGAAGAAGTACGGCAAGGTCCACAAGATCAA
[0718] TATTGAGCTGGCCCGGGAGGTTGGAAAGAACCACTCTCAGAGAGCAAAGATCGAGAAGGAGCAAAACGAGA
[0719] ACTACAAAGCGAAGAAGGACGCCGAACTGGAGTGCGAGAAGCTTGGCCTGAAGATCAACTCTAAGAATATC
[0720] CTGAAACTCAGACTTTTCAAAGAACAGAAGGAATTCTGTGCCTACAGCGGCGAGAAGATCAAAATTTCTGA
[0721] CCTGCAGGATGAAAAGATGCTGGAGATCGACCACATCTACCCTTACTCCAGAAGCTTCGACGACAGCTATA
[0722] TGAACAAAGTGCTGGTGTTCACAAAGCAGAACCAGGAGAAGCTGAATCAGACCCCTTTCGAGGCCTTCGGC
[0723] AATGACTCCGCCAAGTGGCAGAAAATCGAGGTGCTGGCCAAAAACCTGCCAACCAAGAAACAGAAGAGAAT
[0724] CCTCGACAAGAACTACAAGGACAAGGAACAGAAGAACTTCAAGGATCGGAACCTGAACGACACCCGGTACA
[0725] TCGCCAGGCTGGTGTTAAATTACACCAAGGACTACCTGGATTTCCTGCCCCTGAGCGACGACGAGAACACC
[0726] AAGCTGAACGACACACAGAAGGGCAGCAAGGTGCACGTGGAAGCCAAGAGCGGCATGCTGACCAGCGCACT
[0727] GCGCCACACTTGGGGCTTCAGCGCTAAGGACCGGAACAACCACCTGCATCACGCCATCGATGCCGTGATCA
[0728] TAGCCTACGCCAACAACTCAATCGTGAAAGCTTTCAGTGACTTTAAGAAAGAACAGGAGAGCAACTCTGCC
[0729] GAACTGTACGCCAAGAAAATTAGCGAGCTGGACTACAAGAACAAACGGAAGTTCTTCGAACCTTTCTCAGG
[0730] ATTTAGACAGAAGGTGCTGGATAAGATCGATGAAATCTTCGTGTCCAAGCCCGAGAGAAAGAAGCCTAGCG
[0731] GAGCCCTGCACAAGGAAACCTTCAGAAAGGAAGAAGAGTTCTACCAGTCTTATGGGGGCAAAGAAGGCGTG
[0732] CTGAAGGCCCTGGAACTCGGCAAGATCAGAAAGGTGAAGGGAAAGATTGTGAAGAACGGCGACATGTTCAG
[0733] AGTGGACATCTTCAAGCACAAGAAGACCAACAAGTTCTATGCCGTGCCTATCTACACAATGGATTTCGCCC
[0734] TGAAAGTGCTGCCTAACAAGGCTGTCGCCAGATCTAAGAAGGGCGAGATCAAGGACTGGATCCTGATGGAC
[0735] GAAAACTACGAGTTCTGCTTCAGCCTGTACAAGGACAGCCTGATCCTGATCCAAACAAAGAAGATGCAGGA
[0736] GCCTGAATTCGTGTACTACAACGCCTTCACAAGCAGCACCGTGAGCCTGATCGTGTCTAAACATGATAACA
[0737] AGTTCGAAACCCTGTCCAAGAACCAAAAGATCCTGTTCAAGAACGCCAACGAGAAGGAAGTGATCGCCAAG GGCATCGGCATTCAGAATCTGAAGGTGTTCGAGAAATATATCGTGTCCGCTCTGGGAGAGGTTACAAAGGC
[0738] CGAGTTTCGGCAGAGAGAAGATTTTAAGAAGAGCGGCGGGAGCGGCGGGAGCGGGGGGAGCAC
[0739] TAATCTGAGCGACATCATTGAGAAGGAGACTGGGAAACAGCTGGTCATTCAGGAGTCCAT
[0740] CCTGATGCTGCCTGAGGAGGTGGAGGAAGTGATCGGCAACAAGCCAGAGTCTGACATCC
[0741] TGGTGCACACCGCCTACGACGAGTCCACAGATGAGAATGTGATGCTGCTGACCTCTGACG
[0742] CCCCCGAGTATAAGCCTTGGGCCCTGGTCATCCAGGATTCTAACGGCGAGAATAAGATCA
[0743] AGATGCTGTCTGGCGGCTCAAAAAGAACCGCCGACGGCAGCGAATTCGAGCCCAAGAAGAAGAGGAAAGT
[0744] CTAATAGATCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCTCCCCCGTGCCTTCCTTGA
[0745] CCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATGAGGAAATTGCATCGCATTGTCTGAGTAGG
[0746] TGTCATTCTATTCTGGGGGGTGGGGTGGGGCAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCA
[0747] TGCTGGGGATGCGGTGGGCTCTATGGCTCGAGCGGCCCAAAAAAAGCGGTTTTAGGGGATTGTAACCCCGC
[0748] AGAGTCCCGCAAACTCTTTATTCTAGTCCCTTTTCAGGGACTAGAACNNNNNNNNNNNNNNNNNNNN
[0749] NNGGTGTTTCGTCCTTTCCACAAGATATATAAAGCCAAGAAATCGAAATACTTTCAAGTTACGGTAAGCAT
[0750] ATGATAGTCCATTTTAAAACATAATTTTAAAACTGCAAACTACCCAAGAAATTATTACTTTCTACGTCACG
[0751] TATTTTGTACTAATATCTTTGTGTTTACAGTCAAATTAATTCTAATTATCTCTCTAACAGCCTTGTATCGT
[0752] ATATGCAAATATGAAGGAATCATGGGAAATAGGCCCTCAGGAACCCCTAGTGATGGAGTTGGCCACTCCCT
[0753] CTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCG
[0754] GCCTCAGTGAGCGAGCGAGCGCGCAG
[0755] Alternative Promoter sequences that would also work length-wise
[0756] ( instead of P3 , hSyn) :
[0757] Efl-alpha short (SEQ ID NO: 19):
[0758] GGGCAGAGCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGGGTCGGCAATTGAACCGGTGCCTAG
[0759] AGAAGGTGGCGCGGGGTAAACTGGGAAAGTGATGTCGTGTACTGGCTCCGCCTTTTTCCCGAGGGTGGGGG
[0760] AGAACCGTATATAAGTGCAGTAGTCGCCGTGAACGTTCTTTTTCGCAACGGGTTTGCCGCCAGAACACAG
[0761] HlcoreM (SEQ ID NO: 20):
[0762] ATATTTGCATGTCGCTATGTGTTCTGGGAAATCACCATAAACGTGAAATGTCTTTGGATTTGGGAATCCTC
[0763] GAGGTTCTGTATGAGACCACTCTTTCCC
[0764] HlcoreM with chimeric Intron (SEQ ID NO: 21):
[0765] ATATTTGCATGTCGCTATGTGTTCTGGGAAATCACCATAAACGTGAAATGTCTTTGGATTTGGGAATCCTC
[0766] GAGGTTCTGTATGAGACCACTCTTTCCCGGGTAAGTATCAAGGTTACAAGACAGGTTTAAGGAGACCAATA
[0767] GAAACTGGGCTTGTCGAGACAGAGAAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCC
[0768] ACTTTGCCTTTCTCTCCACAG hUPi i (SEQ ID NO: 22):
[0769] CATCGGGGGAGCAGTCCTCCAAGGACTGGCCAGTCTCCAGATGCCCGTGCACACAGGAACACTGCCTTATG
[0770] CACGGGAGTCCCAGAAGAAGGGGTGATTTCTTTCCCCACCTTAGTTACACCATCAAGACCCAGCCAGGGCA
[0771] TCCCCCCTCCTGGCCTGAGGGCCAGCTCCCCATCCTGAAAAACCTGTCTGCTCTCCCCACCCCTTTGAGGC TATAGGGCCCAAGGGGCAGGTTGGACTGGATTCCCCTCCAGCCCCTCCCACCCCCAGGACAAAATCAGCCA
[0772] CCCCAGGGGCAGGGCCTCACTTGCCTCAGGAACCCCAGCCTGCCAGCACCTATTCCACCTCCCAGCCCAGC
[0773] ProA7 ( als o known as SynPVI ) (SEQ ID NO: 23):
[0774] CATCCTGAGAGATGAGCCAGGACAAAGAACCAGTAATAGCTCCTGGAGCAGCACATCTGTTTTGCCAGGAT
[0775] TATCCCTTGGATCTCTTAAAACCGAGACCTTGTAATCTGAAGACTCAACTTGGGCTGTACCCTTAACCTTC
[0776] AGCTCTATGATGCAAGTGAGTCCACAGGACCGGAGGCTTTGAGATGAGCTTTTCAGAAGGGAGGAGTTGGC
[0777] CGCTTGCTCCCAGAGCTCCAGCACCTGCATTCTTCTGGCTATGTCAGAAGCCAGATCATTTCCCTCGTTAA
[0778] AAACAAAAACAAAAAAACAAACAAACAAAATGTTAGTCTTTGCCCTTTATCTGCCTGGCAAAGCTTTTAAT
[0779] TGGCTTGATCTGTCATTCCGCTAGACATAAAGGGGACAATCCCCGGATTAGGAAGGAGCTCTCCAGCTCGG
[0780] GTAAGGAGTCTCAAGGCAAGGTAGGCAAGCACCACCGGTCCGCACTCTCGCCCAGCTTTTACGGGAAGAAG AGA
[0781] Target 1 (Spacer sequence, PAM sequence) (SEQ ID NO: 24):
[0782] 5’-GAGTAGAGGCGGCCACGACCTGGTGNNNNN-3’
[0783] Target 2 (Spacer sequence, PAM sequence) (SEQ ID NO: 25):
[0784] 5’-GTTAGGCAGATTCCTTATCTGGTGANNNNN-3’ sgRNA 1 (Spacer sequence, scaffold sequence) (SEQ ID NO: 26):
[0785] GAGUAGAGGCGGCCACGACCUGGUUUUAGUCCCUGAAAAGGGACUAAAAUAAAGAGUUU
[0786] GCGGGACUCUGCGGGGUUACAAUCCCCUAAAACCGCUUUUAAAA sgRNA 2 (Spacer sequence, scaffold sequence) (SEQ ID NO: 27):
[0787] GUUAGGCAGAUUCCUUAUCUGGGUUUUAGUCCCUGAAAAGGGACUAAAAUAAAGAGUUUG
[0788] CGGGACUCUGCGGGGUUACAAUCCCCUAAAACCGCUUUUAAAA
[0789] Cited prior art documents:
[0790] Nakagawa, R. et al. Engineered Campylobacter jejuni Cas9 variant with enhanced activity and broader targeting range. Commun. Biol. 5, 1-8 (2022).
[0791] All scientific publications and patent documents cited in the present specification are incorporated by reference herein.
Claims
Claims1 . A polypeptide comprising a variant of C / CAS9 (SEQ ID NO 001 ), wherein the polypeptide further comprises a nucleotide deaminase activity, characterized in that relative to the sequence of SEQ ID NO 001 , the following mutations are present:- E789K and S951G;- a substitution of L58 by Y or Q, and- N821 K.
2. A polypeptide according to claim 1 , further comprising the mutation, relative to the sequence of SEQ ID NO 001 , of D900K or D900R.
3. The polypeptide according to claim 1 or 2, further comprising a substitution, relative to the sequence of SEQ ID NO 001 , selected from the group of G323V, D376Y, F434C, K622N, D663E and T671A.
4. The polypeptide according to any one of the preceding claims, further comprising a substitution, relative to the sequence of SEQ ID NO 001 , selected from E871 K, D891 N and E946K.
5. The polypeptide according to any one of the preceding claims, comprising the sequence of SEQ ID NO 002.
6. The polypeptide according to any one of the preceding claims, comprising a sequence having at least 95% identity to SEQ ID NO 002 and bearing the mutations L58Y, E789K, N821 K, D900K, S951G, and having a less stringent PAM requirement than SEQ ID NO 001.
7. The polypeptide according to any one of the preceding claims, wherein the deaminase is an adenosine deaminase activity.
8. The polypeptide according to any one of the preceding claims, wherein the deaminase is a cytidine deaminase activity.
9. A polypeptide comprising a CjCAS9 variant according to any one of the preceding claims, wherein the polypeptide further comprises a reverse transcriptase activity.
10. The polypeptide according to claim 9, wherein the polypeptide does not comprise an RNAseH activity.11 . A polynucleotide sequence encoding a polypeptide comprising a C / CAS9 variant, as specified in any one of the preceding claims.
12. An expression vector comprising a polynucleotide according to claim 11 , under control of a promoter.
13. The expression vector according to claim 12, wherein the expression vector is an AAV virus particle.