Synthetic Cas proteins

JP2024521806A5Inactive Publication Date: 2025-06-02アソシエーション セントロ デ インベスティゲイション クーペレイティヴァ アン ナノシエンシアス シーアイシー ナノグネ
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023573060
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-30
Filing Date
2022-05-25
Publication Date
2025-06-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems face limitations such as undesirable mutations, genetic mosaicism, low efficiency, and immune responses, particularly with SpyCas9, which has stringent PAM recognition requirements, limiting their application in genome editing and therapeutic uses.

Method used

Reconstruct ancestral Cas proteins using phylogenetic tracing from evolutionary trees, creating Cas9 variants like LFCA, LBCA, and LSCA with relaxed PAM requirements, higher nicking rates, and flexibility in gRNA usage, enabling efficient gene editing and reduced immune response.

Benefits of technology

The ancestral Cas proteins exhibit enhanced stability, efficiency, and versatility, allowing for high expression and efficient gene editing in human cells with relaxed PAM recognition, reduced double-strand breaks, and ability to cleave single-stranded nucleic acids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to the use of phylogenetic ancestral sequence reconstruction to generate new Cas enzymes with improved capabilities.This strategy has yielded ancestral variants of currently existing species of Cas9 proteins that can exhibit nickase activity separate from endonuclease activity and relaxed, if not abolished, PAM requirements.The ability to use the tracrRNA components of Cas9 gRNA from a wide variety of existing bacterial species has also been observed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method for obtaining Cas proteins suitable for use as single effector CRISPR system-associated nucleases that are not isolatable from recognized microbial sources, i.e. class II Cas proteins. To that end, the present invention provides reconstructed ancestral sequences obtained by evolutionary tracing from phylogenetic trees compiled with existing species Cas protein sequences. Such reconstructed proteins are therefore synthetic proteins in the sense that they are not isolatable from modern sources, but can be utilized in the same manner as naturally occurring Cas proteins in the now widely used class II CRISPR systems for genome editing. The inventors have coined the term "ancestral Cas" or "AnCas" for such reconstructed sequences. This route to novel Cas proteins has been found to advantageously add to the diversity of Cas proteins available for genome editing, with respect to useful properties including a relaxation of the protospacer adjacent motif (PAM) requirement for type II proteins, compared to the most commonly used type II Cas protein Streptococcus pyogenes (SpyCas9). [Background technology]

[0002] Natural CRISPR-Cas systems provide prokaryotes with immunity to respond to invading nucleic acids from infectious genetic agents. Guided by CRISPR-encoded RNA molecules (gRNAs), Cas proteins recognize specific regions of the foreign genome and cleave it for inactivation. Such CRISPR systems and other class II CRISPR-Cas systems have revolutionized the field of genome engineering, as the first CRISPR-Cas9 system was repurposed as a genome editing tool. Nevertheless, CRISPR is not ready for implementation as a therapeutic tool due to limitations such as the occurrence of undesired mutations at similar loci, production of multiple alleles resulting in genetic mosaicism, low efficiency and possible induction in the host of immune responses. Studies have found that blood samples from human donors show a higher percentage of antibodies to SpyCas9 and Cas9 from Staphylococcus aureus (Charlesworth et al., 2019. Nature Med. 25(2):249-254).

[0003] The number and diversity of known CRISPR-Cas systems have increased dramatically since the first disclosure of in vitro DNA editing studies using the CRISPR-Cas9 system in 2012. A distinctive feature of class II systems is that the nuclease effector of the complex consists of a single multi-domain protein, as exemplified by Cas9, which utilizes a type II system. Target recognition is achieved by a structural non-coding RNA that guides the Cas protein to its target nucleic acid sequence site via base pairing for endonuclease cleavage action. In addition to guide RNA (gRNA) recognition, a sequence motif named protospacer adjacent motif (PAM) is required for Cas-guide RNA target binding and initiation of cleavage. This is important for distinguishing self from non-self in natural antiviral defense systems. However, it is less convincing for some desired applications arising from the repurposing of class II CRISPR-Cas systems as genetic tools.

[0004] Bacterial Cas9 protein was the first Cas protein studied, and SpyCas9 remains the most extensively studied Cas9 and is the most used for genome editing. Such proteins are characterized as type II Cas9 proteins by containing two nuclease domains, HNH-like and RuvC-like nuclease domains, and associated catalytic residues required for double-stranded DNA endonucleolytic cleavage occurring at blunt ends. Since 2012, Cas endonucleases have been isolated from many different bacteria and archaea. The recent classification of class II Cas nucleases includes three types and 17 subtypes, which are reviewed in Makarova et al. (2020) (Nature Rev Microbiol. 18(2):67-83). Thus, class II CRISPR-Cas systems currently include type II, V, and VI systems, with type VI system being the first and so far the only variant of CRISPR-Cas system that exclusively cleaves RNA.

[0005] Type V systems are fundamentally different from type II systems due to the domain structure of their effector Cas proteins. Type II effectors (Cas9 nucleases) contain two nuclease domains, each of which contributes to cleavage of a single strand of the target DNA, with the HNH nuclease inserted inside the RuvC-like nuclease domain sequence, whereas type V effectors (Cas12 nucleases) in contrast contain only a RuvC-like domain that cleaves both strands. Type VI effectors (Cas13 nucleases) are unrelated to type II and type V effectors, as they contain two HEPN domains and apparently target transcripts of the invasive DNA genome in their natural environment. Cas13 proteins also exhibit collateral nonspecific ribonuclease activity triggered by target recognition.

[0006] Among the V-type variants, those with smaller RuvC-like domains are currently classified as V~U subtype effectors. They show high sequence similarity to the TnpB protein (a predicted RuvC-like nuclease) encoded by IS605-like transposons and are thought to be intermediates on the evolutionary pathway from TnpB to fully fledged V-type effectors. CRISPR-Cas systems evolved from different groups of TnpB on multiple independent occasions, as known by phylogenetic analysis of the TnpB family. Analysis of the interference activity of four V~U subtype effectors yielded one such variant that has only recently been upgraded to a distinct V~F subtype. A V~F subtype effector, namely Cas12f (originally designated Cas14), was found to cleave single-stranded DNA (ssDNA). However, phylogenetic analysis of type V Cas enzymes has only been used as a means to classify naturally occurring such nucleases that are isolated by a single RuvC-like nuclease domain.

[0007] Despite the diversity of known naturally occurring class II CRISPR-Cas systems recently added by the extensive search of natural sources for CRISPR-Cas9 orthologues by Gasiunas et al. (2020) (Nature Com. 11(1):5512), still more diversity is desired for the desire for easy and diverse genetic modifications, especially in the therapeutic field. SpyCas9 has been extensively modified for genome editing and as a fusion enzyme for transcriptional regulation, epigenome editing, base editing and prime editing. Despite its versatility, SpyCas9 is still limited for certain such applications by its "NGG" PAM recognition requirement. Attempts have been made to relax this requirement by searching for new orthologues as well as utilizing directed evolution techniques, such as random mutagenesis of the PAM interaction domain (PID) by selection, and structure-guided mutagenesis as reviewed by Collias and Beisel (2021) (Nature Com. 12(1):555). By applying structure-guided mutagenesis, Walton et al. (2020) (Science. 368(6488):290-296) achieved a SpRY nuclease variant that recognizes the consensus NR PAM sequence (where R is A or G) and, with lower efficacy, the RY PAM sequence (where Y is C or T). This is the most relaxed PAM requirement reported for a Cas9 variant to date. Summary of the Invention

[0008] As shown above, in this case, we have employed a novel approach to expand the available toolset of RNA-programmable CRISPR-associated nucleases. This approach, called phylogenetic ancestral sequence reconstruction, has been used to generate variants of bacterial Cas9 predicted to have existed in organisms that lived billions of years ago. The ancestral enzymes have greater stability and efficiency, exhibit chemical promiscuity, and are more versatile than their modern descendants. Furthermore, the advantage of turning to ancestral enzymes for gene therapy is that the host's pre-existing immunity against these proteins can potentially be discarded. We have designed and tested Cas9 forms from, for example, ancestral Firmicutes, Bacillario, and Streptococcus. They show high levels of expression, non-specific tracrRNA binding, and highly efficient gene editing in cells of the human HEK293T cell line.

[0009] Although the present invention is based on the use of phylogenetic information for a diverse population of Cas9 enzymes within the bacterial classes Clostridium and Bacillus, including many species of the genus Streptococcus, including e.g. Streptococcus pyogenes, plus several Cas9 sequences from the phylum Actinobacteria, it will be appreciated that the same approach may be used to obtain ancestral versions of other taxonomic types of Cas single nuclease effectors, e.g. ancestral type V or type VI Cas enzymes. The ancestral versions may be of the same type or of different subtypes, including assignment of new subtypes from any subtype of the current taxonomy as described in Makarova et al. ibid.

[0010] Thus, in one aspect, the invention provides a method for phylogenetic ancestry reconstruction to obtain functional single effector Cas protein nucleases (commonly referred to as class II Cas proteins), such as functional Cas9 variants, comprising: (a) providing a phylogenetic tree from sequence analysis of a population of Cas sequences comprising a population of naturally occurring single effector Cas nuclease sequences of the same taxonomic type, e.g., type II Cas9 sequences, and obtained from a number of existing species, preferably two or more genera, and even more preferably two or more classes (optionally spanning two or more phyla); (b) selecting ancestral variant sequences by tracing an evolutionary path from the phylogenetic tree, determining a high probability amino acid for each amino acid in the selected ancestral variant; (c) producing said variant capable of exhibiting Cas protein endonuclease and / or nickase activity; The present invention provides a method comprising:

[0011] It will be appreciated that the starting population of Cas sequences for providing the phylogenetic tree in step (a) may comprise one or more predefined ancestral variant sequences obtained by prior application of such methods.

[0012] Computer-implemented methods for editing phylogenetic trees by protein sequence alignment of protein orthologs are well known. Further described herein is the use of computer-implemented methods that allow editing evolutionary pathways, thereby predicting and reconstructing naturally occurring ancestral variants of Cas proteins from millions of years ago. The present inventors report for the first time the "resurrection" of the 2-3 billion year old Cas9 enzyme, which shows high production levels and high efficiency of DNA targeting and editing.

[0013] In brief, the computer implementation of step (b) comprises: (i) compiling sequences of ancestral variants that are just ancestral variants for sequences of multiple species, each of which forms a portion of a sequence of the same genus, e.g., the genus Streptococcus, and preferably further (ii) using the sequences achieved in (i) to compile one or more ancestral variant sequences that have been assigned as an ancestral genera, e.g., ancestral sequences that have been compiled for only all or at least a large proportion of the available Streptococcus Cas sequences, and / or one or more ancestral variant sequences that have been assigned as an ancestral class that can be traced back to the sequences of the starting species of multiple genera, e.g., ancestral sequences that have been compiled for only all or at least a large proportion of the available Bacillus sequences across multiple genera, preferably further (iii) Compiling at least one interclass ancestral sequence that can be traced back to the starting species of two or more classes. While such an ancestral variant may be the preferred choice for production, various ancestral variants so edited may be found to have advantageous properties.

[0014] One such evolutionary pathway map leading to the compilation of such interclass ancestral variant sequences (or common ancestral phylum sequences) starting from a population of Cas9 sequences described above, including Cas9 sequences of existing bacterial species of both Bacillus and Clostridia, is shown in FIG. 1. It will be noted that the starting sequences of bacterial species of Bacillus include many (28 in number) known Cas9 sequences of the genus Streptococcus, including SpyCas9. The use of such a diverse population of starting sequences including Cas9 sequences from a diverse range of bacteria belonging to the class Bacillus, including a significant number (e.g., 25 or more) of existing Cas9 sequences of the genus Streptococcus, will be recognized as highly desirable for such evolutionary map construction.

[0015] The production of step (c) is usually by providing the nucleic acid sequence for expression in a suitable host cell, for example E. coli. The coding sequence may be codon optimized.

[0016] A typical example shows how the desired cleavage activity can be tested even in the absence of knowledge of any PAM requirements. Preferably, if such activity is observed initially by in vitro testing, it will be maintained in further testing in human cells. For example, if a selected ancestral variant is achieved starting from a population of Cas9 sequences from multiple existing species, the activity of the selected ancestral variant may be tested under in vitro conditions and in a human cell line known to be suitable for the endonuclease activity of Cas9 sequences from existing species, e.g., SpyCas9. For such testing in human cells, a human codon-optimized sequence for the Cas enzyme will preferably be used in an expression vector suitable for Cas protein expression in the selected cell.

[0017] If the initially selected variant is a Cas endonuclease, it may then be converted to a nickase or deadCas (dcas) by known methods for amino acid mutagenesis of the relevant nuclease catalytic site, and / or fused to a non-nuclease effector. For example, methods are well known for inactivating one or both nuclease sites of the Cas9 endonuclease and binding the Cas9 enzyme to another effector, such as an enzyme for base editing or prime editing, or a transcriptional or epigenetic regulator.

[0018] The novel Cas enzymes obtained by the ancestral reconstruction methods described above, and the nucleic acid sequences encoding same, e.g., provided in an expression vector for expression in a host cell, are also encompassed by the present invention. Within the scope of the present invention, the Cas nucleases or Cas nuclease variants described herein are interchangeably referred to herein as "ancestral Cas" or "AnCas".

[0019] Of particular interest as taught herein is the achievement of ancestral sequence reconstruction of AnCas variant enzymes that exhibit time-separable nickase and endonuclease activities reflected by a higher ratio of nicked template:linearized template compared to SpyCas9 under conditions for linearization of dsDNA plasmid targets with SpyCas as a reference nuclease. That is, the AnCas variant enzymes described herein can produce a greater percentage of nicked plasmid DNA templates compared to SpyCas9 under the same conditions, or can produce a lower percentage of double-strand breaks in plasmid DNA templates compared to SpyCas9 under the same conditions. As is well known in practice, SpyCas9 is not recognized as a nickase enzyme under commonly used conditions, except when one nuclease site is removed. In contrast, herein, there is provided an AnCas enzyme obtained as a Cas9 ancestral variant with a ratio of linearized DNA plasmid template:nicked DNA plasmid template of at least 2.3:1 to at least 1:4, under conditions where SpyCas9 provides a ratio of linearized DNA plasmid target:nicked DNA plasmid template of at least 4:1. That is, after 30 minutes, the AnCas enzyme obtained according to the present invention, such as LFCA, LBCA and LSCA, can nick at least 30% of the DNA template to at least 70%, for example about 80%, of the DNA template, whereas under the same conditions, SpyCas9 nicks about 10% of the DNA template in the same amount of time (see FIG. 11). Thus, in other words, AnCas nuclease has a higher nicking rate and a lower linearization rate for dsDNA plasmid target under conditions where SpyCas9 causes substantially exclusively linearization or almost exclusively linearization, whereas AnCas nuclease and the variant of interest provide observable nicked target under the same conditions.

[0020] This may be combined with relaxed PAM requirements in the AnCas enzyme compared to SpyCas9, i.e., a three-nucleotide or up to seven-nucleotide specific PAM requirement that is in fact unobservable under conditions in which SpyCas9 maintains its recognized three-nucleotide PAM requirement of NGG, e.g. TGG, and exhibits significant DNA cleavage activity.

[0021] Moreover, such relaxed PAM requirements have been observed to be combined with very flexible gRNA usage. Thus, as reported herein, AnCas has been found to be able to utilize sgRNAs with Cas9 tracrRNA components corresponding to any of a wide variety of existing bacterial species that have Cas9 orthologs. Thus, although the targeting sequences may differ, the sgRNAs are otherwise similar to sgRNAs that can be used with multiple known Cas9 orthologs. Such non-specific tracrRNA usage is not a property that has been previously reported for any known Cas9 ortholog.

[0022] Moreover, such advantageous properties were achieved in combination with the ability to cleave single-stranded DNA and RNA, again a property that distinguishes it from SpyCas9.

[0023] As an example of such an AnCas, provided herein is a highly favored ancestor of existing Firmicutes Cas9 enzymes, including SpyCas9, which we have named the Last Firmicutes Common Ancestor (LFCA) (see Figure 1, node 63 in Figure 9 and sequence number 1). SEQ ID NO:1-LFCA

[0024] As a further example, also provided herein are highly advantageous ancestral variants of existing Bacillus Cas9 enzymes, including SpyCas9, which we have named the Last Bacillus Common Ancestor (LBCA) (see FIG. 1, see node 70 in FIG. 9 and SEQ ID NO:2). This AnCas represents an ancestral variant that can be traced in evolution to Cas sequences in a wide range of modern bacterial species of the Bacillus class, including Streptococcus species. SEQ ID NO:2-LBCA

[0025] As an example, a highly advantageous ancestral variant of SpyCas9 is provided herein, which the inventors have named the Last Streptococcus Common Ancestor (LSCA) (see FIG. 1, node 91 in FIG. 9 and SEQ ID NO: 3). This AnCas represents the AnCas considered to be the common ancestral genus. It is the generated common ancestor of all the starting Streptococcus sequences listed in Table 1, including SpyCas9. SEQ ID NO:3-LSCA

[0026] Also provided herein are ancestral variants of SpyCas9 that can be traced evolutionarily back to two or more Streptococcus species, including Streptococcus pyogenes, which the inventors have named the Last Pyopathogenic Common Ancestor (LPCA) and the Last Pyopathogenic / Hemolytic Common Ancestor (LPDCA) (see nodes 92 and 95, respectively, of Figure 9 and sequence numbers 4 and 5). SEQ ID NO:4-LPCA SEQ ID NO:5-LPDCA

[0027] We found that LFCA Cas, LBCA Cas, and LSCA Cas all exhibit higher nicking rates and lower endonuclease (double-strand break) rates than SpyCas9 under conditions favorable for SpyCas9 dsDNA cleavage, and that the ratio of these activities is a function of the age of the ancestor, with LFCA Cas having the highest ratio observed to date.

[0028] LFCA Cas, LBCA Cas and LSCA Cas all also illustrate the utility of the ancestral reconstruction strategy of the present invention to achieve Cas9 variants with the relaxation of PAM requirements described above. Thus, such AnCas nucleases may be considered PAMless.

[0029] It was further found that LFCA Cas, LBCA Cas and LSCA Cas can all cleave single-stranded DNA, as shown by the studies reported herein. LFCA and LBCA Cas can also cleave single-stranded RNA, as shown by the studies reported herein.

[0030] Finally, it was further found that all of the LFCA Cas, LBCA Cas, and LSCA Cas elicited only weak responses to anti-Cas9 antibodies, as shown by the studies reported herein, a feature that may be of particular interest for in vivo applications.

[0031] The present invention is further described below with reference to the following figures and the appended claims. [Brief description of the drawings]

[0032] [Figure 1]Ancestral Cas9 reconstruction and characterization. A phylogenetic tree of Cas9 enzymes from the classes Clostridia and Bacillario of the Firmicutes plus several actinomycetes is shown. The evolutionary path from the last common ancestor of the Firmicutes (LFCA; SEQ ID NO:1) through the Bacillario ancestor (LBCA; SEQ ID NO:2), the genus Streptococcus ancestor (LSCA; SEQ ID NO:3) and several genus Streptococcus species ancestors (LPCA; SEQ ID NO:4 and LPCDA; SEQ ID NO:5) to modern Streptococcus pyogenes is shown by the white dashed arrows. Sequence relationships are also shown in the simplified cladogram in Figure 9. [Diagram 2] Figure 2a-c show a test of AnCas endonuclease activity exemplified by a test of LFCA. Figure 2a shows a DNA library containing seven random nucleotides immediately following the target DNA, where these seven Ns represent all possible PAM sequences. Figure 2b shows a Cas9 activity assay comparing LFCA to SpyCas9. The LFCA Cas cleaves the PCR target amplified from the DNA library, generating two fragments with the expected sizes. Figure 2c shows a Cas9 activity assay using the Streptococcus pyogenes PAM sequence. LFCA is able to recognize the NGG PAM sequence as recognized by SpyCas9. [Diagram 3]Figures 3a-d show a demonstration of nicking and endonuclease activity of LFCA. Figure 3a shows a DNA plasmid containing a TGG PAM sequence after the DNA target. Cas9 can cleave one or both strands of DNA. Figure 3b shows a 1% agarose gel of the DNA plasmid after 1 hour of contact with 30 nM Cas9 resulting in endonuclease activity. LFCA shows nicking activity after 10 minutes of incubation under the same conditions and double strand breaks after 1 hour. In contrast, SpyCas9 shows mainly double strand break activity. Figure 3c shows the total cleavage rate (nicking + endonuclease activity) expressed as a function of time from both LFCA and SpyCas9. Figure 3d shows the cleavage rate expressed as a percentage with a distinction made between nicking and endonuclease activity. LFCA cleaved one strand of DNA and started to cleave the other strand after 1 hour of incubation. SpyCas9 has primarily endonuclease activity (i.e., it cleaves both strands). [Figure 4] Figure 4a-b show the PAM determination for LFCA. Figure 4a shows the PAM wheel graph from LFCA PAM evaluation by providing a 3-nucleotide PAM. LFCA does not show recognition specificity with any particular PAM with 3 nucleotides. Similar results were obtained when using a 7-nucleotide PAM. Figure 4b shows the results of an in vitro cleavage assay of plasmid DNA with different PAM sequences comparing both LFCA and SpyCas9. LFCA (10 nM) nicked all plasmids containing different PAMs within only 10 minutes of reaction. In contrast, SpyCas9 cleaved nearly 100% of the plasmids containing its canonical PAM sequence, i.e., TGG, and showed low or no activity with other PAM sequences. [Diagram 5]Figure 5a-d show the thermal and pH stability test of LFCA. Figure 5a shows the total Cas enzyme activity assay over 1 hour using plasmid DNA containing TGG PAM with 30 nM Cas enzyme at pH 7.9 and different temperatures ranging from 4 °C to 60 °C. LFCA shows higher activity than SpyCas9 at 4 °C and 53-60 °C. Figure 5b shows both nicking and endonuclease activity of Cas enzyme at different temperatures. Figure 5c shows the total Cas enzyme activity assay over 1 hour using plasmid DNA containing TGG PAM with 30 nM Cas enzyme at 37 °C and different pH ranging from 4-9.5. LFCA showed higher activity at acidic pH (4-5.5) in comparison to SpyCas9. Figure 5d shows both nicking and endonuclease activity of Cas enzyme at different pH. [Figure 6] Figure 6a-f show a comparison of LFCA and SpyCas9 genome editing in HEK293T cells. Figure 6a shows the humanized LFCA and SpyCas9 coding sequences cloned into the expression plasmid pCDNA3.1 to transfect HEK293T cells with gRNA for targeting the AAVS1 locus. Figure 6b shows immunofluorescence images from cells transfected with either Cas enzyme. Cells expressing the hCas coding sequence are stained orange-yellow and nuclei are stained with DAPI. Figure 6c shows the results of a T7 assay for Cas enzyme activity. PCR products from cells transfected with gRNA and hCas coding sequence were incubated with T7E1 to measure indel formation. Figure 6d shows hCas9, gRNA, and donor DNA carrying the eGFP gene transfected into HEK293T cells for knock-in of eGFP into the AAVS1 locus. Figure 6e shows confocal microscopy images of cells expressing eGFP. FIG. 6f shows the relative fluorescence measured in images from cells expressing eGFP after transfection with the hCas enzyme. [Figure 7]Comparison of LFCA and SpyCas9 knock-in in HEK293T cells targeting the TTC PAM. Electrophoresis gel with gDNA extracted from cells and amplified locus is also shown. Bands of the expected size are seen on the gel in all samples except for the TTC PAM targeted with Spy Cas9. [Figure 8] Agarose gel test results showing the ability of LFCA to use sgRNA, where the targeting sequence is attached to a tracRNA component that corresponds to the Cas9 tracrRNA component of one of a wide variety of existing bacterial species that have Cas9 orthologs. [Figure 9] We present a cladogram constructed using the sequences listed in Table 1. Each node represents an ancestral state with the sequence shown in the sequence listing. [Figure 10] Figures 10a-c show the nicking and endonuclease Cas9 activity of LFCA, LBCA, and LSCA compared to SpyCas9. Figure 10a shows the total cleavage rate (both nicking and endonuclease activity) for LFCA, LBCA, LSCA, and SpyCas9. All Cas9 enzymes reached the total cleavage rate within about 10 minutes of incubation. Figure 10b shows the plasmid linearization rate of LFCA, LBCA, LSCA, and SpyCas9. Figure 10c shows the nicking rate of LFCA, LBCA, LSCA, and SpyCas9. [Figure 11] The percentage of DNA templates with double-strand breaks (DSBs), i.e., the percentage of linearized templates after 30 minutes of incubation with each Cas nuclease, and the percentage of nicked DNA templates after 30 minutes of incubation with each Cas nuclease, are shown. LFCA, LBCA, and LSCA produce less linearized DNA templates (i.e., lower percentage of DSBs) than SpyCas9, but more nicked templates than SpyCas9. [Figure 12]Figure 11 shows an alternative way of showing the difference in cleavage (endonuclease + nickase) activity of ancestral Cas enzymes compared to SpyCas9. Linearization and nicking rates are plotted against AnCas age for LFCA, LBCA and LSCA compared to SpyCas9. Higher nickase activity measured by nicking results in lower linearity rates, so the higher nicking rate for SpyCas9 is in absolute terms the result of a fit assuming equal initial points, which is true for all enzymes at t=0 min but not at t>0 min. Hence, negative values ​​are shown for nicking rates. This is a conversion of the fraction of cleaved (nicked or linearized) plasmid template shown in Figure 11 into time units, which provides the negative λ parameter of the exponential decay shown in Figure 10. LBCA and LSCA have higher linearization rates than LFCA but still lower than SpyCas9, so the trend is for linearization rates to decrease with ancestral age. In contrast, nicking rates increase with ancestral age. [Figure 13] Figure 13a-b show PAM determination for LBCA and LSCA. Figure 13a shows the PAM wheel graph from LBCA and LSCA PAM sequencing. LBCA and LSCA do not show specificity of recognition for any 3-nucleotide PAM. Similar results were obtained using 7-nucleotides. Figure 13b shows the results of an in vitro cleavage assay of plasmid DNA with different PAM sequences, comparing LFCA, LBCA, LSCA and SpyCas9. LBCA (10 nM) nicked all plasmids with different PAMs within a 10-minute reaction. LSCA showed similar selectivity but with a higher linearization rate cleavage. [Figure 14] Testing the same ancestral Cas9 enzymes for endonuclease activity against single-stranded DNA shows that the three ancestral enzymes cleave single-stranded DNA with or without gRNA. As expected, SpyCas9 was unable to cleave the same single-stranded DNA. [Figure 15]Figures 15A-E show the activity of AnCas endonuclease on supercoiled DNA substrates. Figure 15A shows an in vitro cleavage assay for SpCas9 and all AnCas on a 4007 bp substrate at different reaction times showing the nicked and linear fractions. Figure 15B shows the quantification and exponential fit (lines) of the total cleavage rate at different reaction times. Figure 15C shows the quantification of the nicked fraction for all AnCas and SpCas9 at different times. Figure 15D shows the quantification of the DSB cleavage rate. A single exponential fit was used to obtain the k cleavage and maximum fraction cleaved (amplitude). Figure 15E shows the DSB fraction (left axis) and nicked fraction (right axis) plotted against evolution time. [Figure 16] Figures 16A-C show the PAM determination of AnCas. Figure 16A shows the PAM wheel graph (Krona plot) for all five AnCas and SpCas9 used as a control. Figure 16B shows the percentage of reads containing the NGG PAM sequence 3-4 bp downstream of the cleavage position plotted against evolutionary time. Figure 16C shows the in vitro cleavage assay (DSB and nicked products) with various PAM sequences, represented by TNN and CCC as controls. The incubation time was 10 min. [Figure 17]Figure 17A-G show sgRNA testing and nuclease activity of AnCas on single-stranded substrates. Figure 17A shows in vitro cleavage assays on supercoiled DNA substrates for AnCas and SpCas9 with sgRNAs from different species. LFCA[FCA], LBCA[BCA] and SpCas9 are shown. Figure 17B shows quantification of in vitro cleavage rates for all AnCas and SpCas9 with different sgRNAs. Figure 17C shows in vitro cleavage assays on 85 nt ssDNA fragments at different incubation times for LFCA[FCA], LBCA[BCA] and SpCas9. Figure 17D shows in vitro cleavage assays on 60 nt ssRNA at different incubation times for LFCA[FCA], LBCA[BCA] and SpCas9. Figure 17E shows the exponential fit for the quantification of the fractional cleavage rate of ssDNA at different times and the determination of kinetic parameters. In both Figure 17C and Figure 17D, the control lanes are the same for the three proteins. Figure 17F shows the exponential fit for the quantification of the fractional cleavage rate of ssRNA at different times and the determination of kinetic parameters. All kinetic parameters are summarized in Table 2. Figure 17G shows the results from the ELISA test of anti-Cas9 rabbit antibodies against SpCas9, LFCA [FCA], LBCA [BCA] and BSA used as controls. [Figure 18] FIG. 1 shows in vitro site-specific editing scale in HEK293T cells by NGS targeted sequencing using Illumina technology in three independent targets. [Figure 19]Figures 19A-D show the activity of LFCA[FCA]H838A endonuclease on supercoiled DNA substrates. Figure 19A shows an in vitro cleavage assay for LFCA[FCA]H838A on a 4007 bp substrate at different reaction times showing the nicked and linear fractions. Figure 19B shows the quantification of the total cleaved fraction at different reaction times and the exponential fit (line). Figure 19C shows the quantification of the nicked fraction at different times. Figure 19D shows the quantification of the DSB cleavage rate. A single exponential fit was used to obtain the k cleavage and maximum fraction cleaved (amplitude). [Figure 20] Figure 1 shows the posterior probability distribution for each of the predicted residues of all ancestral AnCas endonucleases. The residue with the highest posterior probability is assigned to each position. In all cases, the average posterior probability is close to 1, except for LFCA[FCA], which shows a mean value of 0.74. [Figure 21] Figures 21A-B show the activity of AnCas endonuclease at different temperatures and pH values. Figure 21A shows the quantification of the total cleavage rate at different temperatures ranging from 5 to 60°C. Figure 21B shows the quantification of the total cleavage rate at different pH values ​​ranging from 4 to 9.5. [Figure 22] PAM wheel graphs (Krona plots) for all five AnCas and SpCas9 with 7-nucleotide PAM analysis. Selectivity for NGG PAM is observed, except for LFCA[FCA]. [Figure 23] Traffic light reporter cleavage assay. Relative NHEJ frequency was estimated by the number of RFP-positive cells and normalized to SpCas9. [Figure 24]Comparative evaluation of PAM selectivity for two AnCas [LFCA and LBCA] against wild-type Streptococcus pyogenes Cas9 [SpCas9], the so-called "ancestral Cas9 protein" of WO 2021 / 084533A1 (sequence number 268 of WO '533) [Anc.Cas], and the so-called "nearly PAMless Cas9 proteins SpG and SpRY" [SpRY and SpG, respectively] of Walton et al. (2020. Science. 368(6488):290-296). The PAM selectivity of each of these nucleases is shown with N=any nucleotide and R=A or G. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0033] The starting sequence population for obtaining functional ancestral Cas variants by the strategy taught herein may preferably be a population of Cas9 sequences from existing bacterial species, as exemplified, by which a phylogenetic tree can be constructed based on sequence alignment information. Computer-implemented methods for constructing phylogenetic trees are well known in the art, requiring sequence alignment and recognition of conserved regions. However, it is not excluded that the phylogenetic tree can be constructed from the sequences of other class II Cas enzyme types, as described above.

[0034] Preferably, the starting sequence spans two or more genera. For example, as shown in the Examples section, in searching for advantageous ancestral Cas9 variants, the multiple Cas9 sequences may be selected from two or more of Streptococcus, Enterococcus, Listeria, Clostridium, Pelagirhabdus, Halolactibacillus, Floricoccus, Vagococcus, Urinacoccus, Vagococcus, Dorea, Ruminococcus, Lachnospira, Anaerostipes, Oisenella, and Bifidobacterium. The two or more species sequences may be selected from sequences from each selected genus, e.g., two, three, four or more, e.g., up to 25 or more, e.g., 28 Streptococcus species. As described above, LSCA is the evolutionarily traceable common ancestor of all 28 Streptococcus species listed in Table 1, including Streptococcus pyogenes, as shown in FIG. 1.

[0035] More preferably, the starting population of sequences spans two or more classes of the phylum of interest. Thus, as shown above, it has been found useful to combine Cas9 sequences from both the Bacillus and Clostridia classes of bacteria in the starting population when searching for advantageous ancestral Cas9 variants. For example, the starting population of sequences may desirably include at least a number of sequences from different species of Streptococcus, a number of sequences from different species of Enterococcus, a number of sequences from different species of Listeria, and a number of sequences from Clostridium species. The diversity of the starting population may be further expanded across phyla, as shown by the inclusion of several Actinomycete sequences in the starting population of Cas9 sequences used in the Examples section. Desirably, the starting population of Cas sequences may span multiple subtypes.

[0036] Starting from a phylogenetic tree compiled from known Cas sequences, an evolutionary path to a predicted ancestral form may be compiled, which may be equivalent to going back millions of years from today. The selected ancestral variant sequences obtained according to the present invention may be equivalent to an evolutionary period of at least 500 million years from the present, such as at least 700-800 million years or even more than 1 billion years. As described above, the evolutionary period may be as long as 2-3 billion years, such as approximately 2.2-2.4 billion years.

[0037] As shown by Figure 1, LFCA Cas is such a reconstructed ancestor of Cas enzymes obtained from evolutionary pathway analysis of a phylogenetic tree of Cas9 sequences from a collection of existing bacterial species spanning the classes Clostridia and Bacillario (both bacterial genera of the Firmicutes phylum) and supplemented with some actinomycetes. It can be considered as an ancestor of earlier ancestral members of the evolutionary pathway to existing Cas9 enzymes from the class Bacillario (reconstructed ancestral genus named LBCA Cas having SEQ ID NO: 2) and more specifically of a reconstructed ancestor of a wide range of Cas enzymes from the genus Streptococcus (reconstructed ancestral genus LSCA Cas having SEQ ID NO: 3). LSCA is then the ancestor of a smaller selection (8 of 28) of Cas enzymes from the genus Streptococcus (reconstructed ancestor designated LPCA Cas having SEQ ID NO:4), which in turn is the ancestor of Streptococcus pyogenes and Streptococcus dysgalactiae sequences (reconstructed ancestral LPDCA Cas having SEQ ID NO:5).

[0038] As used in this specification and the accompanying figures, the terms "Bys," "Bya," and "Gya" are used interchangeably to refer to billions of years.

[0039] As described above, LFCA Cas is represented by node 63 of the illustrated evolutionary pathway and has the amino acid sequence shown in SEQ ID NO: 1. It exhibits high production levels and high efficiency of targeting and editing DNA both in vitro and in human cells.

[0040] Thus, herein provided as an embodiment of the invention is a Cas nuclease comprising or consisting of LFCA Cas having the amino acid sequence of SEQ ID NO: 1, which represents a preferred example of a functional ancestral Cas obtained by employing the strategies taught herein for the identification of such novel Cas enzymes. LFCA Cas is considered to be evolutionarily related to SpyCas9, but has a number of advantageous differences that make it a particularly preferred AnCas nuclease.

[0041] The following interesting properties of LFCA Cas have been reported: (i) LFCA Cas is not known in nature and has only 54% sequence identity with SpyCas9. Nevertheless, it can use sgRNAs with the 3' end of SpyCas sgRNA for guide RNA / Cas protein interactions as shown in the Examples section. (ii) In contrast to SpyCas9, it exhibits time-resolvable nicking activity followed by endonuclease activity on double-stranded plasmid DNA providing a SpyCas9 PAM sequence and under conditions for endonucleolytic cleavage of the same plasmid by SpyCas9. (iii) It showed broad PAM specificity, whereas LFCA Cas showed no specificity for PAM recognition, as shown by the PAM wheel graph in Figure 4a. The cleavage data in Figure 4b show that nicked plasmid DNA provided the proposed PAM sequence at 10 nM within 10 min, regardless of the 3-nucleotide sequence. In contrast, SpyCas9 cleaved nearly 100% of the plasmids containing its canonical PAM sequence, i.e., TGG, under the same conditions, and showed low or no activity with other PAM sequences. Similar results were obtained using a 7-nucleotide variant sequence for PAM provision. Therefore, under the conditions tested, LFCA Cas can be named "PAMless". (iv) It exhibits high flexibility in gRNA requirements; as shown by Figure 8, it has Cas9 tracrRNA components that correspond to any of several existing bacterial species and can utilize sgRNAs across a wide variety of species, including Streptococcus thermophilus, Enterococcus faecium, Clostridium perfringens, and Finegoldia magna, as well as Streptococcus pyogenes. (v) LFCA Cas showed higher cleavage activity, followed by SpyCas9, which was observed as nicking activity at low temperatures (4–20 °C) and pH 7.9. (vi) At higher temperatures (53–60 °C), it exhibited higher thermostability than SpyCas9, with both nickase and endonuclease activities observed. (vii) In pH stability tests at different pHs and 37 °C, LFCA Cas maintained higher activity than SpyCas9 under acidic conditions. At alkaline pH, the activity of LFCA Cas remained the same, where SpyCas9 showed its optimal performance. (viii) LFCA Cas exhibits cleavage activity against single-stranded DNA substrates as shown in Figure 13. As is well known, this is not the activity of SpyCas9 under conditions of normal use in the field of gene modification.

[0042] In human cells (exemplified herein using HEK293T cells), it was found that LFCA Cas can promote indel formation at targeted loci when expressed in such cells using suitable gRNAs (exemplified herein using AAVS1 locus).Furthermore, the ability to promote knock-in gene modification at the same locus was confirmed as shown in Figure 6 and Figure 7.

[0043] As an embodiment of the present invention, Cas nucleases are provided that comprise or consist of LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas, which have the amino acid sequences of SEQ ID NOs: 2, 3, 4 and 5, respectively. LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas each share an interesting property with LFCA Cas that distinguishes them from the Cas9 enzyme known as Cas enzyme. Of particular interest is the higher percentage of nicked plasmid templates produced compared to SpyCas, as further shown for LBCA Cas and LSCA Cas in the Examples section (which is equivalent to a higher nicking rate relative to the plasmid linearization rate compared to SpyCas, as shown by the following exemplary embodiment). As described above, the percentage of nicked plasmid templates (and / or nicking rate) is interestingly found to increase in this group of enzymes with ancestral age, and this feature is most evident in LFCA Cas. In contrast, the linearization rate and / or the percentage of double-strand breaks are found to decrease with ancestral age.

[0044] Thus, the present invention provides (i) a Cas nuclease designated as LFCA Cas and having the amino acid sequence set forth in SEQ ID NO:1; (ii) a Cas nuclease designated as LBCA Cas and having the amino acid sequence set forth in SEQ ID NO:2; (iii) a Cas nuclease designated as LSCA Cas and having the amino acid sequence set forth in SEQ ID NO:3; (iv) a Cas nuclease designated as LPCA Cas and having the amino acid sequence set forth in SEQ ID NO:4; (v) a Cas nuclease designated as LPDCA Cas and having the amino acid sequence set forth in SEQ ID NO:5; or Distinguishing features compared to SpyCas9: (a) a higher percentage of nicked DNA plasmid templates and / or a lower percentage of linearized DNA plasmid templates, under conditions in which SpyCas9 yields substantially exclusively linearized DNA plasmid templates; (b) a higher nicking rate and a lower linearization rate for DNA plasmid targets (preferably a ratio of linearized DNA plasmid target:nicked DNA plasmid template of at least about 4:1), under conditions where the variant provides an observable nicked target while being substantially exclusively linearized or nearly exclusively linearized by SpyCas9; (c) relaxed PAM requirements comparable to either LFCA Cas, LBCA Cas and / or LSCA Cas; (d) the ability to make single-stranded DNA breaks; (e) The ability to use an sgRNA in which the targeting sequence is linked to a tracrRNA component that can be selected from the tracrRNA components of Cas9 gRNAs used by multiple existing bacterial species. A variant of such a Cas nuclease that retains one or more of The present invention relates to a Cas nuclease comprising or consisting of:

[0045] The term "variant" as used herein refers to a Cas nuclease having at least one amino acid mutation (e.g., addition, substitution or deletion) compared to any one of the sequences SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4 and SEQ ID NO:5, respectively. Typically, a Cas nuclease variant shares at least 60% sequence identity, preferably at least 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4 or SEQ ID NO:5. Of course, the amino acid sequence of a Cas variant is not 100% identical to any one of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4 or SEQ ID NO:5.

[0046] In some embodiments, the amino acid sequence of the Cas nuclease variant shares at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or more sequence identity with the amino acid sequence of SEQ ID NO:1.

[0047] In some embodiments, the amino acid sequence of the Cas nuclease variant shares at least 75%, 80%, 85%, 90%, 95% or more sequence identity with the amino acid sequence of SEQ ID NO:2.

[0048] In some embodiments, the amino acid sequence of the Cas nuclease variant shares at least 80%, 85%, 90%, 95% or more sequence identity with the amino acid sequence of SEQ ID NO:3.

[0049] In some embodiments, the amino acid sequence of the Cas nuclease variant shares at least 85%, 90%, 95% or more sequence identity with the amino acid sequence of SEQ ID NO:4.

[0050] In some embodiments, the amino acid sequence of the Cas nuclease variant shares at least 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence of SEQ ID NO:5.

[0051] In some embodiments, the Cas nucleases of the invention have one or more amino acid changes, such as substitutions or deletions, such as one or more conservative substitutions, whereby endonuclease and / or nickase activity is retained with the relaxed PAM specificity of LFCA Cas, LBCA Cas and / or LSCA Cas. In some embodiments, the Cas nucleases of the invention have nickase activity. In some embodiments, the Cas nucleases of the invention have relaxed PAM requirements. In some embodiments, the Cas nucleases of the invention have no PAM requirements, i.e., the Cas nucleases are PAMless.

[0052] In a preferred embodiment, the Cas nuclease of the invention is an LFCA having SEQ ID NO:1 or a variant thereof, or an LBCA having SEQ ID NO:2 or a variant thereof.

[0053] By way of example, LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas have the following additional interesting properties: (a) LBCA Cas is not known in nature and has only 70% identity with SpyCas9. LSCA Cas is not known in nature and has only 75% identity with SpyCas9. LPCA Cas is not known in nature and has 83.5% identity with SpyCas9. LPDCA Cas is not known in nature and has 97.5% identity with SpyCas9. Nevertheless, they can use sgRNA with the 3' end of SpyCas sgRNA for guide RNA / Cas protein interaction shown in the Examples section. (b) In contrast to SpyCas9, both LBCA Cas and LSCA Cas were found to share with LFCA Cas the relaxed PAM requirement (Fig. S3a, S3b and S3c) and its ability to cleave single-stranded DNA (Fig. S3c), as described above.

[0054] However, it will be understood that the novel AnCas enzymes exemplified herein are merely illustrative of the utility of the novel ancestral sequence reconstruction approach of the present invention to achieve such enzymes advantageous for nucleic acid modification. Of course, the teachings herein do not necessarily provide for the same novel properties, e.g., the ability to modify the ancestral sequence of the enzymes, as compared to SpyCas9, e.g., the ability to modify the ancestral sequence of the enzymes, and the ability to modify the ancestral sequence of the enzymes. (a) a higher percentage of nicked DNA plasmid templates and / or a lower percentage of linearized DNA plasmid templates (i.e., percentage of double-stranded breaks), under conditions in which SpyCas9 produces substantially exclusively linearized DNA plasmid templates; and / or (b) a higher nicking rate and a lower linearization rate for a DNA plasmid target (which may be equivalent to a ratio of linearized DNA plasmid target:nicked DNA plasmid template of at least about 4:1), under conditions in which the variant is substantially exclusively linearized or nearly exclusively linearized by SpyCas9 while providing an observable nicked target; and / or (c) relaxed PAM requirements comparable to those of LFCA Cas, LBCA Cas, and LSCA Cas; and / or (d) the ability to cleave single-stranded DNA and / or single-stranded RNA, and / or (e) the ability to use an sgRNA in which the targeting sequence is linked to a tracrRNA component selectable from the tracrRNA components of Cas9 gRNAs used by multiple existing bacterial species, including, for example, all of Streptococcus pyogenes, Streptococcus thermophilus, Enterococcus faecium, Clostridium perfringens, and Finegoldia magna. This paves the way for the achievement of other Cas9 variant enzymes, including variants of the exemplified AnCas enzymes that share one or more of the following:

[0055] As a particular embodiment of the invention, preferred ancestral functional equivalent Cas nucleases are also provided herein, which functional equivalents are represented by nodes on the evolutionary pathway of Figure 9. The amino acid sequences of these functional equivalent Cas nucleases are provided in the attached sequence listing as SEQ ID NOs: 10-236.

[0056] The present invention extends to methods for obtaining an ancestral Cas nuclease, wherein the selected ancestral enzyme is evolutionarily traceable to an existing Cas9 enzyme, preferably, for example, to SpyCas9, and has the properties (a)-(e) described above: (a) a higher percentage of nicked DNA plasmid templates and / or a lower percentage of linearized DNA plasmid templates, under conditions in which SpyCas9 yields substantially exclusively linearized DNA plasmid templates; (b) a higher nicking rate and a lower linearization rate for DNA plasmid targets (preferably a ratio of linearized DNA plasmid target:nicked DNA plasmid template of at least about 4:1), under conditions where the variant provides an observable nicked target while being substantially exclusively linearized or nearly exclusively linearized by SpyCas9; (c) relaxed PAM requirements comparable to either LFCA Cas, LBCA Cas and / or LSCA Cas; (d) the ability to make single-stranded DNA breaks; (e) The ability to use an sgRNA in which the targeting sequence is linked to a tracrRNA component that can be selected from the tracrRNA components of Cas9 gRNAs used by multiple existing bacterial species. has one or more of the following.

[0057] Particularly preferred may be the selection of AnCas with relaxed PAM specificity as described above, optionally in combination with one or both of characteristics (a) and (d) or one or both of characteristics (a) and (e), or optionally in combination with all of (a), (d) and (e). As described above, the selected AnCas may provide a ratio of linearized DNA plasmid target: nicked DNA plasmid template between at least about 2.3: 1 and at least 1: 4, e.g., under conditions where SpyCas9 provides a ratio of linearized DNA plasmid target: nicked DNA plasmid template of at least about 4: 1.

[0058] Thus, applying the method of the present invention, a ratio of linearized DNA plasmid target:nicked DNA plasmid template between at least about 2.3:1 and at least 1:4, under conditions in which SpyCas9 provides a ratio of linearized DNA plasmid target:nicked DNA plasmid template of at least about 4:1; Relaxed PAM requirements comparable to LFCA Cas, LBCA Cas and LSCA Cas, the ability to cleave single-stranded DNA and / or single-stranded RNA, and The ability to use an sgRNA in which the targeting sequence is linked to a tracrRNA component that can be selected from the tracrRNA components of Cas9 gRNAs used by multiple existing bacterial species. It is possible to obtain an ancestral Cas nuclease that has one or more of the following properties:

[0059] Such methods may further include converting such AnCas nucleases to variants that are either nickases only or deadCas with no nuclease activity, and / or providing binding to a non-nuclease effector, for example in a fusion protein. Such variants and / or fusion proteins that are either nickases only or deadCas with no nuclease activity are also considered as products per se in the present invention.

[0060] In some embodiments, the Cas nuclease is a non-nuclease modified deadCas variant of the nuclease of the present invention that has been converted to a deadCas with no nuclease activity by mutagenesis of the catalytic site. In some embodiments, the Cas nuclease is devoid of catalytic activity. In some embodiments, the Cas nuclease is a deadCas.

[0061] In some embodiments, the Cas nuclease or Cas nuclease variant is combined with a non-nuclease effector for genetic modification or regulation. In some embodiments, the non-nuclease effector is a fusion protein comprising the Cas nuclease or Cas nuclease variant and the non-nuclease effector.

[0062] In some embodiments, the ancestral Cas nuclease has the following properties: the ability to make single-stranded DNA breaks, and / or Ability to cleave single-stranded RNA The compound further comprises one or more of:

[0063] In one embodiment, the ancestral Cas nuclease has relaxed PAM requirements comparable to any of LFCA Cas, LBCA Cas, and LSCA Cas. In one embodiment, the ancestral Cas nuclease has no PAM requirements.

[0064] In some embodiments, the ancestral Cas nuclease has the following properties: Ability to break single-stranded DNA, the ability to cleave single-stranded RNA, and / or Relaxed PAM requirements comparable to LFCA Cas, LBCA Cas and LSCA Cas has one or more of the following:

[0065] In some embodiments, the ancestral Cas nuclease has the following properties: Ability to break single-stranded DNA, the ability to cleave single-stranded RNA, and / or · LFCA Cas equivalent to PAM requirement (i.e. PAMless activity) has one or more of the following:

[0066] Variants of the exemplified AnCas nucleases described above that retain one or more of the above distinguishing characteristics (a)-(e) compared to SpyCas9 are also provided. Again, retention of relaxed PAM specificity as exhibited by LFCA Cas may be particularly preferred, for example, optionally one, two or all of the characteristics specified in (a), (b), (d) and (e) above, such as production of more nicked templates and / or less linearized templates (amount of double-stranded breaks) compared to SpyCas9 described above, and / or a higher ratio of nicking rate:linearization rate compared to SpyCas9 described above, and / or the ability to cleave single-stranded DNA. Preferably, all these characteristics are retained.

[0067] Particularly preferred are LFCA Cas and variants thereof that are functional equivalents as indicated above, i.e., maintain all of the above-listed properties (i)-(viii) of LFCA Cas, except that LFCA Cas variants that retain at least the relaxed PAM specificity and / or flexible tracrRNA utilization discussed above are considered highly preferred additions to the Cas enzyme toolbox.

[0068] By convention, the terms "linear activity" and "endonuclease activity" are used interchangeably herein to refer to nuclease activity to cleave both strands of double-stranded DNA provided in the form of a plasmid. The terms "linearization activity rate", "linearization rate" and "linear activity rate" are used interchangeably herein to refer to a measure of the amount of target dsDNA cleaved through both strands as a function of time. As used herein, "nickases" refer to nucleases that cleave only one strand of a dsDNA molecule, such as a plasmid, thereby forming a nick. "Nickase activity rate" and "nick rate" are used interchangeably and refer to a measure of the amount of target dsDNA cleaved through a single strand as a function of time. Nickase and / or linear activity rates are (i) incubating 30 nM Cas nuclease with gRNA in a cleavage buffer (e.g., 100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg BSA, pH 7.9) at a 1:1 ratio and at 37° C. for at least 5 min; (ii) adding target DNA, e.g., a plasmid; (iii) incubating for, e.g., 30 minutes; (iv) stopping the cleavage reaction; and (v) visualizing the final reaction products, for example, by running them on an agarose gel; may be tested by a method including:

[0069] Using methods similar to those described using a 30 minute incubation time, preferred Cas nucleases of the invention may generate a linearized DNA plasmid target:nicked DNA plasmid template ratio of between at least 2.3:1 and at least 1:4, under conditions where SpyCas9 produces a linearized DNA plasmid target:nicked DNA plasmid template ratio of at least 4:1. Thus, after 30 minutes as described above, AnCas enzymes obtained according to the invention, such as LFCA, LBCA and LSCA, may nick at least 30% to at least 70% of the DNA template, for example about 80% of the DNA template, whereas under the same conditions SpyCas9 nicks about 10% of the DNA template in the same amount of time as shown in FIG.

[0070] The percentage of DNA templates (i.e. linearized templates) with double strand breaks (DSBs) formed by preferred Cas nucleases of the invention may be between 10% and about 70%. The percentage of DNA templates with double strand breaks (DSBs) formed by preferred Cas nucleases of the invention may be up to 70%, 60%, 50%, 40%, 30%, 20% or 10%. The percentage of DNA templates with double strand breaks (DSBs) formed by preferred Cas nucleases of the invention may be between 15% and about 65%. The percentage of double strand breaks (DSBs) formed in DNA templates by preferred Cas nucleases of the invention may be between 19% and about 62%.

[0071] The percentage of DNA templates with double-stranded breaks (DSBs) formed by LFCA Cas may be about 19%. The percentage of DNA templates with double-stranded breaks (DSBs) formed by LBCA Cas may be about 36%. The percentage of DNA templates with double-stranded breaks (DSBs) formed by LSCA Cas may be about 62%. Conversely, under the same experimental conditions as used to test any one of the preferred Cas nucleases of the present invention, the percentage of DNA templates with double-stranded breaks (DSBs) formed by SpyCas9 is at least 70%, 75% or 80%.

[0072] The percentage of nicked DNA templates formed by preferred Cas nucleases of the invention (i.e., nicked templates are generated) may be between 20% and about 100%. The percentage of nicked DNA templates formed by preferred Cas nucleases of the invention may be at least 30%, 40%, 50%, 60%, 70%, 80% or 90%. The percentage of nicked DNA templates formed by preferred Cas nucleases of the invention may be between 20% and about 90%. The percentage of nicked DNA templates formed by preferred Cas nucleases of the invention may be between 35% and about 85%.

[0073] The percentage of DNA templates with nicks formed by LFCA Cas may be about 80%. The percentage of DNA templates with nicks formed by LBCA Cas may be about 65%. The percentage of DNA templates with nicks formed by LSCA Cas may be about 35%. Conversely, the percentage of DNA templates with nicks formed by SpyCas9 under the same experimental conditions used to test any one of the preferred Cas nucleases of the present invention is at most about 20% or 10%.

[0074] As shown in Figure 12, the difference in cleavage activity of the preferred ancestral Cas enzymes of the present invention can be shown by the linearization rate and the nick rate. Therefore, the linearization rate of the preferred Cas nucleases of the present invention is about 0.001 to about 0.1 m. -1 The preferred Cas nuclease of the present invention has a nick rate of about -0.4 to about -0.1 m. -1 For example, the straightening rate and the nick rate may be as shown in Table 2.

[0075] In some embodiments, the Cas nuclease variant is modified by mutagenesis of the catalytic site to retain only the nickase activity. In some embodiments, the Cas nuclease variant is a Cas nickase. In some embodiments, the amino acid sequence of the Cas nuclease contains one or more amino acid changes by substitution or deletion, e.g., one or more conservative substitutions, thereby retaining the endonuclease and / or nickase activity with relaxed PAM specificity of the LFCA Cas.

[0076] From the above, it will be clear that the above-discussed LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas and their variants (as well as their functional equivalents, such as those represented by the nodes on the evolutionary pathway in FIG. 9, or others obtained according to the ancestral reconstruction strategy of the present invention) are considered as very useful novel additions to the toolbox of Cas proteins, especially LFCA Cas with the highest observed nick rate. They may be used directly under conditions to promote the action of only nickases (however, modification by mutagenesis of the catalytic site to retain only nickases activity, such as LFCA nickase, LBCA nickase and LSCA nickase, LPCA nickase and LPDCA nickase, is not excluded). Any of the LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas or such variants may be coupled (e.g. fused) with an effector protein for gene modification, such as a base editor, such as a deaminase for base editing or a reverse transcriptase for prime editing.

[0077] It will be appreciated that any variant of LFCA Cas, LBCA Cas, LSCA Cas, LPCA and LPCDA Cas (or the corresponding nickases obtained by mutagenesis of the catalytic site) having one or more amino acid changes by substitution or deletion, e.g. one or more conservative substitutions, may be similarly used as Cas9 endonucleases or Cas9 nickases, provided that the endonuclease and / or nickase activity is retained. Such variants that also retain the relaxed PAM specificity shown for LFCA Cas, LBCA Cas and / or LSCA Cas are of particular interest and form part of the present invention.

[0078] The variants of AnCas nucleases obtained by the ancestral reconstruction strategy of the invention forming part of the present invention or the corresponding variants of nickases obtained by mutagenesis of the catalytic site may have different degrees of sequence identity with the parent enzyme, provided that they retain, for example, one or more desired distinguishing properties: they may have, for example, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity.

[0079] Similar to the naturally occurring Cas9 nuclease, any of the above-described LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas or variants thereof (as well as their functional equivalents, such as those represented by the nodes on the evolutionary pathway in FIG. 9 or others obtained according to the ancestral reconstruction strategy of the present invention) may be converted to a dCas lacking nuclease activity by mutagenesis of the catalytic site and may further be coupled (e.g. fused) to a non-nuclease effector protein.

[0080] Thus, the LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas and LPCDA Cas discussed above, their variants (as well as their functional equivalents, such as those represented by the nodes on the evolutionary pathway in FIG. 9 or others obtained according to the ancestral reconstruction strategy of the present invention) may be used in any kind of genetic modification technique conceivable for naturally occurring Cas9 nucleases of existing species and their modified versions. These extend to use in combination with effectors for genetic modification or regulation, such as base editors, where binding is via RNA extension of guide and RNA binding domains as taught in WO 2017 / 011721 (Rutgers University, licensed to Horizon Discovery). See also Collantes et al., 2021. CRISPR J. 4(1):58-68.

[0081] Preferably, the Cas enzyme variant, e.g. linked or fused to a non-nuclease effector, may be any AnCas or variant thereof obtained according to the present invention having altered nuclease activity by mutagenesis of the conventional catalytic site, e.g. to have only nickase activity or no nuclease activity (i.e. dCas), and exhibits relaxed PAM specificity as observed, e.g., by LFCA Cas, LBCA Cas and / or LSCA Cas.

[0082] The present invention further provides nucleic acids for the expression of the AnCas proteins described herein, including their variants and functional equivalents, such as expression vectors for the expression of such proteins. Such vectors may be used with guide RNA or guide RNA expressed from DNA. Thus, a combination of a vector providing the AnCas nuclease or its variant or functional equivalent as taught herein and a suitable guide RNA may be provided for transfection into cells. Such combinations, including one or more vectors, including vectors for the expression of LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas or LPCDA Cas or its variant or functional equivalent in human cells, may be provided in the form of a pharmaceutical composition, including a pharmaceutically acceptable excipient.

[0083] Preferably, the Cas protein may be an LFCA Cas or a corresponding nickase. Of course, a corresponding nickase of any of the ancestor enzymes taught herein may be provided, such as an LBCA nickase, an LSCA nickase, an LPCA nickase, or an LPCDA nickase.

[0084] The present invention therefore further relates to nucleic acids capable of expressing a Cas nuclease or a Cas nuclease variant according to the invention.

[0085] In some embodiments, the nucleic acid is a DNA or RNA molecule. In some embodiments, the nucleic acid is a DNA molecule, e.g., a complementary DNA molecule. In some embodiments, the nucleic acid is an RNA molecule, e.g., a messenger RNA molecule.

[0086] In some embodiments, the nucleic acid is single-stranded or double-stranded. In some embodiments, the nucleic acid is single-stranded. In some embodiments, the nucleic acid is double-stranded.

[0087] In some embodiments, the nucleic acid comprises naturally occurring nucleotides. In some embodiments, the nucleic acid comprises a combination of naturally occurring and non-naturally occurring nucleotides.

[0088] In some embodiments, the nucleic acid is comprised in a vector. Non-limiting examples of suitable vectors include a plasmid, a fosmid, a cosmid, an artificial chromosome, or a viral vector. In some embodiments, the vector is comprised in a nanoparticle, such as a lipid nanoparticle.

[0089] The present invention further relates to a vector comprising a nucleic acid according to the invention as described herein above in combination with a guide RNA for targeting the Cas nuclease or a variant or functional equivalent thereof to a target DNA sequence, or a vector capable of expressing a guide RNA.

[0090] Alternatively, the novel AnCas nucleases taught herein, such as LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas or LPCDA Cas, or variants or functional equivalents thereof, such as the corresponding nickases discussed above, may be provided as ribonucleoprotein (RNP) complexes comprising guide RNA for transfection into cells, for example by electroporation into isolated cells. Thus, the present invention further refers to ribonucleoprotein complexes comprising the Cas nuclease or Cas nuclease variants according to the present invention, or the Cas nuclease or Cas nuclease variants and guide RNA for targeting the Cas nuclease or Cas nuclease variants to a target DNA sequence.

[0091] It should be understood that the term "guide RNA" as used herein may refer to a single molecule targeting RNA (sgRNA) or, as appropriate for naturally occurring Cas9, an RNA consisting of two sequences including (i) a DNA-targeting segment (crRNA) that includes a nucleotide sequence complementary to the target sequence, and (ii) a protein-binding segment (tracrRNA) that interacts with the Cas protein.

[0092] In a further aspect, the invention relates to a method for modifying or modulating a target nucleic acid sequence, e.g. a target DNA sequence, comprising contacting the target sequence with a complex comprising (i) a taught Cas protein, e.g. an LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas or LPCDA Cas as discussed above or a variant or functional equivalent thereof, and (ii) a guide RNA for targeting the Cas protein to the target sequence, (a) the contacting is with an isolated target nucleic acid sequence in vitro or ex vivo in a cell, with the proviso that methods of altering human germline identity are preferably excluded; and / or (b) The method is not a medical treatment performed on the human or animal body. The present invention provides a method for

[0093] The complex may further comprise a nucleic acid molecule encoding a transgene of interest, for example for introducing said transgene of interest into a target DNA sequence.

[0094] Preferably, the Cas protein may be, for example, an LFCA Cas, LBCA Cas or another exemplary AnCas nuclease described above that retains the same relaxed PAM requirements. Preferably, the Cas protein may be such an AnCas, but modified in the form of a fusion protein to exhibit only nickase activity or no nuclease activity.

[0095] Such methods extend to, for example, ex vivo genetic modification of human or animal cells. The Cas proteins taught herein, such as LFCA Cas, LBCA Cas or other AnCas (including variants and functional equivalents thereof), may have application in the context of genetic modification in plants, for example, but not limited to, optionally by modifying target sequences in protoplasts.

[0096] It will be appreciated that the present invention extends to combinations for use in therapeutic treatment by modifying or modulating a target nucleic acid sequence, e.g. a DNA sequence, wherein the combination comprises: (i) a Cas protein as taught herein, such as the LFCA Cas, LBCA Cas, LSCA Cas, LPCA Cas or LPCDA Cas discussed above, or a variant or functional equivalent thereof, or a polynucleotide capable of expressing same; (ii) a guide RNA for targeting the Cas protein to a target nucleic acid sequence or a polynucleotide capable of expressing the guide RNA; Includes.

[0097] In particular, the therapeutic treatment may include the prevention and / or treatment of a genetic disease. The combination may then further comprise a nucleic acid molecule encoding a transgene of interest, wherein said transgene of interest may for example compensate for a genetic defect involved in the genetic disease.

[0098] As shown above, for example, the low sequence identity of LFCA Cas with SpyCas9 is considered advantageous in the context of envisioning such uses, which may include, for example, Cas action in pathogenic bacteria, or for manipulation of the gut or skin microbiota.

[0099] The following exemplary embodiment illustrates the invention with reference to both obtaining and testing the Cas enzymes LFCA Cas, LBCA Cas and LSCA Cas, however, it is believed that other ancestral Cas proteins with advantageous properties can be obtained by the same strategy, depending on the choice of starting population of Cas enzyme sequences that provide the phylogenetic tree for evolutionary analysis, as described above. Predicted resurrections may be as old as 3 billion years. EXAMPLES

[0100] Example 1 Reconstructing the ancestral sequence of Cas nucleases from existing Cas9 sequences Using the SpyCas9 (Uniprot code: Q99ZW2) sequence as a query, we collected sequences of the gene Cas9 from several Firmicutes bacterial species from the Uniprot database. From this search, we confirmed the presence of hundreds of sequences of Cas9 genes within the classes Bacillario and Clostridia from the Firmicutes. We also found several sequences from Actinobacteria. After downloading the 59 sequences (Table 1), we constructed a sequence alignment, whereby the parts of the sequences showing significant conservation were used to confirm the common origin of the Cas9 sequences. Using Bayesian inference (BEAST software), we compiled a phylogenetic tree to confirm the phylogenetic relationship of these sequences. Using maximum likelihood methods, we reconstructed the five AnCas sequences to the last common ancestor (LFCA) of the Firmicutes phylum, dating back approximately 2.4 billion years (see Figure 1). We traced the evolutionary path from the LFCA Cas to the modern Streptococcus pyogenes. LFCA Cas was found to have only about 50% identity with SpyCas9 (with over 500 mutations). Other evolutionary pathways, such as those leading to Clostridium and Enterococcus species, were constructed. Genes encoding ancestral Cas enzymes were synthesized, cloned into expression vectors, expressed in E. coli, and purified in the laboratory. Five AnCas enzymes (shown in bold in Figure 1) were obtained from the pathway from LFCA to Streptococcus pyogenes. All five were found to be highly expressed, folded, and soluble despite sequence identities ranging from about 50 to almost 95% with SpyCas9. [Table 1] TIFF2024521806000003.tif252159TIFF2024521806000004.tif249159TIFF2024521806000005.tif251159TIFF2024521806000006.tif201159

[0101] For further information regarding evolutionary analysis, see Materials and Methods below.

[0102] Testing the endonuclease activity of AnCas enzymes To test whether the ancestral Cas exhibited endonuclease activity, we designed a DNA library with seven random nucleotides (NNNNNNN) and a gRNA to target the sequence after these nucleotides. This random DNA library was necessary because the LFCA Cas PAM sequence was unknown (Figure 2a). The sequence for providing the random library was cloned into the plasmid pUC18, and an 844 bp sequence was amplified to generate a linear DNA fragment. Two fragments were generated by SpyCas9 cleavage: one with 566 bp and a smaller one with 278 bp that contained the PAM sequence. Incubation of both LFCA Cas and SpyCas9 with gRNA targeting the PCR fragments successfully generated both fragments (Figure 2b), confirming that the ancestral Cas had catalytic activity. Furthermore, a DNA fragment containing the Streptococcus pyogenes PAM (NGG), extracted from the library by PCR, was recognized by both Cas enzymes. LFCA was able to recognize Streptococcus pyogenes PAM and cleave DNA (Fig. 2c).

[0103] To investigate the DNA cleavage kinetics of LFCA Cas, a DNA fragment containing Streptococcus pyogenes PAM (TGG) was cloned and incubated with LFCA Cas or SpyCas9 for different times ranging from 5 to 160 min. Both enzymes were incubated with gRNA and target DNA, and the reaction was stopped by adding loading buffer and EDTA. The samples were run on a 1% agarose gel to detect supercoiled DNA, nicked DNA and linear DNA (Figure 3a). Different DNA higher-order structures were observed after Cas enzyme activity on the agarose gel (Figure 3b). The band intensities were measured and the total cleavage rates by both enzymes were calculated at different times (Figure 3c). Both enzymes cleaved almost 100% of the plasmid DNA after 10 min. However, LFCA Cas cleaves one strand of DNA and the cleavage of the other strand increases with time, i.e., the nicking and endonuclease activities are observed to be separated in time, as shown in Figure 3d, whereas SpyCas9 cleaves both strands of DNA simultaneously.

[0104] In another set of experiments, AnCas genes were synthesized and cloned into the pBAD / gIII expression vector carrying an arabinose-inducible promoter and a gIII-encoded signal that directs AnCas to the periplasmic space. All AnCas were expressed at high levels in E. coli BL21 cells.

[0105] Activity testing began by assuming a simplistic scenario in which AnCas recognizes the sgRNA from Streptococcus pyogenes as well as its canonical 5'-NGG-3' PAM sequence. We designed the sgRNA containing a 20 nt-long spacer region targeted toward the DNA fragment upstream of the TGG PAM, all placed in a 4007 bp supercoiled plasmid. In vitro cleavage assays were performed by incubating AnCas or SpCas9 with the target DNA and sgRNA for different digestion times. Although there were clear differences in cleavage efficiency, all tested enzymes produced relaxed linear products indicative of nickase and DSB activity, respectively. As expected, SpCas9 showed nicked products after short incubation times and linear products after longer incubation times (Figure 15A). However, in the case of AnCas, the behavior changed from the most ancient LFCA AnCas to the more recent enzymes (Figure 15A). LFCA AnCas showed mainly nickase activity, and DSB activity became evident only after more than 60 min. Other AnCas showed a progressive behavior with more intense DSB activity in younger AnCas (Figure 15A). Both nicked and linear fractions for each AnCas and SpCas9 were quantified in three forms, namely total cleavage rate (Figure 15B), nicked fraction (Figure 15A) and linear fraction (Figure 15D), and plotted against incubation time, showing a progressive decrease in nicked fraction and an increase in linear fraction. SpCas9 had a higher proportion of linear products, and the oldest LFCA AnCas had the highest proportion of nicked fraction. The proportion of linear fraction and nicked product was plotted against geological age, showing an evolutionary trend from nickase to DSB activity (Figure 15E). The observed structural differences in LFCA AnCas proteins associated with HNH domain displacements, together with the evolutionary trend from nickase to DSB activity, suggest that the most ancient AnCas, i.e., LFCA AnCas, may display an ancestral HNH domain with reduced or suppressed activity.To investigate this, we tested the in vitro activity of the H838A LFCA AnCas mutant (H840A for the wild-type SpCas9 amino acid sequence). This mutant was able to produce nicked and, surprisingly, linear products, with a profile virtually identical to that obtained with wild-type LFCA AnCas (Figures 19A-D). These results suggest that LFCA AnCas may contain an immature HNH domain along with the RuvC domain that contributes to the nickase and DSB activity observed in LFCA AnCas, as has been previously shown for several V-type effector nucleases lacking the HNH domain, such as Cpf1 (Cas12a), Cas14 (Cas12f) or CasΦ (Cas12j). [Table 2]

[0106] In another set of experiments, the DNA cleavage activity of two AnCas, namely LBCA Cas and LSCA Cas, was compared with that of LFCA Cas and SpyCas9.

[0107] A plasmid containing a TGG PAM sequence after the target sequence was incubated with each enzyme for different times. The cleavage rates of the four enzymes were similar when comparing the total cleavage rates (Fig. 10a). However, the linearization and nick rates differed among the enzymes. LBCA Cas and LSCA Cas were found to have a higher linearization rate than LFCA Cas, but still lower than SpyCas9 (Fig. 10b). In contrast, the nick rates of LBCA Cas and LSCA Cas were lower than those of LFCA Cas, but higher than that of SpyCas9 (Fig. 10c).

[0108] The cleavage activity of AnCas enzyme was found to follow the trend shown in Figure 11. The ratio of double-strand breaks after 30 minutes of incubation time was found to be the highest for SpyCas9 and decreased with the age of the ancestral Cas enzyme (i.e., it can be observed that the ratio of double-strand breaks decreases as the ancestral enzyme gets older, such as %DSB of SpyCas9 > %DSB of LSCA Cas > %DSB of LBCA Cas > %DSB of LFCA Cas).

[0109] The opposite correlation was observed for the ratio of nicked templates formed (i.e., the nick cleavage rate % of SpyCas9 < the nick cleavage rate % of LSCA Cas < the nick cleavage rate % of LBCA Cas < the nick cleavage rate % of LFCA Cas).

[0110] By converting the ratios of DSB and nicked templates shown in Figure 11 into time-dependent values, the linearization rate and the nick rate could be calculated and plotted. The linearization rate and the nick rate seem to follow the trend seen when the rates are plotted against the age of each ancestor (Figure 12).

[0111] PAM determination As described in the following section on PAM library construction, the PAM specificity of LFCA Cas was determined using DNA fragments amplified by PCR from a DNA library. LFCA Cas was incubated with the gRNA and the DNA library for 1 hour, and the reaction products were electrophoresed on a 2% agarose gel. The small 278bp fragment was extracted from the agarose gel and analyzed by Ion Torrent next-generation sequencing (NGS). From the sequencing data, the frequency of each PAM recognized by LFCA Cas was analyzed, and the total ratio of each PAM to the overall frequency of each PAM in the library was calculated. The calculated frequencies were plotted in a PAM wheel graph to visualize the PAM affinity of the ancestral Cas (Figure 4a). This wheel graph shows the loss of PAM selectivity of LFCA Cas compared to SpyCas9.

[0112] To confirm this result, an in vitro PAM determination assay was performed using DNA plasmids carrying target DNA containing different PAM combinations (TNN trinucleotide combination and CCC). LFCA and SpyCas9 (10 nM) incubated with each PAM for 10 min had different cleavage activities. LFCA Cas showed similar nicking activity with all PAMs tested. In contrast, SpyCas9 showed cleavage with TGG PAM (its well-known canonical PAM sequence) and lower cleavage with other PAMs than LFCA Cas (Figure 4b).

[0113] In another set of experiments, we investigated the ability of AnCas endonucleases to recognize different PAMs. To determine the preferred PAM sequence of each AnCas, we designed a DNA library containing the target sequence corresponding to all possible PAMs followed by seven random nucleotides (NNNNNNN). sgRNAs were designed with 20 nucleotides complementary to the scaffold and target sequence of Streptococcus pyogenes. PCR primers were designed to amplify an 844 bp fragment containing both the target and PAM sequences, which was used as a substrate for AnCas and SpCas9. In vitro digestion with purified Cas proteins and transcribed sgRNA was performed with the PCR target. All five AnCas generated two fragments, one of 566 bp and a smaller one of 278 bp that contained the PAM sequence recognized by AnCas. Small fragments were purified, sequenced by next-generation sequencing (NGS), and analyzed, allowing us to determine the PAM sequence diversity from each AnCas and to infer how evolution has changed it. Figure 16A summarizes the results of the PCR cleavage assay in the form of a PAM wheel graph (Krona plot) for the five AnCas and SpCas9. As previously observed, LFCA AnCas showed no preference for any of the PAM sequences tested. For other Cas proteins, preferences for specific nucleotides at target-proximal positions 2 and 3 were detected (Figure 22). For example, in the case of LBCA AnCas, a slight preference for NGG was evident, but additional PAM sequences (NNG) were also detected. For more recent AnCas, the bias towards NGG was more evident (data not shown).

[0114] After all sequences were analyzed, the percentage of reads containing NGG PAM was plotted against the geological age of each AnCa estimated in the phylogenetic analysis. A trend reflecting NGG enrichment over time was observed, indicating that NGG fidelity is an evolutionary feature that describes a gradual progression from PAMless to NGG selective in the more recent Streptococcus ancestor (Figure 16B). This confirms the hypothesis of an evolving adaptive response to PAM recognition, which is expected as the number of spacers acquired by the host cell increases over time. Finally, a strong PAM recognition ability would be required to avoid self-cleavage of CRISPR loci, especially in a scenario where increasing DSB activity over nickase activity (which is deleterious in most prokaryotes) further increases the selection pressure against this ability. Although PAM-tolerant (i.e., "nearly PAMless") Cas9 variants have been described previously, LFCA AnCas is, to the best of our knowledge, the first completely PAMless Cas9 endonuclease reported to date.

[0115] To further explore the PAMless capability of LFCA AnCas, an in vitro PAM determination assay was designed to test the cleavage of target DNA adjacent to all six PAM sequences (TAC, TCC, TAT, TTT, TTC and TAC) within the common TNN PAM. The CCC PAM was also included in the set to confirm the possibility of other than the first T nucleotide. AnCas effectors were incubated with each of the target DNAs and sgRNA for 10 min, and the cleavage products were confirmed by agarose gel (data not shown). Both nicked and linear products were observed, indicating cleavage activity by all TNN PAM sequences. In the case of SpCas9, only the TGG PAM showed double-stranded cleavage of supercoiled DNA substrates. The percentage of nicked and linear products was quantified for each PAM represented in Figure 16C. For the most ancient AnCas (LFCA Cas and LBCA Cas), the percentage of cleavage was similar for all PAM sequences tested, yielding primarily nicked products, as expected given the incubation time. For the younger AnCas and SpCas9, the cleavage fraction reached high levels for the TGG PAM, indicating NGG PAM selectivity. For the CCC control, the cleavage profile was similar to that obtained from the non-NGG PAM sequence.

[0116] In another set of experiments, PAM determination was performed using DNA fragments amplified by PCR from the DNA library, again as described in the PAM library construction section below. LBCA Cas or LSCA Cas was incubated with gRNA and DNA library for 1 hour by running the reaction in a 2% agarose gel. Small fragments of 278 bp were extracted from the agarose gel and analyzed by Ion Torrent next-generation sequencing (NGS). The frequency of each PAM recognized by the AnCas enzyme was determined from the sequencing data, and the total percentage of the overall frequency of each PAM in the library was calculated. The calculated frequencies were plotted in a PAM wheel graph to visualize the PAM affinity for both AnCas (Figure 13a). This wheel graph shows similar PAM selectivity of LBCA and LSCA Cas, as shown by LFCA Cas.

[0117] PAM determination assays were performed in vitro using DNA plasmids carrying target DNA with different PAM nucleotide combinations (TNNs). LBCA Cas and LSCA Cas (10 nM) were incubated with each PAM for 10 min (Figure 13b). LBCA Cas showed similar or even higher cleavage rates with some PAMs than LFCA Cas. LSCA Cas also showed cleavage with all PAM sequences tested, but with lower activity. Higher linear activity was consistently observed for all PAMs tested compared to LFCA Cas, highlighting the higher linear cleavage rates of LBCA Cas and LSCA Cas.

[0118] Finally, the inventors have demonstrated that the Cas9 proteins disclosed in the art, in particular The so-called "ancestral Cas9 protein" (SEQ ID NO: 268 of WO '533), as disclosed in WO 2021 / 084533 A1, and - Walton et al.'s (2020. Science. 368(6488):290-296) nearly PAMless Cas9 protein "SpG" *and "SpRY" ** For this purpose, we wished to compare the PAM selectivity of LFCA Cas (which, as shown above, has no preference for any PAM sequences) and LBCA Cas (which was found to be nearly PAMless due to its selectivity for NNG sequences). * SpG:D1135L / S1136W / G1218K / E1219Q / R1335Q / T1337R ** SpRY:A61R / L1111R / D1135L / S1136W / G1218K / E1219Q / N1317R / A1322R / R1333P / R1335Q / T1337R

[0119] In all cases, a wild-type SpCas9 control was included as a reference.

[0120] A DNA library containing seven random nucleotides was designed and cloned into pUC18 plasmid (Genscript Inc.). The random library was transfected into XL1blue E. coli and amplified for several hours to achieve maximum variability in the PAM sequence.

[0121] PAM determination assays were performed by incubating 3 nM of DNA library plasmid containing 30 nM of each test Cas protein in cleavage buffer with gRNA targeting 20 nucleotides upstream of 7 random nucleotides. The reactions were incubated at 37°C for 1 h, stopped by adding 6x loading dye with EDTA (NEB), and run on a 2% agarose gel. Gels were stained with SYBR gold (ThermoFisher Scientific) and imaged with a ChemiDoc XRS+ system (Bio-Rad). PAM library-specific PCR-based amplification was performed using adapters and specific oligos: [Table 3]

[0122] The fragments were sequenced by Illumina sequencing, and the reads were mapped to the reference sequence using Geneious Prime (2020 version). Illumina miSeq reads were aligned to the amplified sequence using minimap2 for short reads to exclude non-specific sequences. Then, from the aligned reads, reads with three nucleotides before the PAM region were selected. Nucleotides in the region of interest were extracted using a custom script. Finally, logo plots of the PAM region were obtained using ggseqlogo, and PAM wheel graphs for each sample were graphed using KronaTools.

[0123] Our data showed that of the six different Cas proteins tested, only the LFCA Cas was completely PAMless. The LBCA Cas was the second Cas protein with the least restrictive PAM requirements. All other Cas proteins had more restrictive PAM preferences (Figure 24).

[0124] Interestingly, the so-called "ancestral Cas9 protein" disclosed in WO 2021 / 084533 A1 exhibited an NGG PAM requirement identical to that of wild-type SpCas9.

[0125] These data therefore confirm that our AnCas proteins are truly PAMless or at least nearly PAMless, contrary to what has been disclosed in the art.

[0126] gRNA recognition The promiscuity of PAM recognition exhibited by the most ancient AnCas raised the question of whether these AnCas also exhibit promiscuity for gRNA recognition. Although reconstruction of ancestral gRNAs would be ideal, the variability in the sequences of crRNA repeats and tracrRNAs from different species makes this very difficult. To overcome this limitation and still evaluate the promiscuity of AnCas, modern sgRNAs from different species were tested. A total of five sgRNAs were selected from Streptococcus thermophilus, Enterococcus faecium, Clostridium perfringens, Staphylococcus aureus, and Finegoldia magna, covering several classes of the Firmicutes phylum. These sgRNAs were selected according to previous studies on the classification and function of sgRNAs, and here we divided the sgRNAs into seven clusters. These different sgRNAs were contrasted with Streptococcus pyogenes guides that contained two sizes of spacers, 18 and 20 nucleotides long, termed "18sgRNA" and "20nt sgRNA", respectively.

[0127] SpCas9 and five AnCas were incubated with the target plasmid DNA and the TGG PAM recognition site at 37°C for 10 min. From the agarose gel of the cleavage products in Figure 17A, it can be observed that, as expected, SpCas9 linearized the plasmid DNA only when its own sgRNA was used, but was more efficient when the 20 nt spacer version was used, and sgRNAs from other species mainly produced nicked products, leaving most of the supercoiled DNA substrate intact. On the contrary, LFCA Cas and LBCA Cas could nick and linearize the plasmid DNA with all sgRNAs, with the sgRNA from Enterococcus faecium showing better efficiency for LFCA Cas, and the 18 nt sgRNA from Streptococcus pyogenes being preferred for LBCA Cas. Other AnCas were also tested, and it was observed that mainly LFCA Cas and LBCA Cas had a remarkable promiscuity towards sgRNAs. All other AnCas and SpCas9 appeared to work best with the 20 nt sgRNA from Streptococcus pyogenes (Figure 17B).

[0128] Previous studies have shown the contribution of the REC domain to the specificity of sgRNA recognition. As shown in Table 3, this domain shows the highest RMSD difference with a decreasing RMSD trend from the oldest to the newest AnCas. These findings, together with the sgRNA promiscuity observed in the oldest AnCas, suggest that selection pressure may have induced Cas nucleases to improved guide specificity over time. Indeed, this promiscuity has already been observed in type II-C Cas9s, which have been suggested to be an old memory of Cas9 nucleases. In these nucleases, this promiscuity has also been associated with PAM-independent ssDNA cleavage and weaker substrate DNA unwinding ability. [Table 4]

[0129] In another set of experiments, we investigated the ability of LFCA Cas to use sgRNA with targeting sequences bound to various tracrRNA sequences using the same in vitro plasmid cleavage assay described above. We used tracrRNA sequences that correspond to the tracrRNA components used by Cas9 gRNAs of various existing bacterial species. Thus, a plasmid containing Streptococcus pyogenes PAM TGG was provided. It was incubated with either SpyCas9 or LFCA Cas in the presence of sgRNA with tracrRNA derived from the bacterial species Streptococcus thermophilus, Enterococcus faecium, Clostridium perfringens and Finegoldia magna or Streptococcus pyogenes sgRNA, respectively, containing the normal 20 nt spacer or a shortened 18 nt spacer. For further information on the gRNA sequences used, one can refer to Gasiunas et al. (2020) (Nat Commun. 11(1):5512; see Supplementary Data providing gRNA sequences of identified Cas9 orthologs). After incubation, the reaction products were run on an agarose gel to observe the degree of supercoiled, nicked and linear DNA. The gel results are shown in Figure 8. It is clear that LFCA Cas9 has a very flexible gRNA usage. It was able to either nick or linearize the plasmid DNA regardless of the tracrRNA element of the gRNA used. Indeed, improved cleavage was observed with several sgRNAs other than the conventional Streptococcus pyogenes sgRNA. Such gRNA flexibility was not shown for SpyCas9 and is considered to be another novel property of LFCA Cas as shown above.

[0130] Heat and pH stability The thermal stability of LFCA Cas was investigated by performing the cleavage reaction at pH 7.9 and different temperatures ranging from 4 °C to 60 °C for 1 h. LFCA Cas showed higher activity than SpyCas9 at lower temperatures (4 °C and 20 °C) and higher thermal stability at 53 °C to 60 °C (Figure 5a). Nicking and endonuclease activities were calculated and it was observed that LFCA Cas had nicking activity at lower temperatures, while at higher temperatures the two activities were equally distributed (Figure 5b).

[0131] For the evaluation of pH stability, the assay was performed at different pH values ​​(4-9.5) and 37 °C (Figure 5c). At acidic pH, LFCA Cas maintained higher activity compared to SpyCas9. At alkaline pH, the activity of LFCA Cas remained the same, and SpyCas9 showed its optimal performance. With regard to nicking and endonuclease activity, pH affects activity as does temperature (Figure 5d). These results generally indicate the high stability possessed by the ancestral enzyme.

[0132] As mentioned above, AnCas, especially LFCA AnCas, may share some commonalities with V-type effector nucleases lacking HNH domains (e.g., Cpf1 (Cas12a), Cas14 (Cas12f) or CasΦ (Cas12j)27-30). Considering that many of their ancestor enzymes showed the ability to function in wider pH and temperature ranges, thus indicating later adaptation to environmental conditions, in another set of experiments, AnCas nucleases were tested under different temperature and pH conditions. As shown in Figure 21A-B, unlike SpCas9 and newer AnCas, whose activities declined sharply, the most ancient AnCas, such as LFCA Cas, LBCA Cas and LSCA Cas, showed high activity at pH values ​​below 7. With regard to temperature, AnCas endonucleases outperformed SpCas9 at low and high temperatures, i.e., below 10°C and above 50°C.

[0133] Genome editing in HEK293T cells using LFCA Cas HEK293T cells were transfected with an expression plasmid carrying the LFCA Cas humanized gene to examine the effectiveness of the ancestral enzyme in editing genomic DNA. A gRNA was designed to target the AAVS1 locus by Streptococcus pyogenes PAM. The expression plasmid containing the encoding LFCA Cas was co-transfected with another plasmid to express the gRNA (Figure 6a).

[0134] Genomic DNA was then extracted to check for insertion and deletion events (indels). Intracellular LFCA Cas expression was confirmed by immunofluorescence imaging of cells using anti-Cas9 antibody (orange-yellow) (Figure 6b). Cell nuclei were stained blue with DAPI. Similar to SpyCas9, cells that expressed LFCA Cas in the nucleus were observed. 72 hours after transfection, genomic DNA was extracted from HEK293T cells and a fragment of the AAVS1 locus was amplified, where Cas enzyme cleavage was targeted.

[0135] Genome editing was confirmed by performing a T7E1 endonuclease assay with these fragments (Figure 6c). After T7E1 incubation, two expected fragments were observed, confirming indel formation after LFCA Cas transfection. The same was observed with intracellular expression of SpyCas9 as a control.

[0136] We further investigated the knock-in activity after LFCA Cas and SpyCas9 genome cleavage. We followed the same strategy as in the previous experiment, but added a DNA template containing the eGFP gene flanked by sequences homologous to the AAVS1 locus to promote homologous end joining (HDR) (Fig. 6d). 72 hours after transfection, cells showed green fluorescence (Fig. 6e). Quantified fluorescence intensity showed higher values ​​in cells transfected with LFCA Cas (Fig. 6f).

[0137] A similar knock-in experiment was performed targeting the AAVS1 region, but using a different PAM than that of SpyCas9. In this experiment, the TTC PAM was targeted (Figure 7). Cells were transfected with gRNA and LFCA Cas or SpyCas9. Fluorescence was observed after 72 hours in all samples from the DNA template (some transient fluorescence in the TTC sample with SpyCas9). gDNA was extracted and the AAVS1 locus was amplified. PCR amplicons were run on a gel. The expected band was observed in all LFCA Cas samples, but not in the samples with SpyCas9. gDNA was extracted and amplified at the AAVS1 locus. PCR amplicons were run on an electrophoretic gel, and as expected, the expected band was observed in all samples, except for the TTC PAM targeted by SpyCas9.

[0138] Example 2: Endonuclease activity on single-stranded DNA and single-stranded RNA As mentioned above, the most ancient AnCas (LFCA Cas and LBCA Cas) showed notable nickase activity. This nickase activity may be related to ssDNA activity. It was suggested that ssDNA cleavage activity was an ancestral trait present in smaller Cas9s, such as subtype II-C Cas9s. This may also be reflected in the nickase activity of ancestral forms from subtype II-A, such as AnCas. Earlier forms of Cas9 with smaller catalytic domains may have been the origin of this ssDNA cleavage activity that was still present in the larger ancestral nucleases, which then gradually evolved towards DSB activity over time as part of the differentiation process.

[0139] The activity of the ancestral Cas9 enzymes (LFCA Cas, LBCA Cas, and LSCA Cas) on single-stranded DNA was tested. The single-stranded plasmid m13mp18 linearized by EcoRI restriction enzyme was used as the substrate. AnCas enzyme and SpyCas9 were incubated with the plasmid and gRNA designed to target the plasmid, respectively. As a control, DNA and enzymes (but without gRNA) were incubated together (Figure 14). The three ancestral enzymes were found to cleave single-stranded DNA with or without gRNA. Similar activity was observed with SpyCas9 when manganese was present in the reaction system. LSCA Cas showed the highest cleavage rate for single-stranded DNA.

[0140] In another set of experiments, the most ancient AnCas was tested with an 85-nt ssDNA substrate containing a target sequence complementary to the 20-nt spacer region of the Spy-sgRNA. As shown in Figure 17C and Figure 17E, LFCA Cas and LBCA Cas showed the highest levels of ssDNA cleavage, higher than that of SpCas9. Exponential fits to the data showed much faster rates for LFCA Cas and LBCA Cas, with LBCA Cas reaching nearly complete cleavage (Table 4). Activity was tested against a 60-nt ssRNA target, which showed comparable results (Figure 17D and Figure 17F). Exponential fits of cleavage activity showed the maximum rate and amplitude for LBCA Cas, again reaching complete cleavage (Table 4). These results indicate that both ancient LFCA Cas and LBCA Cas, especially LBCA Cas, behave as RNA-guided ribonucleases.

[0141] The activity of the most ancient AnCas on single-stranded substrates suggests that early Cas nucleases may have been active on those substrates, which appears to be an ancient trait as mentioned earlier. Given that the remarkable activity of LBCA Cas on ssDNA and ssRNA resembles that of Cas12a, Cas14 and Cas13a, these abilities may have further important implications, suggesting a relationship among the activities of all class 1 effector nucleases. This functional promiscuity also recommends LBCA Cas as a highly versatile endonuclease for genome editing applications.

[0142] We further investigated whether the most ancient AnCas endonucleases, which show more promiscuous characteristics, may also have different responses to anti-Cas9 antibodies. LBCA Cas and LFCA Cas were incubated with anti-Cas9 rabbit antibodies. ELISA tests showed reduced antibody binding (Figure 17g). This is expected, considering that the host organisms carrying these nucleases have been extinct for a long time and therefore have not been in contact with any organisms. It can be inferred that antibodies against Cas9 may have a weaker response to ancient Cas forms. This lower antibody response may be interesting for potential applications in in vivo editing, where immune responses against SpCas9 and other modern endonucleases represent a current limitation. [Table 5]

[0143] Example 3: In vivo activity of AnCas variants To answer the question of whether these synthetic ancestral Cass could make DNA cleavage, i.e. double-strand breaks (DSBs), and induce editing in cells by non-homologous end joining (NHEJ) under conditions similar to those associated with standard SpCas9, the genome editing activity of these ancestral nucleases was tested in mammalian cells (HEK293T) in culture. These cells were co-transfected with plasmid vectors containing humanized versions of AnCas or SpCas9 as well as the corresponding sgRNA (standard sgRNA from Streptococcus pyogenes carrying a 20-nt spacer target with SEQ ID NOs: 237-239). 72 hours after co-transfection, cells were harvested and genomic DNA was extracted. In vitro site-specific editing was measured in HEK293T cells by next-generation sequencing (NGS) using advanced analysis with Mosaic Finder software.

[0144] As shown in Figure 18, AnCas endonucleases performed robust gene editing in human genomic DNA, except for LFCA Cas. This is expected considering the unique feature of LFCA Cas, which probably does not use the HNH domain for cleavage, and appears to work better on single-stranded substrates, similar to other types of Cas nucleases.

[0145] Site-specific cleavage was tested using traffic light reporters (TLRs) based on RFP reconstruction. This method allows monitoring of DNA repair in HEK293T cells based on fluorescence-activated cell sorting (FACS). Again using conditions optimized for SpCas9, the results are in agreement with those determined by NGS (Figure 23), indicating their robustness.

[0146] material and method Ancestral sequence reconstruction The starting Cas9 sequences described above and listed in Table 1 and Figure 1 and Figure 9 were downloaded from the NCBI database. Sequence alignment was performed using MUSCLE software on the MEGA platform and manually edited. The best evolutionary model was estimated using MEGA, resulting in Jones-Tylor-Thornton (JTT) with gamma distribution model. Phylogenetic tree estimation was performed using BEAST v1.8.4 package software with BEAGLE library for parallel processing and based on Bayesian estimation using Markov Chain Monte Carlo (MCMC). Divergence times were estimated by the uncorrelated log-normal clock model (UCLN) using molecular information from TTOL with default birth and death rates. Calculations were performed on a multi-core server. 25% of them were discarded as burn-in from the generated phylogenetic tree using the LogCombiner utility from BEAST. TRACER was used to check the MCMC log files to ensure that all parameters showed an effective sample size (ESS) > 100. The posterior probabilities of all nodes were above 0.65, and most of them were close to 1. Phylogenetic trees were visualized and edited using Figure Tree v1.4.2. Finally, ancestral sequence reconstruction was performed by maximum likelihood using PAML4.8 with gamma distribution for variable substitution rates across sites and JTT model. Posterior probabilities were calculated for all amino acids, and the residue with the highest posterior probability was selected for each site. From the phylogenetic tree for reconstruction, we selected the last Firmicutes common ancestor (LFCA), the last Bacillario common ancestor (LBCA, the last Streptococcus common ancestor (LSCA), the last pyogenes common ancestor (LPCA), and the last pyogenes / haemolytics common ancestor (LPDCA).

[0147] The node sequences identified by the above methods are set forth in the Sequence Listing provided with this application, which is expressly incorporated herein in its entirety.

[0148] Protein production and purification The LFCA Cas coding sequence was synthesized with codon optimization for E. coli cell expression. The coding sequence was cloned into pBAD / His expression vector (ThermoFisher) and transfected into E. coli BL21(DE3) (Life Technologies) for protein expression. SpyCas9 expression plasmid was purchased from Addgene (plasmid #62934). Cells were incubated at 37°C in LB medium until OD600 reached 0.6. L-arabinose was added to cells to 0.1% for LFCA Cas expression, IPTG was added to cells to 1 mM for SpyCas9 expression, and protein induction was performed overnight at 20°C. Cells were pelleted by centrifugation at 4000 rpm. The pellet was resuspended in extraction buffer (20 mM HEPES (pH 7.5), 300 mM NaCl, 25 mM imidazole, 0.5 mM TCEP). The pellet was added with 100 mg / mL lysozyme (Thermo Scientific) with 15 min incubation. The pellet was then sonicated at 30% amplitude and 3 cycles for 10 min. Cell debris was separated by ultracentrifugation at 33,000 g for 1 h. For purification, the supernatant was mixed with a His GraviTrap affinity column (GE Healthcare) and eluted with elution buffer (20 mM HEPES (pH 7.5), 300 mM NaCl, 500 mM imidazole, 0.5 mM TCEP). The protein was further purified by size exclusion chromatography using a Superdex 200HR column (GE Healthcare) and eluted with 20 mM HEPES (pH 7.5), 1 M KCl, 10 mM MgCl2, 0.5 mM TCEP. To confirm protein purification, SDS-PAGE was used with 8% gels. Protein concentration was calculated by measuring absorbance at 280 nm on a Nanodrop 2000C.

[0149] gRNA synthesis gRNA with a sequence complementary to the target was synthesized and cloned into pUC18 vector. gRNA sequence was amplified by PCR using Phusion® Hot Start Flex DNA polymerase (NEB). PCR product was purified using mi-PCR purification kit (Metabion). gRNA was synthesized using HiScribe T7 High Yield RNA Synthesis Kit (NEB). PCR fragment had T7 promoter at 5' end and sequence from sgRNA of Streptococcus pyogenes at 3' end. Reaction was incubated overnight and sgRNA was purified according to Monarch® RNA Purification Column Kit protocol. gRNA integrity was analyzed by electrophoresis through 2% agarose gel containing TBE buffer.

[0150] In vitro cleavage assay In vitro cleavage assays were performed with purified LFCA Cas and SpyCas9. In all assays, 30 nM Cas nuclease was incubated with 30 nM gRNA at 1:1 ratio in cleavage buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg BSA, pH 7.9) for 15 min at 37 °C. Then, 3 nM target DNA was added and incubated for different times depending on the experiment. The reaction was stopped by adding 6x loading dye with EDTA (NEB), and the final reaction products were run on a 2% agarose gel. The gel was stained with SYBR gold (ThermoFisher) and imaged with a ChemiDoc XRS+ system (Bio-Rad). Cleavage was quantified with ImageJ.

[0151] In vitro thermal and pH stability The assays were performed according to the protocol previously described for in vitro cleavage, except for the modified conditions. The assay for thermostability was performed at pH 7.9 with temperatures varying from 4 to 60 °C. The assay for pH stability was performed at 37 °C with pH varying from 4 to 9.5. After 1 h, the reaction was stopped by adding 6x loading dye with EDTA (NEB), and the final reaction products were run on a 2% agarose gel. The gel was stained with SYBR gold (ThermoFisher) and imaged with a ChemiDoc XRS+ system (Bio-Rad). Cleavage was quantified with ImageJ.

[0152] Building the PAM library A DNA library containing seven random nucleotides was designed and cloned into pUC18 plasmid from Genscript. This random library was transfected into XL1blue E. coli and amplified for several hours to achieve maximum variability in the PAM sequence. Primers from the DNA library containing seven random nucleotides (F'AATAGGCGTATCACGAGGC (SEQ ID NO: 6) and R'AGCGAGTCAGTGAGCGAG (SEQ ID NO: 7)) were used to amplify a 844 bp PCR fragment.

[0153] PAM decision The PAM determination assay was performed by incubating 3 nM PCR fragments from the DNA library with 30 nM LFCA Cas and 30 nM gRNA targeting 20 nucleotides upstream of the 7 random nucleotides in cleavage buffer. The reaction was incubated at 37 °C for 1 h. The reaction was stopped by adding 6x loading dye (NEB) with EDTA, and the final reaction products were run on a 2% agarose gel. The gel was stained with SYBR gold (ThermoFisher) and imaged with a ChemiDoc XRS+ system (Bio-Rad). A small fragment of 278 bp was purified from the agarose gel using a GeneJet gel extraction kit (ThermoFisher). The fragment was sequenced by Ion Torrent, and the resulting reads were mapped in the reference sequence. Reads aligned to the reference with 0 mismatches were selected, and the frequency was calculated for each PAM.

[0154] Genome editing in HEK293T cells Cells were maintained in DMEN+10% FBS medium supplemented with 1% (w / v) L-glutamine and penicillin-streptomycin (100 IU / ml). Humanized coding sequences for LFCA Cas and SpyCas9 were cloned into pCDNA3.1 (ThermoFisher) expression vectors, and gRNAs were cloned into TOPO vectors (ThermoFisher). Plasmids were incubated with Lipofectamine LTX (ThermoFisher) for 5 min and co-transfected into cells. The medium was changed 24 h after transfection, and cells were harvested 72 h later. gDNA was extracted from cells using DNAzol reagent (ThermoFisher) according to the manufacturer's protocol. Primers (F'TATTGTTCCTCCGTGCGTCAG (SEQ ID NO: 8) and R'GACGAGAAACACAGCCCCA (SEQ ID NO: 9)) from gDNA were used to amplify DNA targets by PCR using Phusion® Hot Start Flex DNA polymerase (NEB). T7EI assays were performed using these PCR amplicons as substrates to confirm indel formation. T7E1 endonuclease (NEB) was used according to the manufacturer's protocol. Reactions were stopped by adding 6x loading dye with EDTA (NEB) and the final reaction products were run on a 2% agarose gel. Gels were stained with SYBR gold (ThermoFisher) and imaged on a ChemiDoc XRS+ system (Bio-Rad).

[0155] The same strategy as in the knock-in experiments was used, except that a double-stranded DNA template containing the eGFP gene and the CMV promoter flanked by 500 bp homology arms of the AAVS1 locus was utilized. Immunofluorescence was quantified by confocal microscopy after 72 hours.

[0156] Immunofluorescence studies After 24 h of transfection of LFCA Cas and SpyCas9 plasmids, HEK293T cells were fixed with 4% paraformaldehyde for 30 min. Cells were incubated with 0.2% TritonX-100 / PBS for 30 min at room temperature, then incubated with 3% BSA, 0.05% Tween20 for 1 h for a blocking step. Cells were washed three times with TPBS (0.05% Tween-PBS) and incubated with polyclonal anti-Cas9 antibody (1:100, 600-401-GK0, Thermofisher) for 1 h at 37 °C. Cells were washed three times with TPBS and incubated with secondary antibody (Alexa Fluor 555-labeled goat anti-rabbit, 1:200, A-21428, Thermofisher) for 10 min. DAPI was added at this step, cells were washed once at the end and visualized by confocal microscopy.

[0157] In vitro cleavage assay for gRNA promiscuity For in vitro cleavage for gRNA promiscuity, a DNA plasmid carrying a TGG PAM was used. Cleavage assays were performed at 37°C in cleavage buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / BSA, pH 7.9). 3 nM AnCas and SpCas9 were incubated with 3 nM sgRNA of each bacterial species in a 1:1 ratio in cleavage buffer for 15 min, and 3 nM DNA plasmid was added. The reaction was stopped after 10 min by adding 6x loading dye with EDTA (NEB) and run on a 2% agarose gel. Similarly, the gel was stained with SYBR gold (ThermoFisher Scientific) and imaged with a ChemiDoc XRS+ system (Bio-Rad). Cleavage was quantified with ImageJ.

[0158] In vitro cleavage assays for ssDNA and ssRNA In vitro cleavage assays were performed with purified LFCA (FCA) AnCas, LBCA (BCA) AnCas and SpCas9 endonucleases. In all assays, 30 nM enzyme was incubated with 30 nM sgRNA (Spy-sgRNA 20nt) at a 1:1 ratio in cleavage buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / BSA, pH 7.9) for 15 min at 37 °C. Then, 3 nM target (ssDNA or ssRNA) was added and incubated for different time intervals (0, 5, 10, 30 and 60 min). For ssDNA target, the reaction was stopped by adding 6x loading dye (NEB) with urea. Samples were boiled at 80 °C for 10 min and resolved on a 2.5% denaturing urea agarose gel. For ssRNA targets, reactions were stopped by adding 2x RNA gel loading buffer containing urea (NEB). Samples were boiled for 10 min at 95°C and resolved by 15% denaturing urea polyacrylamide gel electrophoresis. In all cases, gels were stained with SYBR gold (ThermoFisher Scientific) and imaged with a ChemiDoc XRS+ system (Bio-Rad). Cleavage was quantified with ImageJ and fitted to a single exponential decay curve.

[0159] ELISA Testing Elisa testing was performed using a modified protocol described elsewhere. 60Briefly, 1 μg / well of SpCas9, LFCA AnCas, LBCA AnCas and bovine serum albumin (BSA, Sigma Aldrich) were diluted in 1× bicarbonate buffer and coated on a 96-well plate (ThermoFisher Scientific) overnight at 4°C. The plate was washed with 1× wash buffer (TBST, ThermoFisher Scientific) and blocked with 1% BSA blocking solution for 1 h at room temperature. Anti-Cas9 rabbit antibody (Rockland, 600-401-GK0) was diluted 1:25000 in 1% BSA blocking solution and the plate was incubated for 2 h at room temperature. The plate was then washed and HRP-conjugated goat anti-rabbit IgG (H+L) (Invitrogen) diluted 1:2000 in 1% BSA blocking solution was added and incubated for 1 h at room temperature. Finally, 3,3',5,5'-tetramethylbenzidine ELISA substrate solution (ThermoFisher Scientific) was added and incubated at room temperature for 10 min. The reaction was stopped with 1N sulfuric acid. The absorbance was measured at 450 nm using a VICTOR X5 microplate reader (PerkinElmer).

[0160] In vivo cleavage of human HEK293T cells Functional validation of the ancestral Cas nuclease was performed in human HEK293T cells as described elsewhere (Harms, DW et al. Human Genetics 83, 2014). Cells were grown in DMEM medium (Dulbecco's Modified Eagle's Medium, Gibco) supplemented with sterile filtered 10% fetal bovine serum (FBS), 10 mM HEPES (pH 7.4), 2 mM L-glutamine and penicillin (100 IU / ml)-streptomycin (100 μg / ml) and handled under sterile conditions using a sterile hood. HEK293T cells were cultured in an incubator at 37 °C, 95% humidity and 5% CO2. Humanized AnCas was cloned into the pcDNA3.1 plasmid expression vector (ThermoFisher). sgRNA target sequences were analyzed using the Breaking-Cas web tool. 62The plasmid was designed using and cloned into the MLM3636 plasmid vector (Addgene #43860) by the Golden Gate cloning method. SpCas9 from the hCas9 plasmid (Addgene #41815) was used as a positive control. For in vivo genome editing studies, cells were cultured at 4 × 10 in a volume of 0.5 ml of DMEM without antibiotics. 5 Cells were seeded in 24-well plates at a density of 1000 cells / ml. The cells were transfected with 1 μg of hCas / hAnCas plasmid and 0.5 μg of the corresponding sgRNA plasmid in 2 μl of Lipofectamine 2000 (Life Technologies) diluted in 100 μl of Opti-MEM (Gibco) per well. 72 hours after transfection, genomic DNA was isolated using the High Pure Template Preparation Kit (Roche). Indel generation was assessed by T7 endonuclease I assay on PCR-amplified DNA fragments surrounding the targeted DSB.

[0161] Reconstructed ancestral sequence with node identification according to Fig. 1 All were 1340 or 1368 amino acid residues (depending on the reconstruction method). The sequences are disclosed as set forth in SEQ ID NOs: 1-5 and 10-236.

[0162] SEQ ID NOs: 1-5 correspond to the ancestral Cas proteins exemplified above.

Claims

1. (i) An LFCA nuclease having the amino acid sequence set forth in SEQ ID NO: 1, (ii) An LBCA nuclease having the amino acid sequence set forth in SEQ ID NO: 2, (iii) An LSCA nuclease having the amino acid sequence set forth in SEQ ID NO: 3, (iv) An LPCA nuclease having the amino acid sequence set forth in SEQ ID NO: 4, (v) An LPDCA nuclease having the amino acid sequence set forth in SEQ ID NO: 5, or (vi) A Cas nuclease comprising or consisting of a variant of the Cas nuclease described in any one of (i) to (v), wherein the variant - has at least 60% sequence identity with the amino acid sequence of SEQ ID NO: 1, or - has at least 75% sequence identity with the amino acid sequence of SEQ ID NO: 2, or - has at least 80% sequence identity with the amino acid sequence of SEQ ID NO: 3, or - has at least 85% sequence identity with the amino acid sequence of SEQ ID NO: 4, or - has at least 95% sequence identity with the amino acid sequence of SEQ ID NO: 5 and further, the variant has the following prominent characteristics when compared to SpyCas9: (a) A higher percentage of nicked DNA plasmid templates and / or a lower percentage of linearized DNA plasmid templates under the condition that substantially only linearized DNA plasmid templates are obtained by SpyCas9, (b) A higher nicking rate and / or a lower linearization rate for DNA plasmid targets under the condition that only linearization occurs substantially by SpyCas9 while the variant provides an observable nicked target, (c) A relaxed PAM requirement comparable to any of LFCA, LBCA or LSCA, (d) The ability to cleave single-stranded DNA, and (e) The ability to use an sgRNA in which the targeting sequence binds to a tracrRNA component selectable from the tracrRNA components of Cas9 gRNAs used by a plurality of existing bacterial species A Cas nuclease retaining one or several of.

2. The Cas nuclease according to Claim 1, wherein the Cas nuclease is an LFCA nuclease having SEQ ID NO: 1 or a variant thereof having at least 60% sequence identity with SEQ ID NO:

1. **Claim 3**: The Cas nuclease according to claim 1, wherein the Cas nuclease is an LBCA nuclease having SEQ ID NO: 2, or a variant thereof having at least 75% sequence identity with SEQ ID NO:

2. **Claim 4**: The Cas nuclease according to claim 1, wherein the Cas nuclease is an LSCA nuclease having SEQ ID NO: 3, or a variant thereof having at least 80% sequence identity with SEQ ID NO:

3. **Claim 5**: The Cas nuclease according to claim 1, wherein the Cas nuclease is an LPCA nuclease having SEQ ID NO: 4, or a variant thereof having at least 85% sequence identity with SEQ ID NO:

4. **Claim 6**: The Cas nuclease according to claim 1, wherein the Cas nuclease is an LP DCA nuclease having SEQ ID NO: 5, or a variant thereof having at least 95% sequence identity with SEQ ID NO:

5. **Claim 7** The Cas nuclease according to claim 1, wherein the Cas nuclease is modified by mutagenesis of the catalytic site so as to retain only its nickase activity. **Claim 8** The Cas nuclease according to claim 1, which comprises one or several amino acid changes by substitution or deletion, whereby the endonuclease and / or nickase activity of the Cas nuclease is retained, and whereby the relaxed PAM specificity of LFCA Cas is retained. **Claim 9**: The Cas nuclease according to claim 1, which comprises one or several amino acid changes by one or several conservative substitutions, whereby the endonuclease and / or nickase activity of the Cas nuclease is retained, and whereby the relaxed PAM specificity of LFCA Cas is retained. **Claim 10** The Cas nuclease according to claim 1, wherein the Cas nuclease is modified by mutagenesis of the catalytic site so as to inactivate its nuclease activity. **Claim 11** The Cas nuclease according to claim 1, wherein the Cas nuclease is bound or fused to a non-nuclease effector for gene modification or regulation. **Claim 12** A nucleic acid encoding the Cas nuclease according to claim 1. **Claim 13** A combination product comprising (i) a vector comprising the Cas nuclease according to claim 1 or the nucleic acid according to claim 12, and (ii) a guide RNA or a vector expressing the guide RNA, wherein the guide RNA targets the Cas nuclease to a target DNA sequence.

14. A ribonucleoprotein complex comprising the Cas nuclease according to claim 1 and a guide RNA, wherein the guide RNA targets the Cas nuclease to a target DNA sequence.

15. A method for modifying or regulating a target nucleic acid sequence, comprising contacting the target nucleic acid sequence with (i) the Cas nuclease according to claim 1, and (ii) a guide RNA that targets the Cas nuclease to the target sequence, and further (a) the contacting is in vitro with an isolated target nucleic acid or in vivo within a cell, provided that the method is not a method for modifying human germline identity, or (b) the method is not a medical treatment method performed on a human or animal body whichever is the case.

16. The method according to claim 15, wherein the target nucleic acid sequence is a DNA sequence.

17. The method according to claim 15, wherein the target nucleic acid sequence is a target DNA sequence in ex vivo human or animal cells.

18. A pharmaceutical composition for use as a drug, comprising the combination product according to claim 13.

19. A pharmaceutical composition for preventing and / or treating a disease by modifying or regulating a target nucleic acid sequence, comprising the combination product according to claim 13.

20. The pharmaceutical composition for use according to claim 19, wherein the disease is a genetic disease.

21. The pharmaceutical composition for use according to claim 19, wherein the combination product further comprises a nucleic acid molecule encoding a transgene of interest.

22. A phylogenetic ancestor reconstruction method for obtaining a functional monomeric effector Cas protein nuclease, comprising (a) providing a phylogenetic tree from sequence analysis of a population of Cas sequences, comprising naturally occurring monomeric effector Cas nuclease sequences of the same taxonomic type and obtained from a plurality of existing species, and Step (b): a step of selecting an ancestral variant sequence by tracing the evolutionary path from the phylogenetic tree, the step of determining a highly probable amino acid for each amino acid of the selected ancestral variant; Step (c): a step of producing the variant that can exhibit Cas protein endonuclease and / or nickase activity A method comprising the steps.

23. The method according to claim 22, wherein in step (a), the monomer effector Cas nuclease sequence is from two or more genera.

24. The method according to claim 22, wherein in step (a), the monomer effector Cas nuclease sequence is from two or more classes.

25. The method according to claim 22, wherein step (b) includes step (i) of editing the sequence of an ancestral variant, each of which forms a part of a sequence of the same genus and is an ancestral variant for sequences of a plurality of species.

26. The method according to claim 25, further including step (ii) of editing one or more ancestral variant sequences assigned as an ancestral genus and / or one or more ancestral variant sequences assigned as an ancestral class that can trace back to the sequences of the starting species of a plurality of genera, using the sequence achieved in (i).

27. The method according to claim 26, further including step (iii) of editing at least one inter-class ancestral sequence that can trace back to the starting species of two or more classes.

28. The method according to claim 22, wherein the selected ancestral variant sequence is an ancestral variant of the Cas9 sequence of an existing bacterial species.

29. The starting population of the Cas9 sequence includes a plurality of bacterial Cas9 sequences selected from two or more of Streptococcus, Enterococcus, Listeria, Clostridium, Pelagirhabdus, Halolactibacillus, Floricoccus, Vagococcus, Urinococcus, Vagococcus, Dorea, Lachnococcus, Ruminospira, Anaerostipes, Olsenella, and Bifidobacterium. The method according to claim 28.

30. The starting population of sequences is the method according to claim 29, spanning two or more bacterial classes, a plurality of sequences from different species of the genus Enterococcus, a plurality of sequences from different species of Listeria, and a plurality of sequences from species of the genus Clostridium.

31. The method according to claim 30, wherein the two or more bacterial classes are Cas9 sequences from both the Bacillus class and the Clostridium class of bacteria.

32. The method according to claim 30, wherein the starting population of sequences is supplemented with a Cas9 sequence of Actinomycetes.

33. The method according to claim 22, wherein the selected ancestral variant sequence is an inter-class ancestral variant sequence that can be traced back to starting species of two or more classes.

34. The method according to claim 22, wherein the selected ancestral variant is determined to be able to exhibit endonuclease double-stranded DNA cleavage, and is further converted into either a nickase only or a deadCas having no nuclease activity, and / or is bound to a non-nuclease effector.

35. An ancestral variant of the Cas9 sequence has the following characteristics: (a) A linearized DNA plasmid target:nicked DNA plasmid template ratio of at least 2.3:1 to at least 1:4 under the condition that a ratio of at least 4:1 of linearized DNA plasmid target:nicked DNA plasmid template is obtained by SpyCas9, (b) A relaxed PAM requirement comparable to any of LFCA, LBCA, or LSCA, (c) The ability to cleave single-stranded DNA, (d) The ability to use an sgRNA in which the targeting sequence binds to a tracrRNA component selectable from the tracrRNA components of Cas9 gRNAs used by a plurality of existing bacterial species The method according to claim 28, having one or several of them.

36. The method according to claim 35, wherein the selected ancestral variant is further converted into a variant that is either a nickase only or a deadCas having no nuclease activity, and / or provides binding to a non-nuclease effector.