Compositions and methods for targeting C9orf72

By using an improved CRISPR system to target and edit the C9orf72 gene, the genetic regulation challenges of ALS and FTD have been solved, reducing the accumulation of toxic proteins and providing therapeutic potential for ALS and FTD.

JP7847853B2Active Publication Date: 2026-04-20SCRIBE THERAPEUTICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SCRIBE THERAPEUTICS INC
Filing Date
2021-03-17
Publication Date
2026-04-20

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively treat ALS and FTD caused by C9orf72 gene mutations. There is a lack of effective genetic engineering methods to regulate the expression of C9orf72 protein and RNA accumulation, leading to disease progression.

Method used

Using an improved type 2 CRISPR protein and guide nucleic acid (gNA) system, the C9orf72 gene is targeted for editing to achieve gene knockdown or deletion, reducing or eliminating the accumulation of toxic proteins caused by abnormal repetitive sequences, and gene modification is performed using a viral vector or virus-like particle delivery system.

Benefits of technology

It can effectively reduce or eliminate the abnormal expression of C9orf72 gene products and RNA accumulation, reduce disease progression, and provide therapeutic potential for ALS and FTD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007847853000075
    Figure 0007847853000075
  • Figure 0007847853000076
    Figure 0007847853000076
  • Figure 0007847853000077
    Figure 0007847853000077
Patent Text Reader

Abstract

Provided herein is a Class 2 Type V system comprising a nuclease, a guide nucleic acid (gNA), and optionally a donor template nucleic acid, that is useful for modifying the C9orf72 gene. The system is also useful for introduction into cells, e.g., eukaryotic cells, that have a mutation or duplication in the C9orf72 gene. Also provided is a method for using such a system to modify cells that have such a mutation or duplication.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 991,403, filed on 18 March 2020, the contents of which are incorporated herein by reference in their entirety.

[0002] Inclusion by referencing the sequence list The contents of the text files submitted electronically with this specification are incorporated herein by reference in their entirety: a computer-readable copy of the sequence listing (filename: SCRB-025-01WO_SeqList_ST25.txt, date recorded: March 12, 2021, file size: 5.61 megabytes). [Background technology]

[0003] Amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD) are progressive neurological diseases that can have devastating consequences. ALS is a fatal neurodegenerative disease clinically characterized by progressive paralysis that typically leads to death from respiratory failure within 2-3 years of symptom onset, and it is the third most common neurodegenerative disease in Western countries (Rowland and Shneider, N.Engl.J.Med., 2001, 344, 1688-1700; Hirtz et al., Neurology, 2007, 68, 326-337). FTD is the second most common cause of presenile dementia and is characterized by progressive changes in personality, behavior, and language due to degeneration of the frontal and temporal lobes of the brain, with relatively preserved perception and memory (Graff-Radford, N and Woodruff, B. Frontotemporal dementia. Semin. Neurol. 27(1):48(2007)).

[0004] The chromosome 9 open reading frame 72 protein is a protein encoded by the C9orf72 gene (sometimes referred to as the C9orf72-SMCR8 complex subunit). Forms of disease associated with mutations or abnormalities in the C9orf72 gene include FTD and ALS. This protein is found in many regions of the brain, including the cytoplasm of neurons, as well as in presynaptic terminals. Specifically, the C9orf72 gene mutations associated with FTD and ALS occur at either intron 1 between two 5'-untranslated region (5'-UTR) exons, or in the promoter region of the C9orf72 gene, where a hexanucleotide repeat expansion of the nucleotide GGGGCC (DeJesus-Hernandez, M., et al. Expanded GGGGCC hexanucleotide repeat in noncoding region of C9ORF72 causes chromosome 9p-linked FTD and ALS. Neuron 72:245 (2011); Niblock, M., et al. Retention of hexanucleotide repeat-containing intron in C9orf72 mRNA: effects for the pathogenesis of ALS / FTD. Acta Neuropathologica Communications 4:18 (2016)). Therefore, the presence of the hexanucleotide repeat (HRS) expansion does not alter the coding sequence of the resulting C9orf72 protein. In healthy individuals, this hexanucleotide repeat is rare, typically less than 30, but in humans with the disease phenotype, the repeat units range from approximately 700 to 1600 (Mori K., et al. The C9orf72 GGGGCC repeat is translated into aggregating dipeptide-repeat proteins in FTLD / ALS. Science. 339:1335 (2013)).Repeat hexanucleotide expansion leads to the loss of one alternatively spliced ​​C9orf72 transcript, resulting in the formation and accumulation of insoluble dipeptide repeat protein aggregates via repeat-associated non-AUG initiated (RAN) translation, which mostly consist of poly(Gly-Ala) and, to a lesser extent, highly hydrophobic and potentially pathogenic to FTD-ALS patients, poly(Gly-Pro) and poly(Gly-Arg) dipeptide repeat proteins (DPRs) (Mori K., et al. 2013, Niblock, M., et al. 2016). Furthermore, three major disease mechanisms have been proposed: loss of function of the C9orf72 protein, toxic gain of function from C9orf72 repeat RNA (by accumulation of RNA transcripts containing repeats and antisense GGCCCC RNA in the frontal cortex and spinal cord), or accumulation of DPR generated by repeat-associated non-ATG translation (Balendra R, Isaacs AM. C9orf72-mediated ALS and FTD: multiple pathways to disease. Nat Rev Neurol. 14:544 (2018)). The inheritance of C9orf72 mutations is autosomal dominant (Iyer et al. C9orf72, a protein associated with amyotrophic lateral sclerosis (ALS) is a guanine nucleotide exchange factor. Peer J 6:e5815 (2018)).

[0005] The emergence of CRISPR / Cas systems and the programmable nature of these minimal systems have facilitated their use as versatile technologies in genome engineering and genomics. However, efforts to modify C9orf72-related diseases such as FTD and ALS through genetic engineering have received limited attention. Therefore, there is a need for compositions and methods to modulate C9orf72 in subjects with C9orf72-related diseases. To address this need, compositions and methods for targeting the C9orf72 gene are provided herein. [Overview of the project]

[0006] This disclosure provides a composition of modified class 2 type V CRISPR protein and guide nucleic acid used for editing the target nucleic acid sequence of the chromosome 9 open reading frame 72 (C9orf72) gene. The class 2 type V CRISPR protein and guide nucleic acid are modified for passive entry into target cells. The class 2 type V CRISPR protein and guide nucleic acid are useful in various methods for target nucleic acid modification for C9orf72-related diseases, and such methods are also provided.

[0007] In one embodiment, the disclosure relates to a CasX:guide nucleic acid system (CasX:gNA system) and a method used to modify a target nucleic acid, including the C9orf72 gene, which has one or more mutations or contains a hexanucleotide repeat expansion (HRS) in cells. In some embodiments of the disclosure, the CasX:gNA system is useful for knocking down or knocking out the C9orf72 gene, which has one or more mutations or contains a hexanucleotide repeat expansion (HRS), in order to reduce or eliminate DPR in subjects having C9orf72 gene product expression, accumulation of RNA from HRS, and / or C9orf72-related disease. In other embodiments, the CasX:gNA system is useful for modifying the C9orf72 gene, which contains HRS.

[0008] In some embodiments of the system, gNA is gRNA, gDNA, or a chimera of RNA and DNA, and may be a single-molecule gNA or a bimolecule gNA. In other embodiments, the CasX:gNA system gNA has a targeting sequence that is complementary to a target nucleic acid sequence containing a region within the C9orf72 gene. In some embodiments, the targeting sequence of gNA is selected from the group consisting of sequence numbers 309-343, 363-2100, 2295-2185, or sequences having at least about 65%, at least about 75%, at least about 85%, or at least about 95% identity thereto. gNA may contain a targeting sequence containing 14-30 consecutive nucleotides. In some embodiments, the targeting sequence of gNA consists of 21 nucleotides. In other embodiments, the targeting sequence of gNA consists of 20 nucleotides. In some embodiments, the targeting sequence consists of 19 nucleotides, and the targeting sequence for gNA has a sequence selected from the group consisting of SEQ ID NOs. 309-343, 363-2100, and 2295-2185, with a single nucleotide removed from the 3' end of the sequence. In other embodiments, the targeting sequence consists of 18 nucleotides, having a sequence selected from the group consisting of SEQ ID NOs. 309-343, 363-2100, and 2295-2185, with two nucleotides removed from the 3' end of the sequence. In other embodiments, the targeting sequence consists of 17 nucleotides, having a sequence selected from the group consisting of SEQ ID NOs. 309-343, 363-2100, and 2295-2185, with three nucleotides removed from the 3' end of the sequence. In another embodiment, the targeting sequence consists of 16 nucleotides having a sequence selected from the group consisting of SEQ ID NOs: 309-343, 363-2100, and 2295-2185, with four nucleotides removed from the 3' end of the sequence. In yet another embodiment, the targeting sequence consists of 15 nucleotides having a sequence selected from the group consisting of SEQ ID NOs: 309-343, 363-2100, and 2295-2185, with five nucleotides removed from the 3' end of the sequence.

[0009] In some embodiments of the system, the gNA has a scaffold containing a sequence selected from the group consisting of sequence numbers 4-16 and 2101-2294, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.

[0010] In some embodiments of the system, the class 2 V-type CRISPR protein includes a reference CasX protein having any one sequence from sequence numbers 1-3, a CasX variant protein having a sequence selected from the group consisting of sequence numbers 49-150, 233-235, 238-239, 240-242, and 272-281, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the reference CasX protein. In these embodiments, the CasX variant exhibits one or more improved properties with respect to the reference CasX protein. In some embodiments, the CasX protein has binding affinity to protospacer adjacent motif (PAM) sequences selected from the group consisting of TTC, ATC, GTC, and CTC. In some embodiments, the CasX protein has a binding affinity to PAM sequences that is at least 1.5 times higher than the binding affinity of any one of the CasX proteins of SEQ ID NOs. 1 to 3 to PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC.

[0011] In some embodiments of the system, the CasX molecule and the gNA molecule associate together in a ribonucleoprotein complex (RNP). In certain embodiments, if one of the PAM sequences TTC, ATC, GTC, or CTC is located at a single nucleotide at 5' relative to a non-target chain sequence identical to the target sequence of gNA, then the RNP containing the CasX variant and the gNA variant exhibits higher editing efficiency and / or binding of the target sequence on the target DNA in the cell assay system compared to the editing efficiency and / or binding of the RNP containing a reference CasX protein and reference gNA in an equivalent assay system.

[0012] In some embodiments, the system further comprises a donor template comprising a nucleic acid containing at least a portion of the C9orf72 gene, the C9orf72 gene portion being selected from the group consisting of C9orf72 exons, C9orf72 introns, C9orf72 intron-exon junctions, C9orf72 regulatory elements, or combinations thereof, and the donor template is used to knock down or knock out the C9orf72 gene or to correct mutations in the C9orf72 gene. In some embodiments, the donor template comprises hexanucleotide repeats of the GGGGCC sequence, the number of repeats ranging from 10 to about 30, and is used to replace hexanucleotide repeat expansions of the mutant C9orf72 gene. In some cases, the donor sequence is a single-stranded DNA template or a single-stranded RNA template. In other cases, the donor template is a double-stranded DNA template.

[0013] In other embodiments, this disclosure relates to nucleic acids encoding any of the systems described herein, as well as vectors containing nucleic acids. In some embodiments, the vector is selected from the group consisting of retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, herpes simplex virus (HSV) vectors, plasmids, minicircles, nanoplasmides, and RNA vectors. In other embodiments, the vector is an RNP of CasX and gNA from any of the embodiments described herein, and optionally, a virus-like particle (VLP) containing a donor template nucleic acid and a targeted moiety such as a virus-derived glycoprotein.

[0014] In other embodiments, the Disclosure provides a method for modifying a C9orf72 target nucleic acid sequence in a cell population, the method comprising introducing into the cells a) a CasX:gNA system of any embodiment disclosed herein, b) a nucleic acid of any embodiment disclosed herein, c) a vector of any embodiment disclosed herein, d) a VLP of any embodiment disclosed herein, or e) a combination thereof, wherein the C9orf72 target nucleic acid sequence in the cells targeted by the first gNA is modified by the CasX protein to introduce single-strand or double-strand breaks in the target nucleic acid sequence. In some embodiments of the Method, the Method further comprises a second gNA or a nucleic acid encoding the second gNA, the second gNA having a targeting sequence complementary to a different portion of the target nucleic acid sequence. In some embodiments of the Method, the modification comprises introducing one or more nucleotide insertions, deletions, substitutions, duplications, or inversions in the target nucleic acid sequence compared to the wild-type sequence. In some embodiments, the method further comprises contacting a target nucleic acid with a donor template nucleic acid from any of the embodiments disclosed herein. In some embodiments, the target C9orf72 gene for modification comprises more than 30, more than 100, more than 500, more than 700, more than 1000, or more than 1600 copies of the hexanucleotide repeat sequence GGGGCC. In some embodiments of the method, the donor template comprises a nucleic acid comprising at least a portion of the C9orf72 gene for modifying a mutation in the C9orf72 gene (by knock-in), or comprises a sequence comprising a mutation or heterologous sequence for knocking down or knocking out a mutant C9orf72, thereby reducing the expression of HRS or DPR by a cell population by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% compared to cells that are not modified. In some cases, modification of the target nucleic acid sequence occurs in vivo. In some embodiments, the cells are eukaryotic cells selected from the group consisting of rodent cells, mouse cells, rat cells, primate cells, and non-human primate cells.In some embodiments, the cells are human cells. In some embodiments, the cells are selected from the group consisting of Purkinje cells, prefrontal cortical neurons, motor cortical neurons, hippocampal neurons, cerebellar neurons, upper motor neurons, spinal neurons, spinal motor neurons, glial cells, and astrocytes.

[0015] In other embodiments, the Disclosure provides a method for modifying a C9orf72 target nucleic acid in a target cell population, wherein the target cells are contacted using a vector encoding a CasX protein and one or more gNAs containing a targeting sequence complementary to the C9orf72 gene, and optionally further comprising a donor template. In some cases, the vector is an adeno-associated virus (AAV) vector selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44.9, AAV-Rh74, or AAVRh10. In other cases, the vector is a lentiviral vector. In other embodiments, the Disclosure provides a method for contacting target cells using a vector, wherein the vector is a virus-like particle (VLP) containing an RNP of CasX and gNA from any of the embodiments described herein, and optionally a donor template nucleic acid. In some embodiments of the method, the vector is administered to a subject in a therapeutically effective dose. The subject may be a mouse, rat, pig, non-human primate, or human. The dose can be administered by a route of administration selected from the group consisting of subcutaneous, intradermal, intraneuronal, intranodal, intramedullary, intramuscular, intralumbar, intramedullary, subarachnoid, intraventricular, intrasacral, intravenous, lymphatic, or intraperitoneal routes, and the method of administration may be injection, infusion, or transplantation.

[0016] In other embodiments, the Disclosure provides a method for treating a C9orf72-related disorder in a subject, comprising modifying a gene encoding the C9orf72 gene in the cells of the subject, the modification comprising contacting the cells with a) a CasX:gNA system of any embodiment disclosed herein, b) a nucleic acid of any embodiment disclosed herein, c) a vector of any embodiment disclosed herein, d) a VLP of any embodiment disclosed herein, or e) a combination thereof, wherein the C9orf72 gene in the cells targeted by the first gNA is modified with the CasX protein. In some embodiments, the subject is selected from the group consisting of mice, rats, pigs, non-human primates, and humans. In some embodiments, the C9orf72-related disorder is ALS or FTD. In some cases, the method for treating a subject with a C9orf72-related disorder results in improvement of at least one clinically relevant parameter. In other cases, the method for treating a subject with a C9orf72-related disorder results in improvement of at least two clinically relevant parameters.

[0017] In other embodiments, the Disclosure provides compositions for use in methods for treating C9orf72-related disorders in subjects. In some embodiments, the method comprises modifying a gene encoding the C9orf72 gene in cells of a subject, the modification comprising contacting the cells with a composition selected from a) a CasX:gNA system of any embodiment disclosed herein, b) a nucleic acid of any embodiment disclosed herein, c) a vector of any embodiment disclosed herein, d) a VLP of any embodiment disclosed herein, or e) a combination thereof, wherein the C9orf72 gene in cells targeted by the first gNA is modified with the CasX protein. In some embodiments, the subject is selected from the group consisting of mice, rats, pigs, non-human primates, and humans. In some embodiments, the C9orf72-related disorder is ALS or FTD. In some cases, the method for treating a subject with a C9orf72-related disorder results in improvement of at least one clinically relevant parameter. In other cases, the method for treating a subject with a C9orf72-related disorder results in improvement of at least two clinically relevant parameters.

[0018] Built-in by reference All publications, patents, and patent applications referenced herein are incorporated herein by reference to the same extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference. The contents of WO2020 / 247882, filed 5 June 2020, disclosing CasX and gNA variants, as well as U.S. Provisional Applications 63 / 121,196, filed 3 December 2020, and 63 / 162,346, filed 17 March 2021, are incorporated herein by reference in their entirety. [Brief explanation of the drawing]

[0019] Novel features of this disclosure are specifically described in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following embodiments for carrying out the invention, illustrating exemplary embodiments in which the principles of this disclosure are utilized, and the accompanying drawings.

[0020] [Figure 1] As described in Example 1, the SDS-PAGE gel of the StX2 purified fraction visualized by colloidal Coomassi staining is shown. [Figure 2] The chromatogram from the CasX StX2 size exclusion chromatography assay using Superdex 200 16 / 600pg gel filtration, as described in Example 1, is shown. [Figure 3] As described in Example 1, the SDS-PAGE gel of the CasX StX2 purified fraction visualized by colloidal Coomassi staining is shown. [Figure 4] This is a schematic diagram showing the organization of the components of the pSTX34 plasmid used to assemble the CasX construct, as described in Example 2. [Figure 5] This is a schematic diagram illustrating the steps for generating a pSTX34 plasmid containing the CasX 119 variant, as described in Example 2. [Figure 6] As described in Example 2, the SDS-PAGE gel of the purified sample visualized on Bio-Rad Stain-Free® gel is shown. [Figure 7] The chromatogram of Superdex 200 16 / 600pg gel filtration is shown, as described in Example 2. [Figure 8] As described in Example 2, the SDS-PAGE gel of the gel-filtered sample stained with colloidal Coomassie is shown. [Figure 9]This graph shows the results of an assay for quantifying the active fraction of RNP formed by sgRNA174 and CasX variants 119, 457, 488, and 491, as described in Example 13. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time points. The mean and standard deviation of three independent replicas per time point are shown. The two-phase fit of the combined replicas is shown. "2" refers to the reference CasX protein of Sequence ID No. 2. [Figure 10] As described in Example 13, this shows the quantification of the active fraction of RNP formed by CasX2 (reference CasX protein of SEQ ID NO: 2) and modified sgRNA. Equimolar amounts of RNP and target were co-incubated, and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicas at each time point are shown. The two-phase fit of the combined replicas is shown. [Figure 11] As described in Example 13, the quantification of the active fraction of RNP formed by CasX491 and modified sgRNA under guide restriction conditions is shown. Equimolar amounts of RNP and target were co-incubated, and the amount of cleaved target was determined at the indicated time point. Two-phase fit of the data is shown. [Figure 12] As described in Example 13, the quantification of the cleavage rate of RNPs formed by sgRNA174 and CasX variants is shown. Target DNA was incubated with a 20-fold excess of the indicated RNP, and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicas per time point are shown, except for 488 and 491 which show single replicas. Monophase fit of combined replicas is shown. [Figure 13] As described in Example 13, the quantification of the cleavage rate of RNPs formed by CasX2 and sgRNA variants is shown. Target DNA was incubated with a 20-fold excess of the indicated RNP, and the amount of target cleaved was determined at the indicated time points. The mean and standard deviation of three independent replicas per time point are shown. The uniphase fit of the combined replicas is shown. [Figure 14]As described in Example 13, we demonstrate the quantification of the initial cleavage rate of RNPs formed by CasX2 and sgRNA variants. The initial cleavage rate was determined by fitting the first two time points of previous cleavage experiments to a linear model. [Figure 15] As described in Example 13, the quantification of the cleavage rate of RNPs formed by CasX491 and sgRNA variants is shown. Target DNA was incubated at 10°C with a 20-fold excess amount of the indicated RNP, and the amount of target cleaved was determined at the indicated time points. Single-phase fit at those time points is shown. [Figure 16] [Figures 16A-16D] These figures show the quantification of the cleavage rate of CasX variants on NTC PAM, as described in Example 14. Target DNA substrates with the same spacer and indicated PAM sequences were incubated at 37°C with a 20-fold excess of indicated RNP, and the amount of target cleaved was determined at the indicated time point. Single-phase fit of a single replica is shown. Figure 16A shows the results for sequences with TTC PAM. Figure 16B shows the results for sequences with CTC PAM. Figure 16C shows the results for sequences with GTC PAM. Figure 16D shows the results for sequences with ATC PAM. [Figure 17] This is a schematic diagram showing an example of a CasX protein and scaffold DNA sequence for packaging in adeno-associated virus (AAV), as described in Example 23. The DNA segment between the AAV reverse-ended repeat (ITR), which consists of the DNA encoding CasX and its promoter, and the DNA encoding the scaffold and its promoter, is packaged within the AAV capsid during AAV generation. [Figure 18]This report presents the results of an editing assay comparing gRNA scaffolds 229–237 and scaffold 174 in mouse neural progenitor cells (mNPCs) isolated from Ai9-tdtomato transgenic mice. Cells were nucleofected with indicated doses of the p59 plasmid encoding mRHO-targeting CasX 491, scaffolds, and spacer 11.30 (5'AAGGGGCUCCGCACCACGCC3', SEQ ID NO: 361). Editing at the mRHO locus was evaluated by NGS 5 days after transfection, and it was shown that editing with constructs containing scaffolds 230, 231, 234, and 235 showed greater editing compared to constructs containing scaffold 174 at both doses. [Figure 19] The results of an editing assay comparing gRNA scaffolds 229–237 and 174 in mNPC cells are presented. Cells were nucleofected with the indicated doses of CasX 491, scaffolds, and a p59 plasmid encoding spacer 12.7 (5'CUGCAUUCUAGUUGUGGUUU3', SEQ ID NO: 362), targeting repeat elements that prevent the expression of the tdtomato fluorescent protein. Editing was evaluated by FACS 5 days after transfection, and the fraction of tdtomato-positive cells was quantified. Cells nucleofected with scaffolds 231–235 showed approximately 35% higher editing at high doses and approximately 25% higher editing at low doses compared to constructs with scaffold 174. [Figure 20]This report presents the results of editing assays comparing CasX nucleases 2, 119, 491, 515, 527, 528, 529, 530, and 531 in a custom HEK293 cell line, PASS_V1.01. Cells were lipofected with 2 μg of p67 plasmid encoding the indicated CasX proteins. After 5 days, cell genomic DNA was extracted. PCR amplification and next-generation sequencing were performed to isolate and quantify the fraction of cells edited at custom-designed on-target editing sites. For each sample, editing was evaluated at target sites (individual dots) consisting of individual sites of the following PAM sequences: 48 TTC, 14 ATC, 22 CTC, and 11 GTC, and the editing percentage was normalized to the vehicle control. Cells lipofected with any of the nucleases showed higher mean editing at TTC PAM target sites (horizontal bars) than those with the wild-type nuclease CasX 2, except for CasX 528. The relative preferences of any given nuclease for four different PAM sequences are also represented by violin plots. Specifically, CasX nucleases 527, 528, and 529 exhibit substantially different PAM preferences than those of the wild-type nuclease CasX 2. [Figure 21]This report presents the results of an editing assay comparing improved CasX nuclease 491 with improved nucleases 532 and 533 in a custom HEK293 cell line, PASS_V1.01. Cells were lipofected in pairs with 2 μg of p67 plasmid encoding the indicated CasX protein and puromycin resistance gene, and grown under puromycin selectivity. After 3 days, cell genomic DNA was extracted. PCR amplification and next-generation sequencing were performed to isolate and quantify the fraction of cells edited at custom-designed on-target editing sites. For each sample, editing was evaluated at target sites consisting of individual sites of the following PAM sequences: 48 TTC, 14 ATC, 22 CTC, and 11 GTC, and fraction editing was normalized to the vehicle control. Cells lipofected with CasX 532 or 533 showed higher mean editing than Cas 491 at each of the PAM sequences, except for CasX 533 at the TTC PAM target site. The error bars represent the standard error of the mean value for n=2 biological samples. [Figure 22] This is a schematic diagram of a portion of the 5' region of the C9orf72 locus. The upper diagram shows the relative positions of exons 1a and 1b, which are adjacent to a hexanucleotide repeat element (HRE), while the open squares indicate downstream exons. The lower diagram shows the region of the locus that is (complementarily) targeted by the targeting segment (spacer) of the guide RNA in Table 15, as described in Example 18. [Figure 23] This graph shows the results of a single-cleavage experiment in which an edit is introduced to exon 1a using targeting sequence 164, as described in Example 18. The black deletion lines show all locations in the amplicon and the fraction of reads where a deletion is present at that location. The gray bars at the bottom of the graph show the locations of sgRNA binding sites. The quantification window shows the region used for quantification of deletions. The predicted cleavage location is the location of the double-strand break induced by CasX. The deletion lines show the percentage and extent of gene deletions generated by the delivered single guide, resulting in an overall deletion efficiency of 65.4%. The data represent the results observed with single cleavage (Table 15). [Figure 24]This graph shows the results of a double-cleavage experiment using target sequences (spacers) 138 and 151, which are target sequences adjacent to hexanucleotide repeat elements (HREs, sometimes referred to herein as hexanucleotide repeat sequence expansions or HRS) at positions 193–248 of the reference amplicon, as described in Example 18. The black deletion lines show all positions on the amplicon and the fraction of reads where deletions are present at those positions. The gray bars at the bottom of the graph show the locations of sgRNA binding sites. The quantification window shows the region used for quantification of deletions. The predicted cleavage location is the location of the CasX-induced double-strand break. In this experiment, the overall deletion efficiency was 45.4%, which is representative of the results observed with double-cleavage (Table 16), and supports the idea that HREs can be deleted using a double-cleavage design under experimental conditions. [Figure 25] As described in Example 26, these are graphs of experiments testing the effect of spacer length on the ability to edit target nucleic acids in Jurkat cells. The results show that shorter spacers of 18 or 19 indicate increased activity in ex vivo editing with RNP compared to a 20-base spacer. [Modes for carrying out the invention]

[0021] While exemplary embodiments are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Numerous variations, modifications, and substitutions will be made today without departing from the claimed invention as described herein. It should be understood that various alternatives to the embodiments described herein can be used when carrying out the embodiments of this disclosure. The claims define the scope of the invention, and methods and structures within the scope of these claims, as well as their equivalents, are intended to be included within the claims.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art to the extent of this invention. Methods and materials similar or equivalent to those described herein may be used in carrying out or testing these embodiments, but preferred methods and materials are described below. In case of any conflict, this specification, including definitions, shall prevail. In addition, the materials, methods, and examples are illustrative and not intended to limit. Numerous modifications, changes, and substitutions will be made today by those skilled in the art without departing from the invention.

[0023] definition As used interchangeably herein, the terms “polynucleotide” and “nucleic acid” refer to polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Therefore, the terms “polynucleotide” and “nucleic acid” encompass single-stranded DNA, double-stranded DNA, multi-stranded DNA, single-stranded RNA, double-stranded RNA, multi-stranded RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, unnatural, or derivatized nucleotide bases.

[0024] The terms "hybridizable" and "complementary" are used interchangeably and mean that a nucleic acid (e.g., RNA, DNA) contains a sequence of nucleotides that allows it to "anneal" or "hybridize" to another nucleic acid in a sequence-specific antiparallel manner (i.e., the nucleic acid specifically binds to a complementary nucleic acid) in a non-covalent manner, i.e., by forming Watson-Crick base pairs and / or G / U base pairs, under appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. It is understood that the sequence of a polynucleotide does not need to be 100% complementary to its target nucleic acid sequence in order to be specifically hybridizable; a polynucleotide sequence may have at least about 70%, at least about 80%, at least about 90%, or at least about 95% sequence identity with the target nucleic acid sequence and still be able to hybridize to the target nucleic acid sequence. Furthermore, a polynucleotide can hybridize across one or more segments such that intervening or adjacent segments do not participate in the hybridization event (e.g., loop or hairpin structures, "bulges," "bubbles," etc.).

[0025] For the purposes of this disclosure, “gene” includes DNA regions that encode a gene product (e.g., protein, RNA), and all DNA regions that regulate the production of the gene product, whether such regulatory sequences are adjacent to the coding and / or transcription sequences. Thus, a gene may include, but is not limited to, regulatory element sequences such as promoter sequences, terminators, ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix binding sites, and locus regulatory regions. The coding sequence encodes the gene product at transcription or at transcription and translation, and the coding sequence in this disclosure may include fragments and does not need to include a full-length open reading frame. A gene may include both the transcribed strand and a complementary strand containing an anticodon.

[0026] The term "downstream" refers to a nucleotide sequence located 3' relative to a reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence is related to the sequence following the transcription start site. For example, the translation start codon of a gene is located downstream of the transcription start site.

[0027] The term "upstream" refers to a nucleotide sequence located 5' of the reference nucleotide sequence. In certain embodiments, the upstream nucleotide sequence is related to a sequence located 5' of the coding region or transcription start site. For example, most promoters are located upstream of the transcription start site.

[0028] The term “regulatory element” is used herein interchangeably with the term “regulatory sequence” and is intended to include promoters, enhancers, and other expression regulatory elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Exemplary regulatory elements include, but are not limited to, single transcripts, metallothioneins, transcription enhancer elements, transcription termination signals, polyadenylation sequences, sequences for optimizing translation initiation, and transcription promoters that enable translation of multiple genes from translation termination sequences, such as CMV, CMV+intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EF1α), MMLV-ltr, internal ribosome entry sites (IRES), or P2A peptides. In systems used for exon skipping, regulatory elements include exon splicing enhancers. It will be understood that the selection of appropriate regulatory elements depends on the encoded component being expressed (e.g., protein or RNA), whether the nucleic acid requires a different polymerase, or whether it contains multiple components not intended to be expressed as a fusion protein.

[0029] The term “promoter” refers to a DNA sequence that includes an RNA polymerase binding site, a transcription start site, a TATA box, and / or a B recognition element, and that supports or promotes the transcription and expression of an associated transcriptable polynucleotide sequence and / or gene (or transgene). Promoters can be produced synthetically or derived from known or naturally occurring promoter sequences or other promoter sequences. Promoters can be located proximal or distal to the gene being transcribed. Promoters can also include chimeric promoters that include a combination of two or more heterologous sequences to confer a particular characteristic. Promoters of this disclosure may include variants of promoter sequences that are similar in composition to, but not identical to, known or other promoter sequences provided herein. Promoters can be classified according to criteria relating to the expression pattern of associated codes or transcriptable sequences or genes operably linked to the promoter, such as constitutive, developmental, tissue-specific, or inducible.

[0030] The term "enhancer" refers to a regulatory DNA sequence that, when bound to a specific protein called a transcription factor, modulates the expression of an associated gene. Enhancers can be located in the intron of a gene, or at the 5' or 3' of the gene's coding sequence. Enhancers can be located proximal to the gene (i.e., within tens or hundreds of base pairs (bp) of the promoter) or distal to the gene (i.e., thousands, hundreds of thousands, or even millions of bp away from the promoter). A single gene may be regulated by two or more enhancers, all of which are assumed to be within the scope of this disclosure.

[0031] As used herein, “recombinant” means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps that result in a construct having a structural coding or non-coding sequence that is distinguishable from endogenous nucleic acids found in nature. Generally, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide synthetic nucleic acids that can be contained in cells or expressed from recombinant transcription units contained in cell-free transcription and translation systems. Such sequences can be provided in the form of an open reading frame that is not interrupted by internal non-coding sequences or introns, which are typically present in eukaryotic genes. Genomic DNA containing the relevant sequences can also be used to form recombinant genes or transcription units. The non-coding DNA sequence may be located at 5' or 3' of the open reading frame, and such sequence may not interfere with the manipulation or expression of the coding region, but may actually play a role in regulating the production of the desired product by various mechanisms (see “enhancer” and “promoter” above).

[0032] The terms “recombinant polynucleotides” or “recombinant nucleic acids” refer to those that do not exist naturally, for example, those created by the artificial combination of two distinctly separated sequence segments through human intervention. This artificial combination is often achieved by either chemical synthesis or artificial manipulation of isolated nucleic acid segments, such as genetic engineering techniques. Such artificial combinations are typically made to replace codons with redundant codons encoding the same or conserved amino acids, typically while introducing or removing sequence recognition sites. Alternatively, this artificial combination is made to combine nucleic acid segments having a desired function to produce a combination of a desired function. This artificial combination is often achieved by either chemical synthesis or artificial manipulation of isolated nucleic acid segments, such as genetic engineering techniques.

[0033] Similarly, the terms “recombinant polypeptide” or “recombinant protein” refer to polypeptides or proteins that do not exist in nature, created, for example, by the artificial combination of two otherwise separated segments of an amino acid sequence through human intervention. Therefore, for example, a protein containing a heterologous amino acid sequence is a recombinant.

[0034] As used herein, the term “contact” means to establish a physical bond between two or more entities. For example, contacting a target nucleic acid sequence with a guide nucleic acid means that the target nucleic acid sequence and the guide nucleic acid are constructed to share a physical bond, for example, they can hybridize if their sequences share sequence similarity.

[0035] "Dissociation constant" or "K" d The terms "L" and "P" are used interchangeably and refer to the affinity between the ligand "L" and the protein "P," i.e., how closely the ligand binds to a particular protein. This is represented by formula K. d It can be calculated using =[L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of the protein, ligand, and complex, respectively.

[0036] This disclosure provides compositions and methods useful for editing target nucleic acid sequences. As used herein, “editing” is used interchangeably with “modifying” and includes, but is not limited to, cleavage, nicking, deletion, knock-in, knock-out, and the like.

[0037] The term "knockout" refers to the removal or expression of a gene. For example, a gene may be knocked out by either the deletion or addition of a nucleotide sequence that leads to the disruption of the leading frame. As another example, a gene may be knocked out by replacing a portion of the gene with an unrelated sequence. As used herein, the term "knockdown" refers to the reduction of the expression of a gene or its gene product. As a result of gene knockdown, the activity or function of a protein may be attenuated, or the protein level may be reduced or eliminated.

[0038] As used herein, “Homologous-Directed Repair” (HDR) refers to a form of DNA repair that occurs during the repair of double-strand breaks in cells. This process requires homology of nucleotide sequences and uses a donor template to repair or knock out target DNA, resulting in the transfer of genetic information from the donor (e.g., a donor template) to the target, resulting in the desired transgene. If the donor template differs from the target DNA sequence, and some or all of the donor template sequence is incorporated into the target DNA, homologous-directed repair may result in alteration of the target nucleic acid sequence by insertion, deletion, or mutation.

[0039] As used herein, “non-homologous end joining” (NHEJ) refers to the repair of a DNA double-strand break by direct ligation of the break ends toward each other, without requiring a homologous template (as opposed to homology-directed repair, which requires homologous sequences to induce repair). NHEJ often results in the loss (deletion) of a nucleotide sequence near the double-strand break site.

[0040] As used herein, “microhomology-mediated end joining” (MMEJ) refers to a mutagenic DSB repair mechanism that always associates with a deletion adjacent to a cleavage site without requiring a homologous template (in contrast to homology-directed repair, which requires a homologous sequence to induce repair). MMEJs often result in a loss (deletion) of nucleotide sequence near a double-strand break site.

[0041] A polynucleotide or polypeptide (or protein) has a certain percentage of "sequence similarity" or "sequence identity" with respect to another polynucleotide or polypeptide, meaning that when aligned, that percentage of bases or amino acids are the same and, when the two sequences are compared, they are in the same relative positions. Sequence similarity (conversely referred to as similarity percentage, identity percentage, or homology) can be determined in several different ways. Sequences can be aligned using methods and computer programs known in the art, including BLAST, which is available on the World Wide Web at ncbi.nlm.nih.gov / BLAST, to determine sequence similarity. The percentage of complementarity between specific stretches of nucleic acid sequences within a nucleic acid can be determined using any convenient method. Examples of methods include using the BLAST program (a basic local alignment search tool) and the PowerBLAST program (Altschul et al., J.Mol.Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), or using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.) which uses the Smith-Waterman algorithm (Adv.Appl.Math., 1981, 2, 482-489), for example, by using default settings.

[0042] The terms “polypeptide” and “protein” are used interchangeably herein and refer to polymeric forms of amino acids of any length that may include encoded and unencoded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having a modified peptide backbone. The terms include, but are not limited to, fusion proteins having heterologous amino acid sequences.

[0043] A "vector" or "expression vector" is a replicon such as a plasmid, phage, virus, virus-like particle, or cosmid, to which another DNA segment, i.e., an "insert," can bind, resulting in the replication or expression of the bound segment within a cell.

[0044] As used herein, the terms “naturally occurring,” “unmodified,” or “wild-type,” applied to nucleic acids, polypeptides, cells, or organisms, refer to nucleic acids, polypeptides, cells, or organisms as they are found in nature. Therefore, “wild-type” can refer to two or more naturally occurring variants of a nucleic acid, polypeptide, cell, or organism. With respect to genes, “wild-type” may also be used to refer to a naturally occurring gene variant that does not cause disease.

[0045] As used herein, “mutation” means an insertion, deletion, substitution, replication, or inversion of one or more amino acids or nucleotides compared to the wild-type or reference amino acid sequence or the wild-type or reference nucleotide sequence.

[0046] As used herein, the term “isolated” is intended to describe polynucleotides, polypeptides, or cells that exist in an environment different from the environment in which the cells naturally exist. Isolated genetically modified host cells may exist in a mixed population of genetically modified host cells.

[0047] As used herein, “host cell” means a eukaryotic cell, a prokaryotic cell, or a cell derived from a multicellular organism (e.g., a cell line) cultured as a single-cell entity, and these eukaryotic or prokaryotic cells include offspring of the original cell that has been genetically modified by the nucleic acid and used as a recipient of nucleic acid (e.g., an expression vector). It is understood that, due to natural, accidental, or intentional mutations, the morphology or genomic or total DNA complement of the single-cell offspring may not necessarily be completely identical to that of the original parent. A “recombinant host cell” (also referred to as a “genetically modified host cell”) is a host cell into which a different nucleic acid, such as an expression vector, has been introduced.

[0048] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues with similar side chains in proteins. For example, amino acids with aliphatic side chains include glycine, alanine, valine, leucine, and isoleucine; amino acids with aliphatic hydroxyl side chains include serine and threonine; amino acids with amide-containing side chains include asparagine and glutamine; amino acids with aromatic side chains include phenylalanine, tyrosine, and tryptophan; amino acids with basic side chains include lysine, arginine, and histidine; and amino acids with sulfur-containing side chains include cysteine ​​and methionine. Exemplary conservative amino acid substitutions are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.

[0049] As used herein, “treatment” or “to treat” means, interchangeably as used herein, an approach to obtain beneficial or desired outcomes, including but not limited to therapeutic and / or preventive benefits. Therapeutic benefit means the eradication or improvement of the underlying disorder or disease being treated. Therapeutic benefit may also be achieved by the eradication or improvement of one or more symptoms, or by improvement of one or more clinical parameters related to the underlying disease, to which improvement is observed in the subject, even though the subject may still have the underlying disease.

[0050] As used herein, the terms “therapeutic dose” and “therapeutic amount” refer to the amount of a drug or biological preparation, either alone or as part of a composition, that, when administered as a single dose or repeated doses to a subject such as a human or experimental animal, can produce some detectable and beneficial effect on any symptom, aspect, measured parameter, or characteristic of any disease or condition. Such effect does not necessarily have to be absolutely beneficial.

[0051] As used herein, “administer” means a method of giving a subject a certain dose of a compound (e.g., a composition of this disclosure) or a composition (e.g., a pharmaceutical composition).

[0052] The "subjects" are mammals. Mammals include, but are not limited to, livestock, non-human primates, humans, rabbits, mice, rats, and other rodents.

[0053] I. General Methods The implementation of this invention, unless otherwise specified, utilizes conventional techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, as described in Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999), Protein Methods (Bollag et al., John Wiley & Sons 1996), Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999), Viral Vectors (Kaplift & Loewy eds., Academic Press 1995), Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997), and Cell and Tissue Culture: Laboratory Procedures in This information can be found in standard textbooks such as Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), and their disclosures are incorporated herein by reference.

[0054] When a range of values ​​is provided, it should be understood that the endpoint is included, and values ​​between it are between the upper and lower limits of that range, up to one-tenth of the lower limit unit, unless the context explicitly indicates otherwise, and other listed values ​​and values ​​between them within that listed range are included. The upper and lower limits of these smaller ranges may independently be included within smaller ranges, and are also included, subject to any limits specifically excluded in the listed range. If a listed range includes one or both limits, it also includes ranges that exclude one or both of those included limits.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art to which the present invention pertains. All publications referenced herein are incorporated herein by reference to disclose and describe methods and / or materials in relation to those cited herein.

[0056] When used herein and in the appended claims, it should be noted that the singular forms "a," "an," and "the" refer to multiple subjects unless otherwise explicitly indicated by the context.

[0057] For clarity, it will be understood that certain features of the Disclosure described in the context of a separate embodiment may also be provided in combination in a single embodiment. In other instances, for brevity, various features of the Disclosure described in the context of a single embodiment may also be provided separately or in any preferred partial combination. All combinations of embodiments relating to the Disclosure are specifically encompassed by the Disclosure and are intended to be disclosed herein as if each and all combinations were individually and expressly disclosed. In addition, various embodiments and all partial combinations of their elements are also specifically encompassed by the Disclosure and are disclosed herein as if each and all such partial combinations were individually and expressly disclosed herein.

[0058] II. A system for gene editing of the C9orf72 gene In a first aspect, the Disclosure provides a system comprising a class 2 type V CRISPR nuclease protein and one or more guide nucleic acids (gNAs) for use in modifying the C9orf72 gene product, RNA from HRS transcription, and / or DPR (collectively referred to herein as “target nucleic acids,” including coding and non-coding regions) to reduce or eliminate the expression of the C9orf72 gene product, RNA from HRS transcription, and / or DPR (collectively referred herein as “target nucleic acids”).

[0059] The human C9orf72 gene (HGNC:28337) has the sequence MSTLCPPPSPAVAKTEIALSGKSPLLAATFAYWDNILGPRVRHIWAPKTEQVLLSDGEITFLANHTLNGEILRNAESGAIDVKFFVLSEKGVIIVSLIFDGNWNGDRSTYGLSIILPQTELSFYLPLHRVCVDRLTHIIRKGRIWMHKERQENVQKIILEGTERMEDQGQSIIPMLTGEVIPVMELLSSMKSHSVPEEIDIADTVLNDDDIGDSCHEGFLLNAISSHLQTCGCSVVVGSSAEK The C9orf72 gene encodes a protein (Q01453) containing VNKIVRTLCLFLTPAERKCSRLCEAESSFKYESGLFVQGLLKDSTGSFVLPFRQVMYAPYPTTHIDVDVNTVKQMPPCHEHIYNQRRYMRSELTAFWRATSEEDMAQDTIIYTDESFTPDLNIFQDVLHRDTLVKAFLDQVFQLKPGLSLRSTFLAQFLLVLHRKALTLIKYIEDDTQKGKKPFKSLRNLKIDLDLTAEGDLNIIMALAEKIKPGLHSFIFGRPFYTSVQERDVLMTF (Sequence ID 227). The C9orf72 gene is defined as a sequence spanning chr9:27,546,546~27,573,866 on chromosome 9 of the human genome (Homo sapiens Updated Annotation Release 109.20191205,GRCh38.p13(NCBI)). The human C9orf72 gene is partially described in the NCBI database (ncbi.nlm.nih.gov) as reference sequence NC_000009.12, which is incorporated herein by reference. The C9orf72 locus contains 12 exons, including two alternative non-coding first exons (exons 1a and 1b) (DeJesus-Hernandez, M., et al. 2011). In the case of hexanucleotide repeats, the translated DPR protein contains poly(Gly-Ala) and, to less extent, poly(Gly-Pro) and poly(Gly-Arg).The shorter isoform b (NP_659442.2) has the sequence MSTLCPPPSPAVAKTEIALSGKSPLLAATFAYWDNILGPRVRHIWAPKTEQVLLSDGEITFLANHTLNGEILRNAESGAIDVKFFVLSEKGVIIVSLIFDGNWNGDRSTYGLSIILPQTELSFYLPLHRVCVDRLTHIIRKGRIWMHKERQENVQKIILEGTERMEDQGQSIIPMLTGEVIPVMELLSSMKSHSVPEEIDIADTVLNDDDIGDSCHEGFLLK (SEQ ID NO: 228).

[0060] In some embodiments, the disclosure provides systems specifically designed to modify the C9orf72 gene in eukaryotic cells. In some cases, the system is designed to knock down or knock out the C9orf72 gene. In other cases, the system is designed to correct one or more mutations in the C9orf72 gene. In some embodiments, the system is designed to remove a hexanucleotide repeat sequence and restore the ability of cells to express a functional C9orf72 protein. In some embodiments, the system is designed to correct a GGGGCC mutation in the hexanucleotide repeat sequence of the C9orf72 gene encoding the RNA transcripts of HRS and / or DPR, and to restore the ability of cells to express a functional C9orf72 protein.

[0061] In general, any portion of the C9orf72 gene can be targeted using programmable compositions and methods provided herein. In some embodiments, the CRISPR nuclease is a class 2 V nuclease. In some embodiments, the class 2 V nuclease is selected from the group consisting of Cas12a, Cas12b, Cas12c, Cas12d (CasY), Cas12J, and CasX. In some embodiments, the class 2 V nuclease is CasX. In some embodiments, the disclosure provides a CasX:gNA system comprising one or more CasX proteins and one or more guide nucleic acids (gNAs), and optionally one or more donor template nucleic acids. Each of these components and their use in editing the C9orf72 gene are described below herein.

[0062] In some embodiments, this disclosure provides a gene editing pair of CasX and gNA from any of the embodiments described herein, which can be bound together before use for gene editing and thus "pre-complexed" as a ribonucleoprotein complex (RNP). The use of pre-complexed RNPs provides advantages in the delivery of system components to cells or target nucleic acid sequences for editing of target nucleic acid sequences. In some embodiments, functional RNPs can be delivered ex vivo to cells by electrophoresis or by chemical means. In other embodiments, functional RNPs can be delivered ex vivo or in vivo by vectors in their functional form, or expressed and then complexed together as RNPs. gNA can provide target specificity to the complex by including a targeting sequence (or "spacer") having a nucleotide sequence complementary to the sequence of the target nucleic acid sequence, while the CasX protein of a pre-complexed CasX:gNA provides site-specific activity such as cleavage or nicking of the target sequence, which is induced by its association with gNA at a target site within the target nucleic acid sequence (e.g., the modified C9orf72 gene) (e.g., stabilization at the target site). The CasX protein and gNA components of the CasX:gNA system, as well as their sequences, properties, and functions, are described more fully below.

[0063] In some embodiments, the CasX:gNA system used for editing the C9orf72 gene may optionally further include a donor template comprising all or at least a portion of a gene encoding the C9orf72 protein, a non-coding region, or a C9orf72 regulatory element, the donor template comprising one or more mutations compared to the wild-type C9orf72 gene used for insertion to knock out or knock down (more fully described below) a target nucleic acid sequence having one or more mutations or HRS. In other cases, the CasX:gNA system may optionally further include a donor template for introducing (or knocking in) all or a portion of a gene encoding a physiologically normal number of hexanucleotide repeats, or a sequence for the production of the wild-type C9orf72 protein (SEQ ID NO: 227 or 228), or a sequence for the production of a physiologically normal level of C9orf72 in target cells. In some embodiments, the donor template comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 10,000, at least 15,000, or at least 25,000 nucleotides of the wild-type C9orf72 gene, and the C9orf72 gene portion is selected from the group consisting of C9orf72 exons, C9orf72 introns, C9orf72 intron-exon junctions, C9orf72 regulatory elements, C9orf72 coding regions, C9orf72 non-coding regions, or the entire C9orf72 gene. In some embodiments, the C9orf72 gene segment comprises any combination of the following: C9orf72 exon sequence, C9orf72 intron sequence, C9orf72 intron-exon junction sequence, C9orf72 non-coding region, or C9orf72 regulatory element sequence. In certain embodiments, the donor template comprises a sequence having a physiologically normal number of hexanucleotide repeats of the GGGGCC sequence, and upon insertion of the donor template, the hexanucleotide repeat sequence expansion of the C9orf72 gene is replaced.In other embodiments, the donor polynucleotides include at least about 10 to about 15,000 nucleotides, at least about 100 to about 10,000 nucleotides, at least about 400 to about 6,000 nucleotides, at least about 600 to about 4,000 nucleotides, or at least about 1,000 to about 2,000 nucleotides from the wild-type C9orf72 gene. In some embodiments, the donor template is a single-stranded DNA template or a single-stranded RNA template. In other embodiments, the donor template is a double-stranded DNA template.

[0064] III. Guide to Systems for Gene Editing: Nucleic Acids In another embodiment, the disclosure relates to a guide nucleic acid (gNA) comprising a targeting sequence complementary to the target nucleic acid sequence of the C9orf72 gene, the gNA being able to complex with a CRISPR protein specific to a protospacer adjacent motif (PAM) containing a TC motif on a complementary non-target strand, the PAM sequence being located at one nucleotide 5' of the sequence on the non-target strand that is complementary to the target nucleic acid sequence on the target strand of the target nucleic acid. In some embodiments, the gNA can complex with a class 2 type V CRISPR nuclease. In certain embodiments, the gNA can complex with a CasX nuclease.

[0065] In some embodiments, this disclosure provides gNAs for use in a CasX:gNA system that is useful in genome editing in cells, which is useful in editing the C9orf72 gene. This disclosure provides a specifically designed guide nucleic acid ("gNA") having a targeting sequence that is complementary to (and therefore can hybridize to) the C9orf72 gene as a component of a gene editing CasX:gNA system. Representative but non-limiting examples of targeting sequences to the C9orf72 target nucleic acid that can be used in the gNAs of embodiments are presented as SEQ ID NOs.309-343, 363-2100, and 2295-21835. In some embodiments, the gNA is a deoxyribonucleic acid molecule ("gDNA"), in some embodiments, the gNA is a ribonucleic acid molecule ("gRNA"), and in other embodiments, the gNA is a chimera and includes both DNA and RNA. As used herein, the terms gNA, gRNA, and gDNA include naturally occurring molecules as well as sequence variants.

[0066] In some embodiments, multiple gNAs (e.g., multiple gRNAs) are delivered via the CasX:gNA system for modification of one or more regions of the C9orf72 protein, non-coding regions of the C9orf72 gene, or genes encoding C9orf72 regulatory elements. For example, if deletion of a regulatory element or HRS of a gene is desired, a pair of gNAs having targeting sequences to different or overlapping regions of the target nucleic acid sequence can be used to bind and cleave at two different or overlapping sites within the gene. In other cases where a hexanucleotide repeat region is deleted, a pair of gNAs can be used to bind and cleave at two different sites, 5' and 3', of the hexanucleotide repeat within the C9orf72 gene, resulting in the removal of the HRS which is later edited by non-homologous end joining (NHEJ), homology-directed repair (HDR), homology-independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER). Table 16 below shows exemplary pairs of gNAs that may be used to edit HRS, and an exemplary method for achieving editing is described in Example 18 below.

[0067] a. Reference gNA and gNA variants In some embodiments, the gNAs of this disclosure include sequences of naturally occurring gNAs ("reference gNAs"). In other embodiments, the reference gNAs of this disclosure may be subjected to one or more mutagenesis methods, such as deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, to produce one or more gNA variants having improved or altered properties compared to the reference gNA. The gNA variants may also include variants that include one or more exogenous sequences fused to or inserted into either the 5' or 3' end. The activity of the reference gNA can be used as a benchmark against which the activity of the gNA variants can be compared, thereby measuring improvements in the function or other properties of the gNA variants. In other embodiments, the reference gNA may be subjected to one or more intentionally targeted mutations to produce gNA variants, for example, reasonably designed variants.

[0068] The gNAs of this disclosure comprise two segments: a targeting sequence and a protein-binding segment. The targeting segment of the gNA comprises a nucleotide sequence (interchangeably referred to as a guide sequence, spacer, targeter, or targeting sequence) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid sequence (e.g., a target ssRNA, target ssDNA, or the complementary strand of a double-stranded target DNA), and is fully described below. The targeting sequence of the gNA can bind to target nucleic acid sequences including coding sequences, complements of coding sequences, non-coding sequences, and regulatory elements. The protein-binding segment (or “activator” or “protein-binding sequence”) interacts with (e.g., binds to) the CasX protein as a complex to form an RNP (more fully described below).

[0069] In the case of a dual guide RNA (dgRNA), the targeter and activator each have a duplex-forming segment, and the duplex-forming segments of the targeter and activator are complementary to each other and hybridize to form a double-stranded duplex (dsRNA duplex in the case of gRNA). When gNA is gRNA, the terms “targeter” or “targeter RNA” are used herein to refer to the crRNA-like molecule (crRNA: “CRISPR RNA”) of the CasX dual guide RNA (and therefore, for example, the CasX single guide RNA when the “activator” and “targeter” are linked together by an intervening nucleotide). The crRNA has a 5' region that anneals with tracrRNA followed by nucleotides of the targeting sequence. Thus, for example, the guide RNA (dgRNA or sgRNA) includes the guide sequence and the duplex-forming segment of the crRNA, which may also be called the crRNA repeat. The corresponding tracrRNA-like molecule (activator) also includes a duplex-forming stretch of nucleotides that forms the remaining half of the dsRNA duplex of the protein-binding segment of the guide RNA. Thus, the targeter and activator hybridize as a corresponding pair to form a dual guide NA, which is referred to herein as “dual guide NA,” “dual molecular gNA,” “dgNA,” “dual molecular guide NA,” or “two-molecule guide NA.” Site-specific binding and / or cleavage of the target nucleic acid sequence (e.g., genomic DNA) by the CasX protein may occur at one or more locations (e.g., the sequence of the target nucleic acid) determined by the base-pairing complementarity between the targeting sequence of the gNA and the target nucleic acid sequence. For example, the gNA of this disclosure may have a sequence complementary to the target nucleic acid adjacent to a TC PAM motif or a PAM sequence such as ATC, CTC, GTC, or TTC, and thus be able to hybridize. Because the targeting sequence of the guide sequence hybridizes to the sequence of the target nucleic acid sequence, the targeter can be modified by the user to hybridize to a specific target nucleic acid sequence, as long as the position of the PAM sequence is taken into consideration.Therefore, in some cases, the targeter sequence may be a sequence that does not exist in nature. In other cases, the targeter sequence may be a sequence that exists in nature and originates from the gene being edited. In other embodiments, the activator and targeter of gNA are covalently bonded to each other (rather than hybridizing to each other) and comprise a single molecule referred herein as “single-molecule gNA,” “single-molecule guide NA,” “single-guide NA,” “single-guide RNA,” “single-molecule guide RNA,” “single-molecule guide RNA,” “single-guide DNA,” “single-molecule DNA,” or “single-molecule guide DNA” (“sgNA,” “sgRNA,” or “sgDNA”). In some embodiments, sgNA comprises an “activator” or a “targeter,” and therefore may be “activator RNA” and “targeter RNA,” respectively.

[0070] In summary, the gNA assembled in the present disclosure is specific to the target nucleic acid in the embodiments of the present disclosure and comprises four distinct regions or domains located at the 3' end of the gNA: an RNA triplex, a scaffold stem, an elongation stem, and a targeting sequence. The RNA triplex, scaffold stem, and elongation stem together are referred to as the “scaffold” of the reference gNA.

[0071] b.RNA triplex In some embodiments of the guide NAs (including reference sgNAs) provided herein, an RNA triplex is present, the RNA triplex comprising the sequence of a UUU--nX(approximately 4-15)--UUU stem loop (SEQ ID NO: 19) ending at AAAG after two intervening stem loops (a scaffolding stem loop and an elongation stem loop), forming a pseudoknot that can extend beyond the triplex to a duplex pseudoknot. The UU-UUU-AAA sequence of the triplex is formed as a spacer and a nexus between the scaffolding stem and the elongation stem. In the exemplary reference CasX sgNA, the UUU-loop-UUU region is encoded first, followed by the scaffolding stem loop, then the elongation stem loop linked by a tetraloop, and then the triplex is terminated at AAAG before becoming a spacer.

[0072] c. Scaffolding stem loop In some embodiments of the sgNAs of this disclosure, a scaffolding stem-loop follows a triplex region. The scaffolding stem-loop is the region of the gNA to which the CasX protein (such as a reference or CasX variant protein) is bound. In some embodiments, the scaffolding stem-loop is a fairly short and stable stem-loop. In some cases, the scaffolding stem-loop does not tolerate much variation and requires some form of RNA bubble. In some embodiments, the scaffolding stem is required for CasX sgNA function. While perhaps similar to the nexus stem of Cas9 in that it is a critical stem-loop, in some embodiments, the scaffolding stem of the CasX sgNA has a required bulge (RNA bubble) that differs from many other stem-loops found in the CRISPR / Cas system. In some embodiments, the presence of this bulge is conserved across sgNAs interacting with different CasX proteins. An exemplary sequence of a scaffolding stem-loop sequence of a gNA includes the sequence CCAGCGACUAUGUCGUAUGG (SEQ ID NO: 20). In other embodiments, the disclosure provides gNA variants in which the scaffolding stem-loop is replaced with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, a stem-loop sequence selected from MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem-loop. In some cases, the heterologous RNA stem-loop of the gNA can bind to proteins, RNA structures, DNA sequences, or small molecules.

[0073] d. Elongated stem loop In some embodiments of the CasX sgNA of this disclosure, an elongation stem loop follows a scaffolding stem loop. In some embodiments, the elongation stem includes a synthetic tracrRNA and crRNA fusion to which the majority of the CasX protein is not bound. In some embodiments, the elongation stem loop may be highly adaptable. In some embodiments, a single guide gRNA is constructed with a GAAA tetraloop linker or GAGAAA linker between the tracrRNA and crRNA within the elongation stem loop. In some cases, the targeter and activator of the CasX sgNA are linked to each other by intervening nucleotides, and the linker may have a nucleotide length of 3 to 20. In some embodiments of the CasX sgNA of this disclosure, the elongation stem is a large 32 bp loop located outside the CasX protein of the ribonucleoprotein complex. An exemplary sequence of the elongation stem loop sequence of the sgNA includes the sequence GCGCUUAUUUAUCGGAGAGAAAUCCGAUAAAUAAGAAGC (SEQ ID NO: 21). In some embodiments, the elongated stem-loop includes a GAGAAA spacer sequence. In some embodiments, the disclosure provides gNA variants in which the elongated stem-loop is replaced with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, such as, but not limited to, a stem-loop sequence selected from MS2, Qβ, U1 hairpin II, Uvsx, or PP7 stem-loop. In such cases, the heterologous RNA stem-loop increases the stability of the gNA. In other embodiments, the disclosure provides gNA variants having an elongated stem-loop region containing at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides, or at least 10 to 10,000, at least 10 to 100, or at least 10 to 100 nucleotides.

[0074] e. Targeted sequences In some embodiments of the gNA of this disclosure, the elongated stem loop is followed by a region that forms part of a triplex, and then a targeting sequence. The targeting sequence causes the CasX ribonucleoprotein holocomplex to target a specific region of the target nucleic acid sequence of the C9orf72 gene. Thus, for example, the CasX gNA targeting sequence of this disclosure has a sequence that is complementary to and can hybridize to a portion of the C9orf72 gene in a eukaryotic nucleic acid (e.g., eukaryotic chromosomes, chromosome sequences, eukaryotic RNA, etc.) as a component of an RNP, if the TC PAM motif, or one of the PAM sequences TTC, ATC, GTC, or CTC, is located at one nucleotide at the 5' of a non-target strand sequence that is complementary to the target sequence. As long as the position of the PAM sequence is taken into consideration, the targeting sequence of the gNA can be modified so that the gNA can target a desired sequence of any desired target nucleic acid sequence. In some embodiments, the gNA scaffold is located at the 5' end of the targeting sequence, and the targeting sequence is located at the 3' end of the gNA. In some embodiments, the PAM motif sequence recognized by the RNP nuclease is TC. In other embodiments, the PAM sequence recognized by the RNP nuclease is NTC.

[0075] In some embodiments, the gNA targeting sequence is specific to a portion of the gene encoding the C9orf72 protein containing one or more mutations. In some embodiments, the gNA targeting sequence is specific to the C9orf72 exon. In some embodiments, the gNA targeting sequence is specific to the C9orf72 intron. In some embodiments, the gNA targeting sequence is specific to the C9orf72 intron-exon junction. In some embodiments, the gNA targeting sequence has a sequence that hybridizes with a C9orf72 regulatory element, a C9orf72 coding region, a C9orf72 non-coding region, or a combination thereof. In certain embodiments, the gNA targeting sequence hybridizes with a sequence that is 5' of the HRS. In some embodiments where two or more gNAs are used, the first gNA targeting sequence of the gNAs hybridizes with a sequence that is 5' of the HRS, and the second gNA hybridizes with a sequence that is 3' of the HRS. In some embodiments, the gNA targeting sequence is complementary to a sequence containing one or more single nucleotide polymorphisms (SNPs) in the C9orf72 gene or its complement. SNPs within the C9orf72 coding sequence or the C9orf72 non-coding sequence are both within the scope of this disclosure. In other embodiments, the gNA targeting sequence is complementary to a sequence in the intergenetic region of the C9orf72 gene, or to a sequence complementary to the intergenetic region of the C9orf72 gene.

[0076] In some embodiments, the targeting sequence of the gNA is specific to regulatory elements that regulate C9orf72 expression. Such C9orf72 regulatory elements include, but are not limited to, promoter regions, enhancer regions, intergeneric regions, 5' untranslated region (5'UTR), 3' untranslated region (3'UTR), intergeneric regions, gene enhancer elements, conserved elements, and cis-regulatory elements. The promoter region is intended to contain nucleotides within 100 kb of the C9orf72 start site, or, in the case of gene enhancer elements or conserved elements, may be more than 1 Mb distal to the C9orf72 gene. In some embodiments, the disclosure provides a gNA having a targeting sequence that hybridizes with a C9orf72 regulatory element. As described above, the target is intended to knock out or knock down the target coding gene so that a hexanucleotide duplication of the C9orf72 protein or C9orf72 gene product containing a mutation is not expressed in the cell or is expressed at a lower level. In some embodiments, the disclosure provides a CasX:gNA system in which the targeting sequence (or spacer) of gNA is complementary to a nucleic acid sequence encoding a complement of C9orf72, a portion of the C9orf72 protein, a portion of the C9orf72 regulatory element, or a portion of the C9orf72 gene. In some embodiments, the targeting sequence of gNA has 14 to 35 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 18, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 consecutive nucleotides. In some embodiments, the targeting sequence consists of 21 consecutive nucleotides. In some embodiments, the targeting sequence consists of 20 consecutive nucleotides. In some embodiments, the targeting sequence consists of 19 consecutive nucleotides. In some embodiments, the targeting sequence consists of 18 consecutive nucleotides. In some embodiments, the targeting sequence consists of 17 consecutive nucleotides. In some embodiments, the targeting sequence consists of 16 consecutive nucleotides.In some embodiments, the targeting sequence consists of 15 consecutive nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, or 21 consecutive nucleotides, and the targeting sequence contains 0-5, 0-4, 0-3, or 0-2 mismatches with respect to the target nucleic acid sequence, and can maintain sufficient binding specificity so that the RNP containing the gNA containing the targeting sequence can form a complementary bond with the target nucleic acid.

[0077] Representative but non-limiting examples of targeting sequences for wild-type C9orf72 nucleic acid are presented in Tables 3 and 15 as SEQ ID NOs. 309-343, 363-2100, and 2295-21835. In some embodiments, the disclosure provides targeting sequences that are at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identical to the sequences in Tables 3 and 15, SEQ ID NOs. 309-343, 363-2100, and 2295-21835. In some embodiments, the targeting sequences for gNA include the sequences of SEQ ID NOs. 2281-159093, with a single nucleotide removed from the 3' end of the sequence. In other embodiments, the gNA targeting sequence includes sequences 309-343, 363-2100, and ~2295-21835, with two nucleotides removed from the 3' end of the sequence. In other embodiments, the gNA targeting sequence includes sequences 309-343, 363-2100, and ~2295-21835, with three nucleotides removed from the 3' end of the sequence. In other embodiments, the gNA targeting sequence includes sequences 309-343, 363-2100, and ~2295-21835, with four nucleotides removed from the 3' end of the sequence. In other embodiments, the gNA targeting sequence includes sequences 309-343, 363-2100, and ~2295-21835, with five nucleotides removed from the 3' end of the sequence. In the embodiments described above, if the gNA is incorporated into the expression vector in any of the targeting sequences or the coding sequence of the spacer so that it can be gDNA or gRNA, or an RNA and DNA chimera, then thymine (T) nucleotides are used in place of one or more or all of the uracil (U) nucleotides. In some embodiments, the targeting sequences of SEQ ID NOs. 309-343, 363-2100, and 2295-21835 have at least one, two, three, four, five, or six or more thymine nucleotides substituted for uracil nucleotides.In other embodiments, the gNA, gRNA, or gDNA of the Disclosure includes one, two, or three or more targeting sequences of SEQ ID NOs. 309-343, 363-2100, and 2295-21835, or a targeting sequence that is at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to one or more sequences of SEQ ID NOs.

[0078] In some embodiments, the targeting sequence is complementary to a nucleic acid sequence encoding a mutation in the C9orf72 protein of sequence number 227 or 228, or a hexanucleotide duplication that disrupts the function or expression of the C9orf72 protein.

[0079] In some embodiments, CasX:gNA comprises a first gNA and further comprises a second (and optionally, a third, fourth, or fifth) gNA, the second gNA or additional gNA having a targeting sequence complementary to a different portion or complement of the target nucleic acid compared to the targeting sequence of the first gNA, for example, the first gNA targets the 5' of a hexanucleotide repeat, and the second gNA targets the 3' of a hexanucleotide repeat. By selecting the targeting sequences of the gNAs, a defined region of the target nucleic acid sequence can be modified or edited using the CasX:gNA system described herein.

[0080] f.gNA scaffolding In some embodiments, the CasX reference gRNA includes a sequence isolated from or derived from Deltaproteobacter. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated from or derived from Deltaproteobacter may include ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 22) and ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 23). Exemplary crRNA sequences isolated from or derived from Deltaproteobacter may include the sequence CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 24). In some versions, the CasX reference gNA includes a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated from or derived from Deltaproteobacter.

[0081] In some embodiments, the CasX reference guide RNA includes a sequence isolated from or derived from Planctomycetes. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated from or derived from Planctomycetes may include UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 25) and UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGG (SEQ ID NO: 26). Exemplary crRNA sequences isolated from or derived from Planctomycetes may include the sequence UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 27). In some cases, the CasX reference gNA includes a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated from or derived from Planttomycetes.

[0082] In some embodiments, the CasX reference gNA includes a sequence isolated from or derived from Candidatus sungbacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated from or derived from Candidatus sungbacteria may include the sequences GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 28), GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 29), UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 30), and GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 31). In some cases, the CasX reference guide RNA contains a sequence that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to a sequence isolated from or derived from Candidatus Sungbacteria.

[0083] Table 1 provides reference gRNA tracr sequences and scaffold sequences. In some embodiments, the Disclosure provides gNA sequences having a scaffold comprising a sequence having at least one nucleotide modification compared to a reference gNA sequence having one of the sequences of SEQ ID NOs: 4-16 in Table 1. In embodiments where the vector comprises the DNA coding sequence of the gNA, or where the gNA is a gDNA or a chimera of RNA and DNA, it will be understood that a thymine (T) base can be used in place of a uracil (U) base in any of the embodiments of the gNA sequences described herein, including the sequences in Tables 1 and 2. [Table 1]

[0084] g.gNA variant In another aspect, the disclosure relates to guide nucleic acid variants (hereinafter substituted by “gNA variant” or “gRNA variant”) comprising one or more modifications to a reference gRNA scaffold. As used herein, “scaffold” refers to all portions of gNA necessary for gNA function, excluding spacer sequences.

[0085] In some embodiments, the gNA variant comprises one or more nucleotide substitutions, insertions, deletions, or swapped or substituted regions with respect to the reference gRNA sequence of this disclosure. In some embodiments, the mutation can occur in any region of the reference gRNA to produce a gNA variant. In some embodiments, the scaffold of the gNA variant sequence has at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70%, at least 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with respect to the sequence of SEQ ID NO: 4 or SEQ ID NO: 5.

[0086] In some embodiments, the gNA variant includes one or more nucleotide changes within one or more regions of the reference gRNA that improve the properties of the reference gRNA. Exemplary regions include RNA triplexes, pseudoknots, scaffolding stem-loops, and elongation stem-loops. In some cases, the variant scaffolding stem further includes bubbles. In other cases, the variant scaffolding further includes a triplex-loop region. In yet another case, the variant scaffolding further includes a 5' non-structural region. In one embodiment, the gNA variant scaffolding includes a scaffolding stem-loop having at least 60% sequence identity to SEQ ID NO: 14. In another embodiment, the gNA variant includes a scaffolding stem-loop having the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32). In another embodiment, the disclosure provides a gNA scaffold including a C18G substitution, G55 insertion, U1 deletion, and modified elongated stem loop, in which the original 6nt loop and 13 base pairs closest to the loop (32 nucleotides in total) are replaced with a Uvsx hairpin (a 4nt loop and 5 base pairs closest to the loop, 14 nucleotides in total), and the base distal to the loop of the elongated stem is converted into a fully base-paired stem adjacent to the new Uvsx hairpin by A99 deletion and G64U substitution. In the aforementioned embodiment, the gNA scaffold includes the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGUGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG (Sequence ID 33).

[0087] All gNA variants that, when compared to a reference gRNA described herein, have one or more improved functions or properties, or add one or more novel functions, are assumed to be within the scope of this disclosure. A typical example of such a gNA variant is Guide 174 (SEQ ID NO: 2238), whose design is described in the examples. In some embodiments, the gNA variant adds a novel function to the RNP containing the gNA variant. In some embodiments, the gNA variant has improved stability, improved solubility, improved transcription of gNA, improved resistance to nuclease activity, increased folding rate of gNA, reduced byproduct formation during folding, increased productive folding, improved binding affinity to CasX protein, improved binding affinity to target DNA when complexed with CasX protein, improved gene editing when complexed with CasX protein, improved editing specificity when complexed with CasX protein, and improved ability to utilize a broader spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TTC, in editing target DNA when complexed with CasX protein, as well as improved properties selected from any combination thereof. In some cases, one or more of the improved properties of the gNA variant are improved by at least about 1.1 to about 100,000 times compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5. In other cases, one or more improved properties of the gNA variant are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, and at least about 100,000 times or more compared to the reference gNA of SEQ ID NO: 4 or SEQ ID NO: 5.In other cases, one or more of the improved properties of the gNA variant were approximately 1.1 to 100.00 times, approximately 1.1 to 10.00 times, approximately 1.1 to 1,000 times, approximately 1.1 to 500 times, 1.1 to 100 times, approximately 1.1 to 50 times, approximately 1.1 to 20 times, approximately 10 to 100.00 times, approximately 10 to 10.00 times, approximately 10 to 1,000 times, approximately 10 to 500 times, approximately 10 to 100 times, approximately 10 to 50 times, approximately 10 to 20 times, approximately 2 to 70 times, approximately 2 to 50 times, approximately 2 to 30 times, approximately 2 to 20 times, approximately 2 to 10 times, approximately 5 to 50 times, Improvements of approximately 5-30 times, 5-10 times, 100-100.00 times, 100-10.00 times, 100-1,000 times, 100-500 times, 500-100.00 times, 500-10.00 times, 500-1,000 times, 500-750 times, 1,000-100.00 times, 10,000-100.00 times, 20-500 times, 20-250 times, 20-200 times, 20-100 times, 20-50 times, 50-10,000 times, 50-1,000 times, 50-500 times, 50-200 times, or 50-100 times. In other cases, one or more improved properties of a gNA variant were approximately 1.1 times, 1.2 times, 1.3 times, 1.4 times, 1.5 times, 1.6 times, 1.7 times, 1.8 times, 1.9 times, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 11 times, 12 times, 13 times, 14 times, 15 times, 16 times, 17 times, 18 times, 19 times, 20 times, 25 times, 30 times, 40 times, 45 times, 50 times, 55 times, and 60 times compared to the reference gNA of SEQ ID NO: 4 or 5. It is improved by 70 times, 80 times, 90 times, 100 times, 110 times, 120 times, 130 times, 140 times, 150 times, 160 times, 170 times, 180 times, 190 times, 200 times, 210 times, 220 times, 230 times, 240 times, 250 times, 260 times, 270 times, 280 times, 290 times, 300 times, 310 times, 320 times, 330 times, 340 times, 350 times, 360 times, 370 times, 380 times, 390 times, 400 times, 425 times, 450 times, 475 times, or 500 times.

[0088] In some embodiments, gNA variants can be produced by subjecting a reference gRNA to one or more mutagenesis methods, such as deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, as described herein below. The activity of the reference gRNA can be used as a benchmark against which the activity of the gNA variant can be compared, thereby measuring the improvement in the function of the gNA variant. In other embodiments, the reference gRNA can be subjected to one or more intentional targeted mutations, substitutions, or domain swaps to produce gNA variants, e.g., rationally designed variants. Exemplary gRNA variants produced by such methods are described in the examples, and representative sequences of gNA scaffolds are presented in Table 2.

[0089] In some embodiments, the gNA variant includes one or more modifications compared to a reference guide nucleic acid scaffold sequence, the one or more modifications being selected from at least one nucleotide substitution in the gNA variant region, at least one nucleotide deletion in the gNA variant region, at least one nucleotide insertion in the gNA variant region, substitution of all or part of the gNA variant region, deletion of all or part of the gNA variant region, or any combination thereof. In some cases, the modification is the substitution of 1 to 15 consecutive or discontinuous nucleotides in the gNA variant in one or more regions. In other cases, the modification is the deletion of 1 to 10 consecutive or discontinuous nucleotides in the gNA variant in one or more regions. In other cases, the modification is the insertion of 1 to 10 consecutive or discontinuous nucleotides in the gNA variant in one or more regions. In other cases, the modification is the substitution of a scaffold stem-loop or elongated stem-loop in an RNA stem-loop sequence derived from a heterologous RNA source having proximal 5' and 3' ends. In some cases, the gNA variants of this disclosure include two or more modifications in one region. In other cases, the gNA variants of this disclosure include modifications in two or more regions. In other cases, the gNA variants include any combination of the modifications described in this paragraph.

[0090] In some embodiments, a 5'G is added to the gNA variant sequence for in vivo expression because transcription from the U6 promoter is more efficient and consistent with respect to the start site when the +1 nucleotide is G. In other embodiments, two 5'Gs are added to the gNA variant sequence for in vitro transcription to increase production efficiency because T7 polymerase strongly prefers G at position +1 and purine at position +2. In some cases, the 5'G base is added to the reference scaffold in Table 1. In other cases, the 5'G base is added to the variant scaffold in Table 2.

[0091] Table 2 provides exemplary gNA variant scaffold sequences. In Table 2, (-) indicates a deletion at a specified position relative to the reference sequence of SEQ ID NO: 5, (+) indicates an insertion of a specified base at an indicated position relative to SEQ ID NO: 5, and (:) indicates a range of bases at a specified start:end coordinate for a deletion or substitution relative to SEQ ID NO: 5, where multiple insertions, deletions, or substitutions are separated by punctuation, for example, A14C, U17G. In some embodiments, the gNA variant scaffold includes one of the sequences SEQ ID NOs. 2101-2294 listed in Table 2, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% sequence identity thereto. In embodiments where the vector contains a DNA coding sequence of gNA, or where the gNA is a chimera of gDNA or RNA and DNA, it will be understood that a thymine (T) base can be used in place of any of the uracil (U) bases in the embodiments of the gNA sequences described herein. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10] [Table 2-11] [Table 2-12] [Table 2-13]

[0092] In some embodiments, the gNA variant comprises a tracrRNA stem loop containing the sequence -UUU-N4-25-UUU-(SEQ ID NO: 34). For example, the gNA variant comprises a scaffolding stem loop or substitution thereof with two adjacent triplet U motifs contributing to a triplet region. In some embodiments, the scaffolding stem loop or substitution thereof comprises at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, or at least 25 nucleotides.

[0093] In some embodiments, the gNA variant includes a crRNA sequence having -AAAG- at the 5' position relative to the spacer region. In some embodiments, the -AAAG- sequence is immediately 5' relative to the spacer region.

[0094] In some embodiments, at least one nucleotide modification to a reference gNA for producing a gNA variant includes at least one nucleotide deletion in the CasX variant gNA relative to the reference gRNA. In some embodiments, the gNA variant includes deletions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive or discontinuous nucleotides relative to the reference gNA. In some embodiments, at least one deletion includes deletions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotides relative to the reference gNA. In some embodiments, the gNA variant contains deletions of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides relative to the reference gNA, and these deletions are not in consecutive nucleotides. In those embodiments where the gNA variant has two or more discontinuous deletions relative to the reference gRNA, any deletion of any length, and any combination of deletions of any length, as described herein are intended to be within the scope of this disclosure. In some embodiments, the gNA variant contains at least two deletions in different regions of the reference gRNA. In some embodiments, the gNA variant contains at least two deletions in the same region of the reference gRNA. For example, the region may be an elongated stem loop, a scaffolding stem loop, a scaffolding stem bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of the gNA variant. Any deletion of any nucleotide in the reference gRNA is also intended to be within the scope of this disclosure.

[0095] In some embodiments, at least one nucleotide modification of the reference gRNA to generate a gNA variant includes at least one nucleotide insertion. In some embodiments, the gNA variant includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive or discontinuous nucleotide insertions relative to the reference gRNA. In some embodiments, at least one nucleotide insertion includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotide insertions relative to the reference gRNA. In some embodiments, the gNA variant includes two or more insertions compared to the reference gRNA, and these insertions are not consecutive. In those embodiments where the gNA variant has two or more discontinuous insertions compared to the reference gRNA, any length of insertion and any combination of lengths of insertions described herein are intended to be within the scope of this disclosure. For example, in some embodiments, a gNA variant may include a first insertion of one nucleotide and a second insertion of two nucleotides, where these two insertions are not consecutive. In some embodiments, a gNA variant includes at least two insertions in different regions of the reference gRNA. In some embodiments, a gNA variant includes at least two insertions in the same region of the reference gRNA. For example, the region may be an elongation stem loop, a scaffolding stem loop, a scaffolding stem bubble, a triplex loop, a pseudoknot, a triplex, or the 5' end of the gNA variant. Insertions of any A, G, C, U (or T in the corresponding DNA), or any combination thereof, at any position in the reference gRNA are also intended to be within the scope of this disclosure.

[0096] In some embodiments, at least one nucleotide modification of the reference gRNA to generate a gNA variant includes at least one nucleic acid substitution. In some embodiments, the gNA variant includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive or discontinuous substituted nucleotides relative to the reference gRNA. In some embodiments, the gNA variant includes 1 to 4 nucleotide substitutions relative to the reference gRNA. In some embodiments, at least one substitution includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more consecutive nucleotide substitutions relative to the reference gRNA. In some embodiments, the gNA variant includes two or more substitutions relative to the reference gRNA, and these substitutions are not consecutive. In embodiments where the gNA variant has two or more discontinuous substitutions relative to the reference gRNA, any substituted nucleotide of any length, and any combination of substituted nucleotides of any length, as described herein, are intended to be within the scope of this disclosure. For example, in some embodiments, the gNA variant may include a first substitution of one nucleotide and a second substitution of two nucleotides, where these two substitutions are not consecutive. In some embodiments, the gNA variant includes at least two substitutions in different regions of the reference gRNA. In some embodiments, the gNA variant includes at least two substitutions in the same region of the reference gRNA. For example, the region may be a triplex, elongated stem loop, scaffolding stem loop, scaffolding stem bubble, triplex loop, pseudoknot, triplex, or the 5' end of the gNA variant. Any substitution of A, G, C, U (or T in the corresponding DNA), or any combination thereof, at any position in the reference gRNA is intended to be within the scope of this disclosure.

[0097] Any combination of the substitutions, insertions, and deletions described herein can be used to generate the gNA variants of this disclosure. For example, a gNA variant may include, with respect to a reference gRNA, at least one substitution and at least one deletion, with respect to a reference gRNA, at least one substitution and at least one insertion, with respect to a reference gRNA, at least one insertion and at least one deletion, or with respect to a reference gRNA, at least one substitution, one insertion, and one deletion.

[0098] In some embodiments, the gNA variant includes scaffolding regions that are at least 20% identical, at least 30% identical, at least 40% identical, at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to any one of sequence numbers 4 to 16.

[0099] In some embodiments, the gNA variant includes a tracr stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, or at least 95% identical to SEQ ID NO: 14.

[0100] In some embodiments, the gNA variant includes an elongated stem loop that is at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, or at least 95% identical to SEQ ID NO: 15.

[0101] In some embodiments, the gNA variant includes an exogenous extension stem loop, which differs from the reference gNA described below. In some embodiments, the exogenous extension stem loop has little or no identity with respect to the reference stem loop region disclosed herein (e.g., Sequence ID No. 15). In some embodiments, the exogenous stem-loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, at least 1,000 bp, at least 2,000 bp, at least 3,000 bp, at least 4,000 bp, at least 5,000 bp, at least 6,000 bp, at least 7,000 bp, at least 8,000 bp, at least 9,000 bp, at least 10,000 bp, at least 12,000 bp, at least 15,000 bp, or at least 20,000 bp. In some embodiments, the gNA variant includes an elongated stem-loop region containing at least 10, at least 100, at least 500, at least 1000, or at least 10,000 nucleotides. In some embodiments, heterologous stem-loops increase the stability of the gNA. In some embodiments, heterologous RNA stem-loops can bind to proteins, RNA structures, DNA sequences, or small molecules. In some embodiments, the exogenous stem-loop region replacing the stem-loop includes an RNA stem-loop or hairpin that has increased stability and can interact with specific cellular proteins or RNAs depending on the choice of loop.Such exogenous elongated stem-loops can include, for example, thermostable RNA, MS2 (ACAUGAGGAUCACCCAUGU (SEQ ID NO: 35)), Qβ (UGCAUGUCUAAGACAGCA (SEQ ID NO: 36)), U1 hairpin II (AAUCCAUUGCACUCCGGAUU (SEQ ID NO: 37)), Uvsx (CCUCUUCGGAGG (SEQ ID NO: 38)), PP7 (AGGAGUUUCUAUGGAAACCCU (SEQ ID NO: 39)), phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU (SEQ ID NO: 40)), kissing loop_a (UGCUCGCUCCGUUCGAGCA (SEQ ID NO: 41)), kissing loop_b1 (UGCUCGAC Examples include GCGUCCUCGAGCA (SEQ ID NO: 42), kissing loop_b2 (UGCUCGUUUGCGGCUACGAGCA (SEQ ID NO: 43)), G quadriplex M3q (AGGGAGGGAGGGAGAGG (SEQ ID NO: 44)), G quadriplex telomere basket (GGUUAGGGUUAGGGUUAGG (SEQ ID NO: 45)), salsin-lysine loop (CUGCUCAGUACGAGAGGAACCGCAG (SEQ ID NO: 46)), or pseudoknot (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGGAGUUUUAAAAUGUCUCUAAGUACA (SEQ ID NO: 47)). In some embodiments, the exogenous stem-loop includes a long non-coding RNA (lncRNA). As used herein, lncRNA refers to non-coding RNA longer than approximately 200 bp. In some embodiments, the 5' and 3' ends of the exogenous stem-loop form a base pair, i.e., they interact to form a duplex RNA region. In some embodiments, the 5' and 3' ends of the exogenous stem-loop form a base pair, and one or more regions between the 5' and 3' ends of the exogenous stem-loop do not form a base pair.In some embodiments, at least one nucleotide modification includes (a) substitution of 1 to 15 consecutive or discontinuous nucleotides of a gNA variant in one or more regions, (b) deletion of 1 to 10 consecutive or discontinuous nucleotides of a gNA variant in one or more regions, (c) insertion of 1 to 10 consecutive or discontinuous nucleotides of a gNA variant in one or more regions, (d) substitution of a scaffolding stem-loop or elongation stem-loop in an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends, or any combination of (a) to (d).

[0102] In some embodiments, the gNA variant includes a scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32). In some embodiments, the gNA variant includes a scaffold stem-loop sequence of CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 32) having at least 1, 2, 3, 4, or 5 mismatches thereto.

[0103] In some embodiments, the gNA variant includes an elongated stem-loop region containing fewer than 32 nucleotides, fewer than 31 nucleotides, fewer than 30 nucleotides, fewer than 29 nucleotides, fewer than 28 nucleotides, fewer than 27 nucleotides, fewer than 26 nucleotides, fewer than 25 nucleotides, fewer than 24 nucleotides, fewer than 23 nucleotides, fewer than 22 nucleotides, fewer than 21 nucleotides, or fewer than 20 nucleotides. In some embodiments, the gNA variant includes an elongated stem-loop region containing fewer than 32 nucleotides. In some embodiments, the gNA variant further includes a thermally stable stem-loop.

[0104] In some embodiments, the sgRNA variant includes the sequence of SEQ ID NO: 2104, SEQ ID NO: 2106, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO: 2256, SEQ ID NO: 2274, SEQ ID NO: 2275, SEQ ID NO: 2279, or SEQ ID NO: 2281. In some embodiments, the sgRNA variant includes the sequence of SEQ ID NOs: 2238, 2246, 2256, 2274, or 2275.

[0105] In some embodiments, the gNA variant contains any one of the sequences 2236, 2237, 2238, 2241, 2244, 2248, 2249, 2256, or 2259-2294, or has at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the gNA variant contains one or more additional modifications to any one of the sequences 2201-2294. In some embodiments, the gNA variant contains any one of the sequences 2236, 2237, 2238, 2241, 2244, 2248, 2249, or 2259-2294.

[0106] In some embodiments, the sgRNA variant includes one or more additional changes to the sequence of SEQ ID NO: 2104, SEQ ID NO: 2163, SEQ ID NO: 2107, SEQ ID NO: 2164, SEQ ID NO: 2165, SEQ ID NO: 2166, SEQ ID NO: 2103, SEQ ID NO: 2167, SEQ ID NO: 2105, SEQ ID NO: 2108, SEQ ID NO: 2112, SEQ ID NO: 2160, SEQ ID NO: 2170, SEQ ID NO: 2114, SEQ ID NO: 2171, SEQ ID NO: 2112, SEQ ID NO: 2173, SEQ ID NO: 2102, SEQ ID NO: 2174, SEQ ID NO: 2175, SEQ ID NO: 2109, SEQ ID NO: 2176, SEQ ID NO: 2238, SEQ ID NO: 2239, SEQ ID NO: 2240, SEQ ID NO: 2241, SEQ ID NO: 2243, SEQ ID NO: 2256, SEQ ID NO: 2274, SEQ ID NO: 2275, SEQ ID NO: 2279, or SEQ ID NO: 2281.

[0107] In some embodiments of the gNA variants of the present disclosure, the gNA variant comprises at least one modification, the at least one modification compared to the reference guide scaffold of SEQ ID NO: 5, which is selected from one or more of the following: (a) C18G substitution in a triple loop, (b) G55 insertion in a stem bubble, (c) U1 deletion, and (d) extension stem loop modification, wherein (i) a 6nt loop and 13 loop proximal base pairs are replaced by a Uvsx hairpin, and (ii) a deletion of A99 and substitution of G65U resulting in a fully based loop distal base. In such embodiments, the gNA variant comprises any one of the sequences SEQ ID NOs: 2236, 2237, 2238, 2241, 2244, 2248, 2249, 2256, or 2259-2294.

[0108] In embodiments of the gNA variant, the gNA variant further comprises a spacer (or targeted sequence) region located at the 3' end of the gNA, which is specific to the C9orf72 sequence. Exemplary spacers and their homologous PAM sequences are shown in Table 3 below. [Table 3]

[0109] In embodiments of the gNA variant, the gNA variant further comprises a spacer (or targeting sequence) region located at the 3' end of the gNA, as more fully described above, comprising at least 14 to about 35 nucleotides, the spacer being designed using a sequence complementary to the target nucleic acid. In some embodiments, the gNA variant comprises a targeting sequence of at least 10 to 30 nucleotides complementary to the target nucleic acid. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides. In some embodiments, the gNA variant comprises a targeting sequence having 20 nucleotides. In some embodiments, the targeting sequence has 25 nucleotides. In some embodiments, the targeting sequence has 24 nucleotides. In some embodiments, the targeting sequence has 23 nucleotides. In some embodiments, the targeting sequence has 22 nucleotides. In some embodiments, the targeting sequence has 21 nucleotides. In some embodiments, the targeting sequence has 19 nucleotides. In some embodiments, the targeting sequence has 18 nucleotides. In some embodiments, the targeting sequence has 17 nucleotides. In some embodiments, the targeting sequence has 16 nucleotides. In some embodiments, the targeting sequence has 15 nucleotides. In some embodiments, the targeting sequence has 14 nucleotides. In some embodiments, the disclosure provides targeting sequences for inclusion in gNA variants of the disclosure that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to sequences SEQ ID NOs. 309-343, 363-2100, and 2295-21835.In some embodiments, the targeting sequences of the gNA variant include sequences 309-343, 363-2100, and 2295-21835, in which a single nucleotide is removed from the 3' end of the sequence. In other embodiments, the targeting sequences of the gNA variant include sequences 309-343, 363-2100, and ~2295-21835, in which two nucleotides are removed from the 3' end of the sequence. In other embodiments, the targeting sequences of the gNA variant include sequences 309-343, 363-2100, and ~2295-21835, in which three nucleotides are removed from the 3' end of the sequence. In other embodiments, the targeting sequences of the gNA variant include sequences 309-343, 363-2100, and ~2295-21835, in which four nucleotides are removed from the 3' end of the sequence. In other embodiments, the targeting sequences of the gNA variant include sequences SEQ ID NOs. 309-343, 363-2100, and 2295-21835, in which five nucleotides are removed from the 3' end of the sequence.

[0110] In some embodiments, the gNA variant further comprises a spacer (targeting) region located at the 3' end of the gNA, the spacer being designed with a sequence complementary to the target nucleic acid. In some embodiments, the target nucleic acid comprises a PAM sequence located at the 5' end of the spacer, having at least a single nucleotide separating the PAM from the first nucleotide of the spacer. In some embodiments, the PAM is located on the untargeted strand of the target region, i.e., the strand complementary to the target nucleic acid. In some embodiments, the PAM sequence is an ATC. In some embodiments, the targeting sequence for the ATC PAM comprises sequence numbers 363-2100 or 2295-5426, or sequences that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or at least 99% identical to sequence numbers 363-2100 or 2295-5426. In some embodiments, the targeting sequence for ATC PAM is selected from the group consisting of SEQ ID NOs: 363-2100 or 2295-5426. In some embodiments, the PAM sequence is CTC. In some embodiments, the targeting sequence for CTC PAM includes SEQ ID NOs: 16203-21835, or sequences that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or at least 99% identical to SEQ ID NOs: 16203-21835. In some embodiments, the PAM sequence is GTC.In some embodiments, the targeting sequence for the GTC PAM includes sequence numbers 12894-16202, or sequences that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or at least 99% identical to sequence numbers 12894-16202. In some embodiments, the targeting sequence for the GTC PAM is selected from the group consisting of sequence numbers 12894-16202. In some embodiments, the PAM sequence is a TTC. In some embodiments, the targeting sequence for TTC PAM includes sequences selected from the group consisting of SEQ ID NOs. 5427 to 12893, or sequences that are at least 50% identical, at least 55% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 99% identical to SEQ ID NOs.

[0111] In some embodiments, the scaffold for the gNA variant is a portion of an RNP having a reference CasX protein containing SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other embodiments, the scaffold for the gNA variant is a portion of an RNP having a CasX variant protein containing any one of the sequences in Tables 4, 6, 7, 8, or 10, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to it. In the embodiments described above, the gNA further includes a spacer sequence.

[0112] In some embodiments, the scaffold of the gNA variant is a variant that includes one or more additional changes to the sequence of a reference gRNA containing SEQ ID NO: 4 or SEQ ID NO: 5. In those embodiments where the scaffold of the reference gRNA is derived from SEQ ID NO: 4 or SEQ ID NO: 5, one or more improvements or added properties of the gNA variant are improved compared to the same properties of SEQ ID NO: 4 or SEQ ID NO: 5.

[0113] Complex formation with m.CasX protein In some embodiments, the gNA variant has an improved ability to form a complex with the CasX protein (such as the reference CasX or CasX variant protein) compared to the reference gRNA. In some embodiments, the gNA variant has an improved affinity for the CasX protein (such as the reference or variant protein) compared to the reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the examples. By improving ribonucleoprotein complex formation, in some embodiments, the efficiency of assembling functional RNPs can be improved. In some embodiments, more than 90%, more than 93%, more than 95%, more than 96%, more than 97%, more than 98%, or more than 99% of the RNPs, including the gNA variant and its spacer, are competent for gene editing of target nucleic acids.

[0114] Formation of a complex with the h.CasX protein In some embodiments, the gNA variant has an improved ability to form a complex with the CasX protein (such as the reference CasX or CasX variant protein) compared to the reference gRNA. In some embodiments, the gNA variant has an improved affinity for the CasX protein (such as the reference or variant protein) compared to the reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the CasX protein, as described in the examples. By improving ribonucleoprotein complex formation, in some embodiments, the efficiency of assembling functional RNPs can be improved. In some embodiments, more than 90%, more than 93%, more than 95%, more than 96%, more than 97%, more than 98%, or more than 99% of the RNPs, including the gNA variant and its spacer, are competent for gene editing of target nucleic acids.

[0115] Exemplary nucleotide alterations that can improve the ability of gNA variants to complex with the CasX protein may, in some embodiments, involve replacing the scaffold stem with a thermostable stem-loop. While we do not wish to be bound by any theory, replacing the scaffold stem with a thermostable stem-loop may increase the overall binding stability of the gNA variant with the CasX protein. Alternatively, or in addition, by removing a large portion of the stem-loop, the folding dynamics of the gNA variant may be altered, allowing for the easy and rapid production of a functionally folded gNA that can be structurally assembled, for example, by reducing the degree to which the gNA variant is likely to "entangle" itself. In some embodiments, the choice of scaffold stem-loop sequence may vary depending on the different spacers utilized by the gNA. In some embodiments, the scaffold sequence may be tuned to the spacers, and therefore the target sequence. Biochemical assays can be used to evaluate the binding affinity of the CasX protein to the gNA variant for RNP formation, including the assay of the Examples. For example, those skilled in the art can measure the change in the amount of fluorescently labeled gNA bound to the immobilized CasX protein in response to an increase in the concentration of additional unlabeled "cold competitor" gNA. Alternatively, or in addition, how the fluorescence signal changes when different amounts of fluorescently labeled gNA are flowed onto the immobilized CasX protein can be monitored or confirmed. Alternatively, the ability to form RNPs can be assessed using an in vitro cleavage assay against a defined target nucleic acid sequence.

[0116] i.gNA stability In some embodiments, the gNA variant exhibits improved stability compared to the reference gRNA. Increased stability and efficient folding can, in some embodiments, increase the extent to which the gNA variant persists within target cells, thereby increasing the opportunity to form a functional RNP capable of performing CasX functions such as gene editing. Increased stability of the gNA variant can, in some embodiments, allow for similar results in the delivery of smaller amounts of gNA to cells, subsequently reducing the opportunity for off-target effects during gene editing. Guide RNA stability can be assessed in various ways, including, for example, assembling the guide in vitro, incubating it for various periods in a solution mimicking the intracellular environment, and then measuring its functional activity by the in vitro cleavage assay described herein. Alternatively, or in addition, gNA can be collected from cells at various point points after initial transfection / transduction of gNA to determine how long the gNA variant persists relative to the reference gRNA.

[0117] j. Solubility In some embodiments, the gNA variant exhibits improved solubility compared to the reference gRNA. In some embodiments, the gNA variant exhibits improved solubility of the CasX protein:gNA RNP compared to the reference gRNA. In some embodiments, the solubility of the CasX protein:gNA RNP is improved by adding a ribozyme sequence to the 5' or 3' end of the gNA variant, for example, to the 5' or 3' end of the reference sgRNA. Some ribozymes, such as M1 ribozymes, can increase protein solubility through RNA-mediated protein folding. The increased solubility of CasX RNPs containing the gNA variants described herein can be evaluated by various means known to those skilled in the art, for example, by performing concentration measurement readings on a gel of the soluble fraction of lysed E. coli expressing CasX and the gNA variant.

[0118] Resistance to K-nuclease activity In some embodiments, gNA variants exhibit improved resistance to nuclease activity compared to reference gRNA, which may increase the persistence of variant gNA in the intracellular environment, thereby improving gene editing. Resistance to nuclease activity can be assessed by various methods known to those skilled in the art. For example, an in vitro method for measuring resistance to nuclease activity may include, for instance, contacting the reference gNA and variant with one or more exemplary RNA nucleases and measuring their degradation. Alternatively, or in addition, a gNA variant may demonstrate some degree of nuclease resistance by measuring its persistence in the cellular environment using the methods described herein.

[0119] l. Binding affinity to target DNA In some embodiments, the gNA variant has improved affinity for target DNA compared to the reference gRNA. In certain embodiments, the ribonucleoprotein complex containing the gNA variant has improved affinity for target DNA compared to the affinity of the RNP containing the reference gRNA. In some embodiments, the improved affinity of the RNP for target DNA includes improved affinity for the target sequence, improved affinity for the PAM sequence, improved ability of the RNP to locate the DNA of the target sequence, or any combination thereof. In some embodiments, the improved affinity for target DNA is a result of an increase in overall DNA binding affinity.

[0120] While we do not wish to be bound by theory, nucleotide changes in gNA variants affecting the function of OBD in the CasX protein may increase the affinity of the CasX variant protein to protospacer adjacent motifs (PAMs), as well as its ability to bind to or utilize PAM sequences of a widened spectrum other than the standard TTC PAM recognized by the reference CasX protein of Sequence ID No. 2, including PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC. This would increase the affinity and diversity of the CasX variant protein to target DNA sequences, resulting in a substantial increase in the number of target nucleic acid sequences that can be edited and / or bound compared to reference CasX. As will be described more fully below, the increase in the number of target nucleic acid sequences that can be edited compared to reference CasX refers to both PAM and protospacer sequences, as well as their orientation depending on the orientation of the non-target strand. This does not mean that the PAM sequence on the non-target strand, rather than the target strand, determines the cleavage or is mechanically involved in target recognition. For example, when referring to the TTC PAM, it may actually be the complementary GAA sequence required for target cleavage, or it may be a combination of nucleotides derived from both strands. In the case of the CasX protein disclosed herein, the PAM is located at 5' of the protospacer, and at least a single nucleotide separates the PAM from the first nucleotide of the protospacer. Alternatively, or in addition, changes in gNA that affect the function of the helical I and / or helical II domains, which increase the affinity of the CasX variant protein to the target DNA strand, can increase the affinity of the CasX RNP containing the variant gNA to the target DNA.

[0121] Addition or modification of m.gNA function In some embodiments, gNA variants may involve larger structural changes that alter the topology of the gNA variant with respect to the reference gRNA, thereby enabling different gNA functionalities. For example, in some embodiments, the gNA variant swaps the endogenous stem loop of the reference gRNA scaffold with a stem loop that either interacts with a previously identified stable RNA structure or a protein or RNA binding partner to recruit an additional portion to CasX, or recruits CasX to a specific location, such as inside a viral capsid that has a binding partner to the RNA structure. In other situations, RNAs may be recruited to each other, as seen in kissing loops, which may result in the co-localization of two CasX proteins for more efficient gene editing at a target DNA sequence. Such RNA structures may include MS2, Qβ, U1 hairpin II, Uvsx, PP7, phage replication loop, kissing loop_a, kissing loop_b1, kissing loop_b2, G quadriplex M3q, G quadriplex telomere basket, salsin-lysine loop, or pseudoknot.

[0122] In some embodiments, the gNA variant includes a terminal fusion partner. Exemplary terminal fusions may include the fusion of gRNA to a self-cleaving ribozyme or a protein-binding motif. As used herein, “ribozyme” refers to RNA or a segment thereof having one or more catalytic activities similar to those of a protein enzyme. Exemplary ribozyme catalytic activities may include, for example, RNA cleavage and / or ligation, DNA cleavage and / or ligation, or peptide bond formation. In some embodiments, such fusions can improve scaffold folding or mobilize DNA repair mechanisms. For example, in some embodiments, gRNA may be fused to the hepatitis delta virus (HDV) antigenomic ribozyme, HDV genomic ribozyme, hatchet ribozyme (from metagenomic data), env25 pistol ribozyme (a representative one derived from Aliistipes putredinis), HH15 minimal hammerhead ribozyme, tobacco ring spot virus (TRSV) ribozyme, wild-type virus hammerhead ribozyme (and reasonable variants), or the twisted sister 1 or RBMX recruiting motif. Hammerhead ribozymes are RNA motifs that catalyze reversible cleavage and ligation reactions at specific sites within an RNA molecule. Hammerhead ribozymes include type I, type II, and type III hammerhead ribozymes. HDV, pistol, and hatchet ribozymes have autocleavage activity. gNA variants containing one or more ribozymes may enable extended gNA function compared to the gRNA reference. For example, gNAs containing self-cleaving ribozymes may, in some embodiments, be transcribed and processed into mature gNAs as part of a polycistron transcript. Such fusions may occur at either the 5' or 3' end of the gNA. In some embodiments, gNA variants include fusions at both the 5' and 3' ends, each fusion being described independently herein. In some embodiments, gNA variants include a phage replication loop or a tetraloop. In some embodiments, the gNA includes a hairpin loop that can bind to a protein.For example, in some embodiments, the hairpin loop is MS2, Qβ, U1 hairpin II, Uvsx, or PP7 hairpin loop.

[0123] In some embodiments, the gNA variant comprises one or more RNA aptamers. As used herein, “RNA aptamer” refers to an RNA molecule that binds to a target with high affinity and high specificity. In some embodiments, the gNA variant comprises one or more riboswitches. As used herein, “riboswitch” refers to an RNA molecule that changes state upon binding to a small molecule. In some embodiments, the gNA variant further comprises one or more protein-binding motifs. By adding protein-binding motifs to the reference gRNA or gNA variant of this disclosure, in some embodiments, it may be possible to enable the CasX RNP to associate with additional proteins, for example, to add the functionality of those proteins to the CasX RNP.

[0124] n. Chemically modified gNA In some embodiments, this disclosure relates to chemically modified gNA. In some embodiments, this disclosure provides chemically modified gNA having guide RNA functionality and reduced sensitivity to nuclease cleavage. A chemically modified gNA is a gNA comprising any nucleotide or deoxynucleotide other than the four standard ribonucleotides A, C, G, and U. In some cases, the chemically modified gNA comprises any backbone or nucleotide bond other than the natural phosphodiester nucleotide bond. In certain embodiments, the retained functionality includes the ability of the modified gNA to bind to any of the embodiments of CasX described herein. In certain embodiments, the retained functionality includes the ability of the modified gNA to bind to the C9orf72 target nucleic acid sequence. In certain embodiments, the retained functionality includes the ability of a pre-complexed CasX protein-gNA to target the CasX protein or to bind to the target nucleic acid sequence. In certain embodiments, the retained functionality includes the ability of the CasX-gNA to nicking a target polynucleotide. In certain embodiments, the retained functionality includes the ability to cleave a target nucleic acid sequence by CasX-gNA. In certain embodiments, the retained functionality is any other known functionality of gNA in a CasX system having the CasX protein of embodiments of this disclosure.

[0125] In some embodiments, the disclosure describes nucleotide sugar modification as 2′-OC 1-4 Alkyl compounds, for example, 2′-O-methyl(2′-OMe), 2′-deoxy(2′-H), 2′-OC 1-3 Alkyl-OC 1-3Provided is a chemically modified gNA incorporated into a gNA selected from the group consisting of an alkyl, such as 2′-methoxyethyl (“2′-MOE”), 2′-fluoro (“2′-F”), 2′-amino (“2′-NH2”), 2′-arabinosyl (“2′-arabino”) nucleotide, 2′-F-arabinosyl (“2′-F-arabino”) nucleotide, 2′-locked nucleic acid (“LNA”) nucleotide, 2′-unlocked nucleic acid (“ULNA”) nucleotide, L-type sugar (“L-sugar”), and 4′-thioth核糖核苷酸. In other embodiments, the internucleotide linkage modification incorporated into the guide RNA is phosphorothioate “P(S)” (P(S)), phosphonocarboxylate (P(CH2) n COOR), such as phosphonoacetate “PACE” (P(CH2COO - )), thiophosphonocarboxylate ((S)P(CH2) n COOR), such as thiophosphonoacetate “thioPACE” ((S)P(CH2) n COO - )), alkylphosphonate (P(C 1-3 alkyl), such as methylphosphonate - P(CH3), boranophosphonate (P(BH3)), and phosphorodithioate (P(S)2), and is selected from the group consisting of.

[0126] In certain embodiments, the disclosure describes nucleic acid base ("base") modifications as follows: 2-thiouracil ("2-thioU"), 2-thiocytosine ("2-thioC"), 4-thiouracil ("4-thioU"), 6-thioguanine ("6-thioG"), 2-aminoadenine ("2-aminoA"), 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine ("5-methylC"), 5-methyluracil ("5-methylU"), 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dehydrouracil The present invention provides chemically modified gNAs incorporated into gNAs selected from the group consisting of 5-propynylcytosine, 5-propynyluracil, 5-ethinylcytosine, 5-ethinyluracil, 5-allyluracil ("5-allyl U"), 5-allylcytosine ("5-allyl C"), 5-aminoallyluracil ("5-aminoallyl U"), 5-aminoallyl-cytosine ("5-aminoallyl C"), debasalized nucleotides, Z bases, P bases, unstructured nucleic acids ("UNA"), isoguanines ("iso G"), isocytosine ("iso C"), 5-methyl-2-pyrimidine, x (A, G, C, T), and y (A, G, C, T).

[0127] In other embodiments, the disclosure relates to one or more isotopic modifications, one or more 15 N, 13 C, 14 C, deuterium, 3 H, 32 P, 125 I, 131 The present invention provides chemically modified gNAs introduced into nucleotide sugars, nucleic acid bases, phosphodiester bonds, and / or nucleotide phosphates, which include a nucleotide containing an I atom or another atom or element used as a tracer.

[0128] In some embodiments, the “terminal” modifications incorporated into gNA are selected from the group consisting of PEG (polyethylene glycol), hydrocarbon linkers (including heteroatom (O, S, N) substituted hydrocarbon spacers, halo-substituted hydrocarbon spacers, keto-, carboxyl-, amid-, thionyl-, carbamoyl-, and thionocarbama oil-containing hydrocarbon spacers), spermine linkers, e.g., 6-fluorescein-hexyl, quenchers (e.g., dabusil, BHQ), and dyes (e.g., fluorescein, rhodamine, cyanine) containing fluorescent dyes bound to linkers such as other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides, and / or proteins). In some embodiments, the “terminal” modifications include the conjugation (or ligation) of gNA to another molecule containing oligonucleotides of deoxynucleotides and / or ribonucleotides, peptides, proteins, sugars, oligosaccharides, steroids, lipids, folic acid, vitamins, and / or other molecules. In certain embodiments, the disclosure provides a chemically modified gNA in which the “terminal” modification (described above) is incorporated as a phosphodiester bond and can be incorporated anywhere between two nucleotides in the gNA, via a linker such as a 2-(4-butylamidefluorescein)propane-1,3-diolbis(phosphodiester) linker, located within the gNA sequence.

[0129] In some embodiments, the disclosure includes amines, thiols (or sulfhydryls), hydroxyls, carboxyls, carbonyls, thionyls, thiocarbonyls, carbamoyls, thiocarbamoyls, phosphoryls, alkenes, alkynes, halogens, or fluorescent dyes, non-fluorescent labels, tags ( 14 In the case of C, for example, biotin, avidin, streptavidin, or 15 N, 13 C, deuterium, 3 H, 32 P, 125The present invention provides chemically modified gNA having terminal modifications including terminal functional groups such as functional group-terminal linkers that can be subsequently conjugated to a desired portion selected from the group consisting of an isotope-labeled portion such as I, an oligonucleotide (including aptamers, deoxynucleotides and / or ribonucleotides), an amino acid, a peptide, a protein, a sugar, an oligosaccharide, a steroid, a lipid, folic acid, and a vitamin. Conjugation is provided with N-hydroxysuccinimide, isothiocyanate, DCC (or DCI), and / or "Bioconjugate Techniques" by Greg T. Hermanson, Publisher Essevier Science, 3 rd Standard chemistry well known in the art is used, including but not limited to coupling by any other standard methods described in ed. (2013) (the entirety of which is incorporated herein by reference).

[0130] IV. Proteins for modifying target nucleic acids This disclosure provides a system comprising CRISPR nucleases useful for genome editing in eukaryotic cells. In some embodiments, the CRISPR nucleases used in the genome editing system are class 2 type V nucleases. Members of the class 2 type V CRISPR-Cas system have differences but share several common features that distinguish them from the Cas9 system. First, type V nucleases have a single RNA-inducible RuvC domain-containing effector but lack an HNH domain, and they differ from the Cas9 system in that they recognize the T-rich PAM 5' upstream of the target region on the untargeted strand and rely on the G-rich PAM located 3' to the target sequence. Unlike Cas9, type V nucleases produce twisted double-strand breaks distal to the PAM sequence and blunt ends proximal to the PAM sequence. In addition, when activated in cis by target dsDNA or ssDNA binding, type V nucleases degrade the ssDNA in trans. In some embodiments, the V-type nuclease of the embodiment recognizes a 5'-TC PAM motif and produces a twisted end that is cleaved only by the RuvC domain. In some embodiments, the V-type nuclease is selected from the group consisting of Cas12a, Cas12b, Cas12c, Cas12d (CasY), and CasX. In some embodiments, the V-type nuclease is a CasX nuclease. In some embodiments, the disclosure provides a system comprising a CasX protein and one or more guide nucleic acids (gNAs) that is specifically designed to modify a target nucleic acid sequence in a eukaryotic cell.

[0131] As used herein, the term “CasX protein” refers to a family of proteins, encompassing all naturally occurring CasX proteins, proteins that share at least 50% identity with naturally occurring CasX proteins, and CasX variants that exhibit one or more improved properties compared to naturally occurring reference CasX proteins.

[0132] Exemplary improved properties of the CasX variant embodiment include, but are not limited to, improved folding of the variant, improved binding affinity to gNA, improved binding affinity to target nucleic acids, improved ability to utilize a broader spectrum of PAM sequences in editing and / or binding of target DNA, improved unwinding of target DNA, increased editing activity, improved editing efficiency, improved editing specificity, increased percentage of eukaryotic genomes that can be effectively edited, increased nuclease activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target breaks, improved binding of non-target strands of DNA, improved protein stability, improved protein:gNA(RNP) complex stability, improved protein solubility, improved protein:gNA(RNP) complex solubility, improved protein yield, improved protein expression, and improved fusion properties, as fully described below. In some embodiments, the RNP of a CasX variant and a gNA variant exhibits one or more improved properties compared to the RNP of a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 with gNA in Table 1, when assayed in an equivalent manner, being improved by at least about 1.1 to about 100,000 times. In other instances, one or more improved properties of the RNP of a CasX variant and a gNA variant are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000 times, or more compared to the RNP of a reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 with gNA in Table 1.In other cases, one or more of the improved properties of the RNP between the CasX variant and the gNA variant were, when assayed in an equivalent manner, approximately 1.1–100,000, 1.1–10,000, 1.1–1,000, 1.1–500, 1.1–100, 1.1–50, 1.1–20, 10–100,000, 10–10,000, 10–1,000, 10–500, 10–100, 10–50, 10–20, 2–70, and 2–50. , about 2 to 30 times, about 2 to 20 times, about 2 to 10 times, about 5 to 50 times, about 5 to 30 times, about 5 to 10 times, about 100 to 100,00 times, about 100 to 10,00 times , about 100 to 1,000 times, about 100 to 500 times, about 500 to 100,00 times, about 500 to 10,00 times, about 500 to 1,000 times, about 500 to 75 The improvement is 0x, approximately 1,000-100.00x, approximately 10,000-100.00x, approximately 20-500x, approximately 20-250x, approximately 20-200x, approximately 20-100x, approximately 20-50x, approximately 50-10,000x, approximately 50-1,000x, approximately 50-500x, approximately 50-200x, or approximately 50-100x. In other cases, one or more improved properties of the RNPs of CasX variants and gNA variants, when assayed in an equivalent manner, were approximately 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, and 20x compared to the RNPs of the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and gNA in Table 1. It is improved by 25 times, 30 times, 40 times, 45 times, 50 times, 55 times, 60 times, 70 times, 80 times, 90 times, 100 times, 110 times, 120 times, 130 times, 140 times, 150 times, 160 times, 170 times, 180 times, 190 times, 200 times, 210 times, 220 times, 230 times, 240 times, 250 times, 260 times, 270 times, 280 times, 290 times, 300 times, 310 times, 320 times, 330 times, 340 times, 350 times, 360 times, 370 times, 380 times, 390 times, 400 times, 425 times, 450 times, 475 times, or 500 times.

[0133] The term "CasX variant" includes variants that are fusion proteins, meaning that CasX is "fused" to a heterologous sequence. This includes CasX variants that include a CasX variant sequence and an N-terminus, C-terminus, or internal fusion of CasX to a heterologous protein or its domain.

[0134] The CasX protein of this disclosure comprises at least one of the following domains: a non-target chain binding (NTSB) domain, a target chain loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC DNA cleavage domain (the last of these may be modified or deleted in catalytically dead CasX variants), as described in more detail below. Additionally, when complexed with gNA as an RNP, the CasX variant protein of this disclosure has an enhanced ability to efficiently edit and / or bind to target DNA by utilizing a PAM TC motif containing a PAM sequence selected from TTC, ATC, GTC, or CTC, compared to the RNP of the reference CasX protein and reference gNA. As stated above, the PAM sequence is located at at least one nucleotide at 5' of the non-target chain of the protospacer that has identity with the targeting sequence of gNA in the assay system, compared to the editing efficiency and / or binding of the RNP containing the reference CasX protein and reference gNA in an equivalent assay system. In one embodiment, an RNP of a CasX variant and a gNA variant exhibits higher editing efficiency and / or binding of the target sequence in target DNA compared to an RNP containing a reference CasX protein and a reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is TTC. In another embodiment, an RNP of a CasX variant and a gNA variant exhibits higher editing efficiency and / or binding of the target sequence in target DNA compared to an RNP containing a reference CasX protein and a reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is ATC. In yet another embodiment, an RNP of a CasX variant and a gNA variant exhibits higher editing efficiency and / or binding of the target sequence in target DNA compared to an RNP containing a reference CasX protein and a reference gNA in an equivalent assay system, and the PAM sequence of the target DNA is CTC.In another embodiment, the RNP of a CasX variant and a gNA variant exhibits higher editing efficiency and / or binding to the target sequence in the target DNA compared to the RNP containing a reference CasX protein and a reference gNA in an equivalent assay system, where the PAM sequence of the target DNA is GTC. In the above embodiment, the increase in editing efficiency and / or binding affinity to one or more PAM sequences is at least 1.5 times higher than the editing efficiency and / or binding affinity to the PAM sequences of the RNP of one of the CasX proteins of SEQ ID NOs: 1-3 and the gNA in Table 1.

[0135] In some embodiments, the CasX protein can bind to and / or modify (e.g., cleavage, nicking, methylation, demethylation, etc.) a target nucleic acid and / or a polypeptide associated with the target nucleic acid (e.g., histone tail methylation or acetylation). In some embodiments, the CasX protein is catalytically dead but retains the ability to bind to a target nucleic acid. An exemplary catalytically dead CasX protein contains one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, the catalytically dead CasX protein contains substitutions at residues 672, 769, and / or 935 of SEQ ID NO: 1. In one embodiment, the catalytically dead CasX protein contains the D672A, E769A, and / or D935A substitutions in the reference CasX protein of SEQ ID NO: 1. In other embodiments, the catalytically dead CasX protein contains substitutions at amino acids 659, 756, and / or 922 in the reference CasX protein of SEQ ID NO: 2. In some embodiments, the catalytically dead CasX protein includes the D659A, E756A, and / or D922A substitutions in the reference CasX protein of SEQ ID NO: 2. In further embodiments, the catalytically dead CasX protein includes the deletion of all or part of the RuvC domain of the CasX protein. It will be understood that the same aforementioned substitutions can be similarly introduced into the CasX variants of this disclosure to yield the dCasX variant. In one embodiment, all or part of the RuvC domain is deleted from the CasX variant to yield the dCasX variant. The catalytically inactive dCasX variant protein can be used for base editing or epigenetic modification in some embodiments.As affinity for DNA increases, in some embodiments, catalytically inactive dCasX variant proteins can locate their target nucleic acids more quickly, remain bound to the target nucleic acids for longer periods, and bind to the target nucleic acids in a more stable manner, or a combination thereof, compared to catalytically active CasX, thereby improving the function of catalytically dead CasX variant proteins compared to CasX variants that retain their cleavage ability.

[0136] a. Non-target chain binding domain The reference CasX protein in this disclosure contains a non-target strand binding domain (NTSBD). The NTSBD is a domain not previously found in any Cas protein, for example, it is not present in Cas proteins such as Cas9, Cas12a / Cpf1, Cas13, Cas14, CASCADE, CSM, or CSY. While not bound by theory or mechanism, the NTSBD in CasX may enable binding to the non-target DNA strand and assist in the unwinding of the non-target and target strands. The NTSBD is presumed to be involved in the unwinding or capture of the unwinded non-target DNA strand. The NTSBD is in direct contact with the non-target strand in the previously derived CryoEM model structure and may contain a non-standard zinc finger domain. The NTSBD may also play a role in DNA stabilization during unwinding, guide RNA insertion, and R-loop formation. In some embodiments, the exemplary NTSBD includes amino acids 101-191 of SEQ ID NO: 1 or amino acids 103-192 of SEQ ID NO: 2. In some embodiments, the NTSBD of the reference CasX protein contains a quadruple-stranded beta sheet.

[0137] b. Target chain loading domain The reference CasX protein in this disclosure includes a target chain loading (TSL) domain. The TSL domain is a domain not found in certain Cas proteins such as Cas9, CASCADE, CSM, or CSY. While we do not wish to be bound by theory or mechanism, the TSL domain is thought to be involved in assisting the loading of the target DNA strand to the RuvC active site of the CasX protein. In some embodiments, the TSL acts to position or capture the target strand in a folded state, placing cleavable phosphates of the target strand DNA backbone at the RuvC active site. The TSL includes a cys4(CXXC, CXXC zinc finger / ribbon domain (SEQ ID NO: 48) which is largely isolated by the TSL. In some embodiments, the exemplary TSL includes amino acids 825-934 of SEQ ID NO: 1 or amino acids 813-921 of SEQ ID NO: 2.

[0138] c. Helical I domain The reference CasX protein in this disclosure contains a helical I domain. Certain Cas proteins other than CasX have domains that may be named in a similar manner. However, in some embodiments, the helical I domain of a CasX protein contains one or more distinctive structural features or sequences, or a combination thereof, compared to non-CasX proteins. For example, in some embodiments, the helical I domain of a CasX protein contains one or more distinctive secondary structures compared to domains of other Cas proteins that may have similar names. For example, in some embodiments, the helical I domain of a CasX protein contains one or more alpha helices that have structures and sequences that are distinctive in terms of arrangement, number, and length compared to other CRISPR proteins. In certain embodiments, the helical I domain is involved in the interaction of spacers with bound DNA and guide RNA. While we do not wish to be bound by theory, in some cases, the helical I domain is thought to contribute to the binding of protospacer adjacent motifs (PAMs). In some embodiments, the exemplary helical I domain includes amino acids 57-100 and 192-332 of SEQ ID NO: 1, or amino acids 59-102 and 193-333 of SEQ ID NO: 2. In some embodiments, the helical I domain of the reference CasX protein includes one or more alpha helices.

[0139] d. Helical II domain The reference CasX protein of this disclosure includes a helical II domain. Certain Cas proteins other than CasX have domains that may be named in a similar manner. However, in some embodiments, the helical II domain of the CasX protein includes one or more distinctive structural features, or a distinctive sequence, or a combination thereof, compared to the domain of other Cas proteins that may have a similar name. For example, in some embodiments, the helical II domain includes one or more distinctive structural alpha-helical bundles aligned along a target DNA:guide RNA channel. In some embodiments, in CasX including the helical II domain, the target strand and guide RNA interact with helical II (and in some embodiments, the helical I domain) to allow the RuvC domain to approach the target DNA. The helical II domain is involved in the guide RNA scaffold stem loop and binding to the bound DNA. In some embodiments, the exemplary helical II domain includes amino acids 333-509 of SEQ ID NO: 1, or amino acids 334-501 of SEQ ID NO: 2.

[0140] e. Oligonucleotide binding domain The reference CasX protein of this disclosure includes an oligonucleotide-binding domain (OBD). Certain Cas proteins other than CasX have domains that may be named in a similar manner. However, in some embodiments, the OBD includes one or more distinctive functional features, or sequences specific to the CasX protein, or a combination thereof. For example, in some embodiments, a cross-linking helix (BH), helical I domain, helical II domain, and oligonucleotide-binding domain (OBD) together are involved in the binding of the CasX protein to the guide RNA. Thus, for example, in some embodiments, the OBD is specific to the CasX protein in that it functionally interacts with the helical I domain, or the helical II domain, or both, and each of these may be specific to the CasX protein described herein. Specifically, in CasX, the OBD primarily binds to the RNA triplex of the guide RNA scaffold. The OBD may also be involved in binding to the protospacer adjacent motif (PAM). An example OBD domain includes amino acids 1-56 and 510-660 of SEQ ID NO: 1, or amino acids 1-58 and 502-647 of SEQ ID NO: 2.

[0141] f.RuvC DNA cleavage domain The reference CasX protein of this disclosure contains a RuvC domain comprising two partial RuvC domains (RuvC-I and RuvC-II). The RuvC domain is the ancestral domain of all type 12 CRISPR proteins. The RuvC domain is derived from TNPB (transposase B), such as the transposase. Like other RuvC domains, the CasX RuvC domain has a DED catalytic triplicate structure involved in the regulation of magnesium (Mg) ions and DNA cleavage. In some embodiments, RuvC has a DED motif active site involved in cleaving both strands of DNA (likely cleaving one strand at a time, first cleaving the non-target strand at 11-14 nucleotides (nt) in the target sequence, and then cleaving the target strand at 2-4 nucleotides after the target sequence). Particularly in CasX, the RuvC domain is unique in that it is also involved in the binding of the guide RNA scaffold stem-loop, which is important for CasX function. An exemplary RuvC domain includes amino acids 661-824 and 935-986 of SEQ ID NO: 1, or amino acids 648-812 and 922-978 of SEQ ID NO: 2.

[0142] g. Reference CasX protein This disclosure provides a naturally occurring CasX protein (hereinafter referred to as the “reference CasX protein”) that functions as an endonuclease that catalyzes double-strand breaks at specific sequences in targeted double-stranded DNA (dsDNA). Sequence specificity is provided by the targeting sequence of the associated gNA that forms the complex and hybridizes to the target sequence in the target nucleic acid. For example, the reference CasX protein can be isolated from naturally occurring prokaryotes such as Deltaproteobacteria, Plantomycetes, or Candidatus Sungbacteria. The reference CasX protein (also referred to as the reference CasX protein herein) is a type V CRISPR / Cas endonuclease belonging to the CasX (sometimes referred to as Cas12e) family of proteins that can interact with guide NA to form ribonucleoproteins (RNPs). In some embodiments, an RNP complex containing a reference CasX protein can target a specific site within a target nucleic acid by base pairing between a targeting sequence (or spacer) of gNA and a target sequence within the target nucleic acid. In some embodiments, an RNP containing a reference CasX protein can cleave target DNA. In some embodiments, an RNP containing a reference CasX protein can nick target DNA. In some embodiments, an RNP containing a reference CasX protein can edit target DNA, for example, in embodiments where the reference CasX protein can cleave or nick DNA, this is followed by non-homologous end joining (NHEJ), homology-directed repair (HDR), homology-independent targeted integration (HITI), microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), or base excision repair (BER). In some embodiments, the RNP containing the CasX protein is a catalytically dead (catalyzably inactive or substantially lacking any cleavage activity) CasX protein (dCasX), as described more fully above, but retains the ability to bind to target DNA.

[0143] In some cases, the reference CasX protein is isolated from or derived from Deltaproteobacter. In some cases, the CasX protein contains a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequences: [ka]

[0144] In some cases, the reference CasX protein is isolated from or derived from Plantomycetes. In some cases, the CasX protein contains a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequences: [ka]

[0145] In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 2 or at least 60% similarity thereto. In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 2 or at least 80% similarity thereto. In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 2 or at least 90% similarity thereto. In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 2 or at least 95% similarity thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 2. In some embodiments, the CasX protein contains or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 2. These mutations may be insertions, deletions, amino acid substitutions, or any combination thereof.

[0146] In some cases, the reference CasX protein is isolated from or derived from Candidatus Sungbacteria. In some cases, the CasX protein contains a sequence that is at least 50% identical, at least 60% identical, at least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, at least 99.5% identical, or 100% identical to the following sequences: [ka]

[0147] In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 3 or at least 60% similarity thereto. In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 3 or at least 80% similarity thereto. In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 3 or at least 90% similarity thereto. In some embodiments, the CasX protein contains the sequence of SEQ ID NO: 3 or at least 95% similarity thereto. In some embodiments, the CasX protein consists of the sequence of SEQ ID NO: 3. In some embodiments, the CasX protein contains or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations relative to the sequence of SEQ ID NO: 3. These mutations may be insertions, deletions, amino acid substitutions, or any combination thereof.

[0148] h.CasX variant protein This disclosure provides variants of a reference CasX protein (hereinafter interchangeably referred to as “CasX variants” or “CasX variant proteins”), the CasX variants comprising at least one modification to at least one domain of the reference CasX protein, including the sequences of SEQ ID NOs: 1–3. In some embodiments, the CasX variants exhibit at least one improved property compared to the reference CasX protein. All variants that improve one or more functions or properties of the CasX variant protein compared to the reference CasX protein described herein are assumed to be within the scope of this disclosure. In some embodiments, the modification is a mutation of one or more amino acids in the reference CasX. In other embodiments, the modification is the substitution of one or more domains of the reference CasX in one or more domains derived from different CasX. In some embodiments, the insertion includes the insertion of some or all of a domain derived from a different CasX protein. Mutations may occur in one or more domains of the reference CasX protein, and may include, for example, the deletion of part or all of one or more domains, or the substitution, deletion, or insertion of one or more amino acids in any domain of the reference CasX protein. Domains of the CasX protein include the non-target chain binding (NTSB) domain, the targeted chain loading (TSL) domain, the helical I domain, the helical II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain. Any change in the amino acid sequence of the reference CasX protein that results in improved properties of the CasX protein is also considered a CasX variant protein as described herein. For example, a CasX variant may include one or more amino acid substitutions, insertions, deletions, or swapped domains relative to the reference CasX protein sequence, or any combination thereof.

[0149] In some embodiments, the CasX variant protein contains at least one modification in at least each of two domains of the reference CasX protein containing sequences 1-3. In some embodiments, the CasX variant protein contains at least one modification in at least two domains, at least three domains, at least four domains, or at least five domains of the reference CasX protein. In some embodiments, the CasX variant protein contains two or more modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein contains at least two modifications in at least one domain of the reference CasX protein, at least three modifications in at least one domain of the reference CasX protein, or at least four modifications in at least one domain of the reference CasX protein. In some embodiments, the CasX variant contains two or more modifications compared to the reference CasX protein, and each modification is made in a domain independently selected from the group consisting of NTSBD, TSLD, helical I domain, helical II domain, OBD, and RuvC DNA cleavage domain.

[0150] In some embodiments, at least one modification of the CasX variant protein includes a deletion of at least a portion of one domain of the reference CasX protein containing the sequences of SEQ ID NOs.1–3. In some embodiments, the deletion is located in the NTSBD, TSLD, helical I domain, helical II domain, OBD, or RuvC DNA cleavage domain.

[0151] Preferred mutagenesis methods for generating the CasX variant proteins of this disclosure may include, for example, deep mutational evolution (DME), deep mutational scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. In some embodiments, the CasX variant is designed, for example, by selecting one or more desired mutations in reference CasX. In certain embodiments, the activity of the reference CasX protein is used as a benchmark against which the activity of one or more CasX variants is compared, thereby measuring the improvement in the function of the CasX variants. Exemplary improvements in the CasX variant include, but are not limited to, improved folding of the variant, improved binding affinity to gNA, improved binding affinity to target DNA, modified binding affinity to one or more PAM sequences, improved unwinding of target DNA, increased editing activity, improved efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target breaks, improved binding of non-target strands of DNA, improved protein stability, improved protein:gNA complex stability, improved protein solubility, improved protein:gNA complex solubility, improved protein yield, improved protein expression, and improved fusion properties, as fully described below.

[0152] In some embodiments of the CasX variants described herein, at least one modification includes (a) substitution of 1 to 100 consecutive or discontinuous amino acids in the CasX variant compared to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3; (b) deletion of 1 to 100 consecutive or discontinuous amino acids in the CasX variant compared to the reference CasX; (c) insertion of 1 to 100 consecutive or discontinuous amino acids in the CasX compared to the reference CasX; or (d) any combination of (a) to (c). In some embodiments, at least one modification includes (a) substitution of 5 to 10 consecutive or discontinuous amino acids in a CasX variant compared to the reference CasX of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3; (b) deletion of 1 to 5 consecutive or discontinuous amino acids in a CasX variant compared to the reference CasX; (c) insertion of 1 to 5 consecutive or discontinuous amino acids in CasX compared to the reference CasX; or (d) any combination of (a) to (c).

[0153] In some embodiments, the CasX variant protein comprises or consists of a sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, or at least 50 mutations compared to the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. These mutations may be insertions, deletions, amino acid substitutions, or any combination thereof.

[0154] In some embodiments, the CasX variant protein contains at least one amino acid substitution in at least one domain of the reference CasX protein. In some embodiments, the CasX variant protein may contain at least about 1 to 4 amino acid substitutions, 1 to 10 amino acid substitutions, 1 to 20 amino acid substitutions, 1 to 30 amino acid substitutions, 1 to 40 amino acid substitutions, 1 to 50 amino acid substitutions, 1 to 60 amino acid substitutions, 1 to 70 amino acid substitutions, 1 to 80 amino acid substitutions, 1 to 90 amino acid substitutions, 1 to 100 amino acid substitutions, 2 to 10 amino acid substitutions, 2 to 20 amino acid substitutions, 2 to 30 amino acid substitutions, 3 to 10 amino acid substitutions, 3 to 20 amino acid substitutions, 3 to 30 amino acid substitutions, 4 to 10 amino acid substitutions, 4 to 20 amino acid substitutions, 3 to 300 amino acid substitutions, 5 to 10 amino acid substitutions, 5 to 20 amino acid substitutions, 5 to 30 amino acid substitutions, 10 to 50 amino acid substitutions, or 20 to 50 amino acid substitutions compared to the reference CasX protein, and may be continuous or discontinuous or in different domains. As used herein, “contiguous amino acids” refers to adjacent amino acids in the primary sequence of a polypeptide. In some embodiments, the CasX variant protein contains at least about 100 amino acid substitutions relative to the reference CasX protein. In some embodiments, the amino acid substitutions are conserved substitutions. In other embodiments, the substitutions are non-conservative, for example, polar amino acids being substituted for non-polar amino acids, or vice versa.

[0155] In the substitutions described herein, any amino acid may be substituted for any other amino acid. Substitutions may be conservative substitutions (e.g., a basic amino acid being substituted for another basic amino acid). Substitutions may be non-conservative substitutions (e.g., a basic amino acid being substituted for an acidic amino acid, or vice versa). For example, proline in the reference CasX protein may be substituted for any of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine to produce the CasX variant proteins of this disclosure.

[0156] In some embodiments, the CasX variant protein contains at least one amino acid deletion compared to the reference CasX protein. In some embodiments, the CasX variant protein contains deletions of 1 to 4 amino acids, 1 to 10 amino acids, 1 to 20 amino acids, 1 to 30 amino acids, 1 to 40 amino acids, 1 to 50 amino acids, 1 to 60 amino acids, 1 to 70 amino acids, 1 to 80 amino acids, 1 to 90 amino acids, 1 to 100 amino acids, 2 to 10 amino acids, 2 to 20 amino acids, 2 to 30 amino acids, 3 to 10 amino acids, 3 to 20 amino acids, 3 to 30 amino acids, 4 to 10 amino acids, 4 to 20 amino acids, 3 to 300 amino acids, 5 to 10 amino acids, 5 to 20 amino acids, 5 to 30 amino acids, 10 to 50 amino acids, or 20 to 50 amino acids compared to the reference CasX protein. In some embodiments, the CasX protein contains at least about 100 consecutive amino acid deletions compared to the reference CasX protein. In some embodiments, the CasX variant protein contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or 100 consecutive amino acid deletions compared to the reference CasX protein. In some embodiments, the CasX variant protein contains 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 consecutive amino acid deletions.

[0157] In some embodiments, the CasX variant protein contains two or more deletions relative to the reference CasX protein, and these two or more deletions are not consecutive amino acids. For example, the first deletion may be in the first domain of the reference CasX protein, and the second deletion may be in the second domain of the reference CasX protein. In some embodiments, the CasX variant protein contains 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 discontinuous deletions relative to the reference CasX protein. In some embodiments, the CasX variant protein contains at least 20 discontinuous deletions relative to the reference CasX protein. Each discontinuous deletion may be any length of amino acids as described herein, e.g., 1 to 4 amino acids, 1 to 10 amino acids, etc.

[0158] In some embodiments, the CasX variant protein includes one or more amino acid insertions relative to the sequence of SEQ ID NO: 1, 2, or 3. In some embodiments, the CasX variant protein includes one amino acid insertion, two to three consecutive or discontinuous amino acids, two to four consecutive or discontinuous amino acids, two to five consecutive or discontinuous amino acids, two to six consecutive or discontinuous amino acids, two to seven consecutive or discontinuous amino acids, two to eight consecutive or discontinuous amino acids, two to nine consecutive or discontinuous amino acids, two to ten consecutive or discontinuous amino acids, two to twenty consecutive or discontinuous amino acids, two to thirty consecutive or discontinuous amino acids, two to forty consecutive or discontinuous amino acids, two to fifty consecutive or discontinuous amino acids, two to sixty consecutive or discontinuous amino acids, and two to The CasX variant protein includes insertions of 70 consecutive or discontinuous amino acids, 2 to 80 consecutive or discontinuous amino acids, 2 to 90 consecutive or discontinuous amino acids, 2 to 100 consecutive or discontinuous amino acids, 3 to 10 consecutive or discontinuous amino acids, 3 to 20 consecutive or discontinuous amino acids, 3 to 30 consecutive or discontinuous amino acids, 4 to 10 consecutive or discontinuous amino acids, 4 to 20 consecutive or discontinuous amino acids, 3 to 300 consecutive or discontinuous amino acids, 5 to 10 consecutive or discontinuous amino acids, 5 to 20 consecutive or discontinuous amino acids, 5 to 30 consecutive or discontinuous amino acids, 10 to 50 consecutive or discontinuous amino acids, or 20 to 50 consecutive or discontinuous amino acids. In some embodiments, the CasX variant protein includes insertions of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive or discontinuous amino acids. In some embodiments, the CasX variant protein includes at least about 100 consecutive or discontinuous amino acid insertions. Any amino acid, or combination of amino acids, can be inserted using the insertions described herein to generate the CasX variant protein.

[0159] Any permutation of the substitution, insertion, and deletion embodiments described herein can be combined to generate the CasX variant proteins of this disclosure. For example, a CasX variant protein may comprise at least one substitution and at least one deletion relative to a reference CasX protein sequence, at least one substitution and at least one insertion relative to a reference CasX protein sequence, at least one insertion and at least one deletion relative to a reference CasX protein sequence, or at least one substitution, one insertion, and one deletion relative to a reference CasX protein.

[0160] In some embodiments, the CasX variant protein has at least about 60% sequence similarity to SEQ ID NO: 2 or a portion thereof. In some embodiments, the CasX variant protein has substitutions of Y789T, P793, Y789D, T72S, I546V, E552A, A636D, F536S, A708K, Y797L, L792G, A739V, G791M, and position 661 of SEQ ID NO: 2. Insertion of A, substitution of A788W in SEQ ID NO: 2, substitution of K390R in SEQ ID NO: 2, substitution of A751S in SEQ ID NO: 2, substitution of E385A in SEQ ID NO: 2, insertion of P at position 696 in SEQ ID NO: 2, insertion of M at position 773 in SEQ ID NO: 2, substitution of G695H in SEQ ID NO: 2, insertion of AS at position 793 in SEQ ID NO: 2, insertion of AS at position 795 in SEQ ID NO: 2, substitution of C477R in SEQ ID NO: 2, substitution of C477K in SEQ ID NO: 2, substitution of C479A in SEQ ID NO: 2, substitution of C479L in SEQ ID NO: 2, substitution of I55F in SEQ ID NO: 2, K2 Substitution of 10R, substitution of C233S in SEQ ID NO: 2, substitution of D231N in SEQ ID NO: 2, substitution of Q338E in SEQ ID NO: 2, substitution of Q338R in SEQ ID NO: 2, substitution of L379R in SEQ ID NO: 2, substitution of K390R in SEQ ID NO: 2, substitution of L481Q in SEQ ID NO: 2, substitution of F495S in SEQ ID NO: 2, substitution of D600N in SEQ ID NO: 2, substitution of T886K in SEQ ID NO: 2, substitution of A739V in SEQ ID NO: 2, substitution of K460N in SEQ ID NO: 2, substitution of I199F in SEQ ID NO: 2, substitution of G492P in SEQ ID NO: 2, substitution of T153I in SEQ ID NO: 2, Substitution of R591I in SEQ ID NO: 2, insertion of AS at position 795 in SEQ ID NO: 2, insertion of AS at position 796 in SEQ ID NO: 2, insertion of L at position 889 in SEQ ID NO: 2, substitution of E121D in SEQ ID NO: 2, substitution of S270W in SEQ ID NO: 2, substitution of E712Q in SEQ ID NO: 2, substitution of K942Q in SEQ ID NO: 2, substitution of E552K in SEQ ID NO: 2, substitution of K25Q in SEQ ID NO: 2, substitution of N47D in SEQ ID NO: 2, insertion of T at position 696 in SEQ ID NO: 2, substitution of L685I in SEQ ID NO: 2, substitution of N880D in SEQ ID NO: 2, substitution of Q102R in SEQ ID NO: 2,Substitution of M734K in SEQ ID NO: 2, A724S in SEQ ID NO: 2, T704K in SEQ ID NO: 2, P224K in SEQ ID NO: 2, K25R in SEQ ID NO: 2, M29E in SEQ ID NO: 2, H152D in SEQ ID NO: 2, S219R in SEQ ID NO: 2, E475K in SEQ ID NO: 2, G226R in SEQ ID NO: 2, A377K in SEQ ID NO: 2, E480K in SEQ ID NO: 2, K416E in SEQ ID NO: 2, H164R in SEQ ID NO: 2, K767R in SEQ ID NO: 2, I7F in SEQ ID NO: 2, M29R in SEQ ID NO: 2 Substitution of, substitution of H435R in SEQ ID NO: 2, substitution of E385Q in SEQ ID NO: 2, substitution of E385K in SEQ ID NO: 2, substitution of I279F in SEQ ID NO: 2, substitution of D489S in SEQ ID NO: 2, substitution of D732N in SEQ ID NO: 2, substitution of A739T in SEQ ID NO: 2, substitution of W885R in SEQ ID NO: 2, substitution of E53K in SEQ ID NO: 2, substitution of A238T in SEQ ID NO: 2, substitution of P283Q in SEQ ID NO: 2, substitution of E292K in SEQ ID NO: 2, substitution of Q628E in SEQ ID NO: 2, substitution of R388Q in SEQ ID NO: 2, substitution of G791M in SEQ ID NO: 2, substitution of L792K in SEQ ID NO: 2 Substitution of L792E in 2, substitution of M779N in 2, substitution of G27D in 2, substitution of K955R in 2, substitution of S867R in 2, substitution of R693I in 2, substitution of F189Y in 2, substitution of V635M in 2, substitution of F399L in 2, substitution of E498K in 2, substitution of E386R in 2, substitution of V254G in 2, substitution of P793S in 2, substitution of K188E in 2, substitution of QT945KI in 2, substitution of T620P in 2, substitution of T946 in 2 Substitution of P, substitution of TT949PP in SEQ ID NO: 2, substitution of N952T in SEQ ID NO: 2, substitution of K682E in SEQ ID NO: 2, substitution of K975R in SEQ ID NO: 2, substitution of L212P in SEQ ID NO: 2, substitution of E292R in SEQ ID NO: 2, substitution of I303K in SEQ ID NO: 2, substitution of C349E in SEQ ID NO: 2, substitution of E385P in SEQ ID NO: 2, substitution of E386N in SEQ ID NO: 2, substitution of D387K in SEQ ID NO: 2, substitution of L404K in SEQ ID NO: 2, substitution of E466H in SEQ ID NO: 2, substitution of C477Q in SEQ ID NO: 2, substitution of C477H in SEQ ID NO: 2, substitution of C479A in SEQ ID NO: 2,Substitution of D659H in SEQ ID NO: 2, substitution of T806V in SEQ ID NO: 2, substitution of K808S in SEQ ID NO: 2, insertion of AS at position 797 in SEQ ID NO: 2, substitution of V959M in SEQ ID NO: 2, substitution of K975Q in SEQ ID NO: 2, substitution of W974G in SEQ ID NO: 2, substitution of A708Q in SEQ ID NO: 2, substitution of V711K in SEQ ID NO: 2, substitution of D733T in SEQ ID NO: 2, substitution of L742W in SEQ ID NO: 2, substitution of V747K in SEQ ID NO: 2, substitution of F755M in SEQ ID NO: 2, substitution of M771A in SEQ ID NO: 2, substitution of M771Q in SEQ ID NO: 2, substitution of W782Q in SEQ ID NO: 2, sequence Substitution of G791F in sequence number 2, substitution of L792D in sequence number 2, substitution of L792K in sequence number 2, substitution of P793Q in sequence number 2, substitution of P793G in sequence number 2, substitution of Q804A in sequence number 2, substitution of Y966N in sequence number 2, substitution of Y723N in sequence number 2, substitution of Y857R in sequence number 2, substitution of S890R in sequence number 2, substitution of S932M in sequence number 2, substitution of L897M in sequence number 2, substitution of R624G in sequence number 2, substitution of S603G in sequence number 2, substitution of N737S in sequence number 2, substitution of L307K in sequence number 2, substitution of I6 in sequence number 2 Substitution of 58V, insertion of PT at position 688 of SEQ ID NO: 2, insertion of SA at position 794 of SEQ ID NO: 2, substitution of S877R in SEQ ID NO: 2, substitution of N580T in SEQ ID NO: 2, substitution of V335G in SEQ ID NO: 2, substitution of T620S in SEQ ID NO: 2, substitution of W345G in SEQ ID NO: 2, substitution of T280S in SEQ ID NO: 2, substitution of L406P in SEQ ID NO: 2, substitution of A612D in SEQ ID NO: 2, substitution of A751S in SEQ ID NO: 2, substitution of E386R in SEQ ID NO: 2, substitution of V351M in SEQ ID NO: 2, substitution of K210N in SEQ ID NO: 2, substitution of D40A in SEQ ID NO: 2, substitution of E77 This includes substitutions of 3G, substitution of H207L in SEQ ID NO: 2, substitution of T62A in SEQ ID NO: 2, substitution of T287P in SEQ ID NO: 2, substitution of T832A in SEQ ID NO: 2, substitution of A893S in SEQ ID NO: 2, insertion of V at position 14 in SEQ ID NO: 2, insertion of AG at position 13 in SEQ ID NO: 2, substitution of R11V in SEQ ID NO: 2, substitution of R12N in SEQ ID NO: 2, substitution of R13H in SEQ ID NO: 2, insertion of Y at position 13 in SEQ ID NO: 2, substitution of R12L in SEQ ID NO: 2, insertion of Q at position 13 in SEQ ID NO: 2, substitution of V15S in SEQ ID NO: 2, insertion of D at position 17 in SEQ ID NO: 2, or combinations thereof.

[0161] In some embodiments, the CasX variant includes at least one modification in the NTSB domain.

[0162] In some embodiments, the CasX variant includes at least one modification in the TSL domain. In some embodiments, the at least one modification in the TSL domain includes one or more amino acid substitutions of amino acids Y857, S890, or S932 in SEQ ID NO: 2.

[0163] In some embodiments, the CasX variant includes at least one modification in the helical I domain. In some embodiments, the at least one modification in the helical I domain includes one or more amino acid substitutions from among the amino acids S219, L249, E259, Q252, E292, L307, or D318 of SEQ ID NO: 2.

[0164] In some embodiments, the CasX variant includes at least one modification in the helical II domain. In some embodiments, the at least one modification in the helical II domain includes one or more amino acid substitutions from among amino acids D361, L379, E385, E386, D387, F399, L404, R458, C477, or D489 of SEQ ID NO: 2.

[0165] In some embodiments, the CasX variant includes at least one modification in the OBD domain. In some embodiments, the at least one modification in the OBD includes one or more amino acid substitutions from among amino acids F536, E552, T620, or I658 of SEQ ID NO: 2.

[0166] In some embodiments, the CasX variant includes at least one modification in the RuvC DNA cleavage domain. In some embodiments, the at least one modification in the RuvC DNA cleavage domain includes one or more amino acid substitutions or deletions of amino acid P793 from among the amino acids K682, G695, A708, V711, D732, A739, D733, L742, V747, F755, M771, M779, W782, A788, G791, L792, P793, Y797, M799, Q804, S819, or Y857 of SEQ ID NO: 2.

[0167] In some embodiments, the CasX variant comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, and is selected from one or more of the following: (a) L379R amino acid substitution, (b) A708K amino acid substitution, (c) T620P amino acid substitution, (d) E385P amino acid substitution, (e) Y857R amino acid substitution, (f) I658V amino acid substitution, (g) F399L amino acid substitution, (h) Q252K amino acid substitution, (i) L404K amino acid substitution, and (j) P793 amino acid deletion.

[0168] In some embodiments, the CasX variant is a substitution of Y789T in SEQ ID NO: 2, a deletion of P793 in SEQ ID NO: 2, a substitution of Y789D in SEQ ID NO: 2, a substitution of T72S in SEQ ID NO: 2, a substitution of I546V in SEQ ID NO: 2, a substitution of E552A in SEQ ID NO: 2, a substitution of A636D in SEQ ID NO: 2, a substitution of F536S in SEQ ID NO: 2, a substitution of A708K in SEQ ID NO: 2, a substitution of Y797L in SEQ ID NO: 2, a substitution of L792G in SEQ ID NO: 2, a substitution of A739V in SEQ ID NO: 2, a substitution of G791M in SEQ ID NO: 2, an insertion of A at position 661 in SEQ ID NO: 2, a substitution of A788W in SEQ ID NO: 2, an array Substitution of K390R in sequence number 2, substitution of A751S in sequence number 2, substitution of E385A in sequence number 2, insertion of P at position 696 in sequence number 2, insertion of M at position 773 in sequence number 2, substitution of G695H in sequence number 2, insertion of AS at position 793 in sequence number 2, insertion of AS at position 795 in sequence number 2, substitution of C477R in sequence number 2, substitution of C477K in sequence number 2, substitution of C479A in sequence number 2, substitution of C479L in sequence number 2, substitution of I55F in sequence number 2, substitution of K210R in sequence number 2, substitution of C233S in sequence number 2, substitution of D231N in sequence number 2, sequence Substitution of Q338E in sequence number 2, substitution of Q338R in sequence number 2, substitution of L379R in sequence number 2, substitution of K390R in sequence number 2, substitution of L481Q in sequence number 2, substitution of F495S in sequence number 2, substitution of D600N in sequence number 2, substitution of T886K in sequence number 2, substitution of A739V in sequence number 2, substitution of K460N in sequence number 2, substitution of I199F in sequence number 2, substitution of G492P in sequence number 2, substitution of T153I in sequence number 2, substitution of R591I in sequence number 2, insertion of AS at position 795 in sequence number 2, insertion of AS at position 796 in sequence number 2 Insertion of L at position 889, substitution of E121D in SEQ ID NO: 2, substitution of S270W in SEQ ID NO: 2, substitution of E712Q in SEQ ID NO: 2, substitution of K942Q in SEQ ID NO: 2, substitution of E552K in SEQ ID NO: 2, substitution of K25Q in SEQ ID NO: 2, substitution of N47D in SEQ ID NO: 2, insertion of T at position 696 in SEQ ID NO: 2, substitution of L685I in SEQ ID NO: 2, substitution of N880D in SEQ ID NO: 2, substitution of Q102R in SEQ ID NO: 2, substitution of M734K in SEQ ID NO: 2, substitution of A724S in SEQ ID NO: 2, substitution of T704K in SEQ ID NO: 2, substitution of P224K in SEQ ID NO: 2, substitution of K25R in SEQ ID NO: 2,Substitution of M29E in SEQ ID NO: 2, H152D in SEQ ID NO: 2, S219R in SEQ ID NO: 2, E475K in SEQ ID NO: 2, G226R in SEQ ID NO: 2, A377K in SEQ ID NO: 2, E480K in SEQ ID NO: 2, K416E in SEQ ID NO: 2, H164R in SEQ ID NO: 2, K767R in SEQ ID NO: 2, I7F in SEQ ID NO: 2, M29R in SEQ ID NO: 2, H435R in SEQ ID NO: 2, E385Q in SEQ ID NO: 2, E385K in SEQ ID NO: 2, I279F in SEQ ID NO: 2, D489S in SEQ ID NO: 2 Substitution of, substitution of D732N in SEQ ID NO: 2, substitution of A739T in SEQ ID NO: 2, substitution of W885R in SEQ ID NO: 2, substitution of E53K in SEQ ID NO: 2, substitution of A238T in SEQ ID NO: 2, substitution of P283Q in SEQ ID NO: 2, substitution of E292K in SEQ ID NO: 2, substitution of Q628E in SEQ ID NO: 2, substitution of R388Q in SEQ ID NO: 2, substitution of G791M in SEQ ID NO: 2, substitution of L792K in SEQ ID NO: 2, substitution of L792E in SEQ ID NO: 2, substitution of M779N in SEQ ID NO: 2, substitution of G27D in SEQ ID NO: 2, substitution of K955R in SEQ ID NO: 2, substitution of S867R in SEQ ID NO: 2 Substitution of R693I, substitution of F189Y in SEQ ID NO: 2, substitution of V635M in SEQ ID NO: 2, substitution of F399L in SEQ ID NO: 2, substitution of E498K in SEQ ID NO: 2, substitution of E386R in SEQ ID NO: 2, substitution of V254G in SEQ ID NO: 2, substitution of P793S in SEQ ID NO: 2, substitution of K188E in SEQ ID NO: 2, substitution of QT945KI in SEQ ID NO: 2, substitution of T620P in SEQ ID NO: 2, substitution of T946P in SEQ ID NO: 2, substitution of TT949PP in SEQ ID NO: 2, substitution of N952T in SEQ ID NO: 2, substitution of K682E in SEQ ID NO: 2, substitution of K975R in SEQ ID NO: 2, substitution of L212 in SEQ ID NO: 2 Substitution of P, substitution of E292R in SEQ ID NO: 2, substitution of I303K in SEQ ID NO: 2, substitution of C349E in SEQ ID NO: 2, substitution of E385P in SEQ ID NO: 2, substitution of E386N in SEQ ID NO: 2, substitution of D387K in SEQ ID NO: 2, substitution of L404K in SEQ ID NO: 2, substitution of E466H in SEQ ID NO: 2, substitution of C477Q in SEQ ID NO: 2, substitution of C477H in SEQ ID NO: 2, substitution of C479A in SEQ ID NO: 2, substitution of D659H in SEQ ID NO: 2, substitution of T806V in SEQ ID NO: 2, substitution of K808S in SEQ ID NO: 2, insertion of AS at position 797 in SEQ ID NO: 2, substitution of V959M in SEQ ID NO: 2,Substitution of K975Q in SEQ ID NO: 2, W974G in SEQ ID NO: 2, A708Q in SEQ ID NO: 2, V711K in SEQ ID NO: 2, D733T in SEQ ID NO: 2, L742W in SEQ ID NO: 2, V747K in SEQ ID NO: 2, F755M in SEQ ID NO: 2, M771A in SEQ ID NO: 2, M771Q in SEQ ID NO: 2, W782Q in SEQ ID NO: 2, G791F in SEQ ID NO: 2, L792D in SEQ ID NO: 2, L792K in SEQ ID NO: 2, P793Q in SEQ ID NO: 2, P793G in SEQ ID NO: 2 Substitution of, substitution of Q804A in SEQ ID NO: 2, substitution of Y966N in SEQ ID NO: 2, substitution of Y723N in SEQ ID NO: 2, substitution of Y857R in SEQ ID NO: 2, substitution of S890R in SEQ ID NO: 2, substitution of S932M in SEQ ID NO: 2, substitution of L897M in SEQ ID NO: 2, substitution of R624G in SEQ ID NO: 2, substitution of S603G in SEQ ID NO: 2, substitution of N737S in SEQ ID NO: 2, substitution of L307K in SEQ ID NO: 2, substitution of I658V in SEQ ID NO: 2, insertion of PT at position 688 in SEQ ID NO: 2, insertion of SA at position 794 in SEQ ID NO: 2, substitution of S877R in SEQ ID NO: 2, distribution Substitution of N580T in column number 2, substitution of V335G in sequence number 2, substitution of T620S in sequence number 2, substitution of W345G in sequence number 2, substitution of T280S in sequence number 2, substitution of L406P in sequence number 2, substitution of A612D in sequence number 2, substitution of A751S in sequence number 2, substitution of E386R in sequence number 2, substitution of V351M in sequence number 2, substitution of K210N in sequence number 2, substitution of D40A in sequence number 2, substitution of E773G in sequence number 2, substitution of H207L in sequence number 2, substitution of T62A in sequence number 2, substitution of T287P in sequence number 2, The at least two amino acid changes to the sequence of a reference CasX variant protein are selected from the group consisting of substitutions of T832A in SEQ ID NO: 2, substitution of A893S in SEQ ID NO: 2, insertion of V at position 14 in SEQ ID NO: 2, insertion of AG at position 13 in SEQ ID NO: 2, substitution of R11V in SEQ ID NO: 2, substitution of R12N in SEQ ID NO: 2, substitution of R13H in SEQ ID NO: 2, insertion of Y at position 13 in SEQ ID NO: 2, substitution of R12L in SEQ ID NO: 2, insertion of Q at position 13 in SEQ ID NO: 2, substitution of V15S in SEQ ID NO: 2, and insertion of D at position 17 in SEQ ID NO: 2. In some embodiments, the at least two amino acid changes to the reference CasX protein are:The amino acid variations are selected from those disclosed in the sequences of SEQ ID NOs. 49-150 shown in Table 4. In some embodiments, the CasX variant includes any combination of the embodiments described above in this paragraph.

[0169] In some embodiments, the CasX variant protein includes two or more substitutions, insertions, and / or deletions of the amino acid sequence of the reference CasX protein. In some embodiments, the CasX variant protein includes the substitution of S794R and Y797L in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of K416E and A708K in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of A708K and the deletion of P793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the deletion of P793 and the insertion of AS at position 795 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of Q367K and the substitution of I425S in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of A708K, the deletion of P at position 793, and the substitution of A793V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of Q338R and A339E in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of Q338R and A339K in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of S507G and G508R in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of L379R, A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of C477K, A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of L379R, C477K, A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of C477K, A708K, P deletion at position 793, and A739V in SEQ ID NO: 2.In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and M779N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and M771N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, 708K, P deletion at position 793, and D489S in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and A739T in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and D732N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and G791M in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, 708K, P deletion at position 793, and Y797L in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of L379R, C477K, A708K, P deletion at position 793, and M779N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of L379R, C477K, A708K, P deletion at position 793, and M771N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of L379R, C477K, A708K, P deletion at position 793, and D489S in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes substitutions of L379R, C477K, A708K, P deletion at position 793, and A739T in SEQ ID NO: 2.In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and D732N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and G791M in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and Y797L in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and T620P in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of A708K, the deletion of P at position 793, and the substitution of E386S in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of E386R, the substitution of F399L, and the deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitutions of R581I and A739V in SEQ ID NO: 2. In some embodiments, the CasX variant includes any combination of the embodiments described above in this paragraph.

[0170] In some embodiments, the CasX variant protein includes two or more substitutions, insertions, and / or deletions of the amino acid sequence of the reference CasX protein. In some embodiments, the CasX variant protein includes a substitution of A708K, a deletion of P at position 793, and a substitution of A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of L379R, a substitution of A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of L379R, a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of C477K, A708K, P deletion at position 793, and A739 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and T620P in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of M771A in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and D732N in SEQ ID NO: 2. In some embodiments, the CasX variant includes any combination of the embodiments described above in this paragraph.

[0171] In some embodiments, the CasX variant protein includes the substitution of W782Q in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of M771Q in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of R458I and A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and M771N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, P deletion at position 793, and A739T in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and D489S in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and D732N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of V711K in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and Y797L in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, A708K, and P deletion at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, C477K, A708K, P deletion at position 793, and M771N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of A708K, a substitution of P at position 793, and a substitution of E386S in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of L379R, a substitution of C477K, a substitution of A708K, and a deletion of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of L792D in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes a substitution of G791F in SEQ ID NO: 2.In some embodiments, the CasX variant protein includes the substitution of A708K, the deletion of P at position 793, and the substitution of A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, the substitution of A708K, the deletion of P at position 793, and the substitution of A739V in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of C477K, the substitution of A708K, and the substitution of P at position 793 in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L249I and the substitution of M771N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of V747K in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of L379R, the substitution of C477, the substitution of A708K, the deletion of P at position 793, and the substitution of M779N in SEQ ID NO: 2. In some embodiments, the CasX variant protein includes the substitution of F755M. In some embodiments, the CasX variant includes any combination of the embodiments described above in this paragraph.

[0172] In some embodiments, the CasX variant protein comprises at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, the at least one modification being selected from one or more of the following: L379R amino acid substitution, A708K amino acid substitution, T620P amino acid substitution, E385P amino acid substitution, Y857R amino acid substitution, I658V amino acid substitution, F399L amino acid substitution, Q252K amino acid substitution, and [P793] amino acid deletion. In some embodiments, the CasX variant protein includes at least one modification compared to the reference CasX sequence of SEQ ID NO: 2, the at least one modification being selected from one or more of the following: L379R amino acid substitution, A708K amino acid substitution, T620P amino acid substitution, E385P amino acid substitution, Y857R amino acid substitution, I658V amino acid substitution, F399L amino acid substitution, Q252K amino acid substitution, L404K amino acid substitution, and [P793] amino acid deletion. In other embodiments, the CasX variant protein includes any combination of the aforementioned substitutions or deletions compared to the reference CasX sequence of SEQ ID NO: 2. In other embodiments, the CasX variant protein may further include, in addition to the aforementioned substitutions or deletions, a substitution of the helical 1b domain derived from the NTSB and / or the reference CasX sequence of SEQ ID NO: 1.

[0173] In some embodiments, the CasX variant protein contains 400 to 2000 amino acids, 500 to 1500 amino acids, 700 to 1200 amino acids, 800 to 1100 amino acids, or 900 to 1000 amino acids.

[0174] In some embodiments, the CasX variant protein includes one or more modifications in a region of non-adjacent residues that form a channel where gNA:target DNA complex formation occurs. In some embodiments, the CasX variant protein includes one or more modifications in a region of non-adjacent residues that form an interface for binding to gNA. For example, in some embodiments of the reference CasX protein, the helical I, helical II, and OBD domains are all in contact with or adjacent to the gNA:target DNA complex, and one or more modifications to non-adjacent residues within any of these domains can improve the function of the CasX variant protein.

[0175] In some embodiments, the CasX variant protein includes one or more modifications in a region of non-adjacent residues that form a channel for binding to non-target strand DNA. For example, the CasX variant protein may include one or more modifications to non-adjacent residues of the NTSBD. In some embodiments, the CasX variant protein includes one or more modifications in a region of non-adjacent residues that form an interface for binding to the PAM. For example, the CasX variant protein may include one or more modifications to non-adjacent residues of the helical I domain or OBD. In some embodiments, the CasX variant protein includes one or more modifications including a region of non-adjacent surface-exposed residues. As used herein, “surface-exposed residue” means an amino acid on the surface of the CasX protein, or at least a portion of an amino acid, e.g., an amino acid in which part of the backbone or side chain is on the surface of the protein. Surface-exposed residues of cellular proteins such as CasX that are exposed to an aqueous intracellular environment are often selected from positively charged hydrophilic amino acids, such as arginine, asparagine, aspartic acid, glutamine, glutamic acid, histidine, lysine, serine, and threonine. Accordingly, for example, in some embodiments of the variants provided herein, the region of surface-exposed residues includes one or more insertions, deletions, or substitutions compared to the reference CasX protein. In some embodiments, one or more positively charged residues are replaced by one or more other positively charged residues, or negatively charged residues, or uncharged residues, or any combination thereof. In some embodiments, one or more amino acid residues for substitution are nearby bound nucleic acids, for example, residues in the RuvC domain or helical I domain that contact target DNA, or residues in the OBD or helical II domain that bind to gNA, and these may be replaced by one or more positively charged or polar amino acids.

[0176] In some embodiments, the CasX variant protein comprises one or more modifications in a region of non - adjacent residues that form a core by hydrophobic packing in a domain of the reference CasX protein. Without wishing to be bound by any theory, regions that form a core by hydrophobic packing are rich in hydrophobic amino acids such as valine, isoleucine, leucine, methionine, phenylalanine, tryptophan, and cysteine. For example, in some reference CasX proteins, the RuvC domain contains a hydrophobic pocket adjacent to the active site. In some embodiments, 2 - 15 residues of the region are charged, polar, or base - stacking. Charged amino acids (sometimes referred to herein as residues) can include, for example, arginine, lysine, aspartic acid, and glutamic acid, and the side chains of these amino acids can form salt bridges, provided that a cross - linking partner is also present. Polar amino acids can include, for example, glutamine, asparagine, histidine, serine, threonine, tyrosine, and cysteine. Polar amino acids can form hydrogen bonds as proton donors or acceptors, depending on the identity of their side chains, in some embodiments. As used herein, "base - stacking" includes the interaction of aromatic side chains of amino acid residues (such as tryptophan, tyrosine, phenylalanine, or histidine) with stacked nucleotide bases within a nucleic acid. Any modification to a region of non - adjacent amino acids that are spatially very close to form a functional part of the CasX variant protein is envisioned to be within the scope of the present disclosure.

[0177] <照 i. A CasX variant protein having domains from multiple source proteins In certain embodiments, the Disclosure provides a chimeric CasX protein comprising protein domains derived from two or more different CasX proteins, for example, two or more reference CasX proteins, or two or more CasX variant protein sequences described herein. As used herein, “chimeric CasX protein” means a CasX comprising at least two domains isolated from or derived from two different sources, such as two naturally occurring proteins, which in some embodiments may be isolated from different species. For example, in some embodiments, a chimeric CasX protein comprises a first domain derived from a first CasX protein and a second domain derived from a second different CasX protein. In some embodiments, the first domain may be selected from the group consisting of NTSB, TSL, helical I, helical II, OBD, and RuvC domains. In some embodiments, the second domain is selected from the group consisting of NTSB, TSL, helical I, helical II, OBD, and RuvC domains, and the second domain is different from the first domain described herein. For example, a chimeric CasX protein may contain the NTSB, TSL, helical I, helical II, and OBD domains derived from the CasX protein of SEQ ID NO: 2, and the RuvC domain derived from the CasX protein of SEQ ID NO: 1, or vice versa. As a further example, a chimeric CasX protein may contain the NTSB, TSL, helical II, OBD, and RuvC domains derived from the CasX protein of SEQ ID NO: 2, and the helical I domain derived from the CasX protein of SEQ ID NO: 1, or vice versa. Therefore, in certain embodiments, a chimeric CasX protein may contain the NTSB, TSL, helical II, OBD, and RuvC domains derived from a first CasX protein and the helical I domain derived from a second CasX protein. In some embodiments of the chimeric CasX protein, the domain of the first CasX protein is derived from the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, and the domain of the second CasX protein is derived from the sequence of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, and the first and second CasX proteins are not the same.In some embodiments, the domain of the first CasX protein includes a sequence derived from SEQ ID NO: 1, and the domain of the second CasX protein includes a sequence derived from SEQ ID NO: 2. In some embodiments, the domain of the first CasX protein includes a sequence derived from SEQ ID NO: 1, and the domain of the second CasX protein includes a sequence derived from SEQ ID NO: 3. In some embodiments, the domain of the first CasX protein includes a sequence derived from SEQ ID NO: 2, and the domain of the second CasX protein includes a sequence derived from SEQ ID NO: 3. In some embodiments, the CasX variant includes SEQ ID NOs: 130-138 or 141-144, the sequences of which are shown in Table 4. In some embodiments, the CasX variant includes the sequences of SEQ ID NOs: 72, 94, 113, 135, 138, 144, 239, 277, or 280. In some embodiments, the CasX variant includes the sequences of SEQ ID NOs: 94, 72, 138, 144, or 280. In some embodiments, the CasX variant protein comprises at least one chimeric domain comprising a first portion derived from a first CasX protein and a second portion derived from a second different CasX protein. As used herein, “chimeric domain” means a domain comprising at least two portions isolated from or derived from different sources, e.g., two naturally occurring proteins, or a portion of a domain derived from two reference CasX proteins. The at least one chimeric domain may be any of the NTSB, TSL, helical I, helical II, OBD, or RuvC domains described herein. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 1, and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 2. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 1, and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 3. In some embodiments, the first portion of the CasX domain comprises the sequence of SEQ ID NO: 2, and the second portion of the CasX domain comprises the sequence of SEQ ID NO: 3. In some embodiments, at least one chimeric domain includes a chimeric RuvC domain.As an example, the chimeric RuvC domain includes amino acids 661-824 of SEQ ID NO: 1 and amino acids 922-978 of SEQ ID NO: 2. As an alternative example, the chimeric RuvC domain includes amino acids 648-812 of SEQ ID NO: 2 and amino acids 935-986 of SEQ ID NO: 1. In some embodiments, the CasX protein includes a first domain derived from a first CasX protein, a second domain derived from a second CasX protein, and at least one chimeric domain comprising at least two portions isolated from different CasX proteins using the approaches of the embodiments described in this paragraph. In the embodiments described above, the chimeric CasX protein having a domain or portion of a domain derived from SEQ ID NOs: 1, 2, and 3 may further include amino acid insertions, deletions, or substitutions of any of the embodiments disclosed herein.

[0178] In some embodiments, the CasX variant protein comprises the sequences shown in Table 4, Table 6, Table 7, Table 8, or Table 10. In some embodiments, the CasX variant protein consists of the sequences shown in Table 4. In other embodiments, the CasX variant protein has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 86%, at least 87%, at least 88%, at least 89%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the sequences of SEQ ID NOs: 49-150, 233-235, 238-252, or 272-281 shown in Table 4, 6, 7, 8, or 10. In other embodiments, the CasX variant protein comprises the sequences of SEQ ID NOs: 49-150 shown in Table 4 and further comprises one or more of the NLSs disclosed herein at or near either or both of the N-terminus, C-terminus, or both. In some cases, it will be understood that the N-terminal methionine of these table's CasX variants is removed from the CasX variants expressed during post-translational modification.

Table 4-1

Table 4-2

Table 4-3

Table 4-4

Table 4-5

Table 4-6

[0179] In some embodiments, the CasX variant protein contains a sequence selected from the group consisting of SEQ ID NOs: 49-150, 233-235, 238-252, and 272-281.

[0180] In some embodiments, the CasX variant protein has one or more improved properties of a reference CasX protein, such as the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, at least one improved property of the CasX variant is improved by at least about 1.1 to about 100,000 times compared to the reference protein. In some embodiments, at least one improved property of the CasX variant is improved by at least about 1.1 to about 10,000 times compared to the reference CasX protein, at least about 1.1 to about 1,000 times, at least about 1.1 to about 500 times, at least about 1.1 to about 400 times, at least about 1.1 to about 300 times, at least about 1.1 to about 200 times, at least about 1.1 to about 100 times, at least about 1.1 to about 50 times, at least about 1.1 to about 40 times, at least about 1.1 to about 30 times, at least about 1.1 to about 20 times, at least about 1.1 to about 10 times, at least about 1.1 to about 9 times, The improvement is at least approximately 1.1 to 8 times, at least approximately 1.1 to 7 times, at least approximately 1.1 to 6 times, at least approximately 1.1 to 5 times, at least approximately 1.1 to 4 times, at least approximately 1.1 to 3 times, at least approximately 1.1 to 2 times, at least approximately 1.1 to 1.5 times, at least approximately 1.5 to 3 times, at least approximately 1.5 to 4 times, at least approximately 1.5 to 5 times, at least approximately 1.5 to 10 times, at least approximately 5 to 10 times, at least approximately 10 to 20 times, at least 10 to 30 times, at least 10 to 50 times, or at least 10 to 100 times. In some embodiments, at least one improved property of the CasX variant is improved at least approximately 10 to 1000 times compared to the reference CasX protein.

[0181] In some embodiments, one or more improved properties of the CasX variant protein are improved by at least about 1.1, at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1,000, at least about 5,000, at least about 10,000, or at least about 100,000 times compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In other cases, one or more improved properties of a CasX variant are approximately 1.1 to 100.00 times, approximately 1.1 to 10.00 times, approximately 1.1 to 1,000 times, approximately 1.1 to 500 times, approximately 1.1 to 100 times, approximately 1.1 to 50 times, approximately 1.1 to 20 times, approximately 10 to 100.00 times, approximately 10 to 10.00 times, approximately 10 to 1,000 times, approximately 10 to 500 times, approximately 10 to 100 times, approximately 10 to 50 times, approximately 10 to 20 times, approximately 2 to 70 times, approximately 2 to 50 times, approximately 2 to 30 times, approximately 2 to 20 times, approximately 2 to 10 times, approximately 5 to Improvements of 50 times, approximately 5-30 times, approximately 5-10 times, approximately 100-100.00 times, approximately 100-10.00 times, approximately 100-1,000 times, approximately 100-500 times, approximately 500-100.00 times, approximately 500-10.00 times, approximately 500-750 times, approximately 1,000-100.00 times, approximately 10,000-100.00 times, approximately 20-500 times, approximately 20-250 times, approximately 20-200 times, approximately 20-100 times, approximately 20-50 times, approximately 50-10,000 times, approximately 50-1,000 times, approximately 50-500 times, approximately 50-200 times, or approximately 50-100 times.

[0182] Exemplary properties that may be improved in CasX variant proteins compared to the same properties in the reference CasX protein include, but are not limited to, improved folding of the variant, improved binding affinity to gNA, improved binding affinity to a broader spectrum of PAM sequences, improved unwinding of target DNA, increased activity, improved editing efficiency, improved editing specificity, increased nuclease activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target breaks, improved binding of non-target strands of DNA, improved protein stability, improved protein:gNA complex stability, improved protein solubility, improved protein:gNA complex solubility, improved protein yield, improved protein expression, and improved fusion properties. In some embodiments, the variant includes at least one improved property. In other embodiments, the variant includes at least two improved properties. In further embodiments, the variant includes at least three improved properties. In some embodiments, the variant includes at least four improved properties. In yet another embodiment, the variant includes at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, or more improved characteristics. These improved characteristics are described in more detail below.

[0183] j. Protein stability In some embodiments, the disclosure provides CasX variant proteins having improved stability relative to a reference CasX protein. In some embodiments, the improved stability of the CasX variant protein results in a higher steady-state expression of the protein, thereby improving editing efficiency. In some embodiments, the improved stability of the CasX variant protein results in a higher percentage of CasX protein remaining folded in a functional conformation and improving editing efficiency or purification for manufacturing purposes. As used herein, “functional conformation” refers to a CasX protein in a conformation in which the protein can bind to gNA and target DNA. In embodiments in which the CasX variant does not have one or more mutations that catalytically kill the CasX variant, the CasX variant can cleave, nicking, or otherwise modify the target DNA. For example, a functional CasX variant can be used for gene editing in some embodiments, and a functional conformation refers to an “editing-competent” conformation. In some exemplary embodiments, including those embodiments that result in a greater proportion of CasX protein remaining folded in a functional conformation, lower concentrations of the CasX variant compared to the reference CasX protein are required for applications such as gene editing. Therefore, in some embodiments, the CasX variant with improved stability has improved efficiency compared to the reference CasX under one or more gene editing conditions.

[0184] In some embodiments, the disclosure provides CasX variant proteins having improved thermal stability compared to a reference CasX protein. In some embodiments, the CasX variant protein has improved thermal stability over a specific temperature range. While we do not wish to be bound by any theory, some reference CasX proteins function intrinsically in organisms with ecological niches in groundwater and sediments, and therefore some reference CasX proteins may have evolved to function optimally at lower or higher temperatures that may be desirable for certain applications. For example, one application of a CasX variant protein is gene editing in mammalian cells, which is typically performed at around 37°C. In some embodiments, the CasX variant proteins described herein have improved thermal stability compared to the reference CasX protein at temperatures of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher. In some embodiments, the CasX variant proteins have improved thermal stability and functionality compared to the reference CasX protein, resulting in improved gene editing functionality, such as mammalian gene editing applications, which may include human gene editing applications. The improved thermal stability of the nuclease can be evaluated by various methods known to those skilled in the art.

[0185] In some embodiments, the disclosure provides CasX variant proteins having improved stability of the CasX variant protein:gNA complex relative to the reference CasX protein:gNA complex, such that the RNP remains in a functional form. Improved stability may include increased thermal stability, resistance to proteolysis, enhanced pharmacokinetic properties, stability across various pH and salt conditions, and isotonicity. The improved stability of the complex may result in improved editing efficiency in some embodiments. In some embodiments, the RNP of the CasX variant and the gNA variant has at least 5%, at least 10%, at least 15%, or at least 20%, or at least 5–20%, higher percentages of cleavage-competent RNP compared to the RNP of the reference CasX (SEQ ID NOs: 1–3) and any one of the gNAs (SEQ ID NOs: 4–16) in Table 1. Exemplary data of increased cleavage-competent RNP are provided in the examples.

[0186] In some embodiments, the disclosure provides a CasX variant protein having improved thermal stability of the CasX variant protein:gNA complex relative to a reference CasX protein:gNA complex. In some embodiments, the CasX variant protein has improved thermal stability compared to the reference CasX protein. In some embodiments, the CasX variant protein:gNA complex has improved thermal stability relative to the complex containing the reference CasX protein at temperatures of at least 16°C, at least 18°C, at least 20°C, at least 22°C, at least 24°C, at least 26°C, at least 28°C, at least 30°C, at least 32°C, at least 34°C, at least 35°C, at least 36°C, at least 37°C, at least 38°C, at least 39°C, at least 40°C, at least 41°C, at least 42°C, at least 44°C, at least 46°C, at least 48°C, at least 50°C, at least 52°C, or higher. In some embodiments, the CasX variant protein exhibits improved thermal stability of the CasX variant protein:gNA complex compared to the reference CasX protein:gNA complex, thereby resulting in improved functionality for gene editing applications, such as mammalian gene editing applications, which may include human gene editing applications. The improved thermal stability of the RNP can be evaluated by various methods known to those skilled in the art.

[0187] In some embodiments, the improved stability and / or thermal stability of the CasX variant protein includes faster folding kinetics of the CasX variant protein compared to the reference CasX protein, slower unfolding kinetics of the CasX variant protein compared to the reference CasX protein, greater free energy release during folding of the CasX variant protein compared to the reference CasX protein, a higher temperature (Tm) compared to the reference CasX protein at which 50% of the CasX variant protein is unfolded, or any combination thereof. These properties can be improved by a wide range of values, for example, by at least 1.1, at least 1.5, at least 10, at least 50, at least 100, at least 500, at least 1,000, at least 5,000, or at least 10,000 times compared to the reference CasX protein. In some embodiments, the improved thermal stability of the CasX variant protein includes a higher Tm of the CasX variant protein compared to the reference CasX protein. In some embodiments, the Tm of the CasX variant protein is approximately 20°C to 30°C, 30°C to 40°C, 40°C to 50°C, 50°C to 60°C, 60°C to 70°C, 70°C to 80°C, 80°C to 90°C, or 90°C to 100°C. Thermal stability is defined as the "melting temperature" (Tm), which is the temperature at which half of the molecule denatures. mTm is determined by measuring the free energy of unfolding. Methods for measuring protein stability properties such as Tm and unfolding free energy are known to those skilled in the art and can be measured in vitro using standard biochemical techniques. For example, Tm can be measured using differential scanning calorimetry, a thermal analysis technique in which the difference in the amount of heat required to increase the temperature of a sample and a reference is measured as a function of temperature (Chen et al (2003) Pharm Res 20:1952-60, Ghirlando et al (1999) Immunol Lett 68:47-52). Alternatively, or in addition, the Tm of CasX variant proteins can be measured using commercially available methods such as the ThermoFisher Protein Thermal Shift system. Alternatively, or in addition, circular dichroism can be used to measure folding and unfolding dynamics, as well as Tm (Murray et al. (2002) J. Chromatogr Sci 40:343-9). Circular dichroism (CD) depends on the uneven absorption of left-handed and right-handed circular polarization by asymmetric molecules such as proteins. Certain protein structures, such as alpha-helices and beta-sheets, have characteristic CD spectra. Therefore, in some embodiments, CD can be used to determine the secondary structure of CasX variant proteins.

[0188] In some embodiments, the improved stability and / or thermal stability of the CasX variant protein includes improved folding kinetics of the CasX variant protein compared to the reference CasX protein. In some embodiments, the folding kinetics of the CasX variant protein are improved by at least about 5, at least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 2,000, at least about 3,000, at least about 4,000, at least about 5,000, or at least about 10,000 times compared to the reference CasX protein. In some embodiments, the folding dynamics of the CasX variant protein are improved by at least about 1 kJ / mol, at least about 5 kJ / mol, at least about 10 kJ / mol, at least about 20 kJ / mol, at least about 30 kJ / mol, at least about 40 kJ / mol, at least about 50 kJ / mol, at least about 60 kJ / mol, at least about 70 kJ / mol, at least about 80 kJ / mol, at least about 90 kJ / mol, at least about 100 kJ / mol, at least about 150 kJ / mol, at least about 200 kJ / mol, at least about 250 kJ / mol, at least about 300 kJ / mol, at least about 350 kJ / mol, at least about 400 kJ / mol, at least about 450 kJ / mol, or at least about 500 kJ / mol compared to the reference CasX protein.

[0189] Exemplary amino acid changes that can increase the stability of a CasX variant protein relative to a reference CasX protein may include, but are not limited to, amino acid changes that increase the number of hydrogen bonds in the CasX variant protein, amino acid changes that increase the number of disulfide crosslinks in the CasX variant protein, amino acid changes that increase the number of salt bridges in the CasX variant protein, amino acid changes that enhance the interactions between segments of the CasX variant protein, amino acid changes that increase the embedded hydrophobic surface regions of the CasX variant protein, or any combination thereof.

[0190] k. Protein yield In some embodiments, the disclosure provides CasX variant proteins having improved expression and yield during purification compared to a reference CasX protein. In some embodiments, the yield of the CasX variant protein purified from bacterial or eukaryotic host cells is improved compared to the reference CasX protein. In some embodiments, the bacterial host cell is Escherichia coli cell. In some embodiments, the eukaryotic cell is yeast, plant (e.g., tobacco), insect (e.g., Spodoptera frugiperda sf9 cell), mouse, rat, hamster, guinea pig, monkey, or human cell. In some embodiments, eukaryotic host cells include, but are not limited to, human fetal kidney 293 (HEK293) cells, HEK292T cells, baby hamster kidney (BHK) cells, NS0 cells, SP2 / 0 cells, YO myeloma cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, NIH3T3 cells, COS, HeLa, or Chinese hamster ovary (CHO) cells.

[0191] In some embodiments, improved yields of CasX variant proteins are achieved by codon optimization. Cells use 64 different codons, 61 of which encode 20 standard amino acids, while another 3 function as stop codons. In some cases, a single amino acid is encoded by two or more codons. Different organisms tend to use different codons for the same naturally occurring amino acid. Therefore, the selection of codons in the protein coding sequence, and the selection of codons that match the organism in which the protein is expressed, can, in some cases, have a significant impact on protein translation and, consequently, protein expression levels. In some embodiments, CasX variant proteins are encoded by codon-optimized nucleic acids. In some embodiments, the nucleic acids encoding CasX variant proteins are codon-optimized for expression in bacterial cells, yeast cells, insect cells, plant cells, or mammalian cells. In some embodiments, mammalian cells are mouse, rat, hamster, guinea pig, monkey, or human. In some embodiments, CasX variant proteins are encoded by codon-optimized nucleic acids for expression in human cells. In some embodiments, CasX variant proteins are encoded by nucleic acids from which nucleotide sequences that reduce translation rates in prokaryotes and eukaryotes have been removed. For example, a sequence of three or more consecutive thymine residues may reduce translation rates in certain organisms, or internal polyadenylation signals may reduce translation rates.

[0192] In some embodiments, the improvements in solubility and stability described herein result in an improved yield of CasX variant protein relative to the reference CasX protein.

[0193] The improved protein yield during expression and purification can be evaluated by methods known in the art. For example, the amount of CasX variant protein can be determined by electrophoresis on an SDS-Page gel and by comparing the CasX variant protein to a control of known quantity or concentration to determine the absolute level of the protein. Alternatively, or in addition, the relative improvement in the yield of the CasX variant protein can be determined by electrophoresis of the purified CasX variant protein on an SDS-Page gel adjacent to a reference CasX protein undergoing the same purification process. Alternatively, or in addition, the level of the protein can be measured using immunohistochemical methods such as Western blotting or ELISA with an antibody against CasX, or by HPLC. For proteins in solution, the concentration can be determined by measuring the protein's intrinsic UV absorbance, or by methods using protein-dependent color changes such as the Lowry assay, Smith copper / bicinconine assay, or Bradford dye assay. Using such methods, the total protein yield (e.g., total soluble protein) obtained by expression under specific conditions can be calculated. This can be compared, for example, to the protein yield of a reference CasX protein under similar expression conditions.

[0194] l. Protein solubility In some embodiments, the CasX variant protein has improved solubility relative to the reference CasX protein. In some embodiments, the CasX variant protein has improved solubility of the CasX:gNA ribonucleoprotein complex variant relative to the ribonucleoprotein complex containing the reference CasX protein.

[0195] In some embodiments, improved protein solubility results in higher protein yields from protein purification techniques, such as purification from E. coli. Since higher protein solubility reduces the likelihood of aggregation within cells, in some embodiments, improved solubility of CasX variant proteins may enable more efficient activity within cells. While protein aggregates can be harmful or cumbersome to cells in certain embodiments, and we do not wish to be bound by any theory, increased solubility of CasX variant proteins may mitigate this outcome of protein aggregation. Furthermore, improved solubility of CasX variant proteins may enable formulation enhancement, for example, in desired gene editing applications, allowing for the delivery of higher effective amounts of functional proteins. In some embodiments, the improved solubility of the CasX variant protein relative to the reference CasX protein results in an improved yield of the CasX variant protein during purification, at least about 5, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 250, at least about 500, or at least about 1000 times.In some embodiments, the improved solubility of the CasX variant protein relative to the reference CasX protein results in the activity of the CasX variant protein in cells being at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 2.1, at least about 2.2, at least about 2.3, at least about 2.4, at least about 2.5, less The improvement is at least about 2.6, at least about 2.7, at least about 2.8, at least about 2.9, at least about 3, at least about 3.5, at least about 4, at least about 4.5, at least about 5, at least about 5.5, at least about 6, at least about 6.5, at least about 7.0, at least about 7.5, at least about 8, at least about 8.5, at least about 9, at least about 9.5, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, or at least about 15 times. The improved solubility of the nuclease can be evaluated by various methods known to those skilled in the art, including performing concentration measurement readings on a gel of the soluble fraction of dissolved E. coli. Alternatively, or in addition, the improvement in the solubility of the CasX variant protein can be measured by measuring the retention of the soluble protein product throughout the process of complete protein purification. For example, soluble protein products can be measured in one or more steps of gel affinity purification, tag cleavage, cation exchange purification, and protein electrophoresis on a size determination column. In some embodiments, concentration measurements of all bands of the protein on the gel are read after each step in the purification process. CasX variant proteins with improved solubility may maintain higher concentrations in one or more steps of the protein purification process compared to a reference CasX protein in some embodiments, while insoluble protein variants may be lost in one or more steps due to buffer exchange, filtration steps, interactions with the purification column, etc.

[0196] In some embodiments, improving the solubility of the CasX variant protein results in a higher yield in terms of mg / L of protein during protein purification compared to the reference CasX protein.

[0197] In some embodiments, improving the solubility of the CasX variant protein allows for a greater amount of editing events compared to less soluble proteins when evaluated in editing assays such as the EGFP disruption assay described herein.

[0198] Protein affinity for mgNA In some embodiments, the CasX variant protein has improved affinity for gNA relative to the reference CasX protein, resulting in the formation of a ribonucleoprotein complex. The increased affinity of the CasX variant protein for gNA leads to, for example, lower K for the formation of the RNP complex. d This can lead to increased affinity of the CasX variant protein to gNA, and in some cases, to more stable ribonucleoprotein complex formation. In some embodiments, increased affinity of the CasX variant protein to gNA results in increased stability of the ribonucleoprotein complex when delivered to human cells. This increased stability can not only affect the function and utility of the complex in target cells, but may also result in improved pharmacokinetic properties in the blood when delivered to the target. In some embodiments, increased affinity of the CasX variant protein, and the resulting increased stability of the ribonucleoprotein complex, allows for the delivery of lower doses of the CasX variant protein to a target or cell while still possessing the desired activity, e.g., in vivo or in vitro gene editing.

[0199] In some embodiments, a higher affinity (tighter binding) of the CasX variant protein to gNA allows for a greater number of editing events if both the CasX variant protein and gNA remain in the RNP complex. The increase in editing events can be evaluated using editing assays such as the EGFP disruption assay described herein.

[0200] In some embodiments, the K of the CasX variant protein relative to gNA d The affinity for gNA is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100 times compared to the reference CasX protein. In some embodiments, the CasX variant has a binding affinity for gNA that is increased by about 1.1 to about 10 times compared to the reference CasX protein of SEQ ID NO: 2.

[0201] While we do not wish to be constrained by theory, in some embodiments, amino acid changes in the helical I domain can increase the binding affinity of the CasX variant protein to the gNA targeting sequence, changes in the helical II domain can increase the binding affinity of the CasX variant protein to the gNA scaffold stem loop, and changes in the oligonucleotide binding domain (OBD) can increase the binding affinity of the CasX variant protein to the gRNA triplex.

[0202] Methods for measuring the binding affinity of CasX proteins to CasX gNAs include in vitro methods using purified CasX proteins and gNAs. When gNAs or CasX proteins are labeled with fluorophores, the binding affinity to reference CasX and variant proteins can be measured by fluorescence polarization. Alternatively, or in addition, binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or filter binding. Additional standard techniques for quantifying the absolute affinity of RNA-binding proteins, such as the reference CasX and variant proteins of this disclosure, to specific gNAs, such as reference gNAs and their variants, include, but are not limited to, isothermal calorimetry (ITC) and surface plasmon resonance (SPR), as well as the methods of the examples.

[0203] n. Affinity for target nucleic acids In some embodiments, CasX variant proteins have improved binding affinity to the target nucleic acid sequence compared to the affinity of the reference CasX protein to the target nucleic acid sequence. In some embodiments, improved affinity to the target nucleic acid sequence includes improved affinity to the target nucleic acid sequence, improved affinity to a broader spectrum of PAM sequences, improved ability to locate DNA for the target nucleic acid sequence, or any combination thereof. While we do not wish to be bound by theory, it is conceivable that CRISPR / Cas system proteins such as CasX can locate their target nucleic acid sequences by one-dimensional diffusion along the DNA molecule. This process is thought to involve (1) binding of the ribonucleoprotein to the DNA molecule, followed by (2) termination at the target nucleic acid sequence, and any of these may, in some embodiments, be influenced by the improved affinity of the CasX protein to the target nucleic acid sequence, thereby improving the function of the CasX variant protein compared to the reference CasX protein.

[0204] In some embodiments, CasX variant proteins with improved target nucleic acid sequence affinity have increased overall affinity to DNA. In some embodiments, CasX variant proteins with improved target nucleic acid affinity have increased affinity to specific PAM sequences other than the standard TTC PAM recognized by the reference CasX protein of SEQ ID NO: 1 or 2, including binding affinity to PAM sequences selected from the group consisting of TTC, ATC, GTC, and CTC. While we do not wish to be bound by theory, these protein variants have an increased ability to interact more strongly with the entire DNA and to approach and edit sequences within the target DNA due to their ability to bind to additional PAM sequences beyond the ability of wild-type CasX, thereby enabling a more efficient search process for the CasX protein against the target sequence. The higher overall affinity to DNA also increases the frequency with which the CasX protein can efficiently initiate and terminate the binding and unwinding steps, thereby facilitating target strand entry and R-loop formation, and ultimately facilitating cleavage of the target nucleic acid sequence.

[0205] While we do not wish to be bound by theory, amino acid changes in the NTSBD that increase the efficiency of unwinding or capturing unwinded non-target DNA strands may increase the affinity of the CasX variant protein to target DNA. Alternatively, or in addition, amino acid changes in the NTSBD that increase the NTSBD's ability to stabilize DNA during unwinding may increase the affinity of the CasX variant protein to target DNA. Alternatively, or in addition, amino acid changes in the OBD may increase the affinity of the CasX variant protein to the protospacer adjacent motif (PAM), thereby increasing the affinity of the CasX variant protein to the target nucleic acid sequence. Alternatively, or in addition, amino acid changes in the helical I and / or II, RuvC, and TSL domains that increase the affinity of the CasX variant protein to the target nucleic acid strand may increase the affinity of the CasX variant protein to the target nucleic acid sequence.

[0206] In some embodiments, the CasX variant protein has increased binding affinity to a target nucleic acid sequence compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3. In some embodiments, the affinity of the CasX variant protein of this disclosure to a target nucleic acid molecule is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100 times compared to the reference CasX protein.

[0207] In some embodiments, the CasX variant protein has improved binding affinity to the non-target strand of the target nucleic acid. As used herein, the term “non-target strand” refers to a strand of DNA target nucleic acid sequence that does not form a Watson-Crick base pair with the targeting sequence of gNA and is complementary to the target strand.

[0208] Methods for measuring the affinity of a CasX protein (such as a reference or variant) to a target nucleic acid molecule may include electrophoretic mobility shift assays (EMSA), filter binding, isothermal calorimetry (ITC), and surface plasmon resonance (SPR), fluorescence polarization, and biolayer interferometry (BLI). Further methods for measuring the affinity of a CasX protein to a target include in vitro biochemical assays that measure DNA cleavage events over time.

[0209] In some embodiments, CasX variant proteins with higher affinity for the target nucleic acid sequence can cleave the target nucleic acid sequence more rapidly than reference CasX proteins that do not have increased affinity for the target nucleic acid sequence.

[0210] In some embodiments, the CasX variant protein is catalytically dead (dCasX). In some embodiments, the disclosure provides an RNP comprising a catalytically dead CasX protein that retains the ability to bind to target DNA. An exemplary catalytically dead CasX variant protein comprises one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, the catalytically dead CasX variant protein comprises substitutions at residues 672, 769, and / or 935 of SEQ ID NO: 1. In some embodiments, the catalytically dead CasX variant protein comprises substitutions at D672A, E769A, and / or D935A in the reference CasX protein of SEQ ID NO: 1. In some embodiments, the catalytically dead CasX protein comprises substitutions at residues 659, 765, and / or 922 of SEQ ID NO: 2. In some embodiments, the catalytically dead CasX protein comprises substitutions at D659A, E756A, and / or D922A in the reference CasX protein of SEQ ID NO: 2. In further embodiments, the catalytically dead reference CasX protein includes the deletion of all or part of the RuvC domain of the reference CasX protein.

[0211] In some embodiments, the improved affinity of the CasX variant protein to DNA also improves the function of the catalytically inactive version of the CasX variant protein. In some embodiments, the catalytically inactive version of the CasX variant protein contains one or more mutations in the DED motif at RuvC. The catalytically dead CasX variant protein can be used for base editing or epigenetic modification in some embodiments. The higher the affinity to DNA, the more the catalytically dead CasX variant protein can find their target DNA faster, remain bound to the target DNA for longer periods, bind to the target DNA in a more stable manner, or a combination thereof, compared to catalytically active CasX, thereby improving the function of the catalytically dead CasX variant protein.

[0212] o. Improved specificity for target sites In some embodiments, CasX variant proteins exhibit improved specificity for target nucleic acid sequences compared to the reference CasX protein. "Specificity," as interchangeably referred to herein as "target specificity," refers to the degree to which the CRISPR / Cas system ribonucleoprotein complex cleaves off-target sequences that are similar to, but not identical to, the target nucleic acid sequence. For example, a CasX variant RNP with higher specificity exhibits reduced off-target cleavage of the sequence compared to the reference CasX protein. The specificity of CRISPR / Cas system proteins, and the reduction of potentially harmful off-target effects, can be critical to achieving a therapeutic index acceptable for use in mammalian subjects.

[0213] In some embodiments, the CasX variant protein has improved specificity for target sites within the target nucleic acid sequence that are complementary to the targeting sequence of the gNA.

[0214] While we do not wish to be constrained by theory, amino acid changes in the helical I and II domains that increase the specificity of the CasX variant protein to a target nucleic acid chain may also increase the specificity of the CasX variant protein to the entire target nucleic acid sequence. In some embodiments, amino acid changes that increase the specificity of the CasX variant protein to a target nucleic acid sequence may also result in a decrease in the affinity of the CasX variant protein to DNA.

[0215] Methods for testing the target specificity of CasX proteins (such as variants or references) may include guide and circularization for in vitro reporting of cleavage effects by sequencing (CIRCLE-seq), or similar methods. Briefly, in the CIRCLE-seq technique, genomic DNA is sheared and circularized by ligation of a stem-loop adapter, which is then nicked at the stem-loop region to expose a 4-nucleotide palindromic overhang. This is followed by intramolecular ligation and degradation of the remaining linear DNA. Subsequently, the circular DNA molecule containing the CasX cleavage site is linearized with CasX, the adapter adapter is ligated to the exposed end, and high-throughput sequencing is performed to generate paired-end reads containing information about off-target sites. Additional assays that can be used to detect off-target events, and consequently CasX protein specificity, include mismatch detection nuclease assays and next-generation sequencing (NGS) assays used to detect and quantify indels (insertions and deletions) formed at selected off-target sites. An example mismatch detection assay involves a nuclease assay in which genomic DNA from cells treated with CasX and sgRNA is PCR amplified, denatured, and re-hybridized to form heteroduplex DNA containing one wild-type strand and one strand with an indel. Mismatches are recognized and cleaved by mismatch detection nucleases such as Surveyor nuclease or T7 endonuclease I.

[0216] p.DNA rewinding In some embodiments, CasX variant proteins possess improved DNA unwinding ability compared to the reference CasX protein. Previous studies have shown that insufficient dsDNA unwinding impairs or hinders the ability of CRISPR / Cas system proteins AnaCas9 or Cas14s to cleave DNA. Therefore, while we do not wish to be bound by any theory, the increased DNA cleavage activity by some of the CasX variant proteins in this disclosure is likely at least partially attributable to an increased ability to locate and unwind dsDNA at target sites.

[0217] While we do not wish to be constrained by theory, it is thought that amino acid changes in the NTSB domain can produce CasX variant proteins with enhanced DNA unwinding properties. Alternatively, or in addition, amino acid changes in the OBD or helical domain region that interacts with PAM can also produce CasX variant proteins with enhanced DNA unwinding properties.

[0218] Methods for measuring the ability of CasX proteins (variants or references, etc.) to unwind DNA include, but are not limited to, in vitro assays that observe an increase in the rate of a dsDNA target in fluorescence polarization or biolayer interferometry.

[0219] q. Catalytic activity The CasX:gNA-based ribonucleoprotein complexes disclosed herein include a reference CasX protein or variant that binds to and cleaves a target nucleic acid sequence. In some embodiments, the CasX variant protein has improved catalytic activity compared to the reference CasX protein. While we do not wish to be constrained by theory, in some cases, cleavage of the target strand may be a limiting factor for the Cas12-like molecule in causing dsDNA cleavage. In some embodiments, the CasX variant protein improves the bending and cleavage of the target strand of DNA, resulting in improved overall efficiency of dsDNA cleavage by the CasX ribonucleoprotein complex.

[0220] In some embodiments, the CasX variant protein has increased nuclease activity compared to the reference CasX protein. Variants with increased nuclease activity can be generated, for example, by amino acid changes in the RuvC nuclease domain. In some embodiments, the CasX variant includes a nuclease domain having nickase activity. In the above, the CasX nickase in the CasX:gNA system produces a single-strand break within 10 to 18 nucleotides at the 3' end of the PAM site on the non-target strand. In other embodiments, the CasX variant includes a nuclease domain having double-strand break activity. In the above, the CasX in the CasX:gNA system produces double-strand breaks within 18 to 26 nucleotides at the 5' end of the PAM site on the target strand and within 10 to 18 nucleotides at the 3' end of the non-target strand. Nuclease activity can be assayed by various methods, including the method described in the examples. In some embodiments, the CasX variant has a K that is at least 2 times, or at least 3 times, or at least 4 times, or at least 5 times, or at least 6 times, or at least 7 times, or at least 8 times, or at least 9 times, or at least 10 times larger than the reference CasX. cleave It has a constant.

[0221] In some embodiments, CasX variant proteins have increased target chain loading for double-strand cleavage compared to reference CasX. Variants with increased target chain loading activity can be generated, for example, by amino acid changes in the TLS domain.

[0222] While we do not wish to be constrained by theory, amino acid changes in the TSL domain may result in CasX variant proteins with improved catalytic activity. Alternatively, or in addition, amino acid changes around the RNA:DNA duplex binding channel may also improve the catalytic activity of CasX variant proteins.

[0223] In some embodiments, the CasX variant protein has increased incidental cleavage activity compared to the reference CasX protein. As used herein, “incidental cleavage activity” refers to recognition of the target nucleic acid sequence and additional non-targeted cleavage of the nucleic acid after cleavage. In some embodiments, the CasX variant protein has decreased incidental cleavage activity compared to the reference CasX protein.

[0224] In some embodiments, for example, in applications where cleavage of a target nucleic acid sequence is not a desirable outcome, improving the catalytic activity of the CasX variant protein includes modifying, reducing, or inactivating the catalytic activity of the CasX variant protein. In some embodiments, a ribonucleoprotein complex containing the dCasX variant protein binds to the target nucleic acid sequence but does not cleave the target nucleic acid.

[0225] In some embodiments, the CasX ribonucleoprotein complex containing the CasX variant protein binds to target DNA but generates single-strand nicks in the target DNA. In some embodiments, particularly in embodiments where the CasX protein is a nickase, the CasX variant protein has a reduced target chain load for single-strand nicking. Variants with a reduced target chain load can be generated, for example, by amino acid changes in the TSL domain.

[0226] Exemplary methods for characterizing the catalytic activity of the CasX protein may include, but are not limited to, in vitro cleavage assays, including the methods of the following examples. In some embodiments, electrophoresis of DNA products on an agarose gel can be used to investigate the dynamics of strand breaks.

[0227] Affinity for r.C9orf72 target DNA and RNA In some embodiments, a ribonucleoprotein complex containing a reference CasX protein or a CasX variant protein binds to target C9orf72 DNA and cleaves the target nucleic acid sequence. In some embodiments, the ribonucleoprotein complex creates a double-strand break in the target nucleic acid. In other embodiments, the ribonucleoprotein complex creates a single-strand break in the target nucleic acid. In some embodiments, a variant of the reference CasX protein increases the specificity of the CasX variant protein to target C9orf72 RNA and increases the activity of the CasX variant protein to target RNA compared to the reference CasX protein. For example, the CasX variant protein may exhibit increased binding affinity to target RNA or increased cleavage of target RNA compared to the reference CasX protein. In some embodiments, a ribonucleoprotein complex containing the CasX variant protein binds to and / or cleaves target RNA. In some embodiments, the CasX variant has at least about 2 to about 10 times increased binding affinity to C9orf72 target RNA compared to the reference protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3.

[0228] s. Combinations of mutations This disclosure provides CasX variants, which are combinations of mutations derived from distinct CasX variant proteins. In some embodiments, any variant for any domain described herein can be combined with any other variant described herein. In some embodiments, any variant within any domain described herein can be combined with any other variant described herein within the same domain. Different amino acid combinations can, in some embodiments, produce novel optimized variants in which the function is further improved by the combination of amino acid changes. In some embodiments, the effect of combining amino acid changes on CasX protein function is linear. As used herein, a linear combination means a combination in which the effect on function is equal to the sum of the individual effects of each amino acid change when assayed individually. In some embodiments, the effect of combining amino acid changes on CasX protein function is synergistic. As used herein, a synergistic combination means a combination in which the effect on function is greater than the sum of the individual effects of each amino acid change when assayed individually. In some embodiments, combining amino acid changes produces CasX variant proteins in which two or more functions of the CasX protein are improved compared to a reference CasX protein.

[0229] t.CasX fusion protein In some embodiments, this disclosure provides a CasX protein comprising a heterologous protein fused to CasX. In some cases, CasX is a reference CasX protein. In other cases, CasX is a CasX variant of one of the embodiments described herein.

[0230] In some embodiments, a CasX variant protein is fused to one or more proteins or their domains having different desired activities, resulting in a fusion protein. For example, in some embodiments, a CasX variant protein is fused to a protein (or its domain) that inhibits transcription, modifies a target nucleic acid sequence, or modifies a polypeptide associated with a nucleic acid (e.g., histone modification).

[0231] In some embodiments, a heterologous polypeptide (or heterologous amino acid, such as a cysteine ​​residue or a non-natural amino acid) can be inserted at one or more positions within the CasX protein to generate a CasX fusion protein. In other embodiments, a cysteine ​​residue can be inserted at one or more positions within the CasX protein, followed by the conjugation of the heterologous polypeptide described below. In some alternative embodiments, the heterologous polypeptide or heterologous amino acid can be added to the N-terminus or C-terminus of a reference or CasX variant protein. In other embodiments, the heterologous polypeptide or heterologous amino acid can be inserted into the sequence of the CasX protein.

[0232] In some embodiments, the reference CasX or variant fusion protein retains RNA guide sequence-specific target nucleic acid binding and cleavage activity. In some cases, the reference CasX or variant fusion protein has (retains) 50% or more of the activity (e.g., cleavage and / or binding activity) of the corresponding reference CasX or variant protein without heterologous protein insertion. In some cases, the reference CasX fusion or CasX variant fusion protein retains at least about 60%, or at least about 70%, at least about 80%, at least about 90%, at least about 92%, at least about 95%, at least about 98%, or at least about 100% of the activity (e.g., cleavage and / or binding activity) of the corresponding CasX protein without heterologous protein insertion.

[0233] In some cases, the reference CasX or variant fusion protein retains (has) target nucleic acid binding activity compared to the activity of the CasX protein without heterologous amino acids or heterologous polypeptide insertions. In some cases, the reference CasX or variant fusion protein retains at least about 60%, or at least about 70%, at least about 80%, at least about 90%, at least about 92%, at least about 95%, at least about 98%, or at least about 100% of the binding activity of the corresponding CasX protein without heterologous protein insertions.

[0234] In some cases, a reference CasX or variant fusion protein retains (has) target nucleic acid binding and / or cleavage activity compared to the activity of a parent CasX protein without heterologous amino acids or polypeptide insertions. For example, in some cases, a reference CasX or variant fusion protein has (retains) 50% or more of the binding and / or cleavage activity of the corresponding parent CasX protein (CasX protein without insertions). For example, in some cases, a reference CasX or variant fusion protein has (retains) 60% or more (70% or more, 80% or more, 90% or more, 92% or more, 95% or more, 98% or more, or 100%) of the binding and / or cleavage activity of the corresponding parent CasX protein (CasX protein without insertions). Methods for measuring the cleavage and / or binding activity of CasX proteins and / or CasX fusion proteins are known to those skilled in the art, and any simple method can be used.

[0235] Various heterologous polypeptides are suitable for inclusion in the reference CasX or CasX variant fusion proteins of this disclosure. In some cases, the fusion partner can regulate the transcription of target DNA (e.g., inhibit transcription, increase transcription). For example, in some cases, the fusion partner is a transcription-inhibiting protein (or protein-derived domain) (e.g., a protein that functions by transcription repressors, recruitment of transcription-inhibiting proteins, modification of target DNA such as methylation, recruitment of DNA modifiers, regulation of histones associated with target DNA, recruitment of histone modifiers such as those that modify histone acetylation and / or methylation). In some cases, the fusion partner is a transcription-increasing protein (or protein-derived domain) (e.g., a protein that acts by transcription activators, recruitment of transcription-activating proteins, modification of target DNA such as demethylation, recruitment of DNA modifiers, regulation of histones associated with target DNA, recruitment of histone modifiers such as those that modify histone acetylation and / or methylation).

[0236] In some cases, the fusion partner has enzymatic activity that modifies the target nucleic acid sequence, such as nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, dismutase activity, alkylation activity, depurine activity, oxidation activity, pyrimidine dimerization activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, or glycosylase activity.

[0237] In some cases, the fusion partner has enzymatic activity that modifies polypeptides associated with the target nucleic acid (e.g., histones) (e.g., methyltransferase activity, demethylase activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, or demyristoylation activity).

[0238] Examples of proteins (or fragments thereof) that can be used as fusion partners to increase transcription include transcription activators such as VP16, VP64, VP48, VP160, p65 subdomains (e.g., derived from NFκB), and EDLL activating domains and / or TAL activating domains (e.g., for activity in plants); histone lysine methyltransferases such as SET domain-containing 1A, histone lysine methyltransferase (SET1A), SET domain-containing 1B, histone lysine methyltransferase (SET1B), lysine methyltransferase 2A (MLL1-5), ASCL1 (ASH1), achaete-scute family bHLH transcription factor 1 (ASH1), SET and MYND domain-containing 2 (SYMD2), nuclear receptor-binding SET domain protein 1 (NSD1); lysine demethylase 3A (JHDM2a) / lysine-specific demethylase 3 Histone lysine demethylases such as B (JHDM2b), lysine demethylase 6A (UTX), lysine demethylase 6B (JMJD3); histone acetyltransferases such as lysine acetyltransferase 2A (GCN5), lysine acetyltransferase 2B (PCAF), CREB-binding protein (CBP), E1A-binding protein p300 (p300), TATA-box-binding dump protein-related factor 1 (TAF1), lysine acetyltransferase 5 (TIP60 / PLIP), lysine acetyltransferase 6A (MOZ / MYST3), lysine acetyltransferase 6B (MORF / MYST4), SRC proto-oncogene, non-receptor tyrosine kinase (SRC1), nuclear receptor coactivator 3 (ACTR), MYB-binding protein 1a (P160), clock circadian regulator (CLOCK); and Ten-Eleven translocation (Ten-Eleven Examples of DNA demethylases include, but are not limited to, translocation (TET) dioxygenase 1 (TET1CD), tet methylcytosine dioxygenase 1 (TET1), demeter (DME), demeter-like 1 (DML1), demeter-like 2 (DML2), and protein ROS1 (ROS1).

[0239] Examples of proteins (or fragments thereof) that can be used as fusion partners to reduce transcription include transcriptional repressors such as Kruppel-associated boxes (KRAB or SKD); KOX1 repressor domains; Mad mSIN3 interaction domains (SID); ERF repressor domains (ERD), SRDX repressor domains (e.g., for repression in plants); PR / SET domain-containing proteins (Pr-SET7 / 8), histone lysine methyltransferases such as lysine methyltransferase 5B (SUV4-20H1), PR / SET domain 2 (RIZ1); lysine demethylase 4A (JMJD2A / JHDM3A), lysine demethylase 4B (JMJD2B), lysine demethylase 4C (JMJD2C / GASC1), lysine demethylase Histone lysine demethylases such as -ase 4D (JMJD2D), lysine demethylase 5A (JARID1A / RBP2), lysine demethylase 5B (JARID1B / PLU-1), lysine demethylase 5C (JARID1C / SMCX), lysine demethylase 5D (JARID1D / SMCY); histone lysine deacetylases such as histone deacetylase 1 (HDAC1), HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, sirtuin 1 (SIRT1), SIRT2, HDAC11; HhaI DNA methylases such as DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), methyltransferase 1 (MET1), S-adenosyl-L-methionine-dependent methyltransferase superfamily protein (DRM3) (plant), DNA cytosine methyltransferase MET2a (ZMET2), chromometilase 1 (CMT1), chromometilase 2 (CMT2) (plant); and peripheral mobilization elements such as lamin A and lamin B.

[0240] In some cases, the fusion partner possesses enzymatic activity that modifies the target nucleic acid sequence (e.g., ssRNA, dsRNA, ssDNA, dsDNA). Examples of enzymatic activity that may be provided by a fusion partner include nuclease activity, such as that provided by restriction enzymes (e.g., FokI nuclease); methyltransferase activity, such as that provided by methyltransferases (e.g., HhaI DNA m5c-methyltransferase (M.HhaI), DNA methyltransferase 1 (DNMT1), DNA methyltransferase 3a (DNMT3a), DNA methyltransferase 3b (DNMT3b), METI, DRM3 (plant), ZMET2, CMT1, CMT2 (plant), etc.); and demethylases (e.g., 10-11 translocation (TET) dioxygenase 1 (TET1)). Examples of demethylase activity, DNA repair activity, DNA damage activity, etc., provided by CD), TET1, DME, DML1, DML2, ROS1, etc.; deamination activity, dismutase activity, alkylation activity, depurination activity, oxidative activity, pyrimidine dimerization activity, etc., provided by deaminase (e.g., cytosine deaminase enzyme, e.g., APOBEC protein such as rat APOBECl); integrase activity, etc., provided by integrase and / or resolverase (e.g., superactive variants of Gin invertase, GinH106Y, etc.; human immunodeficiency virus type 1 integrase (IN); Tn3 resolverase, etc.); transposase activity; recombinase activity, etc., provided by recombinase (e.g., catalytic domain of Gin recombinase); polymerase activity, ligase activity, helicase activity, photolyase activity, and glycosylase activity).

[0241] In some cases, the reference CasX or CasX variant protein of this disclosure is fused to a polypeptide selected from domains for increasing transcription (e.g., VP16 domain, VP64 domain), domains for decreasing transcription (e.g., KRAB domain, e.g., derived from Kox1 protein), core catalytic domains of histone acetyltransferases (e.g., histone acetyltransferase p300), proteins / domains that provide a detectable signal (e.g., fluorescent proteins such as GFP), nuclease domains (e.g., Fokl nuclease), or base editors (e.g., cytidine deaminases such as APOBEC1).

[0242] In some cases, the fusion partner possesses enzymatic activity that modifies proteins associated with the target nucleic acid (e.g., ssRNA, dsRNA, ssDNA, dsDNA) (e.g., histones, RNA-binding proteins, DNA-binding proteins, etc.). Examples of enzymatic activity that may be provided by the fusion partner (modifying proteins related to the target nucleic acid) include histone methyltransferases (HMTs) (e.g., suppressor of variegated 3-9 homolog 1 (SUV39H1, also known as KMT1A), true chromatin histone lysine methyltransferase 2 (G9A, also known as KMT1C and EHMT2), SUV39H2, ESET / SETDB1, SET1A, SET1B, MLL1-5, ASH1, SYMD2, NSD1, DOT) Methyltransferase activity, such as that provided by 1L, Pr-SET7 / 8, SUV4-20H1, EZH2, RIZ1), histone demethylase (e.g., lysine demethylase 1A (KDM1A, also known as LSD1), JHDM2a / b, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, JARID1A / RBP2, JARID1B / PLU-1, JARID1C / SMCX, JARID1D / SMCY, UTX, JMJD3, etc.) Demethylase activity provided by such catalysts, acetyltransferase activity provided by histone acetylase transferases (e.g., human acetyltransferase p300, GCN5, PCAF, CBP, TAF1, TIP60 / PLIP, MOZ / MYST3, MORF / MYST4, HB01 / MYST2, HMOF / MYST1, SRC1, ACTR, P160, CLOCK, etc.), histone deacetylase activity (e.g., HDAC) Examples of deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristoylation activity, and demyristoylation activity are provided by, but are not limited to, those provided by (1, HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, SIRT1, SIRT2, HDAC11, etc.).

[0243] Examples of additional suitable fusion partners include (i) dihydrofolate reductase (DHFR) destabilization domains (for example, to generate a chemically controllable target RNA guide polypeptide or a conditionally active RNA guide polypeptide), and (ii) chloroplast transport peptides.

[0244] Suitable chloroplast transport peptides include MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGR VKCMQVWPPIGKKKFETLSYLPPLTRDSRA (SEQ ID NO: 151), MASMISSSAVTTVSRASRGQSAAMAPFGGLKSMTGFPVRKVNTDITSITSNGGRVKS (SEQ ID NO: 152), MASSMLSSATMVASPAQATMVAPFNGLKSSAAFPATRKANNDITSITSNGGRVNCMQV WPPIEKKKFETLSYLPDLTDSGGRVNC (SEQ ID NO: 153), MAQVSRICNGVQNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIG SELRPLKVMSSVSTAC (SEQ ID NO: 154), and MAQVSRICNGVWNPSLISNLSKSSQRKSPLSVSLKTQQHPRAYPISSSWGLKKSGMTLIG SELRPLKVMSSVSTAC (SEQ ID NO: 155), MAQINNMAQGIQTLNPNSNFHKPQVPKSSSFLVFGSKKLKNSANSMLVLKKDSIFMQLF CSFRISASVATAC (SEQ ID NO: 156), MAALVTSQLATSGTVLSVTDRFRRPGFQGLRPRNPADAALGMRTVGASAAPKQSRKPH RFDRRCLSMVV (SEQ ID NO: 157), MAALTTSQLATSATGFGIADRSAPSSLLRHGFQGLKPRSPAGGDATSLSVTTSARATPKQ QRSVQRGSRRFPSVVVC (SEQ ID NO: 158), MASSVLSSAAVATRSNVAQANMVAPFTGLKSAASFPVSRKQNLDITSIASNGGRVQC (SEQ ID NO: 159), MESLAATSVFAPSRVAVPAARALVRAGTVVPTRRTSSTSGTSGVKCSAAVTPQASPVIS This includes, but is not limited to, RSAAAA (sequence number 160) and MGAAAATSMQSLKFSNRLVPPSRRLSPVPNNVTCNNLPKSAAPVRTVKCCASSWNSTINGAAATTNGASAASS (sequence number 161).

[0245] In some cases, the reference CasX or variant polypeptide of this disclosure may include an endosomal escape peptide. In some cases, the endosomal escape polypeptide comprises the amino acid sequence GLFXALLXLLXSLWXLLLXA (SEQ ID NO: 162), where each X is independently selected from lysine, histidine, and arginine. In some cases, the endosomal escape polypeptide comprises the amino acid sequence GLFHALLHLLHSLWHLLLHA (SEQ ID NO: 163) or HHHHHHHHH (SEQ ID NO: 164).

[0246] Non-exclusive examples of fusion partners for use when targeting ssRNA target nucleic acid sequences include (but are not limited to) splicing factors (e.g., RS domains); protein translation components (e.g., translation initiation, elongation, and / or release factors, e.g., eIF4G); RNA methylases; RNA editing enzymes (e.g., RNA deaminases including A-to-I and / or C-to-U editing enzymes, e.g., adenosine deaminase (ADAR) acting on RNA); helicases; RNA-binding proteins, etc. It is understood that heterologous polypeptides may contain an entire protein, or in some cases, a fragment of a protein (e.g., a functional domain).

[0247] Fusion partners include endonucleases (e.g., RNase III, CRR22 DYW domain, Dicer, and PIN (PilT N-terminal) domains derived from proteins such as SMG5 and SMG6); proteins and protein domains involved in stimulating RNA cleavage (e.g., CPSF, CstF, CFIm, and CFIIm); exonucleases (e.g., XRN-1 or exonuclease T); derdenylases (e.g., HNT3); proteins and protein domains involved in nonsense-mediated RNA degradation (e.g., UPF1, UPF2, UPF3, UPF3b, RNP) SI, Y14, DEK, REF2, and SRm160); proteins and protein domains involved in RNA stabilization (e.g., PABP); proteins and protein domains involved in translational repression (e.g., Ago2 and Ago4); proteins and protein domains involved in translational stimulation (e.g., Staufen); proteins and protein domains involved in (or regulating) translational regulation (e.g., translation factors such as initiation factors, elongation factors, release factors, e.g., eIF4G); proteins and protein domains involved in RNA polyadenylation (e.g., PAP1, GLD-2, and Star-PAP); proteins and protein domains involved in RNA polyuridine (e.g., CI D1 and terminal uridilate transferases); proteins and protein domains involved in RNA localization (e.g., IMP1, ZBP1, She2p, She3p, and Bicaudal-D); proteins and protein domains involved in nuclear retention of RNA (e.g., Rrp6); proteins and protein domains involved in nuclear export of RNA (e.g., TAP, NXF1, THO, TREX, REF, and Aly); proteins and protein domains involved in repression of RNA splicing (e.g., PTB, Sam68, and hnRNP A1); proteins and protein domains involved in stimulation of RNA splicing (e.g., serine / arginine-rich (SR) domains); proteins and protein domains involved in reducing transcription efficiency (e.g., FUS(TLS));The domain may include, but is not limited to, any domain capable of interacting with ssRNA transiently or irreversibly, directly or indirectly, including effector domains selected from the group including proteins and protein domains involved in transcriptional stimulation (e.g., CDK7 and HIV Tat). Alternatively, effector domains may be selected from the group including endonucleases; protein and protein domains capable of stimulating RNA cleavage; exonucleases; deadenylases; protein and protein domains having nonsense-mediated RNA degradation activity; protein and protein domains capable of stabilizing RNA; protein and protein domains capable of repressing translation; protein and protein domains capable of stimulating translation; protein and protein domains capable of regulating translation (e.g., translation factors such as initiation factors, elongation factors, release factors, eIF4G); protein and protein domains capable of polyadenylation of RNA; protein and protein domains capable of polyuridinedation of RNA; protein and protein domains having RNA localization activity; protein and protein domains capable of nuclear retention of RNA; protein and protein domains having RNA nuclear export activity; protein and protein domains capable of repressing RNA splicing; protein and protein domains capable of stimulating RNA splicing; protein and protein domains capable of reducing transcription efficiency; and protein and protein domains capable of stimulating transcription. Another preferred heterologous polypeptide is the PUF RNA-binding domain, described in detail in WO2012 / 068627, which is incorporated entirely herein by reference.

[0248] Some RNA splicing factors that can be used as fusion partners (whole or in fragments) have a modular mechanism with a separate sequence-specific RNA-binding module and a splicing effector domain. For example, members of the serine / arginine-rich (SR) protein family contain a pre-mRNA that promotes exon inclusion and an N-terminal RNA recognition motif (RRM) that binds to an exon splicing enhancer (ESE) in its C-terminal RS domain. As another example, the hnRNP protein hnRNP A1 binds to an exon splicing silencer (ESS) via its RRM domain and inhibits exon inclusion via its C-terminal glycine-rich domain. Some splicing factors can modulate the selective use of splice sites (ss) by binding to a regulatory sequence between two selective sites. For example, ASF / SF2 can recognize an ESE to promote the use of a proximal intron site, while hnRNP A1 can bind to an ESS to shift splicing to the use of a distant intron site. One application of such factors is the generation of ESFs that regulate the alternative splicing of endogenous genes, particularly disease-related genes. For example, Bcl-x pre-mRNA produces two splicing isoforms, each containing two alternative 5' splice sites, encoding proteins with opposite functions. The long splicing isoform Bcl-xL is a potent apoptosis inhibitor expressed in long-lived, terminating cells, upregulated in many cancer cells, and protects cells from apoptotic signaling. The short isoform Bcl-xS is a pro-apoptotic isoform, expressed at high levels in cells with high turnover rates (e.g., differentiating lymphocytes). The ratio of the two Bcl-x splicing isoforms is regulated by multiple cc elements located either in the core-exon region or the exon-extension region (i.e., between the two alternative 5' splice sites). For further examples, see WO2010 / 075303, which is incorporated herein by reference in its entirety.

[0249] Further suitable fusion partners include, but are not limited to, boundary element proteins (e.g., CTCF) (or their fragments), peripheral recruitment-inducing proteins and their fragments (e.g., lamin A, lamin B, etc.), and protein docking elements (e.g., FKBP / FRB, Pill / Abyl, etc.).

[0250] In some cases, the heterologous polypeptide (fusion partner) provides intracellular localization, i.e., the heterologous polypeptide contains intracellular localization sequences (e.g., nuclear localization signals (NLS) for targeting the nucleus, sequences for maintaining the fusion protein outside the nucleus, e.g., nuclear export sequences (NES), sequences for retaining the fusion protein in the cytoplasm, mitochondrial localization signals for targeting mitochondria, chloroplast localization signals for targeting chloroplasts, ER retention signals, etc.). In some embodiments, the target RNA guide polypeptide or a conditionally active RNA guide polypeptide and / or the target CasX fusion polypeptide does not contain an NLS, and therefore the protein does not target the nucleus (this may be advantageous, for example, if the target nucleic acid sequence is RNA present in the cytosol). In some embodiments, the fusion partner can provide tags for facilitating tracking and / or purification (e.g., fluorescent proteins, e.g., green fluorescent protein (GFP), yellow fluorescent protein (YFP), red fluorescent protein (RFP), cyan fluorescent protein (CFP), mCherry, tdTomato, etc.; histidine tags, e.g., 6XHis tag; hemagglutinin (HA) tag; FLAG tag; Myc tag, etc.) (i.e., the heterologous polypeptide is a detectable label).

[0251] In some cases, the reference or CasX variant polypeptide contains (or is fused to) nuclear localization signals (NLS) (e.g., two or more, three or more, four or more, or five or more, six or more, seven or more, or eight or more NLS). Therefore, in some cases, the reference or CasX variant polypeptide contains one or more NLS (e.g., two or more, three or more, four or more, or five or more NLS). In some cases, one or more NLS (two or more, three or more, four or more, or five or more NLS) is located at or near the N-terminus and / or C-terminus (e.g., within 50 amino acids). In some cases, one or more NLS (two or more, three or more, four or more, or five or more NLS) is located at or near the N-terminus (e.g., within 50 amino acids). In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near the C-terminus (e.g., within 50 amino acids). In some cases, one or more NLSs (three or more, four or more, or five or more NLSs) are located at or near both the N-terminus and C-terminus (e.g., within 50 amino acids). In some cases, one NLS is located at the N-terminus and another NLS is located at the C-terminus. In some cases, the reference or CasX variant polypeptide contains (or is fused to) 1 to 10 NLSs (e.g., 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, or 2 to 5 NLSs). In some cases, the reference or CasX variant polypeptide contains (or is fused to) 2 to 5 NLSs (e.g., 2 to 4 or 2 to 3 NLSs).

[0252] Non-limiting examples of NLS include NLS of the SV40 virus large T antigen with the amino acid sequence PKKKRKV (SEQ ID NO: 165), NLS derived from nucleoplasmin (e.g., bifid nucleoplasmin NLS with sequence KRPAATKKAGQAKKKK (SEQ ID NO: 166), c-myc NLS with amino acid sequences PAAKRVKLD (SEQ ID NO: 167) or RQRRNELKRSP (SEQ ID NO: 168), hRNPAl M9 NLS with sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 169), IBB domain sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 170) derived from importin-alpha, myoma T protein sequences VSRKRPRP (SEQ ID NO: 171) and PPKKARED (SEQ ID NO: 172), human p53 sequence PQPKKKPL (SEQ ID NO: 173), and mouse c-abl The sequences of IV SALIKKKKKMAP (SEQ ID NO: 174), influenza virus NS1 DRLRR (SEQ ID NO: 175) and PKQKKRK (SEQ ID NO: 176), hepatitis virus delta antigen RKLKKKIKKL (SEQ ID NO: 177), mouse Mxl protein REKKKFLKRR (SEQ ID NO: 178), human poly(ADP-ribose) polymerase KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 179), steroid hormone receptor (human) glucocorticoid RKCLQAGMNLEARKTKK (SEQ ID NO: 180), Borna disease virus P protein (BDV-P1) PRPRKIPR (SEQ ID NO: 181), hepatitis C virus nonstructural protein (HCV-NS5A) PPRKKRTVV (SEQ ID NO: 182), LEF1 NLSKKKKRKREK (SEQ ID NO: 183), ORF57 Simirae sequence RRPSRPFRKP (sequence number 184), EBVLANA sequence KRPRSPSS (SEQ ID NO: 185), influenza A protein sequence KRGINDRNFWRGENERKTR (SEQ ID NO: 186), human RNA helicase A (RHA) sequence PRPPKMARYDN (SEQ ID NO: 187), nucleolar RNA helicase II sequence KRSFSKAF (SEQ ID NO: 188), TUS protein sequence KLKIKRPVK (SEQ ID NO: 189), importin-alpha related sequence PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 190), HTLV-1 Rex protein derived sequence PKTRRRPRRSQRKRPPT (SEQ ID NO: 191), Caenorhabditis Sequences derived from the EGL-13 protein of elegans: MSRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 192), and sequences KTRRRPRRSQRKRPPT (SEQ ID NO: 193), RRKKRRPRRKKRR (SEQ ID NO: 194), PKKKSRKPKKKSRK (SEQ ID NO: 195), HKKKHPDASVNFSEFSK (SEQ ID NO: 196), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 197), LSPSLSPLLSPSLSPL (SEQ ID NO: 197). Examples include sequences derived from 198), RGKGGKGLGKGGAKRHRK (SEQ ID NO: 199), PKRGRGRPKRGRGR (SEQ ID NO: 200), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 190), PKKKRKVPPPPKKKRKV (SEQ ID NO: 201), the CPV-derived sequence PAKRARRGYKC (SEQ ID NO: 202), the B19-derived sequence KLGPRKATGRW (SEQ ID NO: 203), and the hBOV-derived sequence PRRKREE (SEQ ID NO: 204). Generally, NLS (or multiple NLS) have sufficient strength to drive the accumulation of reference or CasX variant fusion proteins in the nucleus of eukaryotic cells. Detection of accumulation in the nucleus can be performed by any suitable technique. For example, the detectable marker may be fused to the reference or CasX variant fusion protein, thereby making its intracellular location visible. The cell nucleus may be isolated from the cell, and its contents may then be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blotting, or enzyme activity assay. Accumulation within the nucleus can also be determined.

[0253] In some embodiments, the reference or CasX variant fusion protein includes a "protein transduction domain" or PTD (also known as CPP - cell-permeable peptide), which refers to an organic or inorganic compound that facilitates the translocation of proteins, polynucleotides, carbohydrates, or lipid bilayers, micelles, cell membranes, organelle membranes, or vesicle membranes. PTDs bound to other molecules and / or nanoparticles, ranging from small polar molecules to large macromolecules, facilitate membrane translocation of molecules, for example, from the extracellular space to the intracellular space, or from the cytosol to the organelle. In some embodiments, the PTD is covalently bound to the amino terminus of the reference or CasX variant fusion protein. In some embodiments, the PTD is covalently bound to the carboxyl terminus of the reference or CasX variant fusion protein. In some embodiments, the PTD is inserted into the sequence of the reference or CasX variant fusion protein at a preferred insertion site. In some cases, the reference or CasX variant fusion protein contains (is conjugated to, or is fused to) one or more PTDs (e.g., two or more, three or more, or four or more PTDs). In some cases, the PTD contains one or more nuclear localization signals (NLSs).Examples of PTDs include the peptide transduction domain of HIV TAT containing YGRKKRRQRRR (SEQ ID NO: 205), RKKRRQRR (SEQ ID NO: 206), YARAAARQARA (SEQ ID NO: 207), THRLPRRRRRR (SEQ ID NO: 208), and GGRRARRRRRR (SEQ ID NO: 209); polyarginine sequences containing several arginines sufficient for direct entry into cells (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines (SEQ ID NO: 210)); VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); and cleaved human calcitonin peptide (Trehin et al. (2004) Pharm. Research). This includes, but is not limited to, polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); RRQRRTSKLMKR (SEQ ID NO: 211); transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 212); KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 213); and RQIKIWFQNRRMKWKK (SEQ ID NO: 214). In some embodiments, the PTD is an activatable CPP (ACPP) (Aguilera et al. (2009) Integr Biol (Camb) June; 1(5-6):371-381). ACPP contains a polycationic CPP (e.g., Arg9 or "R9") linked to a matching polyanion (e.g., Glu9 or "E9") via a cleavable linker, which reduces the net charge to nearly zero, thereby inhibiting adhesion and uptake to cells. Upon cleavage of the linker, the polyanion is released, locally unmasking polyarginine and its inherent adhesive properties, and thus "activating" the ACPP to traverse the membrane.

[0254] In some embodiments, the reference or CasX variant fusion protein may comprise a CasX protein linked to heterologous amino acids or heterologous polypeptide (heterologous amino acid sequence) inserted internally via a linker polypeptide (e.g., one or more linker polypeptides). In some embodiments, the reference or CasX variant fusion protein may be linked to a heterologous polypeptide (fusion partner) at the C-terminus and / or N-terminus via a linker polypeptide (e.g., one or more linker polypeptides). The linker polypeptide may have any of a variety of amino acid sequences. The protein may be linked by a generally mobile spacer peptide, but other chemical bonds are not excluded. Preferred linkers include polypeptides 4 to 40 amino acids long, or 4 to 25 amino acids long. These linkers are generally produced by coupling them with the protein using synthetic oligonucleotides encoding the linker. Peptide linkers with some degree of mobility can be used. Given that preferred linkers have sequences that generally result in mobile peptides, the linked peptide may have substantially any amino acid sequence. The use of small amino acids such as glycine and alanine is useful for producing mobile peptides. The preparation of such sequences is customary for those skilled in the art. Various different linkers are commercially available and are considered suitable for use. Examples of linker polypeptides include glycine polymer (G)n, glycine-serine polymers (e.g., (GS)n, GSGGSn (SEQ ID NO: 215), GGSGGSn (SEQ ID NO: 216), and GGGSn (SEQ ID NO: 217), where n is at least an integer of 1), glycine-alanine polymers, alanine-serine polymers, glycine-proline polymers, proline polymers, and proline-alanine polymers.Examples of linkers include, but are not limited to, amino acid sequences such as GGSG (SEQ ID NO: 218), GGSGG (SEQ ID NO: 219), GSGSG (SEQ ID NO: 220), GSGGG (SEQ ID NO: 221), GGGSG (SEQ ID NO: 222), GSSSG (SEQ ID NO: 223), GPGP (SEQ ID NO: 224), GGP, PPP, PPAPPA (SEQ ID NO: 225), and PPPGPPP (SEQ ID NO: 226). Those skilled in the art will recognize that the design of a peptide conjugated to any of the above elements may include a fully or partially mobile linker, thereby allowing the linker to include not only a mobile linker but also one or more parts that give a less mobile structure.

[0255] CasX:gNA system and method for modification of the V.C9orf72 gene The CasX proteins, guide nucleic acids, and their variants provided herein are useful for a variety of applications, including therapeutic, diagnostic, and research purposes. To achieve the methods for editing genes described herein, a programmable CasX:gNA system is provided herein. The programmable nature of the CasX:gNA system provided herein allows for precise targeting to achieve a desired effect (such as nicking or cleavage) in one or more regions of a predetermined target nucleic acid sequence that encodes the C9orf72 protein, the C9orf72 regulatory element, a non-coding region of the C9orf72 gene, or both.In some embodiments, the CasX:gNA systems provided herein are at least 60% identical, at least 70% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, and at least 86% identical to any one of the CasX variants shown in Tables 4, 6, 7, 8, or 10, such as SEQ ID NOs: 49-150, 233-235, 238-252, or 272-281. The gNA scaffold contains variant sequences that are at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical, and the gNA scaffold contains any one of the sequences shown in Table 2SEQ ID NOs. 2101 to 2294, or in relation to it At least 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, The gNA includes sequences that are at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical, and the gNA includes one of the target sequences among sequence numbers 309-343, 363-2100, or 2295-21835, or sequences that are 65% identical, at least 70% identical, at least 75% identical, at least 80% identical, at least 85% identical, at least 90% identical, or at least 95% identical to it, and have 15-30 nucleotides.In some embodiments, the gNA targeting sequence hybridizes to a target nucleic acid sequence encoding one or more mutations in the C9orf72 protein of sequence number 227 or 228, or one or more mutations that disrupt the function or expression of the C9orf72 protein. In another embodiment, the gNA targeting sequence hybridizes to a target nucleic acid sequence comprising a sequence that is 5' or 3' to the hexanucleotide repeat GGGGCC or its complement. In yet another embodiment, the gNA targeting sequence hybridizes to a target nucleic acid sequence comprising a regulatory element of the C9orf72 gene. In some embodiments, the gNA targeting sequence has a sequence that hybridizes with the C9orf72 exon sequence. In some embodiments, the gNA targeting sequence has a sequence that hybridizes with the C9orf72 intron sequence. In some embodiments, the gNA targeting sequence has a sequence that hybridizes with intron 1 of the C9orf72 gene. In some embodiments, the targeting sequences of the gNAs have sequences that hybridize with C9orf72 intron-exon junction sequences, C9orf72 regulatory elements, C9orf72 coding regions, C9orf72 non-coding regions, or combinations thereof. In some embodiments of the method, the gNAs are chemically modified. In other embodiments, the disclosure provides one or more polynucleotides encoding the aforementioned CasX variant protein and gNAs. In some cases, the CasX:gNA system further comprises a donor template nucleic acid, which can be inserted by the host cell's HDR or HITI repair mechanism to knock down or knock out the C9orf72 gene, or in other cases, the mutation can be corrected, for example, by deletion of a variant HRS repeat and insertion of an HRS having 10 to 30 repeats of the GGGGCC sequence.

[0256] In some embodiments, the CasX:gNA system provided herein comprises a CasX protein and gNA, or one or more polynucleotides encoding a CasX protein and gNA, wherein the targeting sequence of gNA is complementary to and can hybridize with a target nucleic acid sequence encoding a C9orf72 protein, a C9orf72 regulatory element, a non-coding region of the C9orf72 gene (e.g., intron 1), a sequence that cleaves these regions, or a sequence that is complementary to them. In certain embodiments, the targeting sequence of gNA is complementary to a sequence within the HRS or a region that is 5' or 3' of the HRS, and can therefore hybridize with it. In another particular embodiment, the targeting sequence of gNA is complementary to a sequence within the promoter of C9orf72, and can therefore hybridize with it. Exemplary but non-limiting targeting sequences that may be used to target the C9orf72 HRS include SEQ ID NOs. 309-343, shown in Table 15. In some embodiments, the targeting sequence includes sequences of sequence numbers 309-343. In some embodiments, the CasX:gNA system includes two targeting sequences selected from sequence numbers 309-343, and the two targeting sequences are not the same. In some embodiments, the CasX:gNA system includes two targeting sequences, the first of which includes sequence number 310, and the second targeting sequence is selected from the group consisting of sequence numbers 321-324. In some embodiments, the CasX:gNA system includes two targeting sequences, the first of which includes sequence number 319, and the second targeting sequence is selected from the group consisting of sequence numbers 321-325. In some embodiments, the CasX:gNA system includes two targeting sequences, the first of which includes sequence number 320, and the second targeting sequence is selected from the group consisting of sequence numbers 321-325.In one embodiment, the two targeted sequences include sequence numbers 310 and 321, 310 and 322, 310 and 323, 310 and 324, 319 and 321, 319 and 322, 319 and 323, 319 and 324, 319 and 325, 320 and 321, 320 and 322, 320 and 323, 320 and 324, or 320 and 325.

[0257] The introduction of a recombinant expression vector containing a sequence encoding the CasX:gNA system of this disclosure (and optionally a donor template sequence) into cells under in vitro conditions can occur in any suitable medium and any suitable culture conditions that promote cell viability and CasX:gNA production. T...

Claims

1. A system comprising a CasX variant protein and a guide ribonucleic acid (gRNA), wherein the CasX variant protein comprises a sequence having at least 90% sequence identity with SEQ ID NO: 138, exhibits one or more improved properties compared to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, the CasX variant protein can form a ribonucleoprotein complex (RNP) with the gRNA, the gRNA comprises a targeting sequence complementary to a target nucleic acid sequence including the chromosome 9 open reading frame 72 (C9orf72) gene, and the CasX variant protein comprises a non-targeting chain binding (NTSB) domain at amino acids 101-191 of SEQ ID NO: 1 and a helical 1b domain at amino acids 192-332 of SEQ ID NO:

1.

2. The system according to claim 1, wherein the C9orf72 gene contains one or more mutations, or the C9orf72 gene mutation contains more than 30, more than 100, more than 500, more than 700, more than 1000, or more than 1600 copies of the hexanucleotide repeat sequence GGGGCC in the hexanucleotide repeat sequence expansion (HRS).

3. (a) The gRNA is a single-molecule guide RNA (gRNA) comprising a scaffold sequence and a targeting sequence, wherein the scaffold sequence comprises the scaffold stem-loop sequence CCAGCGACUAUGUUCGUAGUGG (SEQ ID NO: 32), or a sequence having one or two mismatches thereto; (b) The targeting sequence of the gRNA includes a sequence selected from the group consisting of SEQ ID NOs: 309-343, 363-2100, and 2295-21835, or a sequence having at least 90% identity thereto; (c) The gRNA has a scaffold containing a sequence that has at least 90% sequence identity with respect to the sequence of Sequence ID No. 2238; and / or (d) The gRNA is chemically modified, The system according to claim 1 or 2.

4. The system according to any one of claims 1 to 3, further comprising a second gRNA, wherein the second gRNA has a targeting sequence that is complementary to a different or overlapping portion of the target nucleic acid sequence compared to the targeting sequence of the gRNA.

5. The system according to claim 4, wherein the first gRNA is complementary to the sequence that is 5' of the HRS, and the second gRNA is complementary to the sequence that is 3' of the HRS.

6. The CasX variant protein further comprises one or more nuclear localization signals (NLS), wherein the one or more NLS are selected from the group of sequences consisting of SEQ ID NOs. 165 to 204, and here, (i) One or more NLSs are expressed at or near the C-terminus of the CasX variant protein, (ii) The one or more NLSs are expressed at or near the N-terminus of the CasX variant protein; and / or (iii) The CasX variant protein comprises at least two NLSs at or near the N-terminus and C-terminus of the CasX variant protein, The system according to any one of claims 1 to 5.

7. The system according to claim 4, wherein the CasX variant protein and the gRNA are complexed as a ribonucleoprotein complex (RNP).

8. The RNP comprising the CasX variant protein and the gRNA variant exhibits at least one improved characteristic compared to a reference CasX protein of any one of SEQ ID NOs: 1, SEQ ID NOs: 2, and SEQ ID NOs: 3 and a gRNA of any one of SEQ ID NOs: 4 to 16, wherein the improved characteristics include improved folding of the CasX variant protein, improved binding affinity to the guide ribonucleic acid (gRNA), improved binding affinity to the target DNA, improved ability to utilize a broader spectrum of one or more PAM sequences including ATC, CTC, GTC, or TTC in editing the target DNA, improved unwinding of the target DNA, increased editing activity, improved editing efficiency, improved editing specificity, and increased nucleation The system according to claim 7, wherein the improved properties of the CasX variant protein are selected from the group consisting of ase activity, increased target strand loading for double-strand breaks, decreased target strand loading for single-strand nicking, decreased off-target breaks, improved binding of non-target DNA strands, improved protein stability, improved protein solubility, improved protein:gRNA complex (RNP) stability, improved protein:gRNA complex solubility, improved protein yield, improved protein expression, and improved fusion properties, wherein the improved properties of the RNP of the CasX variant protein are improved by at least about 1.1 to about 100,000 times compared to the gRNA comprising the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3 and any one of SEQ ID NOs: 4 to 16.

9. (a) If any one of the PAM sequences TTC, ATC, GTC, or CTC is located at the 5' first nucleotide relative to a non-target strand sequence identical to the targeting sequence of the gRNA, the RNP comprising the CasX variant protein and the gRNA variant exhibits higher editing efficiency and / or binding of the target sequence in the target DNA in a cell assay system compared to the editing efficiency and / or binding of an RNP comprising a reference CasX protein and a reference gRNA in an equivalent assay system, where (i) The PAM sequence is TTC, and the targeting sequence of the gRNA includes a sequence selected from the group consisting of SEQ ID NOs: 5427 to 12893; (ii) The PAM sequence is an ATC, and the targeting sequence of the gRNA includes a sequence selected from the group consisting of SEQ ID NOs: 363-2100 and 2295-5426; (iii) The PAM sequence is a CTC, and the targeting sequence of the gRNA includes a sequence selected from the group consisting of SEQ ID NOs: 16203 to 21835; or (iv) The PAM sequence is a GTC, and the targeting sequence of the gRNA includes a sequence selected from the group consisting of SEQ ID NOs: 12894 to 16202. and / or (b) The system according to claim 8, wherein the RNP has a percentage of cleavage-competent RNPs that is at least 5%, at least 10%, at least 15%, or at least 20% higher than the RNP of the reference CasX protein and the reference gRNA containing any one of SEQ ID NOs. 4 to 16, or the RNP has a cleavage rate at least twice, at least five times, or at least ten times higher in an in vitro assay compared to the RNP of the reference CasX protein of SEQ ID NOs. 1 to 3.

10. The system according to any one of claims 1 to 9, wherein the CasX variant protein comprises a nuclease domain having double-strand cleavage activity.

11. (a) gRNA containing a targeting sequence complementary to the target nucleic acid sequence containing the chromosome 9 open reading frame 72 (C9orf72) gene, and optionally, (b) A CasX variant protein comprising a sequence having at least 90% sequence identity with SEQ ID NO: 138, exhibiting one or more improved properties with respect to the reference CasX protein of SEQ ID NO: 1, SEQ ID NO: 2, or SEQ ID NO: 3, wherein the CasX variant protein can form a ribonucleoprotein complex (RNP) with the gRNA described in (a), and wherein the CasX variant protein further comprises one or more nuclear localization signals (NLS), A nucleic acid containing a sequence that codes for something.

12. A vector comprising nucleic acid according to claim 11, wherein the vector is selected from the group consisting of nanoparticles that mediate nucleic acid delivery to cells, retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, herpes simplex virus (HSV) vectors, plasmids, minicircles, nanoplasmids, and RNA vectors.

13. The vector according to claim 12, wherein the vector is an AAV vector, where the AAV vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44.9, AAV-Rh74, or AAVRh10.

14. A host cell comprising the vector according to claim 12 or 13, wherein the host cell is selected from the group consisting of BHK, HEK293, HEK293T, NS0, SP2 / 0, YO myeloma cells, P3X63 mouse myeloma cells, PER, PER. C6, NIH3T3, COS, HeLa, CHO, and yeast cells.

15. A method for modifying the C9orf72 target nucleic acid sequence in a population of cells in vitro or ex vivo, Here, the method described above applies to the cells of the population, a. The system according to any one of claims 1 to 10; b. The nucleic acid according to claim 11; c. The vector according to claim 12 or 13; or d. A combination of two or more of (a) to (c), This includes introducing Here, the C9orf72 gene target nucleic acid sequence of the cell targeted by the first gRNA is modified by the CasX variant protein. The above modification is, (i) to introduce single-strand or double-strand breaks in the target nucleic acid sequence of the cells of the population; and / or (ii) including the insertion, deletion, substitution, duplication, or inversion of one or more nucleotides in the target nucleic acid sequence, And a method by which the expression of HRS is reduced.

16. The method according to claim 15, wherein the cells are eukaryotic cells selected from the group consisting of rodent cells, mouse cells, rat cells, pig cells, primate cells, non-human primate cells, and human cells.

17. The method according to claim 16, wherein the cells are selected from the group consisting of Purkinje cells, prefrontal cortical neurons, motor cortical neurons, hippocampal neurons, cerebellar neurons, upper motor neurons, spinal neurons, spinal motor neurons, glial cells, and astrocytes.

18. The method according to claim 15 or 16, further comprising a second gRNA or a nucleic acid encoding the second gRNA, wherein the second gRNA has a targeting sequence that is complementary to a different or overlapping portion of the target nucleic acid sequence compared to the first gRNA.

19. The method according to claim 18, wherein the first gRNA is complementary to the sequence that is 5' of the HRS, and the second gRNA is complementary to the sequence that is 3' of the HRS.

20. a. The system according to any one of claims 1 to 10; b. The nucleic acid according to claim 11; c. The vector according to claim 12 or 13; or d. A combination of two or more of (a) to (c), A population of cells modified by introducing into cells, wherein the cells are (i) At least 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the modified cells are modified not to express the dipeptide repeat protein (DPR) at a detectable level; (ii) The expression of the functional C9orf72 protein is modified to be increased by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% compared to cells in which the C9orf72 gene has not been modified; or (iii) The mutation in the C9orf72 gene is modified in the modified cells of the population to result in the expression of functional C9orf72 protein by the modified cells. And a population of cells in which the cells are selected from the group consisting of Purkinje cells, prefrontal cortical neurons, motor cortical neurons, hippocampal neurons, cerebellar neurons, upper motor neurons, spinal neurons, spinal motor neurons, glial cells, and astrocytes.

21. A composition for use in a method for treating C9orf72-related disorders in subjects requiring the treatment thereof, wherein the composition is in a therapeutically effective dose, a. The system according to any one of claims 1 to 10; b. The nucleic acid according to claim 11; c. The vector according to claim 12 or 13; or d. The combination, A composition comprising, wherein the method comprises modifying the C9orf72 gene in the target cells, wherein the modification comprises bringing the cells into contact with the composition, and wherein the C9orf72 gene in the cells targeted by the first gRNA is modified by the CasX variant protein.

Citation Information

Patent Citations

  • Materials and methods for treatment of amyotrophic lateral sclerosis and / or frontal temporal lobular degeneration

    WO2017109757A1

  • RNA-guided nucleic acid modifying enzymes and methods of use thereof

    WO2018064371A1

  • Methods of treating amyotrophic lateral sclerosis (ALS)

    WO2018208972A1

  • In vitro isolation and enrichment of nucleic acids using site-specific nucleases

    WO2019030306A1